{"thread":{"id":"15449","subject":"[RFC] origin link for cherry-pick and revert","startedAt":"2008-09-09T13:22:12Z","lastAt":"2008-09-23T13:51:17Z","messageCount":137,"participants":["Stephen R. van den Berg","Paolo Bonzini","Jakub Narebski","Steven Grimm","Jeff King","Junio C Hamano","Shawn O. Pearce","Petr Baudis","Linus Torvalds","Theodore Tso","Dmitry Potapov","Miklos Vajna","Nicolas Pitre","A Large Angry SCM","Sam Vilain","Rogan Dawes","Peter Krefting"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"90212","messageId":"20080909132212.GA25476@cuci.nl","threadId":"15449","inReplyTo":null,"subject":"[RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T13:22:12Z","receivedAt":"2008-09-09T13:22:12Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"I've read and digested the old threads about prior and related links.\nHere's a new proposal which should be able to pass muster, if I read all\nthe relevant suggestions and objections in the old threads:\n\nConsider an origin field as such:\n\ncommit bbb896d8e10f736bfda8f587c0009c358c9a8599\ntree b83f28279a68439b9b044bccc313bbeaa3e973f5\nparent ed0f47a8c431f27e0bd131ea1cf9cabbd580745b\norigin d2b9dff8a08cc2037a7ba0463e90791f07cb49dd\norigin a1184d85e8752658f02746982822f43f32316803 2\nauthor Junio C Hamano <gitster@pobox.com> 1220132115 -0700\ncommitter Junio C Hamano <gitster@pobox.com> 1220153445 -0700\n\nThe definition of the origin field reads as follows:\n\n- There can be an arbitrary number of origin fields per commit.\n  Typically there is going to be at most one origin field per commit.\n\n- At the time of creation, the origin field contains a hash B which refers\n  to a reachable commit pair (B, B~1).  If B has multiple parents and the pair\n  being referred to needs to be e.g. (B, B~2), then the hash is followed by\n  a space and followed by an integer (base10, two in this case),\n  which designates the proper parentnr of B (see: mainline in git\n  cherry-pick/revert).\n\n- In an existing repository gc/prune shall not delete commits being\n  referred to by origin links.\n\n- During fetch/push/pull the full commit including the origin fields is\n  transmitted, however, the objects the origin links are referring to\n  are not (unless they are being transmitted because of other reasons).\n\n- When fetching/pulling it is optionally possible to tell git to\n  actually transmit objects referred to by origin links even if it would\n  otherwise not have done so.\n\n- git cherry-pick/revert allow for the creation of origin links only if\n  the object they are referring to is presently reachable.\n\n- git fsck will traverse origin links, but will stay silent if the\n  object an origin link points to is unreachable (kind of like a shallow\n  repository).\n\n- git rev-list --topo-order will take origin links into account to\n  ensure proper ordering.\n\n- gitk allows for (e.g.) dotted lines to show the origin links.\n\n- git log would show something like:\n\n  commit bbb896d8e10f736bfda8f587c0009c358c9a8599\n  Origin: d2b9dff..53d1589\n  Origin: a1184d8..e596cdd\n  Author: Junio C Hamano <gitster@pobox.com>\n  Date:   Sat Aug 30 14:35:15 2008 -0700\n\n  Note that for easy viewing: git diff d2b9dff..53d1589\n  will show the exact diff the origin link is referring to.\n\n- git log --graph will show a dotted line of somesort just like gitk.\n\n- git blame will follow and use the origin link if the object exists.\n\n- git merge disregards the whole origin field entirely, just like all\n  the rest of git-core.\n\nAnything I missed?\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90214","messageId":"48C67C47.6000107@gnu.org","threadId":"15449","inReplyTo":"20080909132212.GA25476@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-09T13:38:15Z","receivedAt":"2008-09-09T13:38:15Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"> - At the time of creation, the origin field contains a hash B which refers\n>   to a reachable commit pair (B, B~1).  If B has multiple parents and the pair\n>   being referred to needs to be e.g. (B, B~2), then the hash is followed by\n>   a space and followed by an integer (base10, two in this case),\n>   which designates the proper parentnr of B (see: mainline in git\n>   cherry-pick/revert).\n\nWhat about just storing *two* hashes?  This way cherry-pick can store\nB~1..B and revert can store B..B~1.  The two cases can be distinguished\nby checking which commit is an ancestor of which.\n\n> - git cherry-pick/revert allow for the creation of origin links only if\n>   the object they are referring to is presently reachable.\n\nWill cherry-pick -x create origin links?  Also, does the origin link\npropagate through multiple cherry picks?  If not, how can the origin\nobject not be reachable?\n\n> [snip good stuff]\n\ngit cherry will use origin links to mark a commit as present, and will\nonly use patch-ids for commits that have no origin links.  Bonus points\nfor an extra command-line/configuration option to only use origin links:\n\n  --source=default\t  << default: get setting from core.cherrysource\n  --source=patch-id\n  --source=origin\n  --source=origin,patch-id\n\n  core.cherrysource = patch-id\n  core.cherrysource = origin\n  core.cherrysource = origin,patch-id\n\nThanks!\n\nPaolo\n"},{"id":"90216","messageId":"20080909134831.GB25476@cuci.nl","threadId":"15449","inReplyTo":"20080909132212.GA25476@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T13:48:31Z","receivedAt":"2008-09-09T13:48:31Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Stephen R. van den Berg wrote:\n>Anything I missed?\n\nI think I forgot two:\n- git rebase will fixup any origin pointers which point back into the\n  strain being rebased.\n\n- git filter-branch will rewrite origin pointers which point to commits\n  that receive a new hash.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90217","messageId":"20080909140453.GC25476@cuci.nl","threadId":"15449","inReplyTo":"48C67C47.6000107@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T14:04:53Z","receivedAt":"2008-09-09T14:04:53Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Paolo Bonzini wrote:\n>> - At the time of creation, the origin field contains a hash B which refers\n>>   to a reachable commit pair (B, B~1).  If B has multiple parents and the pair\n>>   being referred to needs to be e.g. (B, B~2), then the hash is followed by\n>>   a space and followed by an integer (base10, two in this case),\n>>   which designates the proper parentnr of B (see: mainline in git\n>>   cherry-pick/revert).\n\n>What about just storing *two* hashes?  This way cherry-pick can store\n>B~1..B and revert can store B..B~1.  The two cases can be distinguished\n>by checking which commit is an ancestor of which.\n\nValid point, but consider:\nThe new commit to receive hash A.  The diff between A~1..A and B~1..B\nactually defines the relation.  Revert and cherry-pick are symmetrical\noperations as far as git is concerned since git tracks content.\nSo I'm not quite sure if we actually need this extra information, git\nalready knows it all.\n\n>> - git cherry-pick/revert allow for the creation of origin links only if\n>>   the object they are referring to is presently reachable.\n\n>Will cherry-pick -x create origin links?\n\nI'd propose a cherry-pick -o and revert -o for that.\nI wouldn't want to force the text which -x generates into the commit\nwhen origin links are used.\n\n>  Also, does the origin link\n>propagate through multiple cherry picks?\n\nThe origin link is created point-to-point from the object referenced by\ncherry-pick/revert to the new commit.  The link creation specifically\ndoes not follow any existing origin links.  If you want the origin link\nto point to a deeper origin behind the current, then cherry-pick from\nthere, no need to fake it.\n\n>  If not, how can the origin\n>object not be reachable?\n\nThat can only happen during a fetch/pull, which doesn't use origin links\nto determine transmittability by default.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90232","messageId":"m3zlmhnx1z.fsf@localhost.localdomain","threadId":"15449","inReplyTo":"20080909132212.GA25476@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-09T15:44:43Z","receivedAt":"2008-09-09T15:44:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n\n> I've read and digested the old threads about prior and related links.\n> Here's a new proposal which should be able to pass muster, if I read all\n> the relevant suggestions and objections in the old threads:\n> \n> Consider an origin field as such:\n> \n> commit bbb896d8e10f736bfda8f587c0009c358c9a8599\n> tree b83f28279a68439b9b044bccc313bbeaa3e973f5\n> parent ed0f47a8c431f27e0bd131ea1cf9cabbd580745b\n> origin d2b9dff8a08cc2037a7ba0463e90791f07cb49dd\n> origin a1184d85e8752658f02746982822f43f32316803 2\n> author Junio C Hamano <gitster@pobox.com> 1220132115 -0700\n> committer Junio C Hamano <gitster@pobox.com> 1220153445 -0700\n> \n> The definition of the origin field reads as follows:\n> \n> - There can be an arbitrary number of origin fields per commit.\n>   Typically there is going to be at most one origin field per commit.\n\nI understand that multiple origin fields occur if you do a squash\nmerge, or if you cherry-pick multiple commits into single commit.\nFor example:\n $ git cherry-pick -n <a1>\n $ git cherry-pick    <a2>\n $ git commit --amend        #; to correct commit message\n\nI'm not sure if you plan to automatically add 'origin' field for\nrebase, and for interactive rebase...\n \n> - At the time of creation, the origin field contains a hash B which refers\n>   to a reachable commit pair (B, B~1).  If B has multiple parents and the pair\n>   being referred to needs to be e.g. (B, B~2), then the hash is followed by\n>   a space and followed by an integer (base10, two in this case),\n>   which designates the proper parentnr of B (see: mainline in git\n>   cherry-pick/revert).\n\nI think you wanted to use \"(B, B^2)\", which mean B and second parent\nof B.  B~2 means grandparent of B in the straight line:\n\n      ... <--- B~2 <--- B^1 = B^ = B~1 <--- B\n                                           /\n                          ... <--- B^2 <--/\n\nBesides I very much prefer using 'origin <sha1> <sha2>' (as proposed\nin the neighbouring subthread), which would mean together with\n'parent <parent>' (assuming that there are no other parents; if they\nare it gets even more complicated), that the following is true\n\n  <current> ~= <parent> + (<sha2> - <sha1>),\n\nwhere '<rev1> ~= <rev2>' means that <rev1> is based on <rev2> (perhaps\nwith some fixups, corrections or the like).  Perhaps 'origin' should\nbe then called 'changeset'.\n\nIt would also be easier on implementation to check if\n'origin'/'changeset' weak links are not broken, and to get to know\nwhich commits are to be protected against pruning than your proposal\nof\n\n  origin <\"cousin\" id> [<mainline = parent number>]\n\nwhere <mainline> can be omitted if it is 1 (the default).\n\n\nThis can also lead to replacing\n\n  origin <b> <a>\n  origin <c> <b>\n\nby\n\n  origin <c> <a>\n\nfor squash merge, or squash in rebase interactive.\n\n> - In an existing repository gc/prune shall not delete commits being\n>   referred to by origin links.\n> \n> - During fetch/push/pull the full commit including the origin fields is\n>   transmitted, however, the objects the origin links are referring to\n>   are not (unless they are being transmitted because of other reasons).\n> \n> - When fetching/pulling it is optionally possible to tell git to\n>   actually transmit objects referred to by origin links even if it would\n>   otherwise not have done so.\n>\n> - git fsck will traverse origin links, but will stay silent if the\n>   object an origin link points to is unreachable (kind of like a shallow\n>   repository).\n\nThe above means that it is a 'weak' link, i.e. it is protecting\nagainst pruning (perhaps influenced by some configuration variable),\nbut it is not considered an error for it to be broken.\n\n> - git cherry-pick/revert allow for the creation of origin links only if\n>   the object they are referring to is presently reachable.\n\nErrr... shouldn't objects referenced by 'origin' links be reachable in\norder for \"cherry-pick\" or \"revert\" to succeed?\n\nOn the other hand this leads to the following question: what happens\nif you cherry-pick or revert a commit which has its own 'origin'\nlinks?\n\n> - git rev-list --topo-order will take origin links into account to\n>   ensure proper ordering.\n\nWhat do you mean by that?\n \n> - gitk allows for (e.g.) dotted lines to show the origin links.\n> \n> - git log would show something like:\n> \n>   commit bbb896d8e10f736bfda8f587c0009c358c9a8599\n>   Origin: d2b9dff..53d1589\n>   Origin: a1184d8..e596cdd\n>   Author: Junio C Hamano <gitster@pobox.com>\n>   Date:   Sat Aug 30 14:35:15 2008 -0700\n> \n>   Note that for easy viewing: git diff d2b9dff..53d1589\n>   will show the exact diff the origin link is referring to.\n> \n> - git log --graph will show a dotted line of somesort just like gitk.\n\nThat is I guess the whole and main reason for 'origin' links to exist,\nas having this information in free-form part, i.e. in the commit\nmessage might lead to problems (with parsing and extracting, and\nfinding spurious links).\n \n> - git blame will follow and use the origin link if the object exists.\n\nHmmmm... I'm not sure about that.\n\n> - git merge disregards the whole origin field entirely, just like all\n>   the rest of git-core.\n\nUnless of course one uses more complex merge strategy, which doesn't\ntake into account only endpoints (branches to be merged and merge\nbases), but is also affected in some by history...\n \n> Anything I missed?\n\nHow would git-rebase make use of 'origin' links.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"90236","messageId":"10BD65EF-A44B-4A43-BEB5-0D9231332632@midwinter.com","threadId":"15449","inReplyTo":"m3zlmhnx1z.fsf@localhost.localdomain","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2008-09-09T16:38:50Z","receivedAt":"2008-09-09T16:38:50Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"On Sep 9, 2008, at 8:44 AM, Jakub Narebski wrote:\n> This can also lead to replacing\n>\n> origin <b> <a>\n> origin <c> <b>\n>\n> by\n>\n> origin <c> <a>\n>\n> for squash merge, or squash in rebase interactive.\n\n\nAnd, incidentally, the above representation will potentially mesh well  \nwith svn integration, making it possible to cleanly represent svn 1.5  \nmerge-tracking metadata directly in git.\n\n\n> Unless of course one uses more complex merge strategy, which doesn't\n> take into account only endpoints (branches to be merged and merge\n> bases), but is also affected in some by history...\n\n\nIt does intuitively (but perhaps incorrectly) seem like the origin  \ninformation could be used to make more intelligent decisions about  \nautomatic conflict resolution, if nothing else. Though obviously that  \nmight, as you suggest, be a pretty big departure from the way git  \nmerges currently work.\n\n-Steve\n"},{"id":"90250","messageId":"20080909194354.GA13634@cuci.nl","threadId":"15449","inReplyTo":"m3zlmhnx1z.fsf@localhost.localdomain","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T19:43:54Z","receivedAt":"2008-09-09T19:43:54Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jakub Narebski wrote:\n>\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n>> The definition of the origin field reads as follows:\n\n>> - There can be an arbitrary number of origin fields per commit.\n>>   Typically there is going to be at most one origin field per commit.\n\n>I understand that multiple origin fields occur if you do a squash\n>merge, or if you cherry-pick multiple commits into single commit.\n>For example:\n> $ git cherry-pick -n <a1>\n> $ git cherry-pick    <a2>\n> $ git commit --amend        #; to correct commit message\n\nCorrect.\n\n>I'm not sure if you plan to automatically add 'origin' field for\n>rebase, and for interactive rebase...\n\nThat is not part of the plan so far.\nCan you explain what you would be expecting in the best case?\n\n>> - At the time of creation, the origin field contains a hash B which refers\n>>   to a reachable commit pair (B, B~1).  If B has multiple parents and the pair\n>>   being referred to needs to be e.g. (B, B~2), then the hash is followed by\n>>   a space and followed by an integer (base10, two in this case),\n>>   which designates the proper parentnr of B (see: mainline in git\n>>   cherry-pick/revert).\n\n>I think you wanted to use \"(B, B^2)\", which mean B and second parent\n>of B.  B~2 means grandparent of B in the straight line:\n\nCorrect, sorry about the confusion, I meant B^2 instead of B~2.\n\n>Besides I very much prefer using 'origin <sha1> <sha2>' (as proposed\n>in the neighbouring subthread), which would mean together with\n>'parent <parent>' (assuming that there are no other parents; if they\n>are it gets even more complicated), that the following is true\n\n>  <current> ~= <parent> + (<sha2> - <sha1>),\n\n>where '<rev1> ~= <rev2>' means that <rev1> is based on <rev2> (perhaps\n>with some fixups, corrections or the like).  Perhaps 'origin' should\n>be then called 'changeset'.\n\nThe simplicity sounds inviting.  I'd like to hear from others who have\nmore experience (than I have) with the git vs. changeset paradigms about\nthis.  This allows a bit more flexibility in specifying the origin, the\nquestion is if it's needed.\n\n>It would also be easier on implementation to check if\n>'origin'/'changeset' weak links are not broken, and to get to know\n>which commits are to be protected against pruning than your proposal\n>of\n>  origin <\"cousin\" id> [<mainline = parent number>]\n>where <mainline> can be omitted if it is 1 (the default).\n\nOn the contrary, my current proposal only needs to verify the validity\nof a single commit, changing it like this will require the system to\nverify the validity of two commits.  Given the rareness of the origin\nlinks this will hardly present a problem, but it *does* increase\nthe overhead in checking a bit.\n\n>This can also lead to replacing\n>  origin <b> <a>\n>  origin <c> <b>\n>by\n>  origin <c> <a>\n>for squash merge, or squash in rebase interactive.\n\nOk, *that* is not possible with the original proposal.  This might just\nbe the reason why we'd like to go with the dual-hash link.\n\n>> - git cherry-pick/revert allow for the creation of origin links only if\n>>   the object they are referring to is presently reachable.\n\n>Errr... shouldn't objects referenced by 'origin' links be reachable in\n>order for \"cherry-pick\" or \"revert\" to succeed?\n\nTrue.  But sometimes it's necessary to emphasize the obvious; call it a\npreemptive strike against possible objections to the proposal.\n\n>On the other hand this leads to the following question: what happens\n>if you cherry-pick or revert a commit which has its own 'origin'\n>links?\n\nNothing special.  cherry-pick/revert behave as if the existing origin links\nwere not present in the first place.\n\n>> - git rev-list --topo-order will take origin links into account to\n>>   ensure proper ordering.\n\n>What do you mean by that?\n\nThe order in which commits are listed is defined by the fact that\ndescendent commits are shown before any of their parents.  The presence\nof an origin link will make sure that the current commit will always\nappear *before* the origin-commit it is referring to (if the\norigin-commit is in the displayed set, that is).\n\n>> - git log would show something like:\n\n>>   commit bbb896d8e10f736bfda8f587c0009c358c9a8599\n>>   Origin: d2b9dff..53d1589\n>>   Origin: a1184d8..e596cdd\n>>   Author: Junio C Hamano <gitster@pobox.com>\n>>   Date:   Sat Aug 30 14:35:15 2008 -0700\n\n>>   Note that for easy viewing: git diff d2b9dff..53d1589\n>>   will show the exact diff the origin link is referring to.\n\n>> - git log --graph will show a dotted line of somesort just like gitk.\n\n>That is I guess the whole and main reason for 'origin' links to exist,\n>as having this information in free-form part, i.e. in the commit\n>message might lead to problems (with parsing and extracting, and\n>finding spurious links).\n\nQuite.  Also, having them in a well-defined place will allow for easy\nfixups in case of rebase/filter-branch.\n\n>> - git blame will follow and use the origin link if the object exists.\n\n>Hmmmm... I'm not sure about that.\n\nCare to explain your doubts?\nThe reason I want this behaviour, is because it's all about tracking\ncontent, and that part of the content happens to come from somewhere\nelse, and therefore blame should look there to \"dig deeper\" into it.\n\n>> - git merge disregards the whole origin field entirely, just like all\n>>   the rest of git-core.\n\n>Unless of course one uses more complex merge strategy, which doesn't\n>take into account only endpoints (branches to be merged and merge\n>bases), but is also affected in some by history...\n\nQuite, but that is not a part of the definition of the origin field.\nI can only try and make sure that we have a well-defined, well-behaved\nmechanism in core git.  If someone wants to get creative with the\ninformation presented, by all means, be my guest.\n\n>> Anything I missed?\n\n>How would git-rebase make use of 'origin' links.\n\nAs far as I can imagine, git rebase should alter the origin links during\nrebase if they point to a commit within the strain being rebased.\nAre there any other desirable use cases (for rebase)?\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90251","messageId":"20080909195930.GA2785@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080909194354.GA13634@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-09T19:59:31Z","receivedAt":"2008-09-09T19:59:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 09, 2008 at 09:43:54PM +0200, Stephen R. van den Berg wrote:\n\n> >Besides I very much prefer using 'origin <sha1> <sha2>' (as proposed\n> \n> The simplicity sounds inviting.  I'd like to hear from others who have\n> more experience (than I have) with the git vs. changeset paradigms about\n> this.  This allows a bit more flexibility in specifying the origin, the\n> question is if it's needed.\n\nOne thing to keep in mind is that you are not just proposing some new\nbehavior for a command, but rather a new header for the data structure\nthat we will live with from now until eternity. So I think it makes\nsense to allow the general case even if nobody is generating it yet, if\nthere is some chance that it may be useful for somebody to generate in\nthe future.\n\nAnd yes, you can get _too_ general to the point where your semantics\nbecome meaningless. But I don't think that is the case here. You are\ndefining the origin field as \"by the way, the difference between state X\nand state Y was used to make this commit\". cherry-pick just happens to\nmake Y=X^, but something like rebase could use a series.\n\nAs for \"git vs changeset\": this is git. So you have a sequence of tree\nstates whether that is what you want or not. Thus you are specifying\nthe difference between _some_ pair of commits. I don't see any benefit\nto restricting it to a commit and one of its parents.\n\n> On the contrary, my current proposal only needs to verify the validity\n> of a single commit, changing it like this will require the system to\n> verify the validity of two commits.  Given the rareness of the origin\n> links this will hardly present a problem, but it *does* increase\n> the overhead in checking a bit.\n\nActually, it could decrease it. If I tell you that you must have \"X\" and\n\"X^2\", then you could get away with just checking if you have \"X\". But\nyou might also want to check whether \"X\" even _has_ a second parent. And\nthat means not just looking up the object, but accessing it (resolving\ndeltas if need be, uncompressing, parsing the object).  With \"X\" and\n\"Y\", it is just two object lookups.\n\nNow obviously you don't have to be quite so careful in the \"hash plus\nparent\" case. And if you are going to _do_ anything with the origin\nfield, you will end up accessing those objects anyway. But in that case,\nyou end up with the same number of lookups and accesses anyway: 2 of\neach.\n\n> >On the other hand this leads to the following question: what happens\n> >if you cherry-pick or revert a commit which has its own 'origin'\n> >links?\n> \n> Nothing special.  cherry-pick/revert behave as if the existing origin links\n> were not present in the first place.\n\nI think that is smart; if somebody wants to drill down into the history\nof origin links, they can do so at lookup time.\n\n-Peff\n"},{"id":"90256","messageId":"20080909202503.GB13634@cuci.nl","threadId":"15449","inReplyTo":"20080909195930.GA2785@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T20:25:03Z","receivedAt":"2008-09-09T20:25:03Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Tue, Sep 09, 2008 at 09:43:54PM +0200, Stephen R. van den Berg wrote:\n>> >Besides I very much prefer using 'origin <sha1> <sha2>' (as proposed\n\n>> The simplicity sounds inviting.  I'd like to hear from others who have\n>> more experience (than I have) with the git vs. changeset paradigms about\n>> this.  This allows a bit more flexibility in specifying the origin, the\n>> question is if it's needed.\n\n>And yes, you can get _too_ general to the point where your semantics\n>become meaningless. But I don't think that is the case here. You are\n>defining the origin field as \"by the way, the difference between state X\n>and state Y was used to make this commit\". cherry-pick just happens to\n>make Y=X^, but something like rebase could use a series.\n\n>As for \"git vs changeset\": this is git. So you have a sequence of tree\n>states whether that is what you want or not. Thus you are specifying\n>the difference between _some_ pair of commits. I don't see any benefit\n>to restricting it to a commit and one of its parents.\n\nQuite.  I'll drop the old format and adapt my proposal to use the double\nhash.\n\nAs far as the naming of the field is concerned: a changeset is what the\nfield describes, but changeset implies no sense of direction; origin\nmakes it clear that the current commit was derived *from* the changeset\nrepresented by \"origin\".\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90257","messageId":"7vljy159v7.fsf@gitster.siamese.dyndns.org","threadId":"15449","inReplyTo":"20080909195930.GA2785@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-09-09T20:42:52Z","receivedAt":"2008-09-09T20:42:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> And yes, you can get _too_ general to the point where your semantics\n> become meaningless. But I don't think that is the case here. You are\n> defining the origin field as \"by the way, the difference between state X\n> and state Y was used to make this commit\". cherry-pick just happens to\n> make Y=X^, but something like rebase could use a series.\n>\n> As for \"git vs changeset\": this is git. So you have a sequence of tree\n> states whether that is what you want or not. Thus you are specifying\n> the difference between _some_ pair of commits. I don't see any benefit\n> to restricting it to a commit and one of its parents.\n\nAs for \"by the way ... was used to make this commit\": this is git.  So how\nyou arrived at the tree state you record in a commit *does not matter*.\n\nNot only that, it is not just \"the difference between state X and Y\" that\nyou used to come to that tree.  Another thing that is involved is the\nspecific cherry-pick implementation back when the commit was made.  That\nwas what gave you the tree.\n\nTo my ears, it rhymes rather well with a famous quote from $gmane/217:\n\n    You're freezing your (crappy) algorithm at tree creation time, and\n    basically making it pointless to ever create something better later,\n    because even if hardware and software improves, you've codified that\n    \"we have to have crappy information\".\n\nAfter reading the discussion so far, I am still not convinced if this is a\ngood idea, nor this time around it is that much different from what the\nprevious \"prior\" link discussion tried to do.\n"},{"id":"90258","messageId":"20080909204744.GL10015@spearce.org","threadId":"15449","inReplyTo":"7vljy159v7.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-09-09T20:47:44Z","receivedAt":"2008-09-09T20:47:44Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> \n> To my ears, it rhymes rather well with a famous quote from $gmane/217:\n> \n>     You're freezing your (crappy) algorithm at tree creation time, and\n>     basically making it pointless to ever create something better later,\n>     because even if hardware and software improves, you've codified that\n>     \"we have to have crappy information\".\n> \n> After reading the discussion so far, I am still not convinced if this is a\n> good idea, nor this time around it is that much different from what the\n> previous \"prior\" link discussion tried to do.\n\nYup.  Same here.\n\nI didn't see any information about why this \"origin\" link is\nneeded here, just how it might work.\n\nAnd some of that \"how\" scared me because it was doing some sort of\n\"soft\" reachability, where errors aren't noticed but we are expected\nto protect the data from prune/repack forever once it has entered\nthe repository.\n\n-- \nShawn.\n"},{"id":"90259","messageId":"20080909205003.GA3397@coredump.intra.peff.net","threadId":"15449","inReplyTo":"7vljy159v7.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-09T20:50:03Z","receivedAt":"2008-09-09T20:50:03Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 09, 2008 at 01:42:52PM -0700, Junio C Hamano wrote:\n\n> As for \"by the way ... was used to make this commit\": this is git.  So how\n> you arrived at the tree state you record in a commit *does not matter*.\n\nBut it _does_ matter, which is why we have commit messages to explain\nhow you arrived at this tree state.\n\nNow, that being said:\n\n> After reading the discussion so far, I am still not convinced if this is a\n> good idea, nor this time around it is that much different from what the\n> previous \"prior\" link discussion tried to do.\n\nFor the record, I am not convinced it is a good idea either; I was\nhoping to steer it in a direction where somebody could say \"and now this\nis the useful thing we can do now that we could not do before.\" If the\nultimate goal is to put links to other commits into history viewers,\nthen the commit message is a reasonable place to do so. The only thing I\nsee improving with a header is that it makes more sense for pruning and\nobject transfer.\n\n-Peff\n"},{"id":"90261","messageId":"200809092254.30668.jnareb@gmail.com","threadId":"15449","inReplyTo":"20080909194354.GA13634@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-09T20:54:29Z","receivedAt":"2008-09-09T20:54:29Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 9 Sep 2008, Stephen R. van den Berg wrote:\n> Jakub Narebski wrote:\n>>\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n>>>\n>>> The definition of the origin field reads as follows:\n[...] \n>>> - There can be an arbitrary number of origin fields per commit.\n>>>   Typically there is going to be at most one origin field per commit.\n>> \n>> I understand that multiple origin fields occur if you do a squash\n>> merge, or if you cherry-pick multiple commits into single commit.\n[...]\n>> I'm not sure if you plan to automatically add 'origin' field for\n>> rebase, and for interactive rebase...\n> \n> That is not part of the plan so far.\n> Can you explain what you would be expecting in the best case?\n\nAfter thinking about this a bit, I don't think that (recording\norigin(al) commits) for rebased commits would be good idea.  While one\ncan reasonably expect that cherry-picked changes should stay, and\nreverted changes even more so (usually one reverts commit from\na history), usually the original commits being rebased are meant\nto be pruned; the are rebased.\n\n>>> - At the time of creation, the origin field contains a hash B which refers\n>>>   to a reachable commit pair (B, B~1).  If B has multiple parents and the pair\n>>>   being referred to needs to be e.g. (B, B~2), then the hash is followed by\n>>>   a space and followed by an integer (base10, two in this case),\n>>>   which designates the proper parentnr of B (see: mainline in git\n>>>   cherry-pick/revert).\n[...]\n>> Besides I very much prefer using 'origin <sha1> <sha2>' (as proposed\n>> in the neighbouring subthread), which would mean together with\n>> 'parent <parent>' (assuming that there are no other parents; if they\n>> are it gets even more complicated), that the following is true\n>> \n>>  <current> ~= <parent> + (<sha2> - <sha1>),\n> \n>> where '<rev1> ~= <rev2>' means that <rev1> is based on <rev2> (perhaps\n>> with some fixups, corrections or the like).  Perhaps 'origin' should\n>> be then called 'changeset'.\n> \n> The simplicity sounds inviting.  I'd like to hear from others who have\n> more experience (than I have) with the git vs. changeset paradigms about\n> this.  This allows a bit more flexibility in specifying the origin, the\n> question is if it's needed.\n\nIt is the simplicity that it is the most compelling of this solution.\nFor revert we have \"origin B B^\", for cherry-pick we have \"origin A^ A\";\n(or 'changeset') and always we have <rev> =~ <rev>^ + (<r2> - <r1>),\nwhere '-' denote diff operation (<diff> = <tree1> - <tree2>), and '+'\ndenote patch application (<tree1> = <tree2> + <diff>).\n\n\n[ADDED LATER]\nAlso it could be useful for patch management interfaces using Git\nas engine, such as StGIT, Guilt (formerly gq), TopGit, or now defunct,\nobsoleted and no longer maintained Patchy Git aka 'pg'.\n\nThe \"weak\" 'origin'/'changeset' header would allow some sort of\noperating on patches instead of usual operating on tree states.\n\n>> It would also be easier on implementation to check if\n>> 'origin'/'changeset' weak links are not broken, and to get to know\n>> which commits are to be protected against pruning than your proposal\n>> of\n>>\n>>   origin <\"cousin\" id> [<mainline = parent number>]\n>>\n>> where <mainline> can be omitted if it is 1 (the default).\n> \n> On the contrary, my current proposal only needs to verify the validity\n> of a single commit, changing it like this will require the system to\n> verify the validity of two commits.  Given the rareness of the origin\n> links this will hardly present a problem, but it *does* increase\n> the overhead in checking a bit.\n\nErrr... wasn't you proposing to keep/protect against pruning <cousin>\nAND <cousin>^<mainline>? You want to have _diff_ (changeset) protected,\nnot a single tree state.\n\nAnd having \"origin <r1> <r2>\" makes it easier then to check validity; you\ndon't need to get <r1>, check if it has <mainline> parent and what it is,\nand then check if <r1>^<mainline> exists (and is not for example behind\nshallow clone barrier).\n\n>>> - git cherry-pick/revert allow for the creation of origin links only if\n>>>   the object they are referring to is presently reachable.\n> \n>> Errr... shouldn't objects referenced by 'origin' links be reachable in\n>> order for \"cherry-pick\" or \"revert\" to succeed?\n> \n> True.  But sometimes it's necessary to emphasize the obvious; call it a\n> preemptive strike against possible objections to the proposal.\n\nI don't think that it is true in this case.  This sentence _looks_ like\nit offers / requires additional protection, while this \"protection\" is\nalready ensured by the fact of cherry-picking or reverting a commit.\n\n>> On the other hand this leads to the following question: what happens\n>> if you cherry-pick or revert a commit which has its own 'origin'\n>> links?\n> \n> Nothing special.  cherry-pick/revert behave as if the existing origin links\n> were not present in the first place.\n\nO.K.\n\n>>> - git rev-list --topo-order will take origin links into account to\n>>>   ensure proper ordering.\n>> \n>> What do you mean by that?\n> \n> The order in which commits are listed is defined by the fact that\n> descendent commits are shown before any of their parents.  The presence\n> of an origin link will make sure that the current commit will always\n> appear *before* the origin-commit it is referring to (if the\n> origin-commit is in the displayed set, that is).\n\nHmmm... while I think it might be a good idea, I'm not sure about its\noverhead. Should be much, I guess.\n \n>>> - git log would show something like: [...]\n>>> - git log --graph will show a dotted line of somesort just like gitk.\n>> \n>> That is I guess the whole and main reason for 'origin' links to exist,\n>> as having this information in free-form part, i.e. in the commit\n>> message might lead to problems (with parsing and extracting, and\n>> finding spurious links).\n> \n> Quite.  Also, having them in a well-defined place will allow for easy\n> fixups in case of rebase/filter-branch.\n\nTrue.\n\n>>> - git blame will follow and use the origin link if the object exists.\n>> \n>> Hmmmm... I'm not sure about that.\n> \n> Care to explain your doubts?\n> The reason I want this behaviour, is because it's all about tracking\n> content, and that part of the content happens to come from somewhere\n> else, and therefore blame should look there to \"dig deeper\" into it.\n\nBut blame is all about what commit brought some line to currents version.\nSo the cherry-pick itself, or revert of a commit itself would be blamed,\nand should be blamed, not its parents, nor commit which got cherry-picked,\nor commit which got reversed.\n\nIt would be nice to be able to follow 'origin'/'changeset' lines in the\n_graphical_ blame browser like blameview or \"git gui blame\".\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90262","messageId":"7vfxo958tj.fsf@gitster.siamese.dyndns.org","threadId":"15449","inReplyTo":"20080909195930.GA2785@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-09-09T21:05:28Z","receivedAt":"2008-09-09T21:05:28Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> And yes, you can get _too_ general to the point where your semantics\n> become meaningless. But I don't think that is the case here. You are\n> defining the origin field as \"by the way, the difference between state X\n> and state Y was used to make this commit\". cherry-pick just happens to\n> make Y=X^, but something like rebase could use a series.\n\nAnother thing that made me wonder...\n\nTo be consistent, when you are at HEAD and are merging side branch B,\nbecause that merge is to incorporate what happened on the side branch\nwhile you are looking the other way, we should say \"by the way, the\ndifference between state $(git merge-base HEAD B) and state B was used to\nmake this commit.\" in the resulting merge commit, shouldn't we?\n\nWhat happens if there is more than one merge base?\n"},{"id":"90263","messageId":"20080909210917.GA3715@coredump.intra.peff.net","threadId":"15449","inReplyTo":"7vfxo958tj.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-09T21:09:17Z","receivedAt":"2008-09-09T21:09:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Sep 09, 2008 at 02:05:28PM -0700, Junio C Hamano wrote:\n\n> To be consistent, when you are at HEAD and are merging side branch B,\n> because that merge is to incorporate what happened on the side branch\n> while you are looking the other way, we should say \"by the way, the\n> difference between state $(git merge-base HEAD B) and state B was used to\n> make this commit.\" in the resulting merge commit, shouldn't we?\n\nI suppose you could, though in that case you can obviously calculate the\nmerge base yourself from the parents of the merge. The difference with\ncherry-picking (or rebasing) is that you might not otherwise know about\nthe \"by the way\" commits.\n\n-Peff\n"},{"id":"90264","messageId":"20080909211355.GB10544@machine.or.cz","threadId":"15449","inReplyTo":"20080909132212.GA25476@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-09-09T21:13:55Z","receivedAt":"2008-09-09T21:13:55Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Tue, Sep 09, 2008 at 03:22:12PM +0200, Stephen R. van den Berg wrote:\n> - During fetch/push/pull the full commit including the origin fields is\n>   transmitted, however, the objects the origin links are referring to\n>   are not (unless they are being transmitted because of other reasons).\n> \n> - When fetching/pulling it is optionally possible to tell git to\n>   actually transmit objects referred to by origin links even if it would\n>   otherwise not have done so.\n\nI think this is misguided. In general case, cherrypicks can be from\ncompletely unrelated histories, and if you are doing the cherry pick,\nyou are saying that actually, the history *does not matter*. In that\ncase, this kind of link tries to impose a meaning where there is none,\nand in an ill-defined way when whether the commit is actually around\nanywhere is essentially random.\n\nWhy do you actually *follow* the origin link at all anyway? Without its\nparents, the associated tree etc., the object is essentially useless for\nyou; the authorship information and commit message should've been\npreserved by a proper cherry-pick anyway. You're cluttering the object\nstore with invalid objects, which also breaks quite some fundamental\nlogic within Git (which assumes that if an object exists, all its\nreferences are valid - give or take few special cases like shallow\nrepositories, but this would have very different characteristics).\n\nHaving history browsers draw fancy lines is fine but I see nothing wrong\nwith them extracting this from the free-form part of the commit message.\nFor informative purposes, we don't shy away from heuristics anyway, c.f.\nour renames detection (heck, we are even brave enough to use that for\nmerges).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nThe next generation of interesting software will be done\non the Macintosh, not the IBM PC.  -- Bill Gates\n"},{"id":"90282","messageId":"200809100035.23166.jnareb@gmail.com","threadId":"15449","inReplyTo":"20080909205003.GA3397@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-09T22:35:21Z","receivedAt":"2008-09-09T22:35:21Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 9 Sep 2008, Jeff King wrote:\n> On Tue, Sep 09, 2008 at 01:42:52PM -0700, Junio C Hamano wrote:\n> \n> > As for \"by the way ... was used to make this commit\": this is git.  So how\n> > you arrived at the tree state you record in a commit *does not matter*.\n> \n> But it _does_ matter, which is why we have commit messages to explain\n> how you arrived at this tree state.\n\nWell, that is why I was carefull to say that \"origin <rev1> <rev2>\"\n(or 'changeset', or 'cset') means that tree state for given commit\nis created out of parent commit (or parent commits in the case of merge)\nand of (<rev2> - <rev1>) patch.  This is a bit of enhancement to\n\"parent <rev>\" meaning that tree state for current commit is derived\nfrom tree state of <rev>.\n\nThis is nice generalization...\n\n> Now, that being said:\n> \n> > After reading the discussion so far, I am still not convinced if this is a\n> > good idea, nor this time around it is that much different from what the\n> > previous \"prior\" link discussion tried to do.\n> \n> For the record, I am not convinced it is a good idea either; I was\n> hoping to steer it in a direction where somebody could say \"and now this\n> is the useful thing we can do now that we could not do before.\" If the\n> ultimate goal is to put links to other commits into history viewers,\n> then the commit message is a reasonable place to do so. The only thing I\n> see improving with a header is that it makes more sense for pruning and\n> object transfer.\n\nI'm also not all convinced that 'cousin'/'origin'/'changeset'/'cset'\nheader is a good idea.  I only tried to steer discussion in good\ndirection if it is somewhat a good idea.\n\nFirst, if the only goal would be to add extra links (extra edges) to\n[graphical] history viewer, then full sha-1 of a commit which can be\nrecorded in commit message for cherry-picks and reverts should be\nenough.  It does mean parsing commit message, and all possibilities\nfor mistake which are connected to using conventions in free-form part\nof commit object; on the other hand it is not _that_ critical.\n\nIf however 'origin' links are more (perhaps only a tiny bit more),\nfor example discussed \"weak\" links... then I'm not sure if\nthe tradeoffs are worth it. First, if it is full connectivity like\nin 'parent' header case, then a) why not use 'parent' anyway,\nb) it pins the history indefinitely long. Second, if it is \"weak\"\nlink, i.e. local protect it on prune, then a) there are problems\nwith transferring the data, and protecting links on transfer,\nas somewhere in the middle or at the end there might be repository\nwhich uses older git (backwards compatibility strikes again),\nb) git in many, many places assumes that object is valid if it passes,\nand all objects linked to from object are valid; we would have either\nuse some kind of separate 'not strictly checked' packfile/storage,\nor have grafts-like thingy.\n\nSo I'm not sure if 'origin' links are worth the trouble.\n\n\nAbout much, much earlier \"prior\" link discussion: I think the discussion\nabout \"prior\" header link was done before reflogs, or at least before\nreflogs got turned on by default.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90284","messageId":"20080909225603.GA7459@cuci.nl","threadId":"15449","inReplyTo":"20080909211355.GB10544@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T22:56:03Z","receivedAt":"2008-09-09T22:56:03Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Petr Baudis wrote:\n>On Tue, Sep 09, 2008 at 03:22:12PM +0200, Stephen R. van den Berg wrote:\n>> - During fetch/push/pull the full commit including the origin fields is\n>>   transmitted, however, the objects the origin links are referring to\n>>   are not (unless they are being transmitted because of other reasons).\n\n>> - When fetching/pulling it is optionally possible to tell git to\n>>   actually transmit objects referred to by origin links even if it would\n>>   otherwise not have done so.\n\n>I think this is misguided. In general case, cherrypicks can be from\n>completely unrelated histories, and if you are doing the cherry pick,\n>you are saying that actually, the history *does not matter*. In that\n\nThat is a false assumption in general, I'd say.\n\n>case, this kind of link tries to impose a meaning where there is none,\n>and in an ill-defined way when whether the commit is actually around\n>anywhere is essentially random.\n\nThe purpose I'd use the origin links for is to manage software projects\nthat consist of 7 main branches which have branched in (on average) two\nyear intervals, which never get merged anymore.  The only thing that\nhappens is that there are backports amongst the branches about two per\nweek.\n\nThe only way to perform the backports is by using cherry-pick.\nThe history of each backport *is* important though.\nSince all the developers who care about the multiple release branches\nhave all the relevant branches in their repository, the presence of\na origin object is by no means random, it's a certainty.\n\n>Why do you actually *follow* the origin link at all anyway? Without its\n>parents, the associated tree etc., the object is essentially useless for\n>you; the authorship information and commit message should've been\n>preserved by a proper cherry-pick anyway. You're cluttering the object\n>store with invalid objects, which also breaks quite some fundamental\n>logic within Git (which assumes that if an object exists, all its\n>references are valid - give or take few special cases like shallow\n>repositories, but this would have very different characteristics).\n\nI'd prefer to formalise the (weak) relationship of an origin link, instead of\nrelying on vague assumptions when parsing the free-form commit message\nand then guessing what the mentioned hash might mean.\n\n>Having history browsers draw fancy lines is fine but I see nothing wrong\n>with them extracting this from the free-form part of the commit message.\n>For informative purposes, we don't shy away from heuristics anyway, c.f.\n>our renames detection (heck, we are even brave enough to use that for\n>merges).\n\nIt's not just that.  If I make a change to an area that was cherrypicked\nfrom another branch, then I find it rather important to check if any\nchanges to this area need to be backported/forwardported to the branches\nthe origin links are pointing to.\nI.e. the origin link allows me to improve my efficiency as a programmer.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90286","messageId":"20080909230525.GC10360@machine.or.cz","threadId":"15449","inReplyTo":"20080909225603.GA7459@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-09-09T23:05:25Z","receivedAt":"2008-09-09T23:05:25Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Wed, Sep 10, 2008 at 12:56:03AM +0200, Stephen R. van den Berg wrote:\n> The only way to perform the backports is by using cherry-pick.\n> The history of each backport *is* important though.\n> Since all the developers who care about the multiple release branches\n> have all the relevant branches in their repository, the presence of\n> a origin object is by no means random, it's a certainty.\n\nRecording cherry-picks in your workflow certainly makes sense, but I'm\nnot talking about workflow-level issues here. You are adding an extra\nheader to the commit object. I'm talking about the object database and\nlow-level Git model implications this has.\n\nIn other way, I think this is purely a porcelain matter and recording\nthis information in the free-form area is more than enough.\n\n> >Why do you actually *follow* the origin link at all anyway? Without its\n> >parents, the associated tree etc., the object is essentially useless for\n> >you; the authorship information and commit message should've been\n> >preserved by a proper cherry-pick anyway. You're cluttering the object\n> >store with invalid objects, which also breaks quite some fundamental\n> >logic within Git (which assumes that if an object exists, all its\n> >references are valid - give or take few special cases like shallow\n> >repositories, but this would have very different characteristics).\n> \n> I'd prefer to formalise the (weak) relationship of an origin link, instead of\n> relying on vague assumptions when parsing the free-form commit message\n> and then guessing what the mentioned hash might mean.\n\nWhy?\n\n> >Having history browsers draw fancy lines is fine but I see nothing wrong\n> >with them extracting this from the free-form part of the commit message.\n> >For informative purposes, we don't shy away from heuristics anyway, c.f.\n> >our renames detection (heck, we are even brave enough to use that for\n> >merges).\n> \n> It's not just that.  If I make a change to an area that was cherrypicked\n> from another branch, then I find it rather important to check if any\n> changes to this area need to be backported/forwardported to the branches\n> the origin links are pointing to.\n> I.e. the origin link allows me to improve my efficiency as a programmer.\n\nAnd why are the notes created by git cherry-pick -x insufficient for that?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nThe next generation of interesting software will be done\non the Macintosh, not the IBM PC.  -- Bill Gates\n"},{"id":"90287","messageId":"200809100107.44933.jnareb@gmail.com","threadId":"15449","inReplyTo":"200809100035.23166.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-09T23:07:43Z","receivedAt":"2008-09-09T23:07:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jakub Narebski wrote:\n> On Tue, 9 Sep 2008, Jeff King wrote:\n> > On Tue, Sep 09, 2008 at 01:42:52PM -0700, Junio C Hamano wrote:\n\n> > Now, that being said:\n> > \n> > > After reading the discussion so far, I am still not convinced if this is a\n> > > good idea, nor this time around it is that much different from what the\n> > > previous \"prior\" link discussion tried to do.\n> > \n> > For the record, I am not convinced it is a good idea either; I was\n> > hoping to steer it in a direction where somebody could say \"and now this\n> > is the useful thing we can do now that we could not do before.\" If the\n> > ultimate goal is to put links to other commits into history viewers,\n> > then the commit message is a reasonable place to do so. The only thing I\n> > see improving with a header is that it makes more sense for pruning and\n> > object transfer.\n> \n> I'm also not all convinced that 'cousin'/'origin'/'changeset'/'cset'\n> header is a good idea.  I only tried to steer discussion in good\n> direction if it is somewhat a good idea.\n\nBy the way, beside graphical history viewers it would also help rebase\n(and git-cherry) notice when patch was already applied better.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90288","messageId":"20080909230800.GB7459@cuci.nl","threadId":"15449","inReplyTo":"200809092254.30668.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T23:08:00Z","receivedAt":"2008-09-09T23:08:00Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jakub Narebski wrote:\n>On Tue, 9 Sep 2008, Stephen R. van den Berg wrote:\n>> On the contrary, my current proposal only needs to verify the validity\n>> of a single commit, changing it like this will require the system to\n>> verify the validity of two commits.  Given the rareness of the origin\n>> links this will hardly present a problem, but it *does* increase\n>> the overhead in checking a bit.\n\n>Errr... wasn't you proposing to keep/protect against pruning <cousin>\n>AND <cousin>^<mainline>? You want to have _diff_ (changeset) protected,\n>not a single tree state.\n\nActually, making sure that the commit we reference in the origin link\nexists, we implicitly prove that all the parents of that commit exist as\nwell.  Then again, this point is moot since I already conceded (in a\ndifferent thread) that storing two hashes is better.\n\n>>>> - git rev-list --topo-order will take origin links into account to\n>>>>   ensure proper ordering.\n\n>Hmmm... while I think it might be a good idea, I'm not sure about its\n>overhead. Should be much, I guess.\n\nActually, I have already programmed this part, and the overhead is close\nto zero.\n\n>>>> - git blame will follow and use the origin link if the object exists.\n\n>>> Hmmmm... I'm not sure about that.\n\n>> Care to explain your doubts?\n>> The reason I want this behaviour, is because it's all about tracking\n>> content, and that part of the content happens to come from somewhere\n>> else, and therefore blame should look there to \"dig deeper\" into it.\n\n>But blame is all about what commit brought some line to currents version.\n>So the cherry-pick itself, or revert of a commit itself would be blamed,\n>and should be blamed, not its parents, nor commit which got cherry-picked,\n>or commit which got reversed.\n\nWell, it depends, I guess.\nIf you'd go for a \"committer\" based display, then following origin links\nis bad.\nIf you'd go for an \"author\" based display, then following origin links\nshould be the default (IMHO).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90292","messageId":"20080909233245.GC7459@cuci.nl","threadId":"15449","inReplyTo":"20080909230525.GC10360@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T23:32:45Z","receivedAt":"2008-09-09T23:32:45Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Petr Baudis wrote:\n>On Wed, Sep 10, 2008 at 12:56:03AM +0200, Stephen R. van den Berg wrote:\n\n>In other way, I think this is purely a porcelain matter and recording\n>this information in the free-form area is more than enough.\n\n>> I'd prefer to formalise the (weak) relationship of an origin link, instead of\n>> relying on vague assumptions when parsing the free-form commit message\n>> and then guessing what the mentioned hash might mean.\n\n>Why?\n\nUsing special references in the free-form area of a commit is akin to\nusing X-... headerfields in E-mail with all the assorted mess:\n- No strict definition of what it means.\n- Diverging porcelain implementations making use of the field in ever so\n  slightly changing ways over the years.\n- You cannot rely on the field being always available.\n- Automated \"renumbering\" becomes difficult at best.\n\nWhat we want are concise and unambiguous definitions which allow us to\nbuild tools that operate predictably on them now, and will operate\npredictably on them in the future.\n\n>> >Having history browsers draw fancy lines is fine but I see nothing wrong\n\n>> It's not just that.  If I make a change to an area that was cherrypicked\n>> from another branch, then I find it rather important to check if any\n>> changes to this area need to be backported/forwardported to the branches\n>> the origin links are pointing to.\n>> I.e. the origin link allows me to improve my efficiency as a programmer.\n\n>And why are the notes created by git cherry-pick -x insufficient for that?\n\nThings like rebase/filter-branch/stgit mess that up because they don't\nknow if the hash in the free-form should be altered.\nAlso, there is no automated way to actually fetch missing branches we\ncherry-picked from this way.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90294","messageId":"alpine.LFD.1.10.0809091631250.3117@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"20080909194354.GA13634@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-09T23:35:52Z","receivedAt":"2008-09-09T23:35:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 9 Sep 2008, Stephen R. van den Berg wrote:\n\n> Jakub Narebski wrote:\n> >\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n> >> The definition of the origin field reads as follows:\n> \n> >> - There can be an arbitrary number of origin fields per commit.\n> >>   Typically there is going to be at most one origin field per commit.\n> \n> >I understand that multiple origin fields occur if you do a squash\n> >merge, or if you cherry-pick multiple commits into single commit.\n> >For example:\n> > $ git cherry-pick -n <a1>\n> > $ git cherry-pick    <a2>\n> > $ git commit --amend        #; to correct commit message\n> \n> Correct.\n\nQuite frankly, recording the origins for _any_ of the above sounds like a \nhorribly mistake.\n\nAll those operations are commonly used (along with \"git rebase -i\") to \nclean up history in order to show a nicer version.\n\nThe whole point of \"origin\" seems to be to _destroy_ that.\n\nI would refuse to ever touch anything that had an \"origin\" pointer, so if \ngit were to add that feature, it would be a huge disappointment to me. I'd \nhave to have a version that makes sure that anything it pulls hasn't been \ncrapped on by somebody who added a stupid link to some dirty history that \nI'm not at all interested in seeing.\n\nIOW, I'm seeing a _lot_ of downsides, and not any actual upsides. What are \nthe upsides again? \n\n\t\t\tLinus\n"},{"id":"90293","messageId":"20080909233609.GD7459@cuci.nl","threadId":"15449","inReplyTo":"20080909210917.GA3715@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T23:36:09Z","receivedAt":"2008-09-09T23:36:09Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Tue, Sep 09, 2008 at 02:05:28PM -0700, Junio C Hamano wrote:\n>> To be consistent, when you are at HEAD and are merging side branch B,\n>> because that merge is to incorporate what happened on the side branch\n>> while you are looking the other way, we should say \"by the way, the\n>> difference between state $(git merge-base HEAD B) and state B was used to\n>> make this commit.\" in the resulting merge commit, shouldn't we?\n\n>I suppose you could, though in that case you can obviously calculate the\n>merge base yourself from the parents of the merge. The difference with\n>cherry-picking (or rebasing) is that you might not otherwise know about\n>the \"by the way\" commits.\n\nQuite.  The origin link is primarily intended for cherry-picks/reverts\nwhich are otherwise difficult to find.  Anything that can use the normal\nparent mechanism has no business using the origin links.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90299","messageId":"20080909235848.GE7459@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809091631250.3117@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-09T23:58:48Z","receivedAt":"2008-09-09T23:58:48Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>On Tue, 9 Sep 2008, Stephen R. van den Berg wrote:\n>> Jakub Narebski wrote:\n>> >I understand that multiple origin fields occur if you do a squash\n>> >merge, or if you cherry-pick multiple commits into single commit.\n>> >For example:\n>> > $ git cherry-pick -n <a1>\n>> > $ git cherry-pick    <a2>\n>> > $ git commit --amend        #; to correct commit message\n\n>> Correct.\n\n>All those operations are commonly used (along with \"git rebase -i\") to \n>clean up history in order to show a nicer version.\n\nActually, I'd suggest that cherry-pick takes an -o flag which turns on\nthe origin link.  This needs to be a concious decision because one deems\nthe history relevant.  This typically is a Good Thing when the\ndevelopment has several long-term-stable branches which never get merged\nwith each other, yet they receive frequent backports (using cherry-pick)\nbetween them.\n\n>The whole point of \"origin\" seems to be to _destroy_ that.\n\nOnly in the case where the committer thinks the history is of interest,\nand even then, since the origin link is in the header, displaying it or\nnot suddenly is under the control of git.\nHad it been in the free-form textarea, there'd be no way suppress the\ndisplay of it.\n\n>I would refuse to ever touch anything that had an \"origin\" pointer, so if \n>git were to add that feature, it would be a huge disappointment to me. I'd \n>have to have a version that makes sure that anything it pulls hasn't been \n>crapped on by somebody who added a stupid link to some dirty history that \n>I'm not at all interested in seeing.\n\nAs you might have noticed, the actual process of pulling/fetching\nexplicitly does *not* pull in the objects being pointed to.\nThat, in turn, will cause the origin link output to be automatically\nsuppressed.  I.e. you'll never know the difference.\n\nOTOH, if someone adds a free-form link to the commit message, you\nessentially cannot hide that and are just suffering the clutter without\nhaving any use for it.\n\n>IOW, I'm seeing a _lot_ of downsides, and not any actual upsides. What are \n>the upsides again? \n\nThe upsides are:\n- If your repository contains the proper branches, it will show a richer\n  content.\n- If your repository lacks the proper branches, it will show a *reduced*\n  clutter content (because actual free-text references in the commit\n  messages will decrease).\n\nI see a lot of upsides, what were the downsides again?\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90300","messageId":"200809100159.59060.jnareb@gmail.com","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809091631250.3117@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-09T23:59:57Z","receivedAt":"2008-09-09T23:59:57Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n> On Tue, 9 Sep 2008, Stephen R. van den Berg wrote:\n>> Jakub Narebski wrote:\n>>>\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n>>>> The definition of the origin field reads as follows:\n>> \n>>>> - There can be an arbitrary number of origin fields per commit.\n>>>>   Typically there is going to be at most one origin field per commit.\n>> \n>>> I understand that multiple origin fields occur if you do a squash\n>>> merge, or if you cherry-pick multiple commits into single commit.\n>>> For example:\n>>> $ git cherry-pick -n <a1>\n>>> $ git cherry-pick    <a2>\n>>> $ git commit --amend        #; to correct commit message\n>> \n>> Correct.\n> \n> Quite frankly, recording the origins for _any_ of the above sounds like a \n> horribly mistake.\n\nActually the above is _not_ a good example for using 'origin', and why\nusing 'origin'; just a bit convoluted example of multiple 'origin'\nheaders.\n\n> All those operations are commonly used (along with \"git rebase -i\") to \n> clean up history in order to show a nicer version.\n> \n> The whole point of \"origin\" seems to be to _destroy_ that.\n\nIf I understand correctly the point is to record those 'origin' headers\nfor git-revert (when 'origin'-ed commit is somewhere in the history),\nand for git-cherry-pick from other long lived branch and thus require\nadditional option to git-cherry-pick to record 'origin' (denoting that\nyou this is \"true\" cherry-pick, and not reordering of commits and\ncleaning up a history, better done with interactive rebase).\n\n/me is playing advocatus diaboli here, 'cause I'm not that convinced\nto necessity of this feature.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90302","messageId":"20080910001316.GF7459@cuci.nl","threadId":"15449","inReplyTo":"7vljy159v7.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T00:13:16Z","receivedAt":"2008-09-10T00:13:16Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n>As for \"by the way ... was used to make this commit\": this is git.  So how\n>you arrived at the tree state you record in a commit *does not matter*.\n\nThe typical use case for the origin links is in a project with several\nlong-lived branches which use cherry-picks to backport amongst them.\nThere is no real other way to solve this case, except for some rather\nkludgy stuff in the free-form commit message which doesn't mesh well\nwith rebase/filter-branch/stgit etc.\n\nAs to \"does not matter\": then why does git store parent links?\n\n>To my ears, it rhymes rather well with a famous quote from $gmane/217:\n\n>    You're freezing your (crappy) algorithm at tree creation time, and\n>    basically making it pointless to ever create something better later,\n>    because even if hardware and software improves, you've codified that\n>    \"we have to have crappy information\".\n\nI tried to accomodate this approach by overloading the parent link and\nthen making git more intelligent to figure out if it is a cherry-pick or\nnot.  That was deemed undesirable, so using the origin links is the next\nbest thing (IMHO).\n\n>good idea, nor this time around it is that much different from what the\n>previous \"prior\" link discussion tried to do.\n\nIt is well-defined this time, and doesn't bleed across fetch/pull.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Be spontaneous!\"\n"},{"id":"90303","messageId":"alpine.LFD.1.10.0809091722010.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"20080909235848.GE7459@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-10T00:23:27Z","receivedAt":"2008-09-10T00:23:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 10 Sep 2008, Stephen R. van den Berg wrote:\n> \n> As you might have noticed, the actual process of pulling/fetching\n> explicitly does *not* pull in the objects being pointed to.\n\n.. which makes them _local_ data, which in turn means that they should not \nbe in the object database at all.\n\nIOW, i you want this for local reasons, you should use a local database, \nlike the index or the reflogs (and I don't mean \"like the index\" in the \nsense that it would look _anything_ like that file, but in the sense that \nit's a purely local thing and doesn't show up in the object database).\n\n\t\tLinus\n"},{"id":"90305","messageId":"7vej3s223f.fsf@gitster.siamese.dyndns.org","threadId":"15449","inReplyTo":"20080910001316.GF7459@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-09-10T01:59:00Z","receivedAt":"2008-09-10T01:59:00Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n\n> Junio C Hamano wrote:\n>>As for \"by the way ... was used to make this commit\": this is git.  So how\n>>you arrived at the tree state you record in a commit *does not matter*.\n>\n> The typical use case for the origin links is in a project with several\n> long-lived branches which use cherry-picks to backport amongst them.\n> There is no real other way to solve this case, except for some rather\n> kludgy stuff in the free-form commit message which doesn't mesh well\n> with rebase/filter-branch/stgit etc.\n>\n> As to \"does not matter\": then why does git store parent links?\n\nThe parent links describe *where* you came from, not *how*.\n\nAnd if you think the difference is just \"semantics\", then you haven't\ngrokked the first lesson I gave in this thread.  \"parents\" record the\nreference points against which you make \"this resulting commit suits the\npurpose of my branch better than any histories leading to these commits\".\n"},{"id":"90314","messageId":"20080910053838.GA15715@cuci.nl","threadId":"15449","inReplyTo":"7vej3s223f.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T05:38:38Z","receivedAt":"2008-09-10T05:38:38Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n>\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n>> Junio C Hamano wrote:\n>>>As for \"by the way ... was used to make this commit\": this is git.  So how\n>>>you arrived at the tree state you record in a commit *does not matter*.\n\n>> The typical use case for the origin links is in a project with several\n>> long-lived branches which use cherry-picks to backport amongst them.\n>> There is no real other way to solve this case, except for some rather\n>> kludgy stuff in the free-form commit message which doesn't mesh well\n>> with rebase/filter-branch/stgit etc.\n\n>> As to \"does not matter\": then why does git store parent links?\n\n>The parent links describe *where* you came from, not *how*.\n\n>And if you think the difference is just \"semantics\", then you haven't\n>grokked the first lesson I gave in this thread.  \"parents\" record the\n>reference points against which you make \"this resulting commit suits the\n>purpose of my branch better than any histories leading to these commits\".\n\nThe last question of mine was/is a rethorical one.\n\nConsider the typical use case I describe above.  The developer usually\nhas just created a commit in the developmentbranch, tested it, and deems\nthe patch worthwhile enough to backport it to the latest stable branch.\nSo he cherry picks the from the development branch to the latest stable\nbranch.\nThen tests it, and decides to backport it to the older stable branch,\nso he cherry-picks it again, and commits it there too.  The is repeated\nin rapid succession on three older stable branches as well.\n\nBasically that means that for the patch itself, there is a path in\nhistory to follow as well.  I.e. the patch itself evolves over time.\n\nNow, when another developer makes an additional change to this patch\nin one of the stable versions, it is very helpful to actually be able to\nhave git tell you where the original patch came from and to follow back\nthe chain upward.  It allows you to forward/backward port the new change\nmore easily.\n\nBasically, the normal parent links allow you to follow evolving\nsnapshots of the complete source-tree, whereas the origin links allow\nyou to follow evolving snapshots of a patch.\nAs it happens, the shortest way to describe a patch in git is by\nspecifying two commits of which the difference is exactly your patch.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90315","messageId":"20080910054244.GB15715@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809091722010.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T05:42:44Z","receivedAt":"2008-09-10T05:42:44Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>On Wed, 10 Sep 2008, Stephen R. van den Berg wrote:\n\n>> As you might have noticed, the actual process of pulling/fetching\n>> explicitly does *not* pull in the objects being pointed to.\n\n>.. which makes them _local_ data, which in turn means that they should not \n>be in the object database at all.\n\n>IOW, i you want this for local reasons, you should use a local database, \n>like the index or the reflogs (and I don't mean \"like the index\" in the \n>sense that it would look _anything_ like that file, but in the sense that \n>it's a purely local thing and doesn't show up in the object database).\n\nBut then how would someone who clones the repository get at the information?\nThe information is essential to understand backports between the various\nstable branches.\n\nThe origin links describe the evolving state of a patch (i.e. just like\nregular commits/parents store snapshots of the whole tree, the origin\nlinks store snapshots of a patch as it evolves through time).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90326","messageId":"48C780DF.1070605@gnu.org","threadId":"15449","inReplyTo":"200809100107.44933.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T08:10:07Z","receivedAt":"2008-09-10T08:10:07Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"> By the way, beside graphical history viewers it would also help rebase\n> (and git-cherry) notice when patch was already applied better.\n\nI think that rebase had better not trust the origin links in deciding\nwhether a patch was already applied; it already does it well enough.\n\ngit-cherry is another story, as that tool is not faking \"changeset mode\"\nso well (because it cannot attempt merges, and these are what allows\ngit-rebase to fake changesets much better).\n\nPaolo\n"},{"id":"90328","messageId":"48C785C3.9010204@gnu.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809091722010.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T08:30:59Z","receivedAt":"2008-09-10T08:30:59Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"Linus Torvalds wrote:\n> \n> On Wed, 10 Sep 2008, Stephen R. van den Berg wrote:\n>> As you might have noticed, the actual process of pulling/fetching\n>> explicitly does *not* pull in the objects being pointed to.\n> \n> .. which makes them _local_ data, which in turn means that they should not \n> be in the object database at all.\n\nNot really local data.  More like _weakly referenced_ data.  If it is\nthere, cool.  If it is not there, no big deal.\n\nPaolo\n"},{"id":"90329","messageId":"48C7893B.6050001@gnu.org","threadId":"15449","inReplyTo":"20080909211355.GB10544@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T08:45:47Z","receivedAt":"2008-09-10T08:45:47Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> Having history browsers draw fancy lines is fine but I see nothing wrong\n> with them extracting this from the free-form part of the commit message.\n> For informative purposes, we don't shy away from heuristics anyway, c.f.\n> our renames detection (heck, we are even brave enough to use that for\n> merges).\n\n... and it works only because false positives (which make the merge\nactively wrong) are extremely rare.  If there is the occasional false\nnegative, well, it just makes merges more complicated and screws up\nvisualization a bit, so you live with it.\n\nMove/copy detection almost always worked for me, but there are two cases\nwhere it didn't:\n\n1) empty files.  Each of them is marked as copied from a\nseemingly-picked-at-random one.\n\n2) renaming a Gtk+ class.  You rename it (e.g. from gtkclassnamea.c to\ngtkclassnameb.c) and at the same time do\n\n  s/GtkClassNameA/GtkClassNameB/\n  s/GTK_CLASS_NAME_A(?:\\>|?=_)/GTK_CLASS_NAME_B/\n  s/gtk_class_name_a(?:\\>|?=_)/gtk_class_name_b/\n\nand reindent everything.  Guaranteed to have a similarity index around\n30-40%, not more.\n\nI don't care much about it, but face it, it is *not* perfect.\n\nPaolo\n"},{"id":"90335","messageId":"48C794D6.20001@gnu.org","threadId":"15449","inReplyTo":"20080909230525.GC10360@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T09:35:18Z","receivedAt":"2008-09-10T09:35:18Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> Why do you actually *follow* the origin link at all anyway? Without its\n> parents, the associated tree etc., the object is essentially useless for\n> you\n\nStephen posed the origin links as weak, but it is not necessarily true\nthat you don't have the parents and the associated tree.  For example,\nif you download a repository that includes a \"master\" branch and a few\nstable branches, you *will* have the objects cherry-picked into stable\nbranches, because they are commits in the master branch.\n\nJunio explained that the way achieves the same effect in git is by\nforking the topic branch off the \"oldest\" branch where the patch will\npossibly be of interest.  Then he can merge it in that branch and all\nthe newest ones.  That's great, but not all people are as\nforward-looking (he did say that sometimes he needs to cherrypick).\n\nAnother problem is that in some projects actually there are two \"maint\"\nbranches (e.g. currently GCC 4.2 and GCC 4.3), and most developers do\nnot care about what goes in the older \"maint\" branch; they develop for\ntrunk and for the newer \"maint\" branch, and then one person comes and\ncherry-picks into the older \"maint\" branch.  This has two problems:\n\n1) Having to fork topic branches off the older branch would force extra\ntesting on the developers.\n\n2) Besides this, topic branches are not cloned, so if I am the\nintegrator on the older \"maint\" branch, I need to dig manually in the\ncommits to find bugfixes.  True, I could use Bugzilla, but what if I\nwant to use git instead?  There is \"git cherry -v ... | grep -w ^+.*PR\",\nexcept that it has too many false negatives (fixes that have already\nbeen backported, but do show up in the list).\n\n> And why are the notes created by git cherry-pick -x insufficient for that?\n\nFor example, these notes (or the ones created by \"git revert\") are\n*wrong* because they talk about commits instead of changesets (deltas\nbetween two commits).\n\nWhy is only one commit present?  Because these messages are meant for\nusers, not for programs.  That's easy to show: users think of commits as\ndeltas anyway, even though git stores them as snapshots---\"git show\nHEAD\" shows a delta, not a snapshot.\n\nAnd what does this mean for programs?  That they must resort to\ncommit-message scraping to distinguish the two cases. (*)\n\n   (*) A GUI blame program, for example, would need to distinguish\n   whether code added by a commit is taken from commit 4329bd8, or is\n   reverting commit 4329bd8.  (In the first case, the author of that\n   code is whoever was responsible for that code in 4329bd8; in the\n   second case, it is whoever was responsible for that code in\n   4329bd8^).  If recording changesets, you see 4329bd8^..4329bd8 in\n   the first case, and 4329bd8..4329bd8^ in the second, so it is trivial\n   to follow the chain.\n\nAnd scraping is bad.  Imagine people that are writing commit messages in\ntheir native language.  What if they patch git to translate the magic\nnotes created by \"git cherry-pick -x\" or \"git revert\" (maybe a future\nversion of git will do that automatically)?  Should they translate also\nevery program that scrapes the messages?\n\n\nWhenever there is a piece of data that could be useful to programs (no\nmatter if plumbing or porcelain), I consider free form notes to be bad.\n Because data is data, and metadata is metadata.\n\nIf there was a generic way to put porcelain-level metadata in commit\nmessages (e.g. Signed-Off-By and Acknowledged-By can be already\nconsidered metadata), I would not be so much in favor of \"origin\" links\nbeing part of the commit object's format.  Now if you think about it,\ncommit references within this kind of metadata would have mostly the\nproperties that Stephen explained in his first message:\n\n1) they would be rewritten by git-filter-branch\n\n2) these references, albeit weak by default, could optionally be\nfollowed when fetching (either with command-line or configuration options)\n\n3) they would not be pruned by git-gc\n\n4) possibly, git rev-list --topo-order would sort commits by taking into\naccount metadata references too.\n\nSo the implementation effort would be roughly the same.\n\nBut, can you think of any other such metadata?  Personally I can't, so\nwhile I understand the opposition to a new commit header field that\nwould be there from here to eternity (or until the LHC starts), I do\nthink it is the simplest thing that can possibly work.\n\nPaolo\n"},{"id":"90341","messageId":"20080910104424.GH10360@machine.or.cz","threadId":"15449","inReplyTo":"48C794D6.20001@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-09-10T10:44:24Z","receivedAt":"2008-09-10T10:44:24Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Wed, Sep 10, 2008 at 11:35:18AM +0200, Paolo Bonzini wrote:\n> \n> > Why do you actually *follow* the origin link at all anyway? Without its\n> > parents, the associated tree etc., the object is essentially useless for\n> > you\n> \n> Stephen posed the origin links as weak, but it is not necessarily true\n> that you don't have the parents and the associated tree.  For example,\n> if you download a repository that includes a \"master\" branch and a few\n> stable branches, you *will* have the objects cherry-picked into stable\n> branches, because they are commits in the master branch.\n\nBut that is irrelevant. If you already have the objects, whether to\nfollow the origin link does not matter at all.\n\nI argue that the following the origin link by one step is harmful as it\nviolated the internal Git object model and does not have real benefits.\nIf you want to have the origin links, do not follow them at all - the\ncommit objects themselves are not useful. (Or, optionally, follow them\nfully - that of course can make sense.)\n\n> > And why are the notes created by git cherry-pick -x insufficient for that?\n> \n> For example, these notes (or the ones created by \"git revert\") are\n> *wrong* because they talk about commits instead of changesets (deltas\n> between two commits).\n\n(BTW, I don't feel strongly enough about the header-freeform distinction\nto argue about it and some of your and others' points are good. But even\nif we have the origin links, I think we should only follow them not at\nall or fully.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nThe next generation of interesting software will be done\non the Macintosh, not the IBM PC.  -- Bill Gates\n"},{"id":"90343","messageId":"20080910114940.GA14127@cuci.nl","threadId":"15449","inReplyTo":"20080910104424.GH10360@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T11:49:40Z","receivedAt":"2008-09-10T11:49:40Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Petr Baudis wrote:\n>On Wed, Sep 10, 2008 at 11:35:18AM +0200, Paolo Bonzini wrote:\n>But that is irrelevant. If you already have the objects, whether to\n>follow the origin link does not matter at all.\n\n>I argue that the following the origin link by one step is harmful as it\n>violated the internal Git object model and does not have real benefits.\n>If you want to have the origin links, do not follow them at all - the\n>commit objects themselves are not useful. (Or, optionally, follow them\n>fully - that of course can make sense.)\n\nThe origin links are rarely followed, not even by one step.  They are\nonly followed if a certain operation requires them (not a lot do).\n\n>> > And why are the notes created by git cherry-pick -x insufficient for that?\n\n>> For example, these notes (or the ones created by \"git revert\") are\n>> *wrong* because they talk about commits instead of changesets (deltas\n>> between two commits).\n\n>(BTW, I don't feel strongly enough about the header-freeform distinction\n>to argue about it and some of your and others' points are good. But even\n>if we have the origin links, I think we should only follow them not at\n>all or fully.)\n\nMaybe we have a misunderstanding about what \"follow a link\" means and\nwhen it is done.\nDuring most normal git operation, the origin links are just read, but\nnot followed.\nThe only commands that I expect to follow them are log --graph, gitk, fsck\nand blame.  I may have missed some corner use-cases, but this should\ncover most of it; i.e. most of git ignores them or just makes note of\nthe hashvalues provided.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90345","messageId":"20080910122118.GI21071@mit.edu","threadId":"15449","inReplyTo":"20080909225603.GA7459@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-10T12:21:18Z","receivedAt":"2008-09-10T12:21:18Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Sep 10, 2008 at 12:56:03AM +0200, Stephen R. van den Berg wrote:\n> The purpose I'd use the origin links for is to manage software projects\n> that consist of 7 main branches which have branched in (on average) two\n> year intervals, which never get merged anymore.  The only thing that\n> happens is that there are backports amongst the branches about two per\n> week.\n> \n> The only way to perform the backports is by using cherry-pick.\n> The history of each backport *is* important though.\n> Since all the developers who care about the multiple release branches\n> have all the relevant branches in their repository, the presence of\n> a origin object is by no means random, it's a certainty.\n\nI'd argue that the origin link is a bit too general for your proposed\nuse.  One of the problems with the origin link is that it is only a\none way pointer.  Given a newer commit, you know that it is (somehow)\nweekly related to a older commit.  So your proposed workflow only\nworks if cherry-picks only happen in one direction.  That isn't always\ntrue, especially in distributed environments where the bugfix might\nhappen on someone else's development branch, and then it gets pulled\nin, or perhaps rebased in, and you want to know they are related.\n\nI would argue the best way to do that is to store (either in the\nobject or in the free-form text area) not the link, which would have\nto get renumbered but rather the identifier for the bug(s) that this\ncommit fixes.  So for example, consider a convention where in the body\nof the free-form text area, before the Signed-off-by:, Acked-by:, and\nCC: headers for those projects that use them, we add something like\nthe following:\n\nAddresses-Bug: Red_Hat/149480, Sourceforge_Feature/120167\n\nor\n\nAddresses-Bug: Debian/432865, Launchpad/203323, Sourceforge_Bug/1926023\n\nOnce you have this information, it is not difficult to maintain a\nberk_db database which maps a particular Bug identifier (i.e.,\nRed_Hat/149480, or Debian/471977, or Launchpad/203323) to a series of\ncommits.\n\nThe advantage of this scheme is that if a bug has been fixed in\nmultiple branches, you can see the association between two commits in\ntwo different branches very easily.  Furthermore, you get a link back\nto the actual bug in one or more bug tracking systems, which the some\nporcelain program could use to transform into a hot-link which when\nclicked opens up a browser window to the bug in question.\n\nIn contrast, using your proposed origin scheme, if the bug was\noriginally created in some development branch, and then cherry picked\ninto two separate maintenance branches, if you don't have the\ndevelopment branch in your repository (maybe for some reason that\ndevelopment branch wasn't kept for some reason), the origin link in\nthe two maintenance branches would point to a non-existent commit ID,\nand you wouldn't be able to estabish a linkage between them.  By using\nan independent bug identifer as the way of creating the linkage,\nyou're preserving *much* more useful information, and you can reliably\nestablish a relationship between two commits.\n\n\nIn terms of your arguments about why free-form is bad, in another message:\n\n>- No strict definition of what it means.\n>- Diverging porcelain implementations making use of the field in ever so\n>  slightly changing ways over the years.\n\nThis can be a problem regardless of where you store the information.\nWhether you store it in the free-form text or in the git object\nheader, if you don't make sure it is well-defined, you're in trouble.\n\n>- You cannot rely on the field being always available.\n\nThis is true regardless of where you store it; older versions of git\nwon't store the git origin link, for example, unless you plan to break\nbackwards compatibility with all existing git repositories, which\nwould be a bad idea.  :-)\n\nOne nice thing of using text in free-form text fields is that anyone\ncan enter it without needing a new version of git.  The downside is\nthat people could typo the header in some fashion.  But that can be\ndealt with in a newer version of the git porcelain validates the bug\nidentifier and/or checks for obvious spelling mistakes and issues a\nwarning (\"Looks like you may have mispelled 'Adresses-Bug'; perhaps\nyou should fix this via git commit --amend?\").  \n\nIn contrast, if you put it in the git object header, there is no\npossibility of using the field at all until you update to a version of\ngit that supports it.  And some developer on your project is using an\nolder version of git when they rebase or cherry-pick a commit, the\norigin header will be completely lost; but if it is stored in the\nfree-form area, the information will be brought along for the ride for\nfree.\n\n>- Automated \"renumbering\" becomes difficult at best.\n\nThis is actually one of the reasons why I don't like the origin link.\nIf you use the origin link, it's *still* not obvious whether you\nshould rewrite the commit ID or not.  For example, in some workflows,\nyou have two branches pointing to the same commit before you do the\nrebase, where the rebase will only update the current branch pointer,\nbut there is another branch still pointing at the original series of\ncommits.  Worse yet, someone may have done a cherry-pick *before* the\nrebase.  Hence, the only thing you can do is keep *both* commit ID's.\nThis means that over time, you can't get rid of any commit ID's when\nyou do a rebase, which means the number of commit ID's in the origin\nlink will always increase whenever you do a rebase or a cherry-pick.\n\nThis is why for the use case where you are trying to figure out\nwhether a bug exists in a particular branch, it is ***much*** better\nto rendevous using a bug identifier; it provides an extra layer of\nindirection which results in a much more stable identifer that is\nguaranteed to work.\n\n\nI understand it won't work for those cases where you don't have a bug\ntracking identifer, but in fact, if you need this functionality at all\n(and I am not convinced that you do), the ***much*** better approach\nis to use the same approach as the bug tracking identifier, and add a\nlevel of indirection.  How would that work in practice?  Whenever you\ncreate a new commit, create a UUID which is assigned to the patch.\nThis UUID is not modified by git rebase or git cherry pick, and it\nshould be optionally kept or modified on a git commit --amend.\nIdeally, said UUID would exported via git-format-patch, and imported\nvia git-am, and via systems that use patches, such as guilt or stg.\nThis becomes a handy way of recognizing patches even if they aren't\nbeing stored in git --- for example, Andrew Morton's mm patch series.\n\nNow, whether you store this UUID in the free-form text area, or in the\ngit object header, in the long run really doesn't matter.  You can\njust as easily have porcelein suppress a line in the free-form text\narea, as you can have the procelain print the UUID when it is stored\nin the object header.\n\nYes, it means that you have to maintain a separate database so you can\neasily find the list of commits that contain a particular UUID, but I\nsuspect you would need this in the case of the origin link concept\nanyway, since sooner or later some of the more useful uses of said\nlink would require you to be able to find the commits which had origin\nlinks to the original commit, which means you would need to create and\nmaintain this database anyway.  And the maintenance of this database\nis purely optional; you only need it if you care about efficiently\nlooking up UUID's, and given \"time git log > /dev/null\" on the kernel\ntree only takes six seconds on my laptop, and \"git log > /dev/null\"\nonly takes 0.148 seconds for e2fsprogs, for many projects you might\nnot even need the database to accelerate lookups via UUID.\n\n    \t      \t  \t      \t\t \t     - Ted\n"},{"id":"90346","messageId":"20080910123026.GJ10360@machine.or.cz","threadId":"15449","inReplyTo":"20080910114940.GA14127@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-09-10T12:30:26Z","receivedAt":"2008-09-10T12:30:26Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Wed, Sep 10, 2008 at 01:49:40PM +0200, Stephen R. van den Berg wrote:\n> Maybe we have a misunderstanding about what \"follow a link\" means and\n> when it is done.\n> During most normal git operation, the origin links are just read, but\n> not followed.\n> The only commands that I expect to follow them are log --graph, gitk, fsck\n> and blame.  I may have missed some corner use-cases, but this should\n> cover most of it; i.e. most of git ignores them or just makes note of\n> the hashvalues provided.\n\nOh, I'm sorry. By\n\n\t- During fetch/push/pull the full commit including the origin fields is\n\t  transmitted, however, the objects the origin links are referring to\n\t  are not (unless they are being transmitted because of other reasons).\n\nI have understood that you fetch the origin target but not commits\nreferred from it, but instead you meant that you do not follow the\norigin link at all.\n\n\t\t\t\tPetr \"Pasky\" Baudis\n"},{"id":"90347","messageId":"20080910131424.GA7397@cuci.nl","threadId":"15449","inReplyTo":"20080910123026.GJ10360@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T13:14:24Z","receivedAt":"2008-09-10T13:14:24Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Petr Baudis wrote:\n>On Wed, Sep 10, 2008 at 01:49:40PM +0200, Stephen R. van den Berg wrote:\n>> Maybe we have a misunderstanding about what \"follow a link\" means and\n>> when it is done.\n\n>Oh, I'm sorry. By\n\n>\t- During fetch/push/pull the full commit including the origin fields is\n>\t  transmitted, however, the objects the origin links are referring to\n>\t  are not (unless they are being transmitted because of other reasons).\n\n>I have understood that you fetch the origin target but not commits\n>referred from it, but instead you meant that you do not follow the\n>origin link at all.\n\nIndeed.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90351","messageId":"20080910141630.GB7397@cuci.nl","threadId":"15449","inReplyTo":"20080910122118.GI21071@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T14:16:30Z","receivedAt":"2008-09-10T14:16:30Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Theodore Tso wrote:\n>On Wed, Sep 10, 2008 at 12:56:03AM +0200, Stephen R. van den Berg wrote:\n\n>use.  One of the problems with the origin link is that it is only a\n>one way pointer.  Given a newer commit, you know that it is (somehow)\n>weekly related to a older commit.  So your proposed workflow only\n>works if cherry-picks only happen in one direction.  That isn't always\n>true, especially in distributed environments where the bugfix might\n>happen on someone else's development branch, and then it gets pulled\n>in, or perhaps rebased in, and you want to know they are related.\n\nWell, the definition of the origin link (and a back/forwardport) is that:\n- You (as a developer) consider the link relevant for posterity (IOW,\n  you consider it to be a proper back/forwardport which should be\n  recognisable as such).\n- The back/forwardport always has to reference some existing (stable) commit.\n\nEspecially the second condition always holds at the time of creation of\nthe backport (or forwardport, for that matter).  I'm not quite sure\nwhich circumstances you allude to above which would violate this\nrequirement, can you elaborate on that?\n\n>I would argue the best way to do that is to store (either in the\n>object or in the free-form text area) not the link, which would have\n>to get renumbered but rather the identifier for the bug(s) that this\n\nThe renumbering is not a problem, renumbering is a rare operation since\na project's history is supposed to be stable.  And even if renumbering\nis performed, it is a well understood operation of which the renumbering\nof the origin links imposes a negligible overhead on top of the existing\nrenumbering overhead.\n\n>commit fixes.  So for example, consider a convention where in the body\n>of the free-form text area, before the Signed-off-by:, Acked-by:, and\n>CC: headers for those projects that use them, we add something like\n>the following:\n\n>Addresses-Bug: Red_Hat/149480, Sourceforge_Feature/120167\n>or\n>Addresses-Bug: Debian/432865, Launchpad/203323, Sourceforge_Bug/1926023\n\n>Once you have this information, it is not difficult to maintain a\n>berk_db database which maps a particular Bug identifier (i.e.,\n>Red_Hat/149480, or Debian/471977, or Launchpad/203323) to a series of\n>commits.\n\nThis is nice, I admit, but it has the following downsides:\n- It is nontrivial to automate this on execution of \"git cherry-pick\".\n- In a distributed environment this requires a network-reachable bug\n  database.\n- A network-reachable bug database means that suddenly git needs network\n  access for e.g. cherry-pick, revert, gitk, log --graph, blame.\n- Network queries for commits containing references kind of kills\n  performance.\n- Some backports don't have entries in a bug database because they\n  weren't bugs to begin with, in which case it becomes impossible to add\n  an identifier to the commit message after the fact.\n- It relies heavily on tools outside of git-core, which raises the\n  threshold for using it.\n\n>The advantage of this scheme is that if a bug has been fixed in\n>multiple branches, you can see the association between two commits in\n>two different branches very easily.  Furthermore, you get a link back\n>to the actual bug in one or more bug tracking systems, which the some\n>porcelain program could use to transform into a hot-link which when\n>clicked opens up a browser window to the bug in question.\n\nI'm not opposed to links like this, but I consider them a useful extra.\nThe link back is computationally of the same order of magnitude to find\nall existing children of a certain commit; which is well understood and\nwithin reach in most cases.\n\n>In contrast, using your proposed origin scheme, if the bug was\n>originally created in some development branch, and then cherry picked\n>into two separate maintenance branches, if you don't have the\n>development branch in your repository (maybe for some reason that\n>development branch wasn't kept for some reason), the origin link in\n>the two maintenance branches would point to a non-existent commit ID,\n>and you wouldn't be able to estabish a linkage between them.  By using\n\nYes, you would.  You'd notice that either:\n- One origin will point to the other commit (recommended practice,\n  cherry-pick ripple-through, so to speak).\n- Both origin links point to the same non-existent commit.\n\n>In terms of your arguments about why free-form is bad, in another message:\n\n>>- No strict definition of what it means.\n>>- Diverging porcelain implementations making use of the field in ever so\n>>  slightly changing ways over the years.\n\n>This can be a problem regardless of where you store the information.\n\nTrue.  The point is that specifying a definition for a origin\nheaderfield will narrow down how it is and can be used.  Free-form is\njust that, free-form, and merely defines things by convention.\n\n>Whether you store it in the free-form text or in the git object\n>header, if you don't make sure it is well-defined, you're in trouble.\n\nFree form can take the form of plaintext explanations detailing the\nrelationship in a foreign language (worst case example).\n\n>>- You cannot rely on the field being always available.\n\n>This is true regardless of where you store it; older versions of git\n>won't store the git origin link, for example, unless you plan to break\n>backwards compatibility with all existing git repositories, which\n>would be a bad idea.  :-)\n\nTrue.  What I was alluding to, is that if someone includes a\nback/forwardport link in the free-form part of the commit message, then\nyou cannot predict how they'll do that.  In case of the origin link,\n*if* it is used, it will always look the same.\n\n>One nice thing of using text in free-form text fields is that anyone\n>can enter it without needing a new version of git.  The downside is\n\nGit is rather portable, I'd say, so anyone wanting to use the new\nfeature can be bothered to upgrade.\n\n>that people could typo the header in some fashion.  But that can be\n>dealt with in a newer version of the git porcelain validates the bug\n>identifier and/or checks for obvious spelling mistakes and issues a\n>warning (\"Looks like you may have mispelled 'Adresses-Bug'; perhaps\n>you should fix this via git commit --amend?\").  \n\nYou mean you'd prefer some kind of AI solution to aid the user in\nwriting misspelling-free bug identifiers over a simple clean origin link\nin the header of a commit message?\n\n>In contrast, if you put it in the git object header, there is no\n>possibility of using the field at all until you update to a version of\n>git that supports it.  And some developer on your project is using an\n>older version of git when they rebase or cherry-pick a commit, the\n>origin header will be completely lost; but if it is stored in the\n>free-form area, the information will be brought along for the ride for\n>free.\n\nSame as above:\nIf developers care about the backport information, they *can* be\nbothered to upgrade git.  It's not rocketscience.\n\n>>- Automated \"renumbering\" becomes difficult at best.\n\n>This is actually one of the reasons why I don't like the origin link.\n>If you use the origin link, it's *still* not obvious whether you\n>should rewrite the commit ID or not.  For example, in some workflows,\n>you have two branches pointing to the same commit before you do the\n>rebase, where the rebase will only update the current branch pointer,\n>but there is another branch still pointing at the original series of\n>commits.  Worse yet, someone may have done a cherry-pick *before* the\n>rebase.  Hence, the only thing you can do is keep *both* commit ID's.\n>This means that over time, you can't get rid of any commit ID's when\n>you do a rebase, which means the number of commit ID's in the origin\n>link will always increase whenever you do a rebase or a cherry-pick.\n\nThe recommended practice here is quite simple:\n\n- Origin links should only be created pointing to stable commits (i.e.\n  commits which you'd be willing to publish or already have published).\n\n- This implies that pointing an origin link at a commit in a strain that\n  you still want to rebase is asking for trouble.  Doing this is akin to\n  doing a merge between two branches and then you start rebasing 4\n  commits *below* the mergepoint.  Don't do that.\n\n- The only special case I'd allow is if you rebase a strain and the\n  origin link points from one of the commits in the strain to be rebased\n  back *into* the same strain being rebased (most likely a revert).\n  Rebase can be bothered to renumber the origin link in this case.\n\nAnd when you stick to those rules, the problem you're describing doesn't\nhappen.\n\n>This is why for the use case where you are trying to figure out\n>whether a bug exists in a particular branch, it is ***much*** better\n>to rendevous using a bug identifier; it provides an extra layer of\n>indirection which results in a much more stable identifer that is\n>guaranteed to work.\n\nUnless that commit already lies in the past, and you have no way to\nactually add the bugid to the commit.\n\n>(and I am not convinced that you do), the ***much*** better approach\n>is to use the same approach as the bug tracking identifier, and add a\n>level of indirection.  How would that work in practice?  Whenever you\n>create a new commit, create a UUID which is assigned to the patch.\n\nThis only works if you know at time of commit that you want to backport\nit at some later date.\n\n>Now, whether you store this UUID in the free-form text area, or in the\n>git object header, in the long run really doesn't matter.  You can\n>just as easily have porcelein suppress a line in the free-form text\n>area, as you can have the procelain print the UUID when it is stored\n>in the object header.\n\nTrue.  It's almost as much work.  Though it seems rather silly to start\nsuppressing lines in the free-form text area, if one can add a proper\nheaderfield.\n\n>Yes, it means that you have to maintain a separate database so you can\n>easily find the list of commits that contain a particular UUID, but I\n>suspect you would need this in the case of the origin link concept\n>anyway, since sooner or later some of the more useful uses of said\n>link would require you to be able to find the commits which had origin\n>links to the original commit, which means you would need to create and\n>maintain this database anyway.\n\nThat isn't true.  Finding commits which have origin links to a certain\ncommit is just as hard as finding all children of a certain commit.\nIt's not exactly instant, but it is not a big problem, and depending on\nthe amount of repositorytraversal you already are doing, it might even\nbe a negligible amount of extra overhead.\n\n>  And the maintenance of this database\n>is purely optional; you only need it if you care about efficiently\n>looking up UUID's, and given \"time git log > /dev/null\" on the kernel\n>tree only takes six seconds on my laptop, and \"git log > /dev/null\"\n>only takes 0.148 seconds for e2fsprogs, for many projects you might\n>not even need the database to accelerate lookups via UUID.\n\nThe database needs to be available to anyone doing a clone of the\nrepository, which implies that:\n- It needs to be network based.\n- It needs controlled write access (which is a mess).\n- It is slow during blame/gitk operations.\n- It is rather nontrivial to get things setup such that someone (after\n  cloning the repository) is able to run cherry-pick/gitk/blame/revert\n  and have those commands use the database transparently.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90353","messageId":"20080910143329.GE28210@dpotapov.dyndns.org","threadId":"15449","inReplyTo":"48C794D6.20001@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-09-10T14:33:29Z","receivedAt":"2008-09-10T14:33:29Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Wed, Sep 10, 2008 at 11:35:18AM +0200, Paolo Bonzini wrote:\n> \n> Junio explained that the way achieves the same effect in git is by\n> forking the topic branch off the \"oldest\" branch where the patch will\n> possibly be of interest.  Then he can merge it in that branch and all\n> the newest ones.  That's great, but not all people are as\n> forward-looking (he did say that sometimes he needs to cherrypick).\n\nThose who base their work on the newest ones must be very forward-\nlooking :) but, seriously, cherry-picking is *not* a normal workflow\nwith Git. Git is optimized for easy merging while cherry-picking is\na rare operation reserved for correcting some past mistakes.\n\n> \n> Another problem is that in some projects actually there are two \"maint\"\n> branches (e.g. currently GCC 4.2 and GCC 4.3), and most developers do\n> not care about what goes in the older \"maint\" branch; they develop for\n> trunk and for the newer \"maint\" branch, and then one person comes and\n> cherry-picks into the older \"maint\" branch.  This has two problems:\n> \n> 1) Having to fork topic branches off the older branch would force extra\n> testing on the developers.\n\nIf a branch is meant to included in the oldest version, it must be\ntested with that version anyway, and it is better when it is written for\nthe old version, because functions tend to be more backward compatible\nthan forward compatible. In other words, functions may often acquire\nsome extra functionality over time without changing their signature, so\nthe code written for a new version will merge without any conflict to\nthe old one, but it won't work correctly under some conditions. It is\ncertainly possible to have a problem in the opposite direction, but it\nis much less likely, and usually bugs introduced in the development\nversion are not as bad as destabilizing a stable branch. Thus starting\nbranch that is clearly meant for inclusion to the old version from that\nversion is the right thing do.\n\nOf course, if you have more than one stable branch for a long time then\nyou may want some branches forked from the new stable. You can do that\nby merging uninteresting changes from the new stable with the 'ours'\nstrategy (so they will be ignored), and after that merging actually\ninteresting features from the new stable.\n\nIn contrast to cherry-picking, the real merge creates the history that\ncan be easily visualized and understood.\n\n> \n> 2) Besides this, topic branches are not cloned, so if I am the\n> integrator on the older \"maint\" branch, I need to dig manually in the\n> commits to find bugfixes.  True, I could use Bugzilla, but what if I\n> want to use git instead?  There is \"git cherry -v ... | grep -w ^+.*PR\",\n> except that it has too many false negatives (fixes that have already\n> been backported, but do show up in the list).\n\nIf you clearly mark all bugs in the commit message, there will be no\nproblem to find them by grepping log. There is a lot of potentially\nuseful information, and the 'origin' link is just one of many. It may\nbe okay to do some general mechanism for custom commit attributes (if\nit's really necessary), but making a hack for one specific item of\ninformation feels very wrong. In fact, I have not convinced at all\nthat the free-form text is not suitable to store this information.\n\n\nDmitry\n"},{"id":"90355","messageId":"20080910151015.GA8869@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080910141630.GB7397@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-10T15:10:16Z","receivedAt":"2008-09-10T15:10:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 10, 2008 at 04:16:30PM +0200, Stephen R. van den Berg wrote:\n\n> >Once you have this information, it is not difficult to maintain a\n> >berk_db database which maps a particular Bug identifier (i.e.,\n> >Red_Hat/149480, or Debian/471977, or Launchpad/203323) to a series of\n> >commits.\n> \n> This is nice, I admit, but it has the following downsides:\n> - It is nontrivial to automate this on execution of \"git cherry-pick\".\n\nMaybe a cherry-picking hook?\n\n> - In a distributed environment this requires a network-reachable bug\n>   database.\n\nUse a distributed bug tracking system (DBTS).\n\n> - A network-reachable bug database means that suddenly git needs network\n>   access for e.g. cherry-pick, revert, gitk, log --graph, blame.\n\nUse a DBTS.\n\n> - Network queries for commits containing references kind of kills\n>   performance.\n\nUse a DBTS.\n\n> - Some backports don't have entries in a bug database because they\n>   weren't bugs to begin with, in which case it becomes impossible to add\n>   an identifier to the commit message after the fact.\n\nUse a DBTS, since then you can generally make up a new UUID on the spot.\n\n> - It relies heavily on tools outside of git-core, which raises the\n>   threshold for using it.\n\nTrue.\n\nBut maybe Ted is on to something here. Rather than adding the\ninformation to the commit object itself, why not maintain a separate\nmapping, but keep it _within git_. That is how most of the DBTS's work\nthat I have seen. Maybe it is possible to implement some subset of the\nfeatures in a tool that could become part of core git.\n\nThere was a proposal at some point for a \"notes\" feature which would\nallow after-the-fact annotation of commits. I don't recall the exact\ndetails, but I think it stored its information as a git tree of blobs.\nYou could choose whether or not to transfer the notes based on\ntransferring a ref pointing to the notes tree.\n\nI'm not sure how applicable this is to your problem, but if you want to\ninvestigate you can find discussion in the list archive under the name\n\"notes\".\n\n-Peff\n"},{"id":"90357","messageId":"20080910151542.GA10523@cuci.nl","threadId":"15449","inReplyTo":"20080910143329.GE28210@dpotapov.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T15:15:42Z","receivedAt":"2008-09-10T15:15:42Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Dmitry Potapov wrote:\n>On Wed, Sep 10, 2008 at 11:35:18AM +0200, Paolo Bonzini wrote:\n>> Another problem is that in some projects actually there are two \"maint\"\n>> branches (e.g. currently GCC 4.2 and GCC 4.3), and most developers do\n>> not care about what goes in the older \"maint\" branch; they develop for\n>> trunk and for the newer \"maint\" branch, and then one person comes and\n>> cherry-picks into the older \"maint\" branch.  This has two problems:\n\n>> 1) Having to fork topic branches off the older branch would force extra\n>> testing on the developers.\n\n>If a branch is meant to included in the oldest version, it must be\n>tested with that version anyway, and it is better when it is written for\n>the old version, because functions tend to be more backward compatible\n>than forward compatible. In other words, functions may often acquire\n>some extra functionality over time without changing their signature, so\n>the code written for a new version will merge without any conflict to\n>the old one, but it won't work correctly under some conditions. It is\n>certainly possible to have a problem in the opposite direction, but it\n>is much less likely, and usually bugs introduced in the development\n>version are not as bad as destabilizing a stable branch. Thus starting\n>branch that is clearly meant for inclusion to the old version from that\n>version is the right thing do.\n\n>Of course, if you have more than one stable branch for a long time then\n>you may want some branches forked from the new stable. You can do that\n>by merging uninteresting changes from the new stable with the 'ours'\n>strategy (so they will be ignored), and after that merging actually\n>interesting features from the new stable.\n\n>In contrast to cherry-picking, the real merge creates the history that\n>can be easily visualized and understood.\n\nCould you explain how the above mechanisms work based on the following\ncherry-pick action:\n\nA -- B -- C -- D -- L\n      \\            /\n       E -- F -- G -- H -- K\n\nD is the stable branch.\nK is the development branch.\nG is cherry-picked and applied to D producing L.\nThe origin link of L would have contained (G, F).\n\nHow would such a workflow be implemented using the temporary branches\nyou describe?\n\n>If you clearly mark all bugs in the commit message, there will be no\n>problem to find them by grepping log. There is a lot of potentially\n\nSometimes they're not bugs, yet they still are backported and thus carry\nno special marks.\n\n>useful information, and the 'origin' link is just one of many. It may\n\nTrue, but it's one of the few machine-useable ones.\n\n>be okay to do some general mechanism for custom commit attributes (if\n>it's really necessary),\n\nThat's the problem, a general mechanism is undesirable, that we already\nhave the free-form textfield for.\n\n> but making a hack for one specific item of\n>information feels very wrong.\n\nIt's a rather well-defined usefull property (which precludes it from\nbeing a hack, I suppose).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90358","messageId":"48C7E6B3.7030403@gnu.org","threadId":"15449","inReplyTo":"20080910151542.GA10523@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert, and more about porcelain-level metadata","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T15:24:35Z","receivedAt":"2008-09-10T15:24:35Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"> > > fork topic branches off the older branch [and merge them]\n>\n> Could you explain how the above mechanisms work based on the following\n> cherry-pick action:\n> \n> A -- B -- C -- D -- L\n>       \\            /\n>        E -- F -- G -- H -- K\n> \n> D is the stable branch.\n> K is the development branch.\n> G is cherry-picked and applied to D producing L.\n> The origin link of L would have contained (G, F).\n> \n> How would such a workflow be implemented using the temporary branches\n> you describe?\n\nYou don't.  You do everything in topic branches based off the stable\nbranch, and you merge them.  That's the other way round, compared to\nwhat you (and I) are used to.\n\nPaolo\n"},{"id":"90359","messageId":"alpine.LFD.1.10.0809100828360.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"20080910054244.GB15715@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-10T15:30:39Z","receivedAt":"2008-09-10T15:30:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 10 Sep 2008, Stephen R. van den Berg wrote:\n> \n> But then how would someone who clones the repository get at the information?\n\nYou just said it wouldn't get there with fetches.\n\nIf clone acts differently from a \"full\" fetch, something is really really \nwrong.\n\n> The information is essential to understand backports between the various\n> stable branches.\n\nNo it's not. You can mention the backport explicitly in the commit \nmessage, and then you get hyperlinks in the graphical viewers. That works \nwhen people _want_ it to work, instead of in some hidden automatic manner \nthat does entirely the wrong thing in all the common cases.\n\nWhat more do you want?\n\n\t\tLinus\n"},{"id":"90360","messageId":"alpine.LFD.1.10.0809100830570.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"48C785C3.9010204@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-10T15:32:12Z","receivedAt":"2008-09-10T15:32:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 10 Sep 2008, Paolo Bonzini wrote:\n> \n> Not really local data.  More like _weakly referenced_ data.  If it is\n> there, cool.  If it is not there, no big deal.\n\nYou think it's \"cool\".\n\nI think it is \"unreliable, random, and depends on the phase of the moon\".\n\nMy definition of \"cool\" is a totally different thing. What you describe is \nthe very anti-thesis of cool.\n\nIf you want unreliable and random, use CVS. Please.\n\n\t\tLinus\n"},{"id":"90363","messageId":"48C7E9A1.5080409@gnu.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809100830570.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T15:37:05Z","receivedAt":"2008-09-10T15:37:05Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"Linus Torvalds wrote:\n> \n> On Wed, 10 Sep 2008, Paolo Bonzini wrote:\n>> Not really local data.  More like _weakly referenced_ data.  If it is\n>> there, cool.  If it is not there, no big deal.\n> \n> You think it's \"cool\".\n> \n> I think it is \"unreliable, random, and depends on the phase of the moon\".\n\nI think that shallow clones are not any different from this.  If the\nrequired piece of history is there, cool.  If they're not there, no big\ndeal.\n\nI understood the hyperbole, but I think that it's not unreliable,\nbecause all it relies on is the uniqueness of SHA1 values.\n\nPaolo\n"},{"id":"90365","messageId":"alpine.LFD.1.10.0809100841080.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"48C7E9A1.5080409@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-10T15:43:25Z","receivedAt":"2008-09-10T15:43:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 10 Sep 2008, Paolo Bonzini wrote:\n> \n> I think that shallow clones are not any different from this.  If the\n> required piece of history is there, cool.  If they're not there, no big\n> deal.\n\nSure. I don't use them either. But because I don't use them, it doesn't \naffect me. It also doesn't change the core git data structures in any way \nto introduce any new problems.\n\nAlso, if there isn't a required piece of history, things generally break \nvery loudly. IOW, there are only certain things you can do with a shallow \nrepo. In general it's absolutely _not_ a \"no big deal\" issue, quite the \nreverse - it's a deal-breaker.\n\n\t\t\tLinus\n"},{"id":"90366","messageId":"alpine.LFD.1.10.0809100844040.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809100841080.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-10T15:46:19Z","receivedAt":"2008-09-10T15:46:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 10 Sep 2008, Linus Torvalds wrote:\n>\n> Sure. I don't use them either. But because I don't use them, it doesn't \n> affect me. It also doesn't change the core git data structures in any way \n> to introduce any new problems.\n\nBtw, so far nobody has even _explained_ what the advantage of the origin \nlink is. It apparently has no effect for most things, and for other things \nit has some (unspecified) effect when it can be resolved.\n\nApart from the \"dotted line\" in graphical history viewers, I haven't \nactually heard any single concrete example of exactly what it would *do*.\n\nAnd that dotted line really does sound like something you could do with \njust the existing \"hyperlink\" functionality in the commit message.\n\n\t\tLinus\n"},{"id":"90368","messageId":"48C7EE71.8080305@gnu.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809100844040.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T15:57:37Z","receivedAt":"2008-09-10T15:57:37Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> Btw, so far nobody has even _explained_ what the advantage of the origin \n> link is. It apparently has no effect for most things, and for other things \n> it has some (unspecified) effect when it can be resolved.\n> \n> Apart from the \"dotted line\" in graphical history viewers, I haven't \n> actually heard any single concrete example of exactly what it would *do*.\n\nI mentioned git-cherry as an additional use case.  Automatic rename\ndetection works because it might have the occasional false negative, but\nit has practically no false positive, and those are what screws up\nmerges.  But automatic changeset detection a la git-patch-id has too\nmany false negatives to make the current implementation of git-cherry\npractical, and here's when the origin link comes in.  Also, automatic\nchangeset detection does not work with reverts, only with cherry-picks.\n\nBlame could also use the origin link to go backwards in the history and\nfind the origin of the code, without being fooled by reverts.\n\nI'll quote another message I sent in the thread:\n\n>> And why are the notes created by git cherry-pick -x insufficient for that?\n> \n> For example, these notes (or the ones created by \"git revert\") are\n> *wrong* because they talk about commits instead of changesets (deltas\n> between two commits).\n> \n> Why is only one commit present?  Because these messages are meant for\n> users, not for programs.  That's easy to see: users think of commits as\n> deltas anyway, even though git stores them as snapshots---\"git show\n> HEAD\" shows a delta, not a snapshot.\n> \n> And what does this mean for programs?  That they must resort to\n> commit-message scraping to distinguish the two cases. (*)\n> \n>    (*) A GUI blame program, for example, would need to distinguish\n>    whether code added by a commit is taken from commit 4329bd8, or is\n>    reverting commit 4329bd8.  (In the first case, the author of that\n>    code is whoever was responsible for that code in 4329bd8; in the\n>    second case, it is whoever was responsible for that code in\n>    4329bd8^).  If recording changesets, you see 4329bd8^..4329bd8 in\n>    the first case, and 4329bd8..4329bd8^ in the second, so it is trivial\n>    to follow the chain.\n> \n> And scraping is bad.  Imagine people that are writing commit messages in\n> their native language.  What if they patch git to translate the magic\n> notes created by \"git cherry-pick -x\" or \"git revert\" (maybe a future\n> version of git will do that automatically)?  Should they translate also\n> every program that scrapes the messages?\n> \n> \n> Whenever there is a piece of data that could be useful to programs (no\n> matter if plumbing or porcelain), I consider free form notes to be bad.\n> Because data is data, and metadata is metadata.\n> \n> If there was a generic way to put porcelain-level metadata in commit\n> messages (e.g. Signed-Off-By and Acknowledged-By can be already\n> considered metadata), I would not be so much in favor of \"origin\" links\n> being part of the commit object's format.  Now if you think about it,\n> commit references within this kind of metadata would have mostly the\n> properties that Stephen explained in his first message:\n> \n> 1) they would be rewritten by git-filter-branch\n> \n> 2) these references, albeit weak by default\n\n(Note to Linus: reinforcement of your disagreement will be implicitly\nassumed :-)\n\n> could optionally be\n> followed when fetching (either with command-line or configuration options)\n> \n> 3) they would not be pruned by git-gc, unlike notes\n> \n> 4) possibly, git rev-list --topo-order would sort commits by taking into\n> account metadata references too.\n> \n> So the implementation effort would be roughly the same.\n> \n> But, can you think of any other such metadata?  Personally I can't, so\n> while I understand the opposition to a new commit header field that\n> would be there from here to eternity (or until the LHC starts), I do\n> think it is the simplest thing that can possibly work.\n\nPaolo\n"},{"id":"90369","messageId":"20080910161852.GR21071@mit.edu","threadId":"15449","inReplyTo":"20080910141630.GB7397@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-10T16:18:52Z","receivedAt":"2008-09-10T16:18:52Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Sep 10, 2008 at 04:16:30PM +0200, Stephen R. van den Berg wrote:\n> The renumbering is not a problem, renumbering is a rare operation since\n> a project's history is supposed to be stable.  And even if renumbering\n> is performed, it is a well understood operation of which the renumbering\n> of the origin links imposes a negligible overhead on top of the existing\n> renumbering overhead.\n\nWell *you* were the one using this as an argument for using the origin\nlink.  But I'll note that in some workflows, rebasing happens all the\ntime when a patch is being developed and moved around.  Sometimes\npatches are created in git, exported as a patch, and then it re-enters\ngit again later (which is another reason why using an external UUID or\nbug tracking identifier is a good thing).\n\n> >Addresses-Bug: Red_Hat/149480, Sourceforge_Feature/120167\n> >or\n> >Addresses-Bug: Debian/432865, Launchpad/203323, Sourceforge_Bug/1926023\n> \n> >Once you have this information, it is not difficult to maintain a\n> >berk_db database which maps a particular Bug identifier (i.e.,\n> >Red_Hat/149480, or Debian/471977, or Launchpad/203323) to a series of\n> >commits.\n> \n> This is nice, I admit, but it has the following downsides:\n> - It is nontrivial to automate this on execution of \"git cherry-pick\".\n\nIt's trivial if it's in the free-form text.  In fact, it happens\nautomatically.  If it's stored within the git commit object, then it\nwill be done in the C code (if you've updated to the latest git;\nagain, one of the advantages of doing it in free-form text).\n\n> - In a distributed environment this requires a network-reachable bug\n>   database.\n> - A network-reachable bug database means that suddenly git needs network\n>   access for e.g. cherry-pick, revert, gitk, log --graph, blame.\n> - Network queries for commits containing references kind of kills\n>   performance.\n\nNo, because you don't need to look up the bug identifier unless you\nwant to, you know, actually look at the bug.  Otherwise, we are just\nusing something like \"debian/432865\" as an identifier; you only need\nto look them up if you want to look up the bug.  Any time you have a\ncollaborative development environment, you will need either a\ncentralized, network accessible bug tracking system, or use a\ndistributed bug tracking system.  Either way, though, if it's just\nmatter of seeing whether or not a bug fix such as debian/432865 is\nfixed by some commit in some branch, using the bug identifier actually\nmakes this *easier*, not harder.\n\n> - Some backports don't have entries in a bug database because they\n>   weren't bugs to begin with, in which case it becomes impossible to add\n>   an identifier to the commit message after the fact.\n\nThis is true.  The transition is a little easier if you are pointing\nto a pre-existing commit, whereas if you need some kind of rendevous\nidentifer (whether it is a bug ID or some UUID).  On the other hand,\nyou've cherry-picked some bug fix using a git that didn't support the\norigin link, you'd also be screwed, so \n\n> - It relies heavily on tools outside of git-core, which raises the\n>   threshold for using it.\n\nWell, it relies on changes to git --- just like the origin link\nrequires changes to git.  If the it is implemented using free-form\ntext, which is a great way to prototype it, you have the *option* of\nimplementing it via either git porcelain changes or outside tools like\nemacs or vi macros (just as most of us who are kernel developers have\neditor macros that insert Signed-off-by: into git commit messages, as\nwell as changes in git porcelain such that \"git am -s\" automatically\nadds the Signed-off-by header).  But given the wildly successful use\nof Signed-off-by in the kernel sources, this objection seems not very\ncredible, to say the least.\n\n> The recommended practice here is quite simple:\n> \n> - Origin links should only be created pointing to stable commits (i.e.\n>   commits which you'd be willing to publish or already have published).\n> \n> - This implies that pointing an origin link at a commit in a strain that\n>   you still want to rebase is asking for trouble.  Doing this is akin to\n>   doing a merge between two branches and then you start rebasing 4\n>   commits *below* the mergepoint.  Don't do that.\n\nRight.  And if we use a UUID to identify commits, then we don't have\nto have these restrictions.\n\n> - The only special case I'd allow is if you rebase a strain and the\n>   origin link points from one of the commits in the strain to be rebased\n>   back *into* the same strain being rebased (most likely a revert).\n>   Rebase can be bothered to renumber the origin link in this case.\n\nNope, because you might have a branch to the original origin link, and\nsome body else may have already done a cherry-pick to the original\norigin commit.  You've hand-waved around the problem by saying, \"don't\ndo that\", but it just points out how **fragile** the origin link\nscheme really is.  It's just not robust.\n\nIn contrast, generating a UUID per commit is much more robust, since\nyou can now export it out of git in a patch, and then re-import it\nlater, and have the right thing happen.\n\n> >(and I am not convinced that you do), the ***much*** better approach\n> >is to use the same approach as the bug tracking identifier, and add a\n> >level of indirection.  How would that work in practice?  Whenever you\n> >create a new commit, create a UUID which is assigned to the patch.\n> \n> This only works if you know at time of commit that you want to backport\n> it at some later date.\n\nI'm suggesting that all commits (once you upgrade to a version of git\nthat supports this --- and you've already handwaved away the question\non whether you can get all of developers for a project to upgrade to\nthe latest git, remember) would have a UUID generated.  That UUID\ncould be stored internal to git, or (perhaps as an initial prototyping\nas a proof of concept, before we add something into the git commit\nrecord **forever**) could be in the free-form text.\n\n> >Yes, it means that you have to maintain a separate database so you can\n> >easily find the list of commits that contain a particular UUID, but I\n> >suspect you would need this in the case of the origin link concept\n> >anyway, since sooner or later some of the more useful uses of said\n> >link would require you to be able to find the commits which had origin\n> >links to the original commit, which means you would need to create and\n> >maintain this database anyway.\n> \n> That isn't true.  Finding commits which have origin links to a certain\n> commit is just as hard as finding all children of a certain commit.\n> It's not exactly instant, but it is not a big problem, and depending on\n> the amount of repositorytraversal you already are doing, it might even\n> be a negligible amount of extra overhead.\n\nMy point is you'll need this separate database anyway, in order to\ndeal with the cases where you have two commits that point to the same\n(non-existent) origin link, one in maint1, and one in maint2, and\ngiven the commit in the maint1 branch, you want to see if there is\nrelated comit in the maint2 branch you'll need this database anyway\n(or you do a brute force search of the repository, which isn't too bad\nfor modest datbases).  It's identical in both cases --- but having a\nUUID field in the commit is much *cleaner*, since it merely states\nthat these two commits introduce the same semantic change; it doesn't\nimply some kind of parent/child relationship which an origin link\nimplies.\n\n> The database needs to be available to anyone doing a clone of the\n> repository, which implies that:\n> - It needs to be network based.\n> - It needs controlled write access (which is a mess).\n> - It is slow during blame/gitk operations.\n> - It is rather nontrivial to get things setup such that someone (after\n>   cloning the repository) is able to run cherry-pick/gitk/blame/revert\n>   and have those commands use the database transparently.\n\nNo it doesn't, since the database can be inferred from the objects in\nthe repository.  So you can generate it locally if you need it, merely\nas an optimization.  The same is true for the origin link proposal, as\nI've said.\n\n    \t\t    \t    \t     - Ted\n"},{"id":"90370","messageId":"200809101823.22072.jnareb@gmail.com","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809100844040.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-10T16:23:20Z","receivedAt":"2008-09-10T16:23:20Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n> On Wed, 10 Sep 2008, Linus Torvalds wrote:\n> >\n> > Sure. I don't use them either. But because I don't use them, it doesn't \n> > affect me. It also doesn't change the core git data structures in any way \n> > to introduce any new problems.\n> \n> Btw, so far nobody has even _explained_ what the advantage of the origin \n> link is. It apparently has no effect for most things, and for other things \n> it has some (unspecified) effect when it can be resolved.\n> \n> Apart from the \"dotted line\" in graphical history viewers, I haven't \n> actually heard any single concrete example of exactly what it would *do*.\n> \n> And that dotted line really does sound like something you could do with \n> just the existing \"hyperlink\" functionality in the commit message.\n\nAs far as I understand (note: I'm neither for, nor against the proposal;\nalthough I think it has thin chance to be accepted, especially soon),\nit is for graphical history viewers, for git-cherry to make it more\nprecise (to detect duplicated/cherry-picked changes better), and in\nthe future possibly to help history-aware merge strategies. And probably\nhelp patch management interfaces.\n\nOn the theoretical front it looks like extension/generalization of\na parent link, marking given commit do be derivative not only some\nset of trees, or some line of history, but also on some changeset.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90371","messageId":"20080910164045.GL10360@machine.or.cz","threadId":"15449","inReplyTo":"20080910141630.GB7397@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-09-10T16:40:45Z","receivedAt":"2008-09-10T16:40:45Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Wed, Sep 10, 2008 at 04:16:30PM +0200, Stephen R. van den Berg wrote:\n> Theodore Tso wrote:\n> >  And the maintenance of this database\n> >is purely optional; you only need it if you care about efficiently\n> >looking up UUID's, and given \"time git log > /dev/null\" on the kernel\n> >tree only takes six seconds on my laptop, and \"git log > /dev/null\"\n> >only takes 0.148 seconds for e2fsprogs, for many projects you might\n> >not even need the database to accelerate lookups via UUID.\n> \n> The database needs to be available to anyone doing a clone of the\n> repository, which implies that:\n> - It needs to be network based.\n> - It needs controlled write access (which is a mess).\n> - It is slow during blame/gitk operations.\n> - It is rather nontrivial to get things setup such that someone (after\n>   cloning the repository) is able to run cherry-pick/gitk/blame/revert\n>   and have those commands use the database transparently.\n\nThe database can just live in a special branch, with trees organized the\nsame way the object database is, possibly in a more optimized way\n(having the HEAD trees cached around inside Git, etc.).  This should be\nno rocked science if the design is given a little thought, and should be\nfairly fast afterwards.\n\nI'm not endorsing assigning UUIDs to commits now at all (but I don't\nhave time to formulate a comprehensive argument against that either).\n\nHowever, having a commit -> nonessential_volatile_metadata database\nwould be useful for many other things as well! For example amending\ncommit messages later, maintaining general linkage between related\ncommits, tracking explicit rename hints for Git (like the Samba guys\nwould appreciate right now, and me many times in the past - note that\nthis is NOT the same as directly tricking renames within Git history)\nor caching expensive computations with mostly static results (like the\nrename detection or maybe pickaxe indexes - that could be quite large,\nso we might want to actually separate different kinds of data to\nseparate branches).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nThe next generation of interesting software will be done\non the Macintosh, not the IBM PC.  -- Bill Gates\n"},{"id":"90374","messageId":"48C80ADC.60207@gnu.org","threadId":"15449","inReplyTo":"20080910164045.GL10360@machine.or.cz","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-10T17:58:52Z","receivedAt":"2008-09-10T17:58:52Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> I'm not endorsing assigning UUIDs to commits now at all (but I don't\n> have time to formulate a comprehensive argument against that either).\n> \n> However, having a commit -> nonessential_volatile_metadata database\n> would be useful for many other things as well!\n\n100 points to Petr. :-)\n\nPaolo\n"},{"id":"90385","messageId":"20080910203249.GX4829@genesis.frugalware.org","threadId":"15449","inReplyTo":"20080909132212.GA25476@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Miklos Vajna","fromEmail":"vmiklos@frugalware.org","sentAt":"2008-09-10T20:32:49Z","receivedAt":"2008-09-10T20:32:49Z","isPatch":false,"sender":{"key":"vmiklos@frugalware.org","avatar":"https://gravatar.com/avatar/401c1cbbb3a5d13e650c691a2c71d6fd0b80df1a01bc74d9f1972675dd58f2bd?d=mp&s=160"},"body":"On Tue, Sep 09, 2008 at 03:22:12PM +0200, \"Stephen R. van den Berg\" <srb@cuci.nl> wrote:\n> origin a1184d85e8752658f02746982822f43f32316803 2\n> author Junio C Hamano <gitster@pobox.com> 1220132115 -0700\n> committer Junio C Hamano <gitster@pobox.com> 1220153445 -0700\n\nFirst, sorry for joining the thread lately, as far as I see the idea I\nwant to shere here was not mentioned by anybody yet.\n\nSo, git revert already includes the \"origin\" of the commit in the commit\nmessage, and I think that is fine for most people.\n\nWhat about adding an option to cherry-pick to add a similar\n\"commit 7b27718bdb1b70166383dec91391df5534d449ee upstream\" or similar\nstring to the commit message?\n\nAs far as I see the kernel -stable tree already have this, but it is\nadded manually and in many different forms, like:\n\n[ Upstream commit 5f3a9a207f1fccde476dd31b4c63ead2967d934f ]\n\ncommit 7b27718bdb1b70166383dec91391df5534d449ee upstream\n\nAlready in Linus' tree:\nhttp://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.6.git;a=commit;h=b25b791b13aaa336b56c4f9bd417ff126363f80b\n\netc.\n\nOnce git would provide a standard way to do this, that could be used to\navoid this.\n"},{"id":"90386","messageId":"alpine.LFD.1.10.0809101654020.23787@xanadu.home","threadId":"15449","inReplyTo":"20080910203249.GX4829@genesis.frugalware.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-10T20:55:19Z","receivedAt":"2008-09-10T20:55:19Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 10 Sep 2008, Miklos Vajna wrote:\n\n> So, git revert already includes the \"origin\" of the commit in the commit\n> message, and I think that is fine for most people.\n> \n> What about adding an option to cherry-pick to add a similar\n> \"commit 7b27718bdb1b70166383dec91391df5534d449ee upstream\" or similar\n> string to the commit message?\n\nIt's already there: -x\n\nNicolas\n"},{"id":"90387","messageId":"20080910210636.GZ4829@genesis.frugalware.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809101654020.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Miklos Vajna","fromEmail":"vmiklos@frugalware.org","sentAt":"2008-09-10T21:06:36Z","receivedAt":"2008-09-10T21:06:36Z","isPatch":false,"sender":{"key":"vmiklos@frugalware.org","avatar":"https://gravatar.com/avatar/401c1cbbb3a5d13e650c691a2c71d6fd0b80df1a01bc74d9f1972675dd58f2bd?d=mp&s=160"},"body":"On Wed, Sep 10, 2008 at 04:55:19PM -0400, Nicolas Pitre <nico@cam.org> wrote:\n> It's already there: -x\n\nThanks, and sorry for not reading carefully the manpage before sending\nthe mail.\n"},{"id":"90389","messageId":"20080910215045.GA22739@cuci.nl","threadId":"15449","inReplyTo":"20080910151015.GA8869@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T21:50:45Z","receivedAt":"2008-09-10T21:50:45Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Wed, Sep 10, 2008 at 04:16:30PM +0200, Stephen R. van den Berg wrote:\n>> This is nice, I admit, but it has the following downsides:\n>> - It is nontrivial to automate this on execution of \"git cherry-pick\".\n\n>Maybe a cherry-picking hook?\n\nYes, that works, but it is non-trivial, especially since it needs to\nwork for gitk, log --graph, blame and revert as well.\n\n>> - In a distributed environment this requires a network-reachable bug\n>>   database.\n\n>Use a distributed bug tracking system (DBTS).\n\nIf it were part of git-core, that would work.\n\n>But maybe Ted is on to something here. Rather than adding the\n>information to the commit object itself, why not maintain a separate\n>mapping, but keep it _within git_. That is how most of the DBTS's work\n>that I have seen. Maybe it is possible to implement some subset of the\n>features in a tool that could become part of core git.\n\nInteresting thought.\n\n>There was a proposal at some point for a \"notes\" feature which would\n>allow after-the-fact annotation of commits. I don't recall the exact\n>details, but I think it stored its information as a git tree of blobs.\n>You could choose whether or not to transfer the notes based on\n>transferring a ref pointing to the notes tree.\n\nThe idea is nice, but if we were to use it to store the origin link\ninformation, the following happens:\n- Origin link information is rare.\n- Yet during a log/gitk/blame run the information might need to\n  be queried for at every commit.\n- Since in most cases the origin information does not exist, this\n  will cause misses to fill the dentry cache for directory lookups, and\n  thus killing performance.\n- In order to make this efficient, a different database lookup system is\n  needed that is fast for misses.\n\nWhereas if the information is part of the commit, it costs nothing in\nthe typical case (no origin information present).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90390","messageId":"20080910215410.GA24432@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080910215045.GA22739@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-10T21:54:10Z","receivedAt":"2008-09-10T21:54:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 10, 2008 at 11:50:45PM +0200, Stephen R. van den Berg wrote:\n\n> >There was a proposal at some point for a \"notes\" feature which would\n> >allow after-the-fact annotation of commits. I don't recall the exact\n> >details, but I think it stored its information as a git tree of blobs.\n> >You could choose whether or not to transfer the notes based on\n> >transferring a ref pointing to the notes tree.\n> \n> The idea is nice, but if we were to use it to store the origin link\n> information, the following happens:\n> - Origin link information is rare.\n> - Yet during a log/gitk/blame run the information might need to\n>   be queried for at every commit.\n> - Since in most cases the origin information does not exist, this\n>   will cause misses to fill the dentry cache for directory lookups, and\n>   thus killing performance.\n> - In order to make this efficient, a different database lookup system is\n>   needed that is fast for misses.\n\nI think you are misunderstanding what I meant by \"git tree\" here. It is\nliterally a git tree object, so you don't ask the filesystem at all. You\nare looking up within the single object file. If it's a miss, you know\nafter seeing that object. If not, then you dereference the blob object\nthat contains the notes.\n\n-Peff\n"},{"id":"90393","messageId":"20080910223427.GB22739@cuci.nl","threadId":"15449","inReplyTo":"20080910215410.GA24432@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T22:34:27Z","receivedAt":"2008-09-10T22:34:27Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Wed, Sep 10, 2008 at 11:50:45PM +0200, Stephen R. van den Berg wrote:\n>> >There was a proposal at some point for a \"notes\" feature which would\n>> >allow after-the-fact annotation of commits. I don't recall the exact\n>> >details, but I think it stored its information as a git tree of blobs.\n>> >You could choose whether or not to transfer the notes based on\n>> >transferring a ref pointing to the notes tree.\n\n>> The idea is nice, but if we were to use it to store the origin link\n>> information, the following happens:\n>> - Origin link information is rare.\n\n>I think you are misunderstanding what I meant by \"git tree\" here. It is\n>literally a git tree object, so you don't ask the filesystem at all. You\n>are looking up within the single object file. If it's a miss, you know\n>after seeing that object. If not, then you dereference the blob object\n>that contains the notes.\n\nI see.  Indeed.  That's a lot better.\nDid the binary search inside tree objects ever get implemented?\n\nIt is unclear why the latest commit notes proposal didn't make it,\nthough I admit that storing the origin link information in there seems\nfeasible.\n\nThe downsides when doing that are:\n- The lookup cost is small, but still noticable, since it is sometimes\n  done on every commit; using the in-commit origin headerfield solves\n  this at negligible cost.\n- The origin information is no longer cryptographically protected (under\n  certain circumstances this could be considered an advantage and a\n  disadvantage at the same time).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90396","messageId":"20080910224403.GC22739@cuci.nl","threadId":"15449","inReplyTo":"48C80ADC.60207@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T22:44:03Z","receivedAt":"2008-09-10T22:44:03Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Petr wrote:\n>> However, having a commit -> nonessential_volatile_metadata database\n>> would be useful for many other things as well!\n\nWhich brings us back to the \"commit notes\" proposal.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90397","messageId":"20080910225518.GA24534@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080910223427.GB22739@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-10T22:55:18Z","receivedAt":"2008-09-10T22:55:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 11, 2008 at 12:34:27AM +0200, Stephen R. van den Berg wrote:\n\n> I see.  Indeed.  That's a lot better.\n> Did the binary search inside tree objects ever get implemented?\n\nI believe it's still linear (and skimming tree-walk.c:find_tree_entry\nseems to confirm). However, one advantage of such an approach is that it\nwill improve as tree lookup improves (e.g., I believe the pack v4 work\nincluded improvements in this area).\n\n> The downsides when doing that are:\n> - The lookup cost is small, but still noticable, since it is sometimes\n>   done on every commit; using the in-commit origin headerfield solves\n>   this at negligible cost.\n> - The origin information is no longer cryptographically protected (under\n>   certain circumstances this could be considered an advantage and a\n>   disadvantage at the same time).\n\nYes, those are inherent in the scheme, as is the upside that one can\nmake and distribute such annotations separately from commit creation.\n\nI haven't thought enough about it to decide whether there is a scenario\nwhere making such a \"cherry-picked from\" annotation might make use of\nthat property.\n\n-Peff\n"},{"id":"90398","messageId":"20080910230906.GD22739@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809100828360.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T23:09:06Z","receivedAt":"2008-09-10T23:09:06Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>On Wed, 10 Sep 2008, Stephen R. van den Berg wrote:\n>> But then how would someone who clones the repository get at the information?\n\n>You just said it wouldn't get there with fetches.\n\nTrue and still valid.\n\n>If clone acts differently from a \"full\" fetch, something is really really \n>wrong.\n\nIt does not act differently.\n\nLet me elaborate:\n\n- The origin field is part of the commit (and only present if\n  *consciously* added by the committer), and therefore is transmitted\n  along with the rest of a commit upon a fetch.\n\n- The commits being referred to by the origin field are *not*\n  transmitted upon a fetch.\n\n- Given a repository with 4 long lived published branches called A, B, C and D\n  and a backport from development branch D cherry-picked -o into branch A\n  which creates an origin field pointing back to (D^,D^^)\n\n- Now you fetch just branch A from this repository.  This will not cause\n  branch D to be pulled in as well.\n\n- However, if you explicitly pull D, the origin information from A to D can\n  be used.  People doing a generic clone get all four branches, and\n  therefore have all the important commits which normally could contain\n  origin links.  Note that even during a clone, commits pointed to by\n  origin links are not being transmitted (unless there already are other\n  reasons to send them along).\n\n>> The information is essential to understand backports between the various\n>> stable branches.\n\n>No it's not. You can mention the backport explicitly in the commit \n>message, and then you get hyperlinks in the graphical viewers. That works \n>when people _want_ it to work, instead of in some hidden automatic manner \n>that does entirely the wrong thing in all the common cases.\n\nCould you spell out one of the common cases where it would do entirely\nthe wrong thing?\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90399","messageId":"20080910231548.GE22739@cuci.nl","threadId":"15449","inReplyTo":"48C7EE71.8080305@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T23:15:48Z","receivedAt":"2008-09-10T23:15:48Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Paolo Bonzini wrote:\n>> Btw, so far nobody has even _explained_ what the advantage of the origin \n>> link is. It apparently has no effect for most things, and for other things \n>> it has some (unspecified) effect when it can be resolved.\n\n>> Apart from the \"dotted line\" in graphical history viewers, I haven't \n>> actually heard any single concrete example of exactly what it would *do*.\n\nIt allows one to follow and view the evolvement of a patch over time during\nthe various backports.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90400","messageId":"20080910231900.GF22739@cuci.nl","threadId":"15449","inReplyTo":"20080910225518.GA24534@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-10T23:19:00Z","receivedAt":"2008-09-10T23:19:00Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Thu, Sep 11, 2008 at 12:34:27AM +0200, Stephen R. van den Berg wrote:\n>> - The origin information is no longer cryptographically protected (under\n>>   certain circumstances this could be considered an advantage and a\n>>   disadvantage at the same time).\n\n>I haven't thought enough about it to decide whether there is a scenario\n>where making such a \"cherry-picked from\" annotation might make use of\n>that property.\n\nBeing able to subvert the authenticity of git blame by providing fake\norigin information is not very appealing.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Am I paying for this abuse or is it extra?\"\n"},{"id":"90401","messageId":"alpine.LFD.1.10.0809101733050.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"20080910230906.GD22739@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-11T00:39:26Z","receivedAt":"2008-09-11T00:39:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n> \n> - However, if you explicitly pull D, the origin information from A to D can\n>   be used.  People doing a generic clone get all four branches, and\n>   therefore have all the important commits which normally could contain\n>   origin links.  Note that even during a clone, commits pointed to by\n>   origin links are not being transmitted (unless there already are other\n>   reasons to send them along).\n\nIOW, it's not actually transferring them and saving them, since a simple \ndelete of the origin branch will basically make them unreachable.\n\nFine. At least it works the same way as fetch, then. But it's still a huge \nmistake, because it really does mean that it is technically no different \nat all to just mentioning the SHA1 in the commit message, the way we \nalready do for backports.\n\nThe \"origin\" link has no _meaning_ for git, in other words.\n\n> >No it's not. You can mention the backport explicitly in the commit \n> >message, and then you get hyperlinks in the graphical viewers. That works \n> >when people _want_ it to work, instead of in some hidden automatic manner \n> >that does entirely the wrong thing in all the common cases.\n> \n> Could you spell out one of the common cases where it would do entirely\n> the wrong thing?\n\nIt carries along information that is worthless and meaningless and hidden.\n\nI refuse to touch such an obviously braindamaged design. It has no sane \n_semantics_. If it doesn't have semantics, it shouldn't exist, certainly \nnot as some architected feature.\n\nNobody has shown any actual sane meaning for it. The only ones that have \nbeen mentioned have been for things like avoiding re-picking commits \nduring a \"git rebase\", but (a) the patch SHA1 does that already for things \nthat are truly identical an (b) since that information isn't reliable \n_anyway_, and since it's apparently a user choice, it's just \"random\".\n\nI'm sorry, but \"good design\" is a hell of a lot more important than some \nmade-up use case that isn't even reliable, and doesn't match any actual \nreal problems that anybody can explain.\n\n\t\t\tLinus\n"},{"id":"90406","messageId":"alpine.LFD.1.10.0809102244100.23787@xanadu.home","threadId":"15449","inReplyTo":"20080910225518.GA24534@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T02:46:19Z","receivedAt":"2008-09-11T02:46:19Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 10 Sep 2008, Jeff King wrote:\n\n> On Thu, Sep 11, 2008 at 12:34:27AM +0200, Stephen R. van den Berg wrote:\n> \n> > I see.  Indeed.  That's a lot better.\n> > Did the binary search inside tree objects ever get implemented?\n> \n> I believe it's still linear (and skimming tree-walk.c:find_tree_entry\n> seems to confirm). However, one advantage of such an approach is that it\n> will improve as tree lookup improves (e.g., I believe the pack v4 work\n> included improvements in this area).\n\nNo, not yet.  Actually that's the part that still needs serious \nthinking.\n\n\nNicolas\n"},{"id":"90411","messageId":"48C8A9A4.7030906@gnu.org","threadId":"15449","inReplyTo":"20080910231900.GF22739@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-11T05:16:20Z","receivedAt":"2008-09-11T05:16:20Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n>>> - The origin information is no longer cryptographically protected (under\n>>>   certain circumstances this could be considered an advantage and a\n>>>   disadvantage at the same time).\n> \n>> I haven't thought enough about it to decide whether there is a scenario\n>> where making such a \"cherry-picked from\" annotation might make use of\n>> that property.\n> \n> Being able to subvert the authenticity of git blame by providing fake\n> origin information is not very appealing.\n\nYou could use a dummy submodule to ensure that each commit pointed to\nthe right set of notes.  It would force to create a separate commit\nwhenever you modified the notes, which is actually not bad.\n\nAlternatively, the header of the commit can be modified to add a pointer\nto a tree object for the notes; I suppose this is more palatable than\nthe origin link.  The tree could be organized in directories+blobs like\n.git/objects to speed up the lookup.\n\nI actually like the commit notes idea, but then I wonder: why are the\nauthor and committer part of the commit object?  How does the plumbing\nuse them?  Isn't that metadata that could live in the \"notes\"?  And so,\nwhy should the origin link have less privileges?\n\nPaolo\n"},{"id":"90416","messageId":"20080911062242.GA23070@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809101733050.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T06:22:42Z","receivedAt":"2008-09-11T06:22:42Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n>> - However, if you explicitly pull D, the origin information from A to D can\n>>   be used.  People doing a generic clone get all four branches, and\n>>   therefore have all the important commits which normally could contain\n>>   origin links.  Note that even during a clone, commits pointed to by\n>>   origin links are not being transmitted (unless there already are other\n>>   reasons to send them along).\n\n>IOW, it's not actually transferring them and saving them, since a simple \n\nCorrect.\n\n>delete of the origin branch will basically make them unreachable.\n\nFalse.\n\nIf you fetch just branches A, B and C, but not D, the origin link from A\nto D is dangling.  Once you have fetched D as well, the origin link from\nA to D is not dangling anymore.  Subsequently deleting branch D but\nkeeping branch A will keep everything in branch D up till the commits\nthe origin link is pointing to alive and prevent those from being\ndeleted.\n\n>Fine. At least it works the same way as fetch, then. But it's still a huge \n>mistake, because it really does mean that it is technically no different \n>at all to just mentioning the SHA1 in the commit message, the way we \n>already do for backports.\n\nNot quite.\n\n>The \"origin\" link has no _meaning_ for git, in other words.\n\nGit will keep alive commits based on origin links once you (the fetcher)\nhas shown interest by fetching the appropriate branches.\n\nAs to \"meaning\" for git, it's there in the form of:\n- --topo-order uses the information to order the output (but only if the\n  target commits of the link are present in the repository).\n\n>> >No it's not. You can mention the backport explicitly in the commit \n>> >message, and then you get hyperlinks in the graphical viewers. That works \n>> >when people _want_ it to work, instead of in some hidden automatic manner \n>> >that does entirely the wrong thing in all the common cases.\n\n>> Could you spell out one of the common cases where it would do entirely\n>> the wrong thing?\n\n>It carries along information that is worthless and meaningless and hidden.\n\nThe common cases would be:\n\na. \"hidden\": It doesn't need to be hidden.  It can be hidden if you want it\n   to be.  We can decide if git hides it sometimes, always or never.\n   So this point is moot.\n\nb. \"meaningless\": Git is all about taking snapshots of sourcetrees and\n   linking them in an orderly fashion.  The origin link is all about\n   taking snapshots of patches and linking them in an orderly fashion.\n   This allows you to see the patch evolve over time, and it allows for\n   diffs between patches.  We're not actually storing patches, we merely\n   store snapshots.  As it happens, the snapshot of a patch is defined\n   by two commit hashes.\n   Doesn't sound meaningless to me.  Just as one needs normal history\n   between commits in a branch to follow development, there is a history\n   of a backport as it \"travels\" from stable branch to stable branch.\n\nc. \"worthless\": Without the tracking of a backport through a series of\n   well-defined patch-snapshots, it becomes kind of haphazard to\n   actually figure out which piece of code came from where.  Having this\n   information in the form of a series of origin links increases the\n   efficiency of a developer maintaining the backports between branches.\n   Maybe you consider that worthless, I consider anything that improves\n   code quality because having access to a concise history of how the\n   code evolved a Good Thing.  Having history of how code evolved is\n   actually one of the main reasons why people use git.  It's just that\n   git lacks support in the tracking of backports.  The origin link\n   fills that gap.  If you don't do backports in your trees, then fine,\n   the origin link will never materialise in your repositories.\n\n>I refuse to touch such an obviously braindamaged design. It has no sane \n>_semantics_. If it doesn't have semantics, it shouldn't exist, certainly \n>not as some architected feature.\n\nIt does have sane semantics, quite well defined, actually.  I'm just not\ngood at explaining them apparently.  Try reading the explanation I gave\nabove.\n\n>Nobody has shown any actual sane meaning for it. The only ones that have \n>been mentioned have been for things like avoiding re-picking commits \n>during a \"git rebase\", but (a) the patch SHA1 does that already for things \n>that are truly identical an (b) since that information isn't reliable \n>_anyway_, and since it's apparently a user choice, it's just \"random\".\n\nQuite frankly I don't see the application for rebase either (yet).\nI'm focusing on sane semantics first, any implications that has for\nusability by rebase will follow from that.  The origin links track\ncontent (patches), nothing else.  They assist the developer in\nunderstanding how patches evolve, any use cases follow from that.\n\n>I'm sorry, but \"good design\" is a hell of a lot more important than some \n>made-up use case that isn't even reliable, and doesn't match any actual \n>real problems that anybody can explain.\n\nPlease focus on the semantics and on the *non*-made up use case of\ndevelopment of several stable branches with backports between them.\nDiscussing made-up use cases is wasting energy at this point.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90420","messageId":"20080911075539.GA27089@cuci.nl","threadId":"15449","inReplyTo":"48C8A9A4.7030906@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T07:55:39Z","receivedAt":"2008-09-11T07:55:39Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Paolo Bonzini wrote:\n>> Being able to subvert the authenticity of git blame by providing fake\n>> origin information is not very appealing.\n\n>You could use a dummy submodule to ensure that each commit pointed to\n>the right set of notes.  It would force to create a separate commit\n>whenever you modified the notes, which is actually not bad.\n\nPossibly, yes.  But we'd have to be careful not to incur too much\noverhead because every indirection will cost, especially since the\norigin link sometimes is checked for on every commit during a treewalk.\nThe fact that it rarely exists means that it should be fast to find out\nthat there are no origin links (which obviously is the common case).\n\n>Alternatively, the header of the commit can be modified to add a pointer\n>to a tree object for the notes; I suppose this is more palatable than\n>the origin link.\n\nThis won't work for the original notes concept, because it makes the\nnotes immutable after commit.  For the origin links this would be fine,\nsince they don't change once committed.\nThe problem with fitting the origin links in the notes is twofold:\n- They become mutable, which is undesirable, I'd like to preserve\n  history as is (just like parent links).\n- There is a performance hit, since origin links need to be found not to\n  exist on every commit (sometimes, depending on the operation of course).\n\n>  The tree could be organized in directories+blobs like\n>..git/objects to speed up the lookup.\n\nYes, that was already in the latest proposal for notes, I believe.\n\n>I actually like the commit notes idea, but then I wonder: why are the\n>author and committer part of the commit object?  How does the plumbing\n>use them?  Isn't that metadata that could live in the \"notes\"?  And so,\n\nIt would fit with a non-mutable version of the notes.  Then again, we\nalready *have* the non-mutable version of the notes, it's called the\nheader of the commit message.\n\n>why should the origin link have less privileges?\n\nThey both belong in the non-mutable notes, and those happen to live in\nthe header of the commit (which *is* the most efficient spot, of course).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90424","messageId":"200809111020.55115.jnareb@gmail.com","threadId":"15449","inReplyTo":"20080911062242.GA23070@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-11T08:20:53Z","receivedAt":"2008-09-11T08:20:53Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Stephen R. van den Berg wrote:\n\n[...]\n> Please focus on the semantics and on the *non*-made up use case of\n> development of several stable branches with backports between them.\n> Discussing made-up use cases is wasting energy at this point.\n\nBy the way, I would really consider trying first to host 'origin' links \nnot in repository database itself, but in some extra database inside \ngit repository, like reflog or index.  Git community is _very_ \nreluctant to modifying / extending format of persistent objects.  From \nall the proposals to add some extra header to a 'commit' object: \nthe 'prior' link to previous version of rebased, cherry-picked or redone \ncommit (superceded somewhat by local reflog, on by default in modern \ngit); the generic 'note' header, with examples of usage including \n_non-linking_ cherry-pick and reverted commit-id, merge strategy used, \nand hints for rename detection, i.e. something like #pragma in C \n(rejected on the grounds that it was too generic and didn't have well \ndefined semantic); the 'generation' header which was meant to help and \nspeed up sorting commits, with root (parentless) commit having \ngeneration of 1, and each commit having generation being 1 more than \nmaximum of generations of its parents (I think that backwards \ncompatibility killed it, and the fact that date-based heuristics was \nimproved); only the 'encoding' header was accepted.\n\nSo I think you should go the route of externally (outside 'commit' \nobjects) maintaing 'origin'/'changeset'/'cset' links (like XLink \nextended links ;-)) as a prototype to examine consequences of the idea. \nThat was the way _submodule_ support was added to Git, by the way.  \nFirst there were (at least) two implementations maintaining submodules \noutside object database (see http://git.or.cz/gitwiki/SubprojectSupport\nespecially \"References\" section), then it was officially added first at \nthe level of plumbing support, as extension of a 'tree' object (and \nindex format, I think).\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90428","messageId":"48C8DABD.40201@gnu.org","threadId":"15449","inReplyTo":"20080911075539.GA27089@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-11T08:45:49Z","receivedAt":"2008-09-11T08:45:49Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n>> I actually like the commit notes idea, but then I wonder: why are the\n>> author and committer part of the commit object?  How does the plumbing\n>> use them?  Isn't that metadata that could live in the \"notes\"?  And so,\n> \n> we already *have* the non-mutable version of the notes, it's called the\n> header of the commit message.\n\nYes, that was my point.  I don't see how the author and committer fit in\nthe header of the commit message, if the origin does not.\n\nPaolo\n"},{"id":"90434","messageId":"48C90F06.4000309@gmail.com","threadId":"15449","inReplyTo":"20080911062242.GA23070@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2008-09-11T12:28:54Z","receivedAt":"2008-09-11T12:28:54Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Stephen R. van den Berg wrote:\n> If you fetch just branches A, B and C, but not D, the origin link from A\n> to D is dangling. \n\nI do not understand how this can be considered an acceptable behavior. \nIf an object ID is referenced in an object header, particularly commit \nobjects, fetch must gather those objects also because to do otherwise \nbreaks the cryptographic authentication in git.\n"},{"id":"90435","messageId":"20080911123148.GA2056@cuci.nl","threadId":"15449","inReplyTo":"200809111020.55115.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T12:31:48Z","receivedAt":"2008-09-11T12:31:48Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jakub Narebski wrote:\n>Stephen R. van den Berg wrote:\n>[...]\n>> Please focus on the semantics and on the *non*-made up use case of\n>> development of several stable branches with backports between them.\n>> Discussing made-up use cases is wasting energy at this point.\n\n>By the way, I would really consider trying first to host 'origin' links \n>not in repository database itself, but in some extra database inside \n>git repository, like reflog or index.  Git community is _very_ \n>reluctant to modifying / extending format of persistent objects.  From \n\nRightfully so, of course.\n\n>So I think you should go the route of externally (outside 'commit' \n>objects) maintaing 'origin'/'changeset'/'cset' links (like XLink \n>extended links ;-)) as a prototype to examine consequences of the idea. \n>That was the way _submodule_ support was added to Git, by the way.  \n>First there were (at least) two implementations maintaining submodules \n>outside object database (see http://git.or.cz/gitwiki/SubprojectSupport\n>especially \"References\" section), then it was officially added first at \n>the level of plumbing support, as extension of a 'tree' object (and \n>index format, I think).\n\nWell, the train of thought here goes as follows:\n1. Sure, why not add a field (zero or more) at the bottom of the free-form\n   commit message reading like:\n\n   Origin: bbb896d8e10f736bfda8f587c0009c358c9a8599 ee837244df2e2e4e9171f508f83f353730db9e53\n\n2. Add support to cherry-pick/revert to actually generate the field upon\n   demand.\n\n3. Then add support to prune/gc/fsck/blame/log --graph to take the field\n   into account.\n\n4. Add support to filter-branch/rebase to renumber the field if necessary.\n\n5. Add support to --topo-order to use the field if present and reachable.\n\n6. For bonus points: add support to log to suppress the display of the\n   field at the end of the commit message, and redisplay the field\n   as Origin: bbb896d..ee83724\n   next to the Parent/Merge fields.\n\nWell, and after having done steps 1 to 5, the net result is that it\nworks almost as if the field is present in the header, except that:\n- It is now at the end of the body in the commit message.\n- It takes more time to find and parse it.\n\nSo that gives two minuses, and no pluses.\nSo short-circuiting the reasoning suggests that since the only thing\nthat actually changes now is the position of the field (at the top or\nend of the commit message), we might as well do it right and put it in\nthe top, that gets rid of the two minuses.\n\nAnything I missed?\n\nBasically it means that:\n\na. If there is a better solution to tracking the backports, I'll gladly\n   use that instead, but simply using the current really freeform\n   approach doesn't cut it (it currently refers to a single commit,\n   instead of a pair of commits, and takes too long to parse out in a\n   --top-order or blame command).  Better solutions I haven't heard so\n   far.\n\nb. I need the integrity protection of a commit to make sure that the\n   origin fields cannot be altered later; blame would be too easy to fool\n   otherwise.  So using the notes solution seems to be out (it would also\n   be quite a performance hit again).\n\nc. I consider the Origin: field at the end of the commit message a\n   workable solution, but it smells like X-header-extension-messes as in\n   E-mail headers, and it incurs a small performance hit (in case of\n   --topo-order/blame/prune/fsck), but maybe this performance hit can be\n   minimised by making sure that the fields are *always* at the end\n   of the commit message.\n\nd. Using the proposed origin header in the standard commit header has\n   close to zero overhead (in most commits the field is not present), yet\n   codecomplexitywise it is almost identical with the Origin: field at\n   the end of the commit message.\n\nI find it remarkable though that people are dragging their feet at\nsolution d, yet are quite ok with solution c.  IMO solution c and d are\nalmost identical, except that solution c is ugly, and solution d is\nelegant.  But if it makes it easier to prove the usefulness by\nimplementing the ugly solution first, that's fine.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90436","messageId":"48C91013.4070907@gmail.com","threadId":"15449","inReplyTo":"20080911075539.GA27089@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2008-09-11T12:33:23Z","receivedAt":"2008-09-11T12:33:23Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Stephen R. van den Berg wrote:\n> It would fit with a non-mutable version of the notes.  Then again, we\n> already *have* the non-mutable version of the notes, it's called the\n> header of the commit message.\n\nAlmost correct. Remove \"header of\" from the above and you'd be correct.\n"},{"id":"90437","messageId":"20080911123902.GB2056@cuci.nl","threadId":"15449","inReplyTo":"48C90F06.4000309@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T12:39:02Z","receivedAt":"2008-09-11T12:39:02Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"A Large Angry SCM wrote:\n>Stephen R. van den Berg wrote:\n>>If you fetch just branches A, B and C, but not D, the origin link from A\n>>to D is dangling. \n\n>I do not understand how this can be considered an acceptable behavior. \n>If an object ID is referenced in an object header, particularly commit \n>objects, fetch must gather those objects also because to do otherwise \n>breaks the cryptographic authentication in git.\n\nNo it does not.\nThe cryptographic seal is calculated over the content of the commit,\nwhich includes the hashes of all referenced objects, but doesn't include\nthe objects themselves.\nThe content of the commit is not violated.\n\nDo not forget though:\n- origin links are a rare occurrence.\n- When they occur, they usually were made to point into other (deemed)\n  important public branches.\n- Due to the fact that the branches they are pointing into are important\n  and public, in most cases the origin links *will* point to objects you\n  actually already have (even if you fetched from someone else).\n- The only time you're going to have dangling origin links is when\n  they were pointing at someone's private branches, in which case it was\n  not very prudent of the committer to actually record the link in the\n  first place.  But nothing breaks if you don't have his private branch\n  locally.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90438","messageId":"20080911135146.GE5082@mit.edu","threadId":"15449","inReplyTo":"20080911123148.GA2056@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-11T13:51:46Z","receivedAt":"2008-09-11T13:51:46Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Sep 11, 2008 at 02:31:48PM +0200, Stephen R. van den Berg wrote:\n> \n> Well, the train of thought here goes as follows:\n> 1. Sure, why not add a field (zero or more) at the bottom of the free-form\n>    commit message reading like:\n> \n>    Origin: bbb896d8e10f736bfda8f587c0009c358c9a8599 ee837244df2e2e4e9171f508f83f353730db9e53\n> \n> 2. Add support to cherry-pick/revert to actually generate the field upon\n>    demand.\n\n\"git cherry-pick -x\" already generates the field you want.\n\n> \n> 3. Then add support to prune/gc/fsck/blame/log --graph to take the field\n>    into account.\n> \n\nUm, why should \"git fsck\", or \"git prune\" or \"git gc\" need to\nunderstand about this field?  What were you saying about unclean\nsemantics, again?  I thought you claimed that dangling origin links\nwere OK?  So why the heck should git fsck care?  And why shouldn't\ngc/prune drop objects that are only referenced via the origin link.\n\n> 4. Add support to filter-branch/rebase to renumber the field if necessary.\n\nAs we discussed earlier in some cases renumbering the field is not the\nright thing to do, especially if the commit in question has already\nbeen cherry-picked --- and you don't know that.  Again, this is why\nprototyping it outside of the core git is so useful; it will show up\nsome of these fundamental flaws in the origin link proposal.\n\n> Well, and after having done steps 1 to 5, the net result is that it\n> works almost as if the field is present in the header, except that:\n> - It is now at the end of the body in the commit message.\n> - It takes more time to find and parse it.\n\nA proof of concept, even if it isn't fully performant, is useful to\nprove that an idea actually has merit --- which clearly not everyone\nbelieves at this point.\n\nI'll also note that having a ***local*** database to cache the origin\nlink is a great way of short-circuiting the performance difficulties.\nIf it works, then it will be a lot easier to convince people that\nperhaps it should be done git-core, and by modifying core git functions.\n\nAlternatively, if you think this is such a great idea, why don't you\ngrab a copy of the git repository, and start hacking the idea\nyourself?  If you have running code, it tends to make the idea much\nmore concrete, and much easier to evaluate.  Or were you hoping to\nconvince other people to do all of this programming for you?\n\n\t\t\t\t\t\t- Ted\n"},{"id":"90440","messageId":"alpine.LFD.1.10.0809111047380.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911123148.GA2056@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T15:02:54Z","receivedAt":"2008-09-11T15:02:54Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> So short-circuiting the reasoning suggests that since the only thing\n> that actually changes now is the position of the field (at the top or\n> end of the commit message), we might as well do it right and put it in\n> the top, that gets rid of the two minuses.\n> \n> Anything I missed?\n\nA good convincing demonstration that this is actually worth doing in the \nfirst place.  And here I'm talking about the _feature_ and not the \n_implementation_.\n\n> Basically it means that:\n> \n> a. If there is a better solution to tracking the backports, I'll gladly\n>    use that instead, but simply using the current really freeform\n>    approach doesn't cut it (it currently refers to a single commit,\n>    instead of a pair of commits, and takes too long to parse out in a\n>    --top-order or blame command).  Better solutions I haven't heard so\n>    far.\n> \n> b. I need the integrity protection of a commit to make sure that the\n>    origin fields cannot be altered later; blame would be too easy to fool\n>    otherwise.  So using the notes solution seems to be out (it would also\n>    be quite a performance hit again).\n> \n> c. I consider the Origin: field at the end of the commit message a\n>    workable solution, but it smells like X-header-extension-messes as in\n>    E-mail headers, and it incurs a small performance hit (in case of\n>    --topo-order/blame/prune/fsck), but maybe this performance hit can be\n>    minimised by making sure that the fields are *always* at the end\n>    of the commit message.\n> \n> d. Using the proposed origin header in the standard commit header has\n>    close to zero overhead (in most commits the field is not present), yet\n>    codecomplexitywise it is almost identical with the Origin: field at\n>    the end of the commit message.\n> \n> I find it remarkable though that people are dragging their feet at\n> solution d, yet are quite ok with solution c.  IMO solution c and d are\n> almost identical, except that solution c is ugly, and solution d is\n> elegant.  But if it makes it easier to prove the usefulness by\n> implementing the ugly solution first, that's fine.\n\nTechnically speaking, implementation d is obviously the most efficient.  \nbut, as mentioned above, the actual need for this feature has not been \nconvincing so far.  Until then, it is not wise to add random stuff to \nthe very structure of a commit object, while c can be done even \nexternally from git which is a good way to demonstrate and convince \npeople about the usefulness of such feature.\n\n\nNicolas\n"},{"id":"90441","messageId":"20080911153202.GD2056@cuci.nl","threadId":"15449","inReplyTo":"20080911135146.GE5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T15:32:02Z","receivedAt":"2008-09-11T15:32:02Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Theodore Tso wrote:\n>On Thu, Sep 11, 2008 at 02:31:48PM +0200, Stephen R. van den Berg wrote:\n>> Well, the train of thought here goes as follows:\n>> 2. Add support to cherry-pick/revert to actually generate the field upon\n>>    demand.\n\n>\"git cherry-pick -x\" already generates the field you want.\n\nWell, sort of.  In order for swift parsing it should be a real field,\ni.e. it should not be an English sentence (in order to avoid people\naccidentally translating it); and it should list a pair of hashes\n(patches/changesets are defined by the difference between two tree\nsnapshots).  So it would be a -o option most likely, in order to provide\nbackward compatibility to the users of -x.\n\n>> 3. Then add support to prune/gc/fsck/blame/log --graph to take the field\n>>    into account.\n\n>Um, why should \"git fsck\", or \"git prune\" or \"git gc\" need to\n>understand about this field?  What were you saying about unclean\n>semantics, again?  I thought you claimed that dangling origin links\n>were OK?  So why the heck should git fsck care?  And why shouldn't\n>gc/prune drop objects that are only referenced via the origin link.\n\nDangling origin links are ok only if the developer in charge of the\nrepository doesn't care about the commits/branches they point to.\nThe definition of a \"caring developer\" is formalised by the fact that\nthe offending commits are already present in the repository or not.\n\nThis implies that fsck will skip the field if the hashes in question are\nunreachable in the current repository.\nIf they are reachable though, fsck will follow the link and check the\nwhole tree referenced by the origin link.  Obviously there are only two\nconditions for an origin link: either the hash points to an unreachable\nobject or the hash points to a reachable object of type commit (and all\nassociated checks that go with any commit).\n\ngc will preserve the commits the origin links point to once they are\nreachable.  I.e. if the developer doesn't care about the commits the\norigin links point to (i.e. if the branches are not reachable) then gc\njust skips them, if the developer *does* care, the origin links are used\nto keep those objects alive (and, of course, all their parenthood).\n\n>> 4. Add support to filter-branch/rebase to renumber the field if necessary.\n\n>As we discussed earlier in some cases renumbering the field is not the\n>right thing to do, especially if the commit in question has already\n>been cherry-picked --- and you don't know that.  Again, this is why\n>prototyping it outside of the core git is so useful; it will show up\n>some of these fundamental flaws in the origin link proposal.\n\nI agree that the behaviour of especially rebase with respect to the\norigin links is still something that needs to be thought through.\nI'm not convinced you are right, but I'm not convinced you are wrong\neither.\n\n>> Well, and after having done steps 1 to 5, the net result is that it\n>> works almost as if the field is present in the header, except that:\n>> - It is now at the end of the body in the commit message.\n>> - It takes more time to find and parse it.\n\n>A proof of concept, even if it isn't fully performant, is useful to\n>prove that an idea actually has merit --- which clearly not everyone\n>believes at this point.\n\nQuite.\n\n>I'll also note that having a ***local*** database to cache the origin\n>link is a great way of short-circuiting the performance difficulties.\n>If it works, then it will be a lot easier to convince people that\n>perhaps it should be done git-core, and by modifying core git functions.\n\nCreating local databases for these kinds of structures feels kludgy\nsomehow, since the git hash objects essentially *are* a working\ndatabase.  I have not checked yet if git already has some kind of\nready-to-use local database lib inside which I could reuse for that.\n\n>Alternatively, if you think this is such a great idea, why don't you\n>grab a copy of the git repository, and start hacking the idea\n>yourself?\n\nActually, in the first hour after posting the initial mail/proposal I\nalready had altered a local version of git to support the origin links\nin commit.[ch], --topo-order and fsck.  Before hacking further I decided\nto get some feedback first to see if someone would come up with\nsomething better.  And they did, instead of the mainline number, I\ndecided that using two hashes is better.  Once the dust has settled,\nI'll fill in the rest of the code.\n\n>  If you have running code, it tends to make the idea much\n>more concrete, and much easier to evaluate.\n\nAgreed, but then again, most of the programming is done without touching\nany code (the design phase), which is where we are now.  Once the design\nis scrutinised (as far as possible), the coding can begin (continue).\nThe feedback so far was very helpful, and caused me to explore (and\ndismiss) some of the alternate avenues to achieve the desired\nfunctionality.\n\n>  Or were you hoping to\n>convince other people to do all of this programming for you?\n\nI've never needed that so far, and will not need that here either.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90446","messageId":"alpine.LFD.1.10.0809110835070.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"20080911062242.GA23070@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-11T15:39:20Z","receivedAt":"2008-09-11T15:39:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n> \n> >delete of the origin branch will basically make them unreachable.\n> \n> False.\n\nStephen, here's a f*cking clue:\n\n - I know how git works.\n\n> If you fetch just branches A, B and C, but not D, the origin link from A\n> to D is dangling.  Once you have fetched D as well [..]\n\nSo I just said we deleted beanch 'D', so there's no way to ever fetch it \nagain.\n\nGet it?\n\nThe fact is, a big part of git is temporary branches. It's one of the \n*best* features of git. Throw-away stuff. Those throw-away branches are \noften done for initial development, and then the final result is often a \ncleaned-up version. Often using rebase or cherry-picking or any number of \nthings.\n\nAnd this is why \"git cherry-pick\" DOES NOT PUT THE ORIGINAL SHA1 IN THE \nCOMMENT FIELD BY DEFAULT.\n\n(Although you can use \"-x\" to make it do so for when you actually _want_ \nto say \"cherry-picked from xyzzy\")\n\nCan you not understand that? The \"origin\" field is _garbage_. It's garbage \nfor all normal cases. The original commit will not ever even EXIST in the \nresult, because it has long since been thrown away and will never exist \nanywhere else.\n\nGarbage should be _avoided_, not added.\n\n\t\t\tLinus\n"},{"id":"90447","messageId":"20080911160040.GE2056@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111047380.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T16:00:40Z","receivedAt":"2008-09-11T16:00:40Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> Anything I missed?\n\n>Technically speaking, implementation d is obviously the most efficient.  \n>but, as mentioned above, the actual need for this feature has not been \n>convincing so far.  Until then, it is not wise to add random stuff to \n>the very structure of a commit object, while c can be done even \n>externally from git which is a good way to demonstrate and convince \n>people about the usefulness of such feature.\n\nThe actual need for the feature seems to be dependent on one's workflow\nhabits.  This is also the problem I sense throughout the thread: some\npeople know exactly what I'm talking about, and would come up with the\nalmost identical design specs for the feature independent of myself, and\nothers need to be explained every tiny detail of the spec because they\nare not familiar with the concept and can't imagine why/how it would be\nused.\n\nLet me try and describe once more the typical environment this origin field\nis vital in:\n\nImagine a repository with:\n- 33774 commits total\n- 13 years of history\n- 1 development branch\n- 9 stable branches (forked off of the development branch at regular\n  intervals during the past 13 years).\n- The stable branches are never merged with each other or with the\n  development branch.\n- 2787 individual back/forward ports between the development and stable\n  branches.\n\nIn order to have meaningful output for git-blame, it needs to follow the\nchain across cherry-picks reliably.\nOnce you alter a piece of code, in order to figure out what more to alter,\nyou need to verify if this piece of code was or wasn't forward/backported.\nReliable and fast reporting of this, and actual comparison of the\ndifferent forward/backports between the 9 branches is essential.  It\nbasically means that you need to view the diffs of the patches across 9\nbranches on a regular basis.\n\nWithout the origin links, this workflow will cost a lot more time to\npursue (I know it, because I'm living it at the moment, and no, I'm not\nthe only developer, it's a development team).\n\nThis development model is not unique to my situation, it occurs at more\nplaces.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90448","messageId":"48C940C8.6040407@gnu.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809110835070.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-11T16:01:12Z","receivedAt":"2008-09-11T16:01:12Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n>> If you fetch just branches A, B and C, but not D, the origin link from A\n>> to D is dangling.  Once you have fetched D as well [..]\n> \n> So I just said we deleted beanch 'D', so there's no way to ever fetch it \n> again.\n> \n> Get it?\n\nYes, but you should not have used Stephen's proposed new option to git\ncherry-pick, just like you shouldn't have used the existing -x option.\n\"-x\" would not have created a dangling reference, but it would have\ncreated a puzzling commit message.\n\n> The fact is, a big part of git is temporary branches. It's one of the \n> *best* features of git. Throw-away stuff. Those throw-away branches are \n> often done for initial development, and then the final result is often a \n> cleaned-up version. Often using rebase or cherry-picking or any number of \n> things.\n\nThese days I doubt people would use cherry-pick, they would probably use\ninteractive rebase.  But anyway, exactly for the same reason...\n\n> \"git cherry-pick\" DOES NOT PUT THE ORIGINAL SHA1 IN THE \n> COMMENT FIELD BY DEFAULT.\n\n... neither should cherry-picking create the origin link by default.\nOnly if requested by the user, using a new option that is basically \"-x\"\ndone in a different way.  Just like \"-x\", it should not be used when\ncherry-picking from private branches.\n\nBut say someone does it, then what happens?  If people clone the branch,\nthe reference will be basically unusable.  But since \"git gc\" does not\ndelete the referenced commit, at least the origin commit is still\navailable in the repository where the cherry-pick was made.  It is\ndebatable whether it is better or worse than \"-x\".\n\nCan we discuss instead a generic way to have porcelain-level metadata,\nimmutable or at least versioned, for the commit objects?  (This is the\nsame kind of metadata as the author or committer, which clearly have\nnothing to do with the git plumbing.)  Do you have any proposal of saner\nsemantics, not for the origin link but for commit references within this\nkind of metadata in general?\n\nPaolo\n"},{"id":"90449","messageId":"alpine.LFD.1.10.0809110910430.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"48C940C8.6040407@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-11T16:23:29Z","receivedAt":"2008-09-11T16:23:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 11 Sep 2008, Paolo Bonzini wrote:\n> \n> Yes, but you should not have used Stephen's proposed new option to git\n> cherry-pick, just like you shouldn't have used the existing -x option.\n> \"-x\" would not have created a dangling reference, but it would have\n> created a puzzling commit message.\n\nBut my point is, _none_ of what Stephen proposes has _any_ advantage over \nthe already existing functionality.\n\nIOW, absolutely *everything* is actually done better with existing data \nstructures, and then just adding tools to perhaps follow those SHA1's in \nthe commit message.\n\nThe whole \"origin\" field doesn't have any semantics that make sense for \ncore git. It's basically ignored by all normal git operations, and the \n_only_ things that people seem to point out as being features are things \nthat can - and obviously in my opinion should - be done by much higher \nlevels.\n\nFor example, the claim was that it's hard to follow the chain of \ncherry-picks. That's not _true_. Use gitweb and gitk, and you can already \nsee them. Sure, you need to use \"-x\", BUT YOU'D HAVE TO USE THAT WITH \nSteven's MODEL TOO!\n\nExactly because it would be a frigging _disaster_ if that \"origin\" field \nwas done by default.\n\nAnd the only thing that \"origin\" does is:\n\n - hide the information\n\n - make it easier to make mistakes (either enable the feature by default, \n   or not notice that you didn't enable it when you wanted to)\n\n - add a requirement for a backwards-incompatible field that is just \n   guaranteed to confuse any old git binaries.\n\n - make it _harder_ to do things like send revert/cherry-pick information \n   by email.\n\nSee? There are only downsides.\n\nLook at the kernel -stable trees. They explicitly add that cherry-pick \ninformation, and can add *more*. For example, they go look at\n\n\thttp://git.kernel.org/?p=linux/kernel/git/stable/linux-2.6.26.y.git;a=commit;h=cb09de4542ad75cc3b66d0cf1a86217bf5633416\n\nand then go to its parent commit (just click on the parent SHA). And \nnotice how the stable kernel tree commits talk about where they were \nback-ported from, or _why_ they aren't back-ports at all!\n\nIOW, there are really two main cases:\n\n - the common case for cherry-picking: you do not want any origin \n   information, because it's irrelevant, pointless, and *wrong*.\n\n - you _do_ want origin information, but you actually want to _explain_ \n   explicitly why it's not irrelevant, pointless, or wrong.\n\nAnd yes, the latter case is about a lot more than \"this was \ncherry-picked\". It's about \"this fixes that other commit we did\", or it's \nabout \"this was anti-cherry-picked - ie reverted\". They are all \"origins\" \nfor the commit in the sense that they are relevant to the commit, but they \nall need some explanation of what _kind_ of origins they are.\n\n\t\t\tLinus\n"},{"id":"90452","messageId":"200809111853.53065.jnareb@gmail.com","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809110835070.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-11T16:53:52Z","receivedAt":"2008-09-11T16:53:52Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n> On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n> \n>> If you fetch just branches A, B and C, but not D, the origin link from A\n>> to D is dangling.  Once you have fetched D as well [..]\n> \n> So I just said we deleted beanch 'D', so there's no way to ever fetch it \n> again.\n> \n> Get it?\n> \n> The fact is, a big part of git is temporary branches. It's one of the \n> *best* features of git. Throw-away stuff. Those throw-away branches are \n> often done for initial development, and then the final result is often a \n> cleaned-up version. Often using rebase or cherry-picking or any number of \n> things.\n> \n> And this is why \"git cherry-pick\" DOES NOT PUT THE ORIGINAL SHA1 IN THE \n> COMMENT FIELD BY DEFAULT.\n> \n> (Although you can use \"-x\" to make it do so for when you actually _want_ \n> to say \"cherry-picked from xyzzy\")\n\nAnd that is why the proposal was to use \"-o\" option to git-cherry-pick\nto add 'origin'/'changeset' header, exactly because git-cherry-pick is\n_abused_ to clean up branches and reorder commits; although I think that\n\"git rebase --interactive\" (and patch management interfaces) do replace\nusing git-cherry-pick for that purpose.\n\ngit-revert would add 'origin'/'changeset' header unconditionally,\njust like by default it seeds commit message with SHA-1 id of reverted\ncommit.\n\n> Can you not understand that? The \"origin\" field is _garbage_. It's garbage \n> for all normal cases. The original commit will not ever even EXIST in the \n> result, because it has long since been thrown away and will never exist \n> anywhere else.\n> \n> Garbage should be _avoided_, not added.\n\nHmmm... the difference between having 'origin' in a commit object header,\nand having it in commit mesage is like difference between 'Signed-off-by:'\nconvention and 'author' header.  First is the matter of workflow, second\nis inherent, required and non-avoidable part of revision information.\n\nOn the other hand git-cherry and git-blame would then have rely on\nparsing correctly free-form part of a commit object, to take advantage\nof 'origin' information: something what 'origin' info is for.\n\n\nP.S. 'generation' header was not added... just saying... :-)\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90453","messageId":"alpine.LFD.1.10.0809111222170.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911160040.GE2056@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T17:02:36Z","receivedAt":"2008-09-11T17:02:36Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> Let me try and describe once more the typical environment this origin field\n> is vital in:\n> \n> Imagine a repository with:\n> - 33774 commits total\n> - 13 years of history\n> - 1 development branch\n> - 9 stable branches (forked off of the development branch at regular\n>   intervals during the past 13 years).\n> - The stable branches are never merged with each other or with the\n>   development branch.\n> - 2787 individual back/forward ports between the development and stable\n>   branches.\n> \n> In order to have meaningful output for git-blame, it needs to follow the\n> chain across cherry-picks reliably.\n> Once you alter a piece of code, in order to figure out what more to alter,\n> you need to verify if this piece of code was or wasn't forward/backported.\n> Reliable and fast reporting of this, and actual comparison of the\n> different forward/backports between the 9 branches is essential.  It\n> basically means that you need to view the diffs of the patches across 9\n> branches on a regular basis.\n> \n> Without the origin links, this workflow will cost a lot more time to\n> pursue (I know it, because I'm living it at the moment, and no, I'm not\n> the only developer, it's a development team).\n> \n> This development model is not unique to my situation, it occurs at more\n> places.\n\nOK.  I think I might be able to believe you.\n\nWhere I feel uncomfortable is with the real semantics of your \"origin\" \nlink proposal.\n\nFirst, its name.  The word \"origin\" probably has a too narrow meaning \nthat creates confusion.  I'd suggest something like a \n\"may-be-related-to\" field that would be like a weak link.\n\nThe format of a may-be-related-to field would be the same as the parent \nfield, except that the object pointed to by the sha1 could have its type \nrelaxed, i.e. it could be anything like a blob or a tag.\n\nThe semantics of a \"may-be-related-to\" link would be defined for object \nreachability only:\n\n- If the may-be-related-to link is dangling then it is ignored.\n\n- If it is not dangling then usual reachability rules apply.\n\nThat's all the core git might care about, and the only real argument for \nnot having this information in the free form commit message.\n\nStill, in your case, you probably won't get rid of your stable branches, \nhence the reachability argument is rather weak for your usage scenario, \nmeaning that you could as well have that info in the free form text \n(like cherry-pick -x), and even generate a special graft file from that \nlocally for visualization/blame purposes.  Sure the indirection will add \nsome overhead, but I doubt it'll be measurable.\n\nPeople fetching your main branch won't have to carry the whole \nrepository because those weak links would otherwise be followed if \nthey're formally part of the commit header.  And if they want \nto benefit from the information those weak links carry then they just \nhave to also fetch the branch(es) where those links are pointing.  At \nthat point it is trivial to regenerate the special graft file locally \nwhich would also have the benefit of only containing links to actually \nreachable commits, hence you'd never have dangling \"origin\" links.\n\nConclusion: the only fundamental reason for having this weak link \ninformation in the commit header is for reachability convenience for \nwhen the actual branch that contained the referenced commits is gone, \nwhich IMHO is a bad justification.  Having lines of developments hanging \noff of a weak link alone is just plain stupid if you can't reach it via \nproper branches or tags.\n\nSo I think that your usage scenario is a valid one, but I think that you \nshould implement it some other ways, like this special graft file I \nmentioned above, which can be generated and updated from custom \ninformation found in the free form comment text of a commit object, and \nthat proper reachability issues should continue to be handled through \nproper branches and tags as it is done today.\n\n\nNicolas\n"},{"id":"90457","messageId":"20080911180037.GH5082@mit.edu","threadId":"15449","inReplyTo":"20080911153202.GD2056@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-11T18:00:37Z","receivedAt":"2008-09-11T18:00:37Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Sep 11, 2008 at 05:32:02PM +0200, Stephen R. van den Berg wrote:\n> gc will preserve the commits the origin links point to once they are\n> reachable.  I.e. if the developer doesn't care about the commits the\n> origin links point to (i.e. if the branches are not reachable) then gc\n> just skips them, if the developer *does* care, the origin links are used\n> to keep those objects alive (and, of course, all their parenthood).\n\nThis seems wrong.  OK, suppose you have branches A, B, C, and D, while\nyou are on branch C, you cherry pick commit 'p' from branch B, so that\nthere is a new commit q on branch C which has an origin link\ncontaining the commit ID's p^ and 'p.    \n\nNow suppose branch B gets deleted, and you do a \"git gc\".  All of the\ncommits that were part of branch B will vanish except for p^ and p,\nwhich in your model will stick around because they are origin links\ncommit q on branch C.  But what good is are these two commits?  They\nrepresent two snapshots in time, with no context now that branch B has\nbeen deleted.  99% of the time, the diff between p^ and p will result\nin the equivalent of the diff between q^ and q.  But even if they\naren't, what use are these isolated, disconnected commits?  So having\n\"git gc\" retain them commits that are pointed to be this proposed\norigin link doesn't seem to make any sense, and doesn't seem to be\nwell thought through.\n\nOh, BTW, suppose you then further do a \"git cherry-pick -o\" of commit\nq while you are on branch D.  Presumably this will create a new\ncommit, r.  But will the origin-link of commit r be p^ and p, or q^\nand q?  And will this change depending on whether or not -o is\nspecified?\n\n> >I'll also note that having a ***local*** database to cache the origin\n> >link is a great way of short-circuiting the performance difficulties.\n> >If it works, then it will be a lot easier to convince people that\n> >perhaps it should be done git-core, and by modifying core git functions.\n> \n> Creating local databases for these kinds of structures feels kludgy\n> somehow, since the git hash objects essentially *are* a working\n> database.  I have not checked yet if git already has some kind of\n> ready-to-use local database lib inside which I could reuse for that.\n\nGitk already keeps a cache (.git/gitk.cache) to speed up some of its\noperations.  And in some ways the index file is a cache, although it\ndoes far more than that.\n\n     \t      \t       \t \t\t    - Ted\n"},{"id":"90458","messageId":"20080911184405.GA1451@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111222170.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T18:44:05Z","receivedAt":"2008-09-11T18:44:05Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> Let me try and describe once more the typical environment this origin field\n>> is vital in:\n\n>> Without the origin links, this workflow will cost a lot more time to\n>> pursue (I know it, because I'm living it at the moment, and no, I'm not\n\n>First, its name.  The word \"origin\" probably has a too narrow meaning \n>that creates confusion.  I'd suggest something like a \n>\"may-be-related-to\" field that would be like a weak link.\n\nWell, the important properties of the name/field would be:\n- It should be as specific as possible, in order to minimise the\n  potential for abuse in the future.  I distill the desirability of this\n  requirement out of the various earlier discussions about commitheaders\n  in the past on this mailinglist held by others.\n- It should convey a sense of direction (it's a directed graph).\n\nAny generic may-be-related-to field is therefore probably a non-starter.\n\n>The semantics of a \"may-be-related-to\" link would be defined for object \n>reachability only:\n>- If the may-be-related-to link is dangling then it is ignored.\n>- If it is not dangling then usual reachability rules apply.\n>That's all the core git might care about, and the only real argument for \n>not having this information in the free form commit message.\n\nThe origin field as currently proposed tightens the requirements that\nit either is dangling and ignored or points to a commit.\nrev-list --topo-order should use the origin links to order the output.\ngc/prune won't delete commits referenced *by* an origin link.\n\nThe only two other arguments one might give to actually keep the field\nin the header of the commit as opposed to the trailer is that the \nphysical field can be kept machine readable, and the actual display can be\nbeautified like:  Origin: 2abcdef..1234567\nThe output of the field could be suppressed (if so desired) if the\ntarget commit isn't reachable.\nAll this is of course possible for a trailer field in the free-form\narea as well, but it seems a bit silly to have two places for \"headers\".\n\n>Still, in your case, you probably won't get rid of your stable branches, \n\nTrue.\n\n>hence the reachability argument is rather weak for your usage scenario, \n\nThen again, I don't want to be bothered by stupid free-form origin links\nmade to local branches by a developer.  If the developer creates them\nusing cherry-pick -o which creates an origin link, I'll never have to\nsee his silly commit hashes where he is referring to commits in his\nlocal branch (and never waste time wondering where those commits are).\n\n>meaning that you could as well have that info in the free form text \n>(like cherry-pick -x), and even generate a special graft file from that \n>locally for visualization/blame purposes.  Sure the indirection will add \n>some overhead, but I doubt it'll be measurable.\n\nThe free-form equivalent looks like:\nOrigin: df85f7855da44c730f942b330ada181209d09d7a ff1e8bfcd69e5e0ee1a3167e80ef75b611f72123\nYou need a pair of hashes, which is, a bit bulky, for my taste.\n\nWhat special graft file would I need to visualise?  Isn't having the\norigin link information enough?\n\n>People fetching your main branch won't have to carry the whole \n>repository because those weak links would otherwise be followed if \n>they're formally part of the commit header.  And if they want \n>to benefit from the information those weak links carry then they just \n>have to also fetch the branch(es) where those links are pointing.  At \n>that point it is trivial to regenerate the special graft file locally \n>which would also have the benefit of only containing links to actually \n>reachable commits, hence you'd never have dangling \"origin\" links.\n\nYou lost me here somewhere.  Could you give a concrete example with one\ncommit, one origin link (your style) and a special graftfile entry?\n\n>Conclusion: the only fundamental reason for having this weak link \n>information in the commit header is for reachability convenience for \n>when the actual branch that contained the referenced commits is gone, \n\nErm.  Quite the opposite, actually.\nThe practical use for the origin link in case the target is unreachable\nis zero to none, so it can gleefully be ignored in that case.\nBut maybe the semantics of your \"related\" link and my origin link are\nsufficiently distinct.\nFor the arguments why it should be in the header of a commit, see above.\n\n>which IMHO is a bad justification.  Having lines of developments hanging \n>off of a weak link alone is just plain stupid if you can't reach it via \n>proper branches or tags.\n\nAgreed.  But this is in reference to your \"related\" link proposal.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90459","messageId":"20080911190335.GB1451@cuci.nl","threadId":"15449","inReplyTo":"20080911180037.GH5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T19:03:35Z","receivedAt":"2008-09-11T19:03:35Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Theodore Tso wrote:\n>On Thu, Sep 11, 2008 at 05:32:02PM +0200, Stephen R. van den Berg wrote:\n>> gc will preserve the commits the origin links point to once they are\n>> reachable.  I.e. if the developer doesn't care about the commits the\n>> origin links point to (i.e. if the branches are not reachable) then gc\n>> just skips them, if the developer *does* care, the origin links are used\n>> to keep those objects alive (and, of course, all their parenthood).\n\n>This seems wrong.  OK, suppose you have branches A, B, C, and D, while\n>you are on branch C, you cherry pick commit 'p' from branch B, so that\n>there is a new commit q on branch C which has an origin link\n>containing the commit ID's p^ and 'p.    \n\nOk.\n\n>Now suppose branch B gets deleted, and you do a \"git gc\".  All of the\n>commits that were part of branch B will vanish except for p^ and p,\n\nNot quite.  Obviously all parents of p and p^ will continue to exist.\nI.e. deleting branch B will cause all commits from p till the tip of B\n(except p itself) to vanish.  Keeping p implies that the whole chain of\nparents below p will continue to exist and be reachable.  That's the way\na git repository works.\n\n>which in your model will stick around because they are origin links\n>commit q on branch C.  But what good is are these two commits?  They\n>represent two snapshots in time, with no context now that branch B has\n\nThe context are all their ancestors, which continue to exist, and that\nis all you need.\n\n>been deleted.  99% of the time, the diff between p^ and p will result\n>in the equivalent of the diff between q^ and q.  But even if they\n>aren't, what use are these isolated, disconnected commits?  So having\n>\"git gc\" retain them commits that are pointed to be this proposed\n>origin link doesn't seem to make any sense, and doesn't seem to be\n>well thought through.\n\nI beg to differ, but I presume you agree with me now?\n\n>Oh, BTW, suppose you then further do a \"git cherry-pick -o\" of commit\n>q while you are on branch D.  Presumably this will create a new\n>commit, r.  But will the origin-link of commit r be p^ and p, or q^\n>and q?\n\nIt will be q^..q, and specifically not p^..p, using ^p..p would be\nlying.  We aim to document the evolvement of the patch in time.\nCherry-pick itself will always ignore the origin links present on the\nold commit, it simply creates new ones as if the old ones didn't exist.\n\n>  And will this change depending on whether or not -o is\n>specified?\n\nNo.  Actually, cherry-pick will never generate origin links unless -o is\nspecified.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90460","messageId":"20080911192356.GC1451@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809110835070.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T19:23:56Z","receivedAt":"2008-09-11T19:23:56Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> >delete of the origin branch will basically make them unreachable.\n\n>> False.\n\n>Stephen, here's a f*cking clue:\n> - I know how git works.\n\nI'd presume you do, but that doesn't mean you always accurately express\nyourself.\n\n>> If you fetch just branches A, B and C, but not D, the origin link from A\n>> to D is dangling.  Once you have fetched D as well [..]\n\n>So I just said we deleted beanch 'D', so there's no way to ever fetch it \n>again.\n\nYou did not state you deleted branch 'D' on the repository being fetched\n*FROM*.  I assumed you meant you deleted branch 'D' on the repository\ndoing the fetching (after having fetched 'D' in the past).\n\n>Get it?\n\n\"You stupid git\".\n\n>The fact is, a big part of git is temporary branches. It's one of the \n>*best* features of git. Throw-away stuff. Those throw-away branches are \n>often done for initial development, and then the final result is often a \n>cleaned-up version. Often using rebase or cherry-picking or any number of \n>things.\n\nIndeed, features I value in git very much, and use every day, thanks.\n\n[...portions of man git-cherry-pick stripped...]\n\n>Can you not understand that? The \"origin\" field is _garbage_. It's garbage \n>for all normal cases. The original commit will not ever even EXIST in the \n>result, because it has long since been thrown away and will never exist \n>anywhere else.\n\nThe origin field will *not* be created on regular cherry-picks, this\n*would* create garbage.  The origin field is not meant to be generated\nwhen doing things with temporary branches.  The origin field is meant to\nbe filled *ONLY* when cherry-picking from one permanent branch to\nanother permanent branch.  This is a *rare* operation.\n\n>Garbage should be _avoided_, not added.\n\nQuite.\n\nI do understand that \"normal cases\" in your case mean cherry-picks among\ntemporary branches.\nWell, you are completely right that *your* normal cases should not (and\nwill not) generate an origin field.\nThe origin field is intended for the *abnormal* cases, which means\ncherry-picking between permanent branches (which, apparently, you rarely\ndo, if ever), this is something that (depending on your workflow) can be\na more frequent event.  For *those* cases, the origin field will not\ncontain garbage.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90461","messageId":"alpine.LFD.1.10.0809111527030.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911190335.GB1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T19:33:00Z","receivedAt":"2008-09-11T19:33:00Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> Theodore Tso wrote:\n> >On Thu, Sep 11, 2008 at 05:32:02PM +0200, Stephen R. van den Berg wrote:\n> >> gc will preserve the commits the origin links point to once they are\n> >> reachable.  I.e. if the developer doesn't care about the commits the\n> >> origin links point to (i.e. if the branches are not reachable) then gc\n> >> just skips them, if the developer *does* care, the origin links are used\n> >> to keep those objects alive (and, of course, all their parenthood).\n> \n> >This seems wrong.  OK, suppose you have branches A, B, C, and D, while\n> >you are on branch C, you cherry pick commit 'p' from branch B, so that\n> >there is a new commit q on branch C which has an origin link\n> >containing the commit ID's p^ and 'p.    \n> \n> Ok.\n> \n> >Now suppose branch B gets deleted, and you do a \"git gc\".  All of the\n> >commits that were part of branch B will vanish except for p^ and p,\n> \n> Not quite.  Obviously all parents of p and p^ will continue to exist.\n> I.e. deleting branch B will cause all commits from p till the tip of B\n> (except p itself) to vanish.  Keeping p implies that the whole chain of\n> parents below p will continue to exist and be reachable.  That's the way\n> a git repository works.\n\nAnd that's what I called stupid in my earlier reply to you.  Either you \nhave proper branches or tags keeping P around, or deleting B brings \neverything not reachable through other branches or tags (or reflog) \naway too.  Otherwise there is no point making a dangling origin link \nvalid.\n\n\nNicolas\n"},{"id":"90462","messageId":"20080911194447.GD1451@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111527030.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T19:44:47Z","receivedAt":"2008-09-11T19:44:47Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> Not quite.  Obviously all parents of p and p^ will continue to exist.\n>> I.e. deleting branch B will cause all commits from p till the tip of B\n>> (except p itself) to vanish.  Keeping p implies that the whole chain of\n>> parents below p will continue to exist and be reachable.  That's the way\n>> a git repository works.\n\n>And that's what I called stupid in my earlier reply to you.  Either you \n>have proper branches or tags keeping P around, or deleting B brings \n>everything not reachable through other branches or tags (or reflog) \n>away too.  Otherwise there is no point making a dangling origin link \n>valid.\n\nWell, the principle of least surprise dictates that they should be kept\nby gc as described above, however...\nI can envision an option to gc say \"--drop-weak-links\" which does\nexactly what you describe.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90463","messageId":"alpine.LFD.1.10.0809111534300.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911192356.GC1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T19:45:33Z","receivedAt":"2008-09-11T19:45:33Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> The origin field will *not* be created on regular cherry-picks, this\n> *would* create garbage.  The origin field is not meant to be generated\n> when doing things with temporary branches.  The origin field is meant to\n> be filled *ONLY* when cherry-picking from one permanent branch to\n> another permanent branch.  This is a *rare* operation.\n\n... and therefore you might as well just have a separate file (which \nmight or might not be tracked by git like the .gitignore files are) \nto keep that information?  Since this is a rare operation, modifying the \ncore database structure for this doesn't appear that appealing to most \nso far.\n\nAnd, while recording this origin link is optional, you are likely to \nmake mistakes like forgotting to record it, or you might even wish to \nfix it with better links after the facts.  Having it versionned also \nmeans that older git versions will be able to carry that information \neven if they won't make any use of it, and that also solves the \ncryptographic issue since that data is part of the top commit SHA1.\n\n\nNicolas\n"},{"id":"90464","messageId":"20080911195516.GE1451@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111534300.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T19:55:16Z","receivedAt":"2008-09-11T19:55:16Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> when doing things with temporary branches.  The origin field is meant to\n>> be filled *ONLY* when cherry-picking from one permanent branch to\n>> another permanent branch.  This is a *rare* operation.\n\n>... and therefore you might as well just have a separate file (which \n>might or might not be tracked by git like the .gitignore files are) \n>to keep that information?  Since this is a rare operation, modifying the \n>core database structure for this doesn't appear that appealing to most \n>so far.\n\nFor various reasons, the best alternate place would be at the trailing\nend of the free-form field.  Using a separate structure causes\n(performance) problems (mostly).\n\n>And, while recording this origin link is optional, you are likely to \n>make mistakes like forgotting to record it,\n\nThat is just as likely filling in the wrong commit message.\n\n> or you might even wish to \n>fix it with better links after the facts.\n\nThat is not possible for commit messages, and should not be possible for\norigin links either (same reasons).\n\n>  Having it versionned also \n>means that older git versions will be able to carry that information \n>even if they won't make any use of it, and that also solves the \n>cryptographic issue since that data is part of the top commit SHA1.\n\nIt would allow the data to be faked, that is undesirable for \"git blame\".\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90465","messageId":"alpine.LFD.1.10.0809111546270.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911184405.GA1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T20:00:55Z","receivedAt":"2008-09-11T20:00:55Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> Nicolas Pitre wrote:\n> >First, its name.  The word \"origin\" probably has a too narrow meaning \n> >that creates confusion.  I'd suggest something like a \n> >\"may-be-related-to\" field that would be like a weak link.\n> \n> Well, the important properties of the name/field would be:\n> - It should be as specific as possible, in order to minimise the\n>   potential for abuse in the future.  I distill the desirability of this\n>   requirement out of the various earlier discussions about commitheaders\n>   in the past on this mailinglist held by others.\n\nWell, sure.  But being too specific sometimes limits its usefulness.  \nThere could be other usages for such a link which IMHO should be defined \nin terms of graph connectivity semantics rather than high level purpose.\n\n> - It should convey a sense of direction (it's a directed graph).\n\nWell, isn't that already obvious?\n\n> Any generic may-be-related-to field is therefore probably a non-starter.\n\nWell, my whole argument is that if it has no generic purpose then it \nprobably doesn't belong in the commit header.\n\n> The origin field as currently proposed tightens the requirements that\n> it either is dangling and ignored or points to a commit.\n> rev-list --topo-order should use the origin links to order the output.\n> gc/prune won't delete commits referenced *by* an origin link.\n\nAnd I disagree on the gc/prune point, as mentioned previously.\n\nAs to rev-list --topo-order, it doesn't need for this link to actually \nbe part of the commit object to accomplish the desired effect.\n\n> The only two other arguments one might give to actually keep the field\n> in the header of the commit as opposed to the trailer is that the \n> physical field can be kept machine readable, and the actual display can be\n> beautified like:  Origin: 2abcdef..1234567\n> The output of the field could be suppressed (if so desired) if the\n> target commit isn't reachable.\n> All this is of course possible for a trailer field in the free-form\n> area as well, but it seems a bit silly to have two places for \"headers\".\n\nAnd I think you should simply create a file within the repository with \nthat info instead of either thecommit header or the free form text.  It \ngives all the usability advantages you wish for and more.\n\n\nNicolas\n"},{"id":"90466","messageId":"alpine.LFD.1.10.0809111601280.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911194447.GD1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T20:03:09Z","receivedAt":"2008-09-11T20:03:09Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> Nicolas Pitre wrote:\n> >On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n> >> Not quite.  Obviously all parents of p and p^ will continue to exist.\n> >> I.e. deleting branch B will cause all commits from p till the tip of B\n> >> (except p itself) to vanish.  Keeping p implies that the whole chain of\n> >> parents below p will continue to exist and be reachable.  That's the way\n> >> a git repository works.\n> \n> >And that's what I called stupid in my earlier reply to you.  Either you \n> >have proper branches or tags keeping P around, or deleting B brings \n> >everything not reachable through other branches or tags (or reflog) \n> >away too.  Otherwise there is no point making a dangling origin link \n> >valid.\n> \n> Well, the principle of least surprise dictates that they should be kept\n> by gc as described above, however...\n> I can envision an option to gc say \"--drop-weak-links\" which does\n> exactly what you describe.\n\nDon't you think this starts to look silly at that point?\n\n\nNicolas\n"},{"id":"90467","messageId":"20080911200452.GM5082@mit.edu","threadId":"15449","inReplyTo":"20080911190335.GB1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-11T20:04:53Z","receivedAt":"2008-09-11T20:04:53Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Sep 11, 2008 at 09:03:35PM +0200, Stephen R. van den Berg wrote:\n> >This seems wrong.  OK, suppose you have branches A, B, C, and D, while\n> >you are on branch C, you cherry pick commit 'p' from branch B, so that\n> >there is a new commit q on branch C which has an origin link\n> >containing the commit ID's p^ and 'p.    \n> \n> >Now suppose branch B gets deleted, and you do a \"git gc\".  All of the\n> >commits that were part of branch B will vanish except for p^ and p,\n> \n> Not quite.  Obviously all parents of p and p^ will continue to exist.\n> I.e. deleting branch B will cause all commits from p till the tip of B\n> (except p itself) to vanish.  Keeping p implies that the whole chain of\n> parents below p will continue to exist and be reachable.  That's the way\n> a git repository works.\n\nThat's still not very useful, since you still don't have a label for\nthis anonymous series of commit chain that just dead ends at commit p.\nHow would anyone find this useful?\n\n> >which in your model will stick around because they are origin links\n> >commit q on branch C.  But what good is are these two commits?  They\n> >represent two snapshots in time, with no context now that branch B has\n> \n> The context are all their ancestors, which continue to exist, and that\n> is all you need.\n\nNeed for what?  What useful information would you devine from it?\n\n> >Oh, BTW, suppose you then further do a \"git cherry-pick -o\" of commit\n> >q while you are on branch D.  Presumably this will create a new\n> >commit, r.  But will the origin-link of commit r be p^ and p, or q^\n> >and q?\n> \n> It will be q^..q, and specifically not p^..p, using ^p..p would be\n> lying.  We aim to document the evolvement of the patch in time.\n> Cherry-pick itself will always ignore the origin links present on the\n> old commit, it simply creates new ones as if the old ones didn't exist.\n\nSo if you never pull branch C (where commit q resides), there is no\nway for you to know that commits p and r are related.  How.... not\nuseful.\n\nIf the scenario was being able to tell which stable branches had a\nparticular bug fixes, I think my proposal of attaching a bug\nidentifier is a far superior solution.\n\nAgain, what's the use case of \"trying to document the development of\nthe patch in time?\"  Aside from drawing pretty dotted lines\neverywhere, what *good* does this actually achieve?  How would it\naffect other git commands' behavior, and how would this change in\nbehavior actually be considered a net improvement over what we have\nnow?\n\n\t\t\t\t\t\t- Ted\n"},{"id":"90468","messageId":"200809112205.16928.jnareb@gmail.com","threadId":"15449","inReplyTo":"20080911194447.GD1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-11T20:05:15Z","receivedAt":"2008-09-11T20:05:15Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Stephen R. van den Berg wrote:\n> Nicolas Pitre wrote:\n>>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>>> Not quite.  Obviously all parents of p and p^ will continue to exist.\n>>> I.e. deleting branch B will cause all commits from p till the tip of B\n>>> (except p itself) to vanish.  Keeping p implies that the whole chain of\n>>> parents below p will continue to exist and be reachable.  That's the way\n>>> a git repository works.\n> \n>>And that's what I called stupid in my earlier reply to you.  Either you \n>>have proper branches or tags keeping P around, or deleting B brings \n>>everything not reachable through other branches or tags (or reflog) \n>>away too.  Otherwise there is no point making a dangling origin link \n>>valid.\n> \n> Well, the principle of least surprise dictates that they should be kept\n> by gc as described above, however...\n> I can envision an option to gc say \"--drop-weak-links\" which does\n> exactly what you describe.\n\nWell, IIRC the need for this was one of the causes of \"death\" of 'prior'\nheader link proposal...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"90470","messageId":"20080911201639.GF1451@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809110910430.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T20:16:39Z","receivedAt":"2008-09-11T20:16:39Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>But my point is, _none_ of what Stephen proposes has _any_ advantage over \n>the already existing functionality.\n\nI think you're missing some of the advantages because you don't have a\nlot of experience with cherry-pick workflows between multiple permanent\nbranches.\n\n>IOW, absolutely *everything* is actually done better with existing data \n>structures, and then just adding tools to perhaps follow those SHA1's in \n>the commit message.\n\nThe best way to explain the difference is probably by implementing the\nfree-form support, so I think I'll do that.\n\n>For example, the claim was that it's hard to follow the chain of \n>cherry-picks. That's not _true_. Use gitweb and gitk, and you can already \n>see them. Sure, you need to use \"-x\", BUT YOU'D HAVE TO USE THAT WITH \n>Steven's MODEL TOO!\n\nThe existing cherry-pick -x option doesn't cut it, it helps for the\nsimple cases, yes, but there are cherry-pick situations where it just\nadds to the confusion.\n\n>Exactly because it would be a frigging _disaster_ if that \"origin\" field \n>was done by default.\n\nThat never was the intention, and never will be happening.\n\n>And the only thing that \"origin\" does is:\n\n> - hide the information\n\nOnly if you want to hide it, you control if it does, this point is moot.\n\n> - make it easier to make mistakes (either enable the feature by default, \n>   or not notice that you didn't enable it when you wanted to)\n\nThe same holds for -x, so this point is moot as well.\n\n> - add a requirement for a backwards-incompatible field that is just \n>   guaranteed to confuse any old git binaries.\n\nThis is a problem, I admit, but maybe this can be solved in the future.\nThen again, since use of the feature is a *very* conscious decision, anyone\nusing the feature can advise their users to use git version xxx at least.\n\n> - make it _harder_ to do things like send revert/cherry-pick information \n>   by email.\n\nNot necessarily, adding an Origin field in the patch sent by mail is\neasy.  I don't see how it would be more difficult otherwise.  Please\nexplain.\n\n>See? There are only downsides.\n\nI think I just neutralised all but one of the mentioned downsides, and\nthe backward compatibility issue is at least mitigated.\n\n>and then go to its parent commit (just click on the parent SHA). And \n>notice how the stable kernel tree commits talk about where they were \n>back-ported from, or _why_ they aren't back-ports at all!\n\nAnd this is impossible when using the origin link?  The usage with an\norigin link would be just as flexible, even more so.\n\n>IOW, there are really two main cases:\n\n> - the common case for cherry-picking: you do not want any origin \n>   information, because it's irrelevant, pointless, and *wrong*.\n\nQuite, and my proposal is not generating those anyway.\n\n> - you _do_ want origin information, but you actually want to _explain_ \n>   explicitly why it's not irrelevant, pointless, or wrong.\n\n>And yes, the latter case is about a lot more than \"this was \n>cherry-picked\". It's about \"this fixes that other commit we did\", or it's \n>about \"this was anti-cherry-picked - ie reverted\". They are all \"origins\" \n>for the commit in the sense that they are relevant to the commit, but they \n>all need some explanation of what _kind_ of origins they are.\n\nYes, and that *extra* information can and should go into the free-form\ncommit message, alongside of the origin field inside the header (or\ntrailer), just edit the commit message before committing after a\ncherry-pick -o.  What's your point?\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90471","messageId":"20080911202228.GG1451@cuci.nl","threadId":"15449","inReplyTo":"200809112205.16928.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T20:22:28Z","receivedAt":"2008-09-11T20:22:28Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jakub Narebski wrote:\n>Stephen R. van den Berg wrote:\n>> Well, the principle of least surprise dictates that they should be kept\n>> by gc as described above, however...\n>> I can envision an option to gc say \"--drop-weak-links\" which does\n>> exactly what you describe.\n\n>Well, IIRC the need for this was one of the causes of \"death\" of 'prior'\n>header link proposal...\n\nAs I understood it, one of the causes of death of the \"prior\" link\nproposal was that it was unclear if it pulled in the linked-to commits\nupon fetch.  In the \"origin\" case, the default is *not* to fetch them.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90472","messageId":"20080911202431.GH1451@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111601280.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T20:24:31Z","receivedAt":"2008-09-11T20:24:31Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> Well, the principle of least surprise dictates that they should be kept\n>> by gc as described above, however...\n>> I can envision an option to gc say \"--drop-weak-links\" which does\n>> exactly what you describe.\n\n>Don't you think this starts to look silly at that point?\n\nNo, it's the developers vote controlling his own repository saying:\nOk, I expressed interest in the other branches and their\nbackport/forwardport relationships, but I changed my mind.  Drop all\nbackport/forwardport information on branches I don't explicitly have.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90474","messageId":"alpine.LFD.1.10.0809111604040.23787@xanadu.home","threadId":"15449","inReplyTo":"20080911195516.GE1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-09-11T20:27:00Z","receivedAt":"2008-09-11T20:27:00Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n\n> Nicolas Pitre wrote:\n> >On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n> >> when doing things with temporary branches.  The origin field is meant to\n> >> be filled *ONLY* when cherry-picking from one permanent branch to\n> >> another permanent branch.  This is a *rare* operation.\n> \n> >... and therefore you might as well just have a separate file (which \n> >might or might not be tracked by git like the .gitignore files are) \n> >to keep that information?  Since this is a rare operation, modifying the \n> >core database structure for this doesn't appear that appealing to most \n> >so far.\n> \n> For various reasons, the best alternate place would be at the trailing\n> end of the free-form field.  Using a separate structure causes\n> (performance) problems (mostly).\n\nDid you try it?  I don't particularly buy this performance argument, and \nthe bulk of my contributions to git so far were about performances.  It \nis quite easy to load a flat file with sorted commit SHA1s, and given \nthat origin links are the result of a rare operation, then there \nshouldn't be too many entries to search through.  Hell, doing 213647 \nlookups (and many other things like inflating zlib deflated data)  with \neach of them for commit objects in my Linux repository which has 1355167 \ntotal entries takes only 6 seconds here, or about a quarter of a \nmilisecond for each lookup.  I doubt doing an extra lookup in a much \nsmaller table would show on the radar.\n\n\nNicolas\n"},{"id":"90476","messageId":"20080911210118.GO5082@mit.edu","threadId":"15449","inReplyTo":"20080911195516.GE1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-11T21:01:18Z","receivedAt":"2008-09-11T21:01:18Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Sep 11, 2008 at 09:55:16PM +0200, Stephen R. van den Berg wrote:\n> >  Having it versionned also \n> >means that older git versions will be able to carry that information \n> >even if they won't make any use of it, and that also solves the \n> >cryptographic issue since that data is part of the top commit SHA1.\n> \n> It would allow the data to be faked, that is undesirable for \"git blame\".\n\nWhy would this matter?  The information is largely\nself-authenticating.  If a commit claims to have come from some other\ncherry-pick, a human taking a quick look at it would know instantly\nthat this wasn't true.  So what's the harm done if some incorrect\ninformation gets introduced?  \"git blame\" is something which is\ngenerally used by humans, not by automated programs.\n\nAlso, what's the attack scenario?  The person who originally makes the\ncommit can easily fake the origin link information.  They can hack git\nto fill on some other commit ID, for example.  So what you are\nprotecting against is someone after the fact adding the annotation\nthat this commit was related to this other commit.  When would this be\na bad thing to do?  If they are adding correct information, it's a\ngood thing.  If they add incorrect information, what's the harm they\ncan as a result of being able to add the incorrect information.\n(Noting that if this annotation file is kept under git control, you\ncan use what ever access controls and/or process controls that verify\nthat a new cherry-pick --- or a commit claiming to be a cherry-pick\n--- is valid and should be accepted into the master git repository for\nthat project.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"90478","messageId":"7vy71ys8a7.fsf@gitster.siamese.dyndns.org","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111222170.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-09-11T21:05:20Z","receivedAt":"2008-09-11T21:05:20Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> Still, in your case, you probably won't get rid of your stable branches, \n> hence the reachability argument is rather weak for your usage scenario, \n> meaning that you could as well have that info in the free form text \n> (like cherry-pick -x), and even generate a special graft file from that \n> locally for visualization/blame purposes.  Sure the indirection will add \n> some overhead, but I doubt it'll be measurable.\n\nI keep hearing \"blame\" in this discussion, but I do not understand why\npeople think blame should _follow_ this \"origin\" information (in the usual\nsense of \"following\").\n\nSuppose you cherry-pick an existing commit from unrelated context:\n\n         ...---A---B\n                    . (origin)\n                     .\n        ...---o---X---Y---Z\n\ni.e. on top of X the difference to bring A to B is applied to produce Y,\nand a new development Z is made on top.  You start digging from Z.\n\nWithout any \"origin\", here is how blame works:\n\n * What Z did is blamed on Z; what Z did not change is passed to Y;\n\n * Y needs to:\n \n   (1) take responsibility for what it changed; and/or\n\n   (2) the remaining contents came from X --- pass the blame to it.\n\nLet's see how we would want \"origin\" get involved.  Instead of the above,\nwhat Y would do would be:\n\n   (1) if the contents (excluding the part Z changed) is different from X,\n       instead of taking the blame itself, give the _final_ blame to B.\n\n   (2) the remainder is passed to X as usual.\n\nThis is different from the normal \"following\" in that B is not allowed to\npass the blame to its parents (should it be allowed to pass it to its\n\"origin\"?), because the _only thing_ cherry-pick did was to transport what\nB did (relative to A) to the unrelated history that led to X.\n\nIOW, you did not look at the contents outside \"diff A..B\" when you made\nthe cherry-pick.  There could well be parts of the content that are common\nacross all of A, Y, X and Z, but as far as Y and Z are concerned, they did\nnot get any part of that common common content from A (otherwise \"origin\"\nis no different from \"parent\", but you did not merge).\n\nThe output from \"origin\" aware blame would be identical to the normal\nblame, except that lines that usually are labeled with Y are labeled with\nB.  However:\n\n   (1) If you _are_ interested in the line that says Y, you can look at\n       the commit object Y and see \"cherry-pick -x\" information to learn\n       it came from B already; and\n\n   (2) More importantly, if you want to dig deeper by peeling the blamed\n       line (I think gitweb allows this, and probably git-gui), you\n       shouldn't peel that line blamed on B to start running blame at A.\n       That would continue digging the history of A, which is wrong when\n       you are examining the history that led to Z.\n\nSo please leave \"blame\" out of this discussion.\n"},{"id":"90483","messageId":"20080911214650.GB3187@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080911200452.GM5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-11T21:46:50Z","receivedAt":"2008-09-11T21:46:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 11, 2008 at 04:04:53PM -0400, Theodore Tso wrote:\n\n> > It will be q^..q, and specifically not p^..p, using ^p..p would be\n> > lying.  We aim to document the evolvement of the patch in time.\n> > Cherry-pick itself will always ignore the origin links present on the\n> > old commit, it simply creates new ones as if the old ones didn't exist.\n> \n> So if you never pull branch C (where commit q resides), there is no\n> way for you to know that commits p and r are related.  How.... not\n> useful.\n\nThat is a good point. Stephen has explained his workflow, and I can see\nwhy he wants to reference the cherry-picked commits, and how he thinks\nthat the referenced commits will always be available in that workflow.\nAnd obviously in Linus's workflow such references are basically useless,\nand they should just not be generated.\n\nBut what about workflows in between? When I pull from some developer who\nhas added a weak reference to a particular commit SHA1, but I _don't_\nhave that commit, my next question \"OK, so what was in that commit?\".\nWhat is the mechanism by which I find out more information on that SHA1?\n\nUsing a key that is meaningful to an external database (like a bug\ntracker) means that you can go to that database to look up more\ninformation.\n\n-Peff\n"},{"id":"90488","messageId":"20080911223203.GA29559@cuci.nl","threadId":"15449","inReplyTo":"7vy71ys8a7.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T22:32:03Z","receivedAt":"2008-09-11T22:32:03Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n>I keep hearing \"blame\" in this discussion, but I do not understand why\n>people think blame should _follow_ this \"origin\" information (in the usual\n>sense of \"following\").\n\n>Suppose you cherry-pick an existing commit from unrelated context:\n\n>         ...---A---B\n>                    . (origin)\n>                     .\n>        ...---o---X---Y---Z\n\n>i.e. on top of X the difference to bring A to B is applied to produce Y,\n>and a new development Z is made on top.  You start digging from Z.\n\n>Let's see how we would want \"origin\" get involved.  Instead of the above,\n>what Y would do would be:\n\n>   (1) if the contents (excluding the part Z changed) is different from X,\n>       instead of taking the blame itself, give the _final_ blame to B.\n\n>   (2) the remainder is passed to X as usual.\n\nSounds reasonable.\n\n>This is different from the normal \"following\" in that B is not allowed to\n>pass the blame to its parents (should it be allowed to pass it to its\n>\"origin\"?), because the _only thing_ cherry-pick did was to transport what\n>B did (relative to A) to the unrelated history that led to X.\n\nWell, I'd expect:\na. That B should be able to pass blame onto it's origin.\nb. That B should be able to pass blame onto A (and deeper).\n\nLet me show another example:\n\n...-C---D---E---F---G\n                 . (origin)\n                  .\n         ...---A---B\n                    . (origin)\n                     .\n        ...---o---X---Y---Z\n\n\nNow suppose there is a piece of sourcecode which evolves from C to F,\nthen when I dig into G using blame I get something like:  CCCFFEGGDDDCC\n(Every letter represents a line in the sourcecode)\n\nDigging into Z I'd expect to see the following:  ZZCCCFFEDDYDCCB\n\nAll this assumes that there were minimal changes to the patch when\ncreating B, and also minimal changes to the patch when creating Y.\n\nI.e. large parts of that code where developed during C, D, E and F, so\nthat is what I expect to see; is that illogical?\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90489","messageId":"20080911224027.GB29559@cuci.nl","threadId":"15449","inReplyTo":"20080911223203.GA29559@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T22:40:27Z","receivedAt":"2008-09-11T22:40:27Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Stephen R. van den Berg wrote:\n>Junio C Hamano wrote:\n>>This is different from the normal \"following\" in that B is not allowed to\n>>pass the blame to its parents (should it be allowed to pass it to its\n>>\"origin\"?), because the _only thing_ cherry-pick did was to transport what\n>>B did (relative to A) to the unrelated history that led to X.\n\n>Well, I'd expect:\n>a. That B should be able to pass blame onto it's origin.\n>b. That B should be able to pass blame onto A (and deeper).\n\n>Let me show another example:\n\n>....-C---D---E---F---G\n>                 . (origin)\n>                  .\n>         ...---A---B\n>                    . (origin)\n>                     .\n>        ...---o---X---Y---Z\n\n>Now suppose there is a piece of sourcecode which evolves from C to F,\n>then when I dig into G using blame I get something like:  CCCFFEGGDDDCC\n>(Every letter represents a line in the sourcecode)\n\n>Digging into Z I'd expect to see the following:  ZZCCCFFEDDYDCCB\n\n>All this assumes that there were minimal changes to the patch when\n>creating B, and also minimal changes to the patch when creating Y.\n\n>I.e. large parts of that code where developed during C, D, E and F, so\n>that is what I expect to see; is that illogical?\n\nI'm sorry, you're right, I'm confusing things here.  The case I'm describing\nhere can only happen when you do this:\n\n....-C---D---E---F---G\n      \\...\\...\\..\\ (origin)\n                  .\n         ...---A---B\n                    . (origin)\n                     .\n        ...---o---X---Y---Z\n\nI.e. the first cherry-pick needs to cherry-pick C, D, E *and* F into B,\nthat will result in four origin fields there.\nAnd yes, that means that:\n- blame follows origin links (repeatedly).\n- blame does *not* travel to parents of commits found through an origin\n  link.\n\nDoes that mean that blame uses origin fields?  Yes, it does, and it has\nto check for origin links at every commit it traverses.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90492","messageId":"20080911225648.GC29559@cuci.nl","threadId":"15449","inReplyTo":"20080911214650.GB3187@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T22:56:48Z","receivedAt":"2008-09-11T22:56:48Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Thu, Sep 11, 2008 at 04:04:53PM -0400, Theodore Tso wrote:\n>> > It will be q^..q, and specifically not p^..p, using ^p..p would be\n>> > lying.  We aim to document the evolvement of the patch in time.\n>> > Cherry-pick itself will always ignore the origin links present on the\n>> > old commit, it simply creates new ones as if the old ones didn't exist.\n\n>> So if you never pull branch C (where commit q resides), there is no\n>> way for you to know that commits p and r are related.  How.... not\n>> useful.\n\n>But what about workflows in between? When I pull from some developer who\n>has added a weak reference to a particular commit SHA1, but I _don't_\n>have that commit, my next question \"OK, so what was in that commit?\".\n>What is the mechanism by which I find out more information on that SHA1?\n\nWell, the usual way to fix this is to actually startup fetch and tell it\nto try and fetch all the weak links (or just fetch a single hash (the\noffending origin link)) from upstream; this is by no means the default\noperatingmode of fetch, but I don't see any harm in allowing to fetch\nthose if one really wants to.\n\n>Using a key that is meaningful to an external database (like a bug\n>tracker) means that you can go to that database to look up more\n>information.\n\nTrue.  And also a Good Thing, I concur.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90493","messageId":"20080911230117.GA4194@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080911225648.GC29559@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-11T23:01:17Z","receivedAt":"2008-09-11T23:01:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 12, 2008 at 12:56:48AM +0200, Stephen R. van den Berg wrote:\n\n> Well, the usual way to fix this is to actually startup fetch and tell it\n> to try and fetch all the weak links (or just fetch a single hash (the\n> offending origin link)) from upstream; this is by no means the default\n> operatingmode of fetch, but I don't see any harm in allowing to fetch\n> those if one really wants to.\n\nMaybe I am misremembering the details of fetching, but I believe you\ncannot fetch an arbitrary SHA-1, and that is by design. So:\n\n  1. You would have to argue the merits of changing that design. I\n     believe the rationale relates to exposing some subset of the\n     content via refs, but I have personally never felt that is very\n     compelling.\n\n  2. Even if we did make a change, that means that _both_ sides need the\n     upgraded version.\n\n-Peff\n"},{"id":"90495","messageId":"alpine.LFD.1.10.0809111533110.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"20080911214650.GB3187@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-11T23:10:26Z","receivedAt":"2008-09-11T23:10:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 11 Sep 2008, Jeff King wrote:\n>\n> And obviously in Linus's workflow such references are basically useless,\n> and they should just not be generated.\n\nThis has _nothing_ to do with workflows or anything else.\n\nWhy are people claiming these total red herrings?\n\nI have asked several times what it is that makes it so important that the \n\"origin\" information be in the headers. Nobody has been able to explain \nwhy it's so different from just doing it in the free-form part. NOBODY.\n\nIf somebody has a workflow where they want to track \"origin\" commits, then \nthey can do it today with the in-body approach. But that has nothing \nwhat-so-ever to do with the question of \"let's change object file format \nto some odd special-case that we just made up and is only apparently \nuseful for some special workflow that uses special tools and special \nrules\".\n\nI want the git object database to have really clear semantics. The fields \nwe have now, we have because we _require_ them. There is nothing unclear \nwhat-so-ever about the semantics of author/commiter-ship, parenthood, \ntrees, or anything else.\n\nAnd there are _zero_ issues about \"workflow\". The workflow doesn't matter, \nthe objects always make sense, and they always work exactly the same way. \nThere are no special magic cases that are in the least questionable in any \nway.\n\nSo this argument is about more than just \"minimalism\", although I'll also \nadmit to that being an issue - I want to be able to basically explain how \ngit data structures work to any CS student, and not have any extra fat or \nany gray areas. It's about everything having a clear design, and a clear \nmeaning, and there never being any question what-so-ever about what the \nreal \"meaning\" of something is.\n\nThen, if you have some special use case or rules for your particular \nproject, well that's where you can have things like formatting rules for \nhow the commit messages should look like. If somebody wants to use fixed \nformat rules for their project, that's fine. And THAT is where \"workflow\" \nissues come up. \n\nBut \"workflow\" has nothing to do with core git data structures. They were \ndesigned for speed, stability, simplicity and good taste. The _workflow_ \npart has been designed on separately on top of that (example: the whole \nthing with a single-line top summary of a commit so that we can have \"git \nshortlog\" and the \"gitk\" single-line commit view etc).\n\nOf course, good and generally useful workflows can then be reflected in \nhow tools work, where that single line commit summary is an example of \nthat: it's not something that git data structures _enforce_ or even care \nabout, but it's obviously something that a lot of the porcelain expects, \nand without it, lots of tools will output less useful information.\n\nThe same goes for the existing SHA1-in-comment support: some tools already \nsupport it and help view it in certain ways, even though it is in no way a \ncore data structure issue. And _extending_ on that kind of helpful \nporcelain support certainly makes sense.\n\nThe only thing I have ever argued against is adding commit headers that \nhave no sane semantics and don't make sense as internal git data \nstructures.\n\n\t\t\tLinus\n"},{"id":"90496","messageId":"20080911231742.GD29559@cuci.nl","threadId":"15449","inReplyTo":"20080911230117.GA4194@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T23:17:42Z","receivedAt":"2008-09-11T23:17:42Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Fri, Sep 12, 2008 at 12:56:48AM +0200, Stephen R. van den Berg wrote:\n>> Well, the usual way to fix this is to actually startup fetch and tell it\n>> to try and fetch all the weak links (or just fetch a single hash (the\n>> offending origin link)) from upstream; this is by no means the default\n>> operatingmode of fetch, but I don't see any harm in allowing to fetch\n>> those if one really wants to.\n\n>Maybe I am misremembering the details of fetching, but I believe you\n>cannot fetch an arbitrary SHA-1, and that is by design. So:\n\nI see, didn't know that.\n\n>  1. You would have to argue the merits of changing that design. I\n>     believe the rationale relates to exposing some subset of the\n>     content via refs, but I have personally never felt that is very\n>     compelling.\n\nWell, I can understand why it is done this way, I think.\n\n>  2. Even if we did make a change, that means that _both_ sides need the\n>     upgraded version.\n\nIf you're using origin links, you'd need that anyway, so that's a given.\nI could imagine the minimum would be something like:\n\n Allow direct SHA1 fetches (which obviously pull in all parents as well)\n if the ref is part of one of the public branches (either as a commit,\n or as an origin link).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90497","messageId":"20080911232610.GA4279@coredump.intra.peff.net","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111533110.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-11T23:26:10Z","receivedAt":"2008-09-11T23:26:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 11, 2008 at 04:10:26PM -0700, Linus Torvalds wrote:\n\n> > And obviously in Linus's workflow such references are basically useless,\n> > and they should just not be generated.\n> \n> This has _nothing_ to do with workflows or anything else.\n> \n> Why are people claiming these total red herrings?\n> \n> I have asked several times what it is that makes it so important that the \n> \"origin\" information be in the headers. Nobody has been able to explain \n> why it's so different from just doing it in the free-form part. NOBODY.\n\nThe message you are responding to has nothing to do with an origin\nheader versus putting it in the free-form part. It is equally a problem\nwith both approaches.\n\nI was purely commenting on the \"if I mention an arbitrary sha-1, what is\nthe person reading it supposed to _do_ with it, if they may never have\nseen that sha-1\" issue.\n\nSo yes, it has _everything_ to do with workflows. In Stephen's case, he\nclaims that all references will be to commits on long-lived branches. In\nwhich case, it is a non-issue because they will have the referenced\ncommits.\n\nBut in the general case, people will not have them, and there is\npotential head-scratching. My point is that even if a feature works for\nStephen's workflow, it may not be a good feature for everyone, since\nother solutions handle the general case (as well as his case) much\nbetter.\n\n> [ranting about how the origin header is bad]\n> The only thing I have ever argued against is adding commit headers that \n> have no sane semantics and don't make sense as internal git data \n> structures.\n\nYes, and I totally agree with everything you said. If you read the mail\nyou are responding to carefully, you will see that I never mention an\norigin header versus the free-form commit.\n\n-Peff\n"},{"id":"90498","messageId":"48C9A9A4.8090703@vilain.net","threadId":"15449","inReplyTo":"200809101823.22072.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2008-09-11T23:28:36Z","receivedAt":"2008-09-11T23:28:36Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Jakub Narebski wrote:\n>> And that dotted line really does sound like something you could do with \n>> just the existing \"hyperlink\" functionality in the commit message.\n> \n> As far as I understand (note: I'm neither for, nor against the proposal;\n> although I think it has thin chance to be accepted, especially soon),\n> it is for graphical history viewers, for git-cherry to make it more\n> precise (to detect duplicated/cherry-picked changes better), and in\n> the future possibly to help history-aware merge strategies. And probably\n> help patch management interfaces.\n\nCan I suggest,\n\n 1. bury this origin link idea\n\n 2. make git-cherry-pick have a similar option to '-x', but instead of\n    recording the original commit ID, record the original *patch* ID,\n    *if* there was a merge conflict for that cherry pick.\n\n 3. tools can build indexes from patch ID => (commit IDs) to make this\n    other form of history navigation fast.\n\nSam\n"},{"id":"90500","messageId":"20080911233634.GE29559@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111533110.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-11T23:36:34Z","receivedAt":"2008-09-11T23:36:34Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>I have asked several times what it is that makes it so important that the \n>\"origin\" information be in the headers. Nobody has been able to explain \n>why it's so different from just doing it in the free-form part. NOBODY.\n\nThat's because the difference is small:\nIn the header is slightly faster and more elegant (both designwise and\ndisplaywise), that's it.\nOther than that, it hardly matters.\n\n>The only thing I have ever argued against is adding commit headers that \n>have no sane semantics and don't make sense as internal git data \n>structures.\n\nOf course.\n\nIn any case, I think I got enough feedback from the list to create\na working implementation/concept which is going to use the free-form\ntrailer to implement the origin field.\n-- \nSincerely,\n           Stephen R. van den Berg.\n"},{"id":"90502","messageId":"alpine.LFD.1.10.0809111641110.3384@nehalem.linux-foundation.org","threadId":"15449","inReplyTo":"48C9A9A4.8090703@vilain.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-11T23:44:36Z","receivedAt":"2008-09-11T23:44:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 12 Sep 2008, Sam Vilain wrote:\n> \n>  2. make git-cherry-pick have a similar option to '-x', but instead of\n>     recording the original commit ID, record the original *patch* ID,\n>     *if* there was a merge conflict for that cherry pick.\n\nActually, don't make it dependent on merge conflicts. Just make it depend \non whether the patch ID is _different_.\n\nIt can happen even without any conflicts, just because the context \nchanged. So it really isn't about merge conflicts per se, just the fact \nthat a patch can change when it is applied in a new area with a three-way \ndiff - or because it got applied with fuzz.\n\nYou could add it as a \n\n\tOriginal-patch-id: <sha1>\n\nor something. And then you just need to teach \"git cherry/rebase\" to take \nboth the original ID and the new one into account when deciding whether it \nhas already seen that patch.\n\n\t\t\tLinus\n"},{"id":"90505","messageId":"48C9B1C8.9070007@gmail.com","threadId":"15449","inReplyTo":"20080911123902.GB2056@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2008-09-12T00:03:20Z","receivedAt":"2008-09-12T00:03:20Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Stephen R. van den Berg wrote:\n> A Large Angry SCM wrote:\n>> Stephen R. van den Berg wrote:\n>>> If you fetch just branches A, B and C, but not D, the origin link from A\n>>> to D is dangling. \n> \n>> I do not understand how this can be considered an acceptable behavior. \n>> If an object ID is referenced in an object header, particularly commit \n>> objects, fetch must gather those objects also because to do otherwise \n>> breaks the cryptographic authentication in git.\n> \n> No it does not.\n> The cryptographic seal is calculated over the content of the commit,\n> which includes the hashes of all referenced objects, but doesn't include\n> the objects themselves.\n> The content of the commit is not violated.\n\nThe fetch MUST gather the referenced objects ALWAYS or I can't verify \nthe history. To do otherwise means that ID strings on the origin lines \nare nothing more than an arbitrary text tag and not pointer to a \nspecific history.\n\n> \n> Do not forget though:\n> - origin links are a rare occurrence.\n> - When they occur, they usually were made to point into other (deemed)\n>   important public branches.\n> - Due to the fact that the branches they are pointing into are important\n>   and public, in most cases the origin links *will* point to objects you\n>   actually already have (even if you fetched from someone else).\n> - The only time you're going to have dangling origin links is when\n>   they were pointing at someone's private branches, in which case it was\n>   not very prudent of the committer to actually record the link in the\n>   first place.  But nothing breaks if you don't have his private branch\n>   locally.\n\nHow do I verify (think git-fsck) that what the origin lines refer to \nare, in fact, commits with the proper relationships? Either they HAVE to \nbe in the repository or the references do not belong in the header.\n"},{"id":"90506","messageId":"20080912001323.GG29559@cuci.nl","threadId":"15449","inReplyTo":"48C9B1C8.9070007@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T00:13:23Z","receivedAt":"2008-09-12T00:13:23Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"A Large Angry SCM wrote:\n>Stephen R. van den Berg wrote:\n>>No it does not.\n>>The cryptographic seal is calculated over the content of the commit,\n>>which includes the hashes of all referenced objects, but doesn't include\n>>the objects themselves.\n>>The content of the commit is not violated.\n\n>The fetch MUST gather the referenced objects ALWAYS or I can't verify \n>the history. To do otherwise means that ID strings on the origin lines \n>are nothing more than an arbitrary text tag and not pointer to a \n>specific history.\n\nTo fetch, by default, the origin lines *are* nothing more than arbitrary\ntext and not a pointer to a specific history.\n\n>How do I verify (think git-fsck) that what the origin lines refer to \n>are, in fact, commits with the proper relationships? Either they HAVE to \n>be in the repository or the references do not belong in the header.\n\nIf the origin hashes are not reachable, then fsck is required to silently\nskip them, according to spec.\nIf the origin hashes *are* reachable, then fsck is required to verify\nthat they refer to proper commits with a normal history.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\"There are three types of people in the world;\n those who can count, and those who can't.\"\n"},{"id":"90507","messageId":"48C9B830.2060903@gmail.com","threadId":"15449","inReplyTo":"20080911202228.GG1451@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2008-09-12T00:30:40Z","receivedAt":"2008-09-12T00:30:40Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Stephen R. van den Berg wrote:\n> Jakub Narebski wrote:\n>> Stephen R. van den Berg wrote:\n>>> Well, the principle of least surprise dictates that they should be kept\n>>> by gc as described above, however...\n>>> I can envision an option to gc say \"--drop-weak-links\" which does\n>>> exactly what you describe.\n> \n>> Well, IIRC the need for this was one of the causes of \"death\" of 'prior'\n>> header link proposal...\n> \n> As I understood it, one of the causes of death of the \"prior\" link\n> proposal was that it was unclear if it pulled in the linked-to commits\n> upon fetch.  In the \"origin\" case, the default is *not* to fetch them.\n\nAnd that's WRONG. Both prior and origin must fetch them if they're \nreference in the header.\n"},{"id":"90510","messageId":"48C9D2E8.1000506@vilain.net","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111641110.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2008-09-12T02:24:40Z","receivedAt":"2008-09-12T02:24:40Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> \n> On Fri, 12 Sep 2008, Sam Vilain wrote:\n>>  2. make git-cherry-pick have a similar option to '-x', but instead of\n>>     recording the original commit ID, record the original *patch* ID,\n>>     *if* there was a merge conflict for that cherry pick.\n> \n> Actually, don't make it dependent on merge conflicts. Just make it depend \n> on whether the patch ID is _different_.\n> \n> It can happen even without any conflicts, just because the context \n> changed. So it really isn't about merge conflicts per se, just the fact \n> that a patch can change when it is applied in a new area with a three-way \n> diff - or because it got applied with fuzz.\n> \n> You could add it as a \n> \n> \tOriginal-patch-id: <sha1>\n> \n> or something. And then you just need to teach \"git cherry/rebase\" to take \n> both the original ID and the new one into account when deciding whether it \n> has already seen that patch.\n\nYes, right - it's the patch ID changing that's the problem for\ngit-cherry / rev-list --cherry-pick to be able to spot changes as the\n'same'.\n\nSomeone else pointed out that git-rebase -i might want to have this as well.\n\nI actually looked into coding this, but there was a little problem with\nthe way git-revert worked - it builds the commit message before the diff\nis calculated.  So there would probably need to be a little trivial\nrefactoring first before this can be implemented.\n\nSam.\n"},{"id":"90515","messageId":"20080912053910.GA22228@cuci.nl","threadId":"15449","inReplyTo":"48C9B830.2060903@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T05:39:10Z","receivedAt":"2008-09-12T05:39:10Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"A Large Angry SCM wrote:\n>Stephen R. van den Berg wrote:\n>>Jakub Narebski wrote:\n>>>Stephen R. van den Berg wrote:\n>>>>Well, the principle of least surprise dictates that they should be kept\n>>>>by gc as described above, however...\n>>>>I can envision an option to gc say \"--drop-weak-links\" which does\n>>>>exactly what you describe.\n\n>>>Well, IIRC the need for this was one of the causes of \"death\" of 'prior'\n>>>header link proposal...\n\n>>As I understood it, one of the causes of death of the \"prior\" link\n>>proposal was that it was unclear if it pulled in the linked-to commits\n>>upon fetch.  In the \"origin\" case, the default is *not* to fetch them.\n\n>And that's WRONG. Both prior and origin must fetch them if they're \n>reference in the header.\n\nBy definition of the origin headerfield that is not wrong, there are no\nother rules.  But the point is moot at the moment, since I'm going to\ncreate a proof of concept which puts the field in the free-form trailer.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90516","messageId":"20080912054739.GB22228@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111641110.3384@nehalem.linux-foundation.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T05:47:39Z","receivedAt":"2008-09-12T05:47:39Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n>On Fri, 12 Sep 2008, Sam Vilain wrote:\n>It can happen even without any conflicts, just because the context \n>changed. So it really isn't about merge conflicts per se, just the fact \n>that a patch can change when it is applied in a new area with a three-way \n>diff - or because it got applied with fuzz.\n\nQuite.\n\n>You could add it as a \n\n>\tOriginal-patch-id: <sha1>\n\nThat will probably work fine when operating locally on (short) temporary\nbranches.\n\nIt would probably become computationally prohibitive to use it between\nlong lived permanent branches.  In that case it would need to be\naugmented by the sha1 of the originating commit.  Which gives you two\nhashes as reference, and in that case you might as well use the two\ncommit hashes of which the difference yields the patch.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90517","messageId":"48CA09E7.8090102@dawes.za.net","threadId":"15449","inReplyTo":"20080912054739.GB22228@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2008-09-12T06:19:19Z","receivedAt":"2008-09-12T06:19:19Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Stephen R. van den Berg wrote:\n> Linus Torvalds wrote:\n>> On Fri, 12 Sep 2008, Sam Vilain wrote:\n>> It can happen even without any conflicts, just because the context \n>> changed. So it really isn't about merge conflicts per se, just the fact \n>> that a patch can change when it is applied in a new area with a three-way \n>> diff - or because it got applied with fuzz.\n> \n> Quite.\n> \n>> You could add it as a \n> \n>> \tOriginal-patch-id: <sha1>\n> \n> That will probably work fine when operating locally on (short) temporary\n> branches.\n> \n> It would probably become computationally prohibitive to use it between\n> long lived permanent branches.  In that case it would need to be\n> augmented by the sha1 of the originating commit.  Which gives you two\n> hashes as reference, and in that case you might as well use the two\n> commit hashes of which the difference yields the patch.\n\nPardon my confusion, but why include two commit hashes? Surely the \ncommit already has its parent, so there is no need to include that in \nyour \"cherry pick\". And if the commit has more than one parent, then I \ndoubt you could/should really cherry-pick it anyway.\n\nBesides, you could always augment your local repo with a mapping of \npatch ids to commits/commit pairs to reduce lookup time.\n\nRogan\n"},{"id":"90519","messageId":"20080912065658.GA15391@cuci.nl","threadId":"15449","inReplyTo":"48CA09E7.8090102@dawes.za.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T06:56:58Z","receivedAt":"2008-09-12T06:56:58Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Rogan Dawes wrote:\n>Stephen R. van den Berg wrote:\n>>It would probably become computationally prohibitive to use it between\n>>long lived permanent branches.  In that case it would need to be\n>>augmented by the sha1 of the originating commit.  Which gives you two\n>>hashes as reference, and in that case you might as well use the two\n>>commit hashes of which the difference yields the patch.\n\n>Pardon my confusion, but why include two commit hashes? Surely the \n>commit already has its parent, so there is no need to include that in \n>your \"cherry pick\". And if the commit has more than one parent, then I \n>doubt you could/should really cherry-pick it anyway.\n\nWell, actually, sometimes cherry-pick does pick just one of the\n(multiple) parents to diff with; also, some people (not I) envisioned\nusing two commits which were not a direct parent and child of one\nanother (I'm not quite sure how that would work, but the model would\nsupport it).\n\n>Besides, you could always augment your local repo with a mapping of \n>patch ids to commits/commit pairs to reduce lookup time.\n\nYes, possible.  But then after cloning, this mapping-cache needs to be\nrecreated, and that would mean that one would have to walk through all\ncommits and calculate all patch-id's, of which then only those few which are\nreferenced need to be stored.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90523","messageId":"20080912084027.GB15391@cuci.nl","threadId":"15449","inReplyTo":"20080911210118.GO5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T08:40:27Z","receivedAt":"2008-09-12T08:40:27Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Theodore Tso wrote:\n>On Thu, Sep 11, 2008 at 09:55:16PM +0200, Stephen R. van den Berg wrote:\n>> >  Having it versionned also \n>> >means that older git versions will be able to carry that information \n>> >even if they won't make any use of it, and that also solves the \n>> >cryptographic issue since that data is part of the top commit SHA1.\n\n>> It would allow the data to be faked, that is undesirable for \"git blame\".\n\n>Why would this matter?  The information is largely\n>self-authenticating.  If a commit claims to have come from some other\n\nAttack-wise, you're right, it's not a big deal.\nI think the comforting feeling one gets about the hashes protecting\nintegrity is what matters more for me here.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90524","messageId":"20080912085021.GC15391@cuci.nl","threadId":"15449","inReplyTo":"alpine.LFD.1.10.0809111604040.23787@xanadu.home","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T08:50:21Z","receivedAt":"2008-09-12T08:50:21Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n>On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> Nicolas Pitre wrote:\n>> >On Thu, 11 Sep 2008, Stephen R. van den Berg wrote:\n>> >> when doing things with temporary branches.  The origin field is meant to\n>> >> be filled *ONLY* when cherry-picking from one permanent branch to\n>> >> another permanent branch.  This is a *rare* operation.\n\n>> >... and therefore you might as well just have a separate file (which \n>> >might or might not be tracked by git like the .gitignore files are) \n>> >to keep that information?  Since this is a rare operation, modifying the \n>> >core database structure for this doesn't appear that appealing to most \n>> >so far.\n\n>> For various reasons, the best alternate place would be at the trailing\n>> end of the free-form field.  Using a separate structure causes\n>> (performance) problems (mostly).\n\n>Did you try it?\n\nNo.\n\n>  I don't particularly buy this performance argument, and \n>the bulk of my contributions to git so far were about performances.  It \n>is quite easy to load a flat file with sorted commit SHA1s, and given \n>that origin links are the result of a rare operation, then there \n>shouldn't be too many entries to search through.  Hell, doing 213647 \n\nTrue.\n\n>lookups (and many other things like inflating zlib deflated data)  with \n>each of them for commit objects in my Linux repository which has 1355167 \n>total entries takes only 6 seconds here, or about a quarter of a \n>milisecond for each lookup.  I doubt doing an extra lookup in a much \n>smaller table would show on the radar.\n\nMaybe you're right.  The reason why my first knee-jerk reaction is\n\"performance problem\" is because:\n- The field is rarely present.\n- When it is used, we look for it on every commit we traverse.\n- This means that finding out the field does *not* exist is the most\n  common operation, and that effort rises linearly with the number of\n  commits visited.\n\nWhereas if the information is present in the header or trailer of the\ncommit, finding out that the field does not exist there is rather\ncheap.  But you could very well be right, that the absolute extra time\nspent might be negligible for all intents and purposes.\n\nNonetheless, the data-integrity argument still holds, i.e. placing it in\nthe commit (header or trailer) automatically protects it.  External\nfiles need extra care if you want the same integrity protection.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90562","messageId":"20080912145802.GV5082@mit.edu","threadId":"15449","inReplyTo":"20080912054739.GB22228@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-12T14:58:02Z","receivedAt":"2008-09-12T14:58:02Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Sep 12, 2008 at 07:47:39AM +0200, Stephen R. van den Berg wrote:\n> >You could add it as a \n> \n> >\tOriginal-patch-id: <sha1>\n> \n> That will probably work fine when operating locally on (short) temporary\n> branches.\n> \n> It would probably become computationally prohibitive to use it between\n> long lived permanent branches.  In that case it would need to be\n> augmented by the sha1 of the originating commit.\n\nNope, as Sam suggested in his original message (but which got clipped\nby Linus when he was replying) all you have to do is to have a\nseparate local database which ties commits and patch-id's together as\na cache/index.\n\nI know you seem to be resistent to caches, but caches are **good**\nbecause they are local information, which by definition can be\nimplementation-dependent; you can always generate the cache from the\ngit repository if for some reason you need to extend it.  It also\nmeans that if it turns out you need to index reationships a different\nway, you can do that without having to make fundamental (incompatible)\nchanges in the git object.  \n\nIt's much like SQL databases; you have your database tables, where\nmaking changes to the database schema is painful --- and indexes,\nwhich can be added and dropped with much less effort.  Think of these\nlocal caches are database indexes.  Just because you need an index in\na particular direction to optimize a query or loopup operation does\n***not*** imply that you need to make a fundamental, globally visible,\ndatabase schema change or git object layout which breaks compatibility\nfor everybody.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"90563","messageId":"48CA8544.8090000@gnu.org","threadId":"15449","inReplyTo":"20080912145802.GV5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-12T15:05:40Z","receivedAt":"2008-09-12T15:05:40Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"Theodore Tso wrote:\n> On Fri, Sep 12, 2008 at 07:47:39AM +0200, Stephen R. van den Berg wrote:\n>>> You could add it as a \n>>> \tOriginal-patch-id: <sha1>\n>> That will probably work fine when operating locally on (short) temporary\n>> branches.\n>>\n>> It would probably become computationally prohibitive to use it between\n>> long lived permanent branches.  In that case it would need to be\n>> augmented by the sha1 of the originating commit.\n> \n> Nope, as Sam suggested in his original message (but which got clipped\n> by Linus when he was replying) all you have to do is to have a\n> separate local database which ties commits and patch-id's together as\n> a cache/index.\n\nYeah, I must admit I am okay with *this* cache.\n\nPaolo\n"},{"id":"90565","messageId":"200809121711.32448.jnareb@gmail.com","threadId":"15449","inReplyTo":"20080912145802.GV5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-09-12T15:11:31Z","receivedAt":"2008-09-12T15:11:31Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Theodore Tso wrote:\n> On Fri, Sep 12, 2008 at 07:47:39AM +0200, Stephen R. van den Berg wrote:\n>>>\n>>>You could add it as a \n>>>\n>>>\tOriginal-patch-id: <sha1>\n>> \n>> That will probably work fine when operating locally on (short) temporary\n>> branches.\n>> \n>> It would probably become computationally prohibitive to use it between\n>> long lived permanent branches.  In that case it would need to be\n>> augmented by the sha1 of the originating commit.\n> \n> Nope, as Sam suggested in his original message (but which got clipped\n> by Linus when he was replying) all you have to do is to have a\n> separate local database which ties commits and patch-id's together as\n> a cache/index.\n> \n> I know you seem to be resistent to caches, but caches are **good**\n> because they are local information, which by definition can be\n> implementation-dependent; you can always generate the cache from the\n> git repository if for some reason you need to extend it. [...]\n\nBut it is not true that \"you can always generate the cache from the\ngit repository\" in this case; the patch-id that is to be saved is\n_original_ patch-id of cherry-picked (or reverted) changeset.\n\nOTOH it is not much different from reflog information, which also\ncannot be regenerated from object database.\n-- \nJakub Narebski\nPoland\n"},{"id":"90570","messageId":"48CA8D6A.4000303@gnu.org","threadId":"15449","inReplyTo":"200809121711.32448.jnareb@gmail.com","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-09-12T15:40:26Z","receivedAt":"2008-09-12T15:40:26Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n>> I know you seem to be resistent to caches, but caches are **good**\n>> because they are local information, which by definition can be\n>> implementation-dependent; you can always generate the cache from the\n>> git repository if for some reason you need to extend it. [...]\n> \n> But it is not true that \"you can always generate the cache from the\n> git repository\" in this case; the patch-id that is to be saved is\n> _original_ patch-id of cherry-picked (or reverted) changeset.\n\nHe's proposing storing the original patch id in the commit message, and\ncaching the commit SHA->patch id association on the side.\n\nPaolo\n"},{"id":"90572","messageId":"20080912155427.GB2915@cuci.nl","threadId":"15449","inReplyTo":"20080912145802.GV5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T15:54:27Z","receivedAt":"2008-09-12T15:54:27Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Theodore Tso wrote:\n>On Fri, Sep 12, 2008 at 07:47:39AM +0200, Stephen R. van den Berg wrote:\n>> It would probably become computationally prohibitive to use it between\n>> long lived permanent branches.  In that case it would need to be\n>> augmented by the sha1 of the originating commit.\n\n>Nope, as Sam suggested in his original message (but which got clipped\n>by Linus when he was replying) all you have to do is to have a\n>separate local database which ties commits and patch-id's together as\n>a cache/index.\n\nTrue.  But repopulating this cache after cloning means that you have to\ncalculate the patch-id of *every* commit in the repository.  It sounds\nlike something to avoid, but maybe I'm overly concerned, I have only a\nvague idea on how computationally intensive this is.\n\n>I know you seem to be resistent to caches, but caches are **good**\n>because they are local information, which by definition can be\n>implementation-dependent; you can always generate the cache from the\n>git repository if for some reason you need to extend it.  It also\n>means that if it turns out you need to index reationships a different\n>way, you can do that without having to make fundamental (incompatible)\n>changes in the git object.  \n\nI fully agree that caches are good.\nAnd yes I seem to resist the idea to create a cache at every whim, but\nthat mostly is because I want to avoid that everyone invents their own\nmini-database for each and every data access they want to accellerate.\n\nI mean, ideally, any database/index/accellerator structure you'd need\ncan reuse the SHA1 object database index, or maybe one or two other\nsemi-standard index types, and git would provide suitable library\nfunctions for all three solutions.  And if that would be the case, I'll\ngladly throw in an extra cache or index at anytime to speed up the\nparticular access pattern I'm trying to make useable.  But as far as I\ncan see, those library functions have not materialised yet, so I'm\nhesitant to create yet another private database structure just for my\naccess patterns; and simply pulling in libdb or sqlite without agreement\nthat those libs are (re)used in a lot of places in git seems a bit\nbloat-prone.\n\n>local caches are database indexes.  Just because you need an index in\n>a particular direction to optimize a query or loopup operation does\n>***not*** imply that you need to make a fundamental, globally visible,\n>database schema change or git object layout which breaks compatibility\n>for everybody.\n\nIt's not a certainty that changing the git object layout has to break\ncompatibility (it should be reasonably possible to add columns to the\nschema without breaking anything, to stay with the database paradigm),\nbut I agree that creating another index can be considered better than\nextending the schema.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90575","messageId":"20080912160059.GY5082@mit.edu","threadId":"15449","inReplyTo":"48CA8D6A.4000303@gnu.org","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-12T16:00:59Z","receivedAt":"2008-09-12T16:00:59Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Sep 12, 2008 at 05:40:26PM +0200, Paolo Bonzini wrote:\n> > But it is not true that \"you can always generate the cache from the\n> > git repository\" in this case; the patch-id that is to be saved is\n> > _original_ patch-id of cherry-picked (or reverted) changeset.\n> \n> He's proposing storing the original patch id in the commit message, and\n> caching the commit SHA->patch id association on the side.\n> \n\nActually its the association in the other direction which you'd want\nto cache.  It's fast given the commit SHA to dig the original patch id\nout of the commit message.  What is harder is given a patch id X, to\nfind all of the commits which either (a) have a patch id of X, or (b)\nhave a commit message indicating that the original patch-id was X.  So\nhaving a database which caches this information, so given a patch-id,\nyou can quickly look up the related commits, is what I believe Sam was\nproposing, and which I think would solve the problem quite nicely.\n\n\t       \t       \t     \t   \t     - Ted\n"},{"id":"90579","messageId":"20080912161911.GA12096@coredump.intra.peff.net","threadId":"15449","inReplyTo":"20080912155427.GB2915@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-09-12T16:19:12Z","receivedAt":"2008-09-12T16:19:12Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 12, 2008 at 05:54:27PM +0200, Stephen R. van den Berg wrote:\n\n> True.  But repopulating this cache after cloning means that you have to\n> calculate the patch-id of *every* commit in the repository.  It sounds\n> like something to avoid, but maybe I'm overly concerned, I have only a\n> vague idea on how computationally intensive this is.\n\nFor a rough estimate, try:\n\n  time git log -p | git patch-id >/dev/null\n\n-Peff\n"},{"id":"90582","messageId":"20080912164348.GC2915@cuci.nl","threadId":"15449","inReplyTo":"20080912161911.GA12096@coredump.intra.peff.net","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T16:43:48Z","receivedAt":"2008-09-12T16:43:48Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jeff King wrote:\n>On Fri, Sep 12, 2008 at 05:54:27PM +0200, Stephen R. van den Berg wrote:\n\n>> True.  But repopulating this cache after cloning means that you have to\n>> calculate the patch-id of *every* commit in the repository.  It sounds\n>> like something to avoid, but maybe I'm overly concerned, I have only a\n>> vague idea on how computationally intensive this is.\n\n>For a rough estimate, try:\n\n>  time git log -p | git patch-id >/dev/null\n\nOn my system that results in 2ms per commit on average.  Not huge, but\nnot small either, I guess.  Running it results in real waiting time, it\nall depends on how patient the user is.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90589","messageId":"20080912184406.GB5082@mit.edu","threadId":"15449","inReplyTo":"20080912164348.GC2915@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-09-12T18:44:06Z","receivedAt":"2008-09-12T18:44:06Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Sep 12, 2008 at 06:43:48PM +0200, Stephen R. van den Berg wrote:\n> >> True.  But repopulating this cache after cloning means that you have to\n> >> calculate the patch-id of *every* commit in the repository.  It sounds\n> >> like something to avoid, but maybe I'm overly concerned, I have only a\n> >> vague idea on how computationally intensive this is.\n> \n> >For a rough estimate, try:\n> \n> >  time git log -p | git patch-id >/dev/null\n> \n> On my system that results in 2ms per commit on average.  Not huge, but\n> not small either, I guess.  Running it results in real waiting time, it\n> all depends on how patient the user is.\n\nFor a local clone, git could be taught to copy the cache file.  For a\nnetwork-based clone, the percentage of time needed to download is\nroughly 2-3 times that (although that will obviously depend on your\nnetwork connectivity).  Building this cache can be done in the\nbackground, though, or delayed until the first time the cache is\nneeded.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"90593","messageId":"20080912205618.GA8711@cuci.nl","threadId":"15449","inReplyTo":"20080912184406.GB5082@mit.edu","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-09-12T20:56:18Z","receivedAt":"2008-09-12T20:56:18Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Theodore Tso wrote:\n>On Fri, Sep 12, 2008 at 06:43:48PM +0200, Stephen R. van den Berg wrote:\n>> On my system that results in 2ms per commit on average.  Not huge, but\n>> not small either, I guess.  Running it results in real waiting time, it\n>> all depends on how patient the user is.\n\n>For a local clone, git could be taught to copy the cache file.  For a\n>network-based clone, the percentage of time needed to download is\n>roughly 2-3 times that (although that will obviously depend on your\n>network connectivity).  Building this cache can be done in the\n>background, though, or delayed until the first time the cache is\n>needed.\n\nFair enough.  If noone beats me to it, I'll probably take a stab at\nimplementing something like this and see how it fares for my own\napplication.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Father's Day: Nine months before Mother's Day.\"\n"},{"id":"90733","messageId":"1221481308.29145.21.camel@maia.lan","threadId":"15449","inReplyTo":"20080912155427.GB2915@cuci.nl","subject":"Re: [RFC] origin link for cherry-pick and revert","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2008-09-15T12:21:48Z","receivedAt":"2008-09-15T12:21:48Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On Fri, 2008-09-12 at 17:54 +0200, Stephen R. van den Berg wrote:\n> Theodore Tso wrote:\n> >Nope, as Sam suggested in his original message (but which got clipped\n> >by Linus when he was replying) all you have to do is to have a\n> >separate local database which ties commits and patch-id's together as\n> >a cache/index.\n> \n> True.  But repopulating this cache after cloning means that you have to\n> calculate the patch-id of *every* commit in the repository.  It sounds\n> like something to avoid, but maybe I'm overly concerned, I have only a\n> vague idea on how computationally intensive this is.\n\nYou don't necessarily need to do that.  If the tool decides that the\nsha1 it finds in the message is a patch-id reference, well it can just\nstart hunting around, caching the patch-ids it calculates as it finds\nthem, until it either finds one that matches, or determines you don't\nhave it.  You can probably find it first try just based on the author\nname and date 90% of the time anyway.\n\nMaybe the machinery could be adequately tilted such that if someone is\nreally desperate to make sure they are found quickly they can put the\ninformation at refs/patches/PATCHID/COMMITID, but that sounds a bit\nabusive.\n\nSam.\n"},{"id":"91393","messageId":"Pine.LNX.4.64.0809231447030.28506@ds9.cixit.se","threadId":"15449","inReplyTo":"20080909211355.GB10544@machine.or.cz","subject":"Recording \"partial merges\" (was: Re: [RFC] origin link for cherry-pick and revert)","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-09-23T13:51:17Z","receivedAt":"2008-09-23T13:51:17Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Petr Baudis:\n\n> I think this is misguided. In general case, cherrypicks can be from\n> completely unrelated histories, and if you are doing the cherry pick,\n> you are saying that actually, the history *does not matter*.\n\nAs my workflow sometimes make me do cherry-picks where history does\nmatter, in form of me doing a \"partial merge\" of one or more than one\ncommit from branch A into branch B, which does not necessarily have to\nbe directly related,\n\n is there a way to perform something like that, while keeping history?\n\n\nPerhaps I'm damaged by having used CVS for too long, and merging just\nsome files, or abusing CVS internals to make some files on branch A\nalso be part of branch B by having their branches point to the same RCS\nbranch revision number, but sometimes I find that I miss being able to\ndo it in Git.\n\nNot really a big deal, just curious.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"}]}