{"thread":{"id":"52509","subject":"Feature request: add a metadata in the commit: the \"commited in branch\" information","startedAt":"2019-12-23T12:56:54Z","lastAt":"2019-12-30T15:40:28Z","messageCount":6,"participants":["Arnaud Bertrand","Junio C Hamano","Theodore Y. Ts'o","Paul Smith"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"388804","messageId":"CAEW0o+jV+r1UMZReRXa3g_fyqCYxHTVYVf6pWvjB7_isofbBaw@mail.gmail.com","threadId":"52509","inReplyTo":null,"subject":"Feature request: add a metadata in the commit: the \"commited in branch\" information","fromName":"Arnaud Bertrand","fromEmail":"xda@abalgo.com","sentAt":"2019-12-23T12:56:41Z","receivedAt":"2019-12-23T12:56:54Z","isPatch":false,"sender":{"key":"xda@abalgo.com","avatar":null},"body":"Hello,\n\nGit is a nice tool but one of the most important missing information\nis the branch in which a commit was done.\nI understood that in git philosophy, once it is merged, a branch can\ndisappear. But for a lot of companies, a SCM is also a guardian of the\nhistory.\nWith this point of view, keeping track of the branch name when the\ncommit was done should be a very very big improvement (and a Major\nargument to switch to git)\nI speak just about a meta-data, exactly as the committer username,\nemail or date... no more.\nIf the branch is removed in the future or is renamed... so what, we\nhave at least its name at the time of the commit (better than\nnothing).\n\nToday, all my git repositories are using hooks to add the name of the\nbranch as header of the comment. But it would be so better to have it\nofficially and automatically and accessible as a git log meta-data.\nIt does not imply any constrains, simply a few characters more in the commit.\nWe can also imagine a core.branchInCommit parameter (true by default\n;-) ) that could be set to false for those that don't one it.\nThe only commands affected should be git commit, git merge --no-ff and\ngit log that should be able to show this metadata.\n\n\nBest regards,\n\nArnaud Bertrand\n"},{"id":"389049","messageId":"xmqqd0c6iuw0.fsf@gitster-ct.c.googlers.com","threadId":"52509","inReplyTo":"CAEW0o+jV+r1UMZReRXa3g_fyqCYxHTVYVf6pWvjB7_isofbBaw@mail.gmail.com","subject":"Re: Feature request: add a metadata in the commit: the \"commited in branch\" information","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-29T23:17:51Z","receivedAt":"2019-12-29T23:19:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Arnaud Bertrand <xda@abalgo.com> writes:\n\n> I understood that in git philosophy, once it is merged, a branch can\n> disappear. But for a lot of companies, a SCM is also a guardian of the\n> history.\n\nA lot more important point than \"once it is merged\" is that the\nbranch identity is strictly local to your repository.  Contaminating\nthe object header, which is cast in stone and cannot be modified\nafter the fact, with such a piece of information will not mix well\nwith the rest of Git, so ...\n\n\n"},{"id":"389052","messageId":"CAEW0o+g7vXj841h+4nNK8iSoO758Uh9fLKMCN87RE2w2Nd=CRg@mail.gmail.com","threadId":"52509","inReplyTo":"xmqqd0c6iuw0.fsf@gitster-ct.c.googlers.com","subject":"Re: Feature request: add a metadata in the commit: the \"commited in branch\" information","fromName":"Arnaud Bertrand","fromEmail":"xda@abalgo.com","sentAt":"2019-12-29T23:53:56Z","receivedAt":"2019-12-30T00:14:55Z","isPatch":false,"sender":{"key":"xda@abalgo.com","avatar":null},"body":"Hi Junio,\n\nIt really depends how git is used. With big collaborative project\n(like git or linux kernel), you are totally right.\nfor development limited to a company that has developments with team\nof 10-20 developers and that uses\na correct SCM plan, the name of the branch is regulated and is\nmeaningful, mostly  linked to a bug tracking system\nsystem. For audits and  traceability, the branch name is really\nimportant... certainly more than the email of the developer ;-)\nSo the \"contamination\" is negligible compare to the bentefit in this context.\nIt will also helps the graphical tools to have a comprehensive\nrepresentation which can do git even better.\n\nIf you think it is a bad idea to have it by default, what about an\noption to activate this functionality ? Today with the patch I've\ndone, it is not a problem if there is no branchname in the commit. The\nonly thing is the \"%Xb\" placeholder.\n\nI would like to have your advice about the name because I have added\nthe \"branch\" metadata but, even it is really the name of the branch, I\nthink it too \"hard\". I preferred \"BranchOfCommit\" or something similar\nthat is more explicit... I think this one is too heavy. Do you have\nother suggestions ?\n\nThanks for your feedback\n.\n\n\nLe lun. 30 déc. 2019 à 00:20, Junio C Hamano <gitster@pobox.com> a écrit :\n>\n> Arnaud Bertrand <xda@abalgo.com> writes:\n>\n> > I understood that in git philosophy, once it is merged, a branch can\n> > disappear. But for a lot of companies, a SCM is also a guardian of the\n> > history.\n>\n> A lot more important point than \"once it is merged\" is that the\n> branch identity is strictly local to your repository.  Contaminating\n> the object header, which is cast in stone and cannot be modified\n> after the fact, with such a piece of information will not mix well\n> with the rest of Git, so ...\n>\n>\n"},{"id":"389055","messageId":"20191230041517.GA84036@mit.edu","threadId":"52509","inReplyTo":"CAEW0o+g7vXj841h+4nNK8iSoO758Uh9fLKMCN87RE2w2Nd=CRg@mail.gmail.com","subject":"Re: Feature request: add a metadata in the commit: the \"commited in branch\" information","fromName":"Theodore Y. Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2019-12-30T04:15:17Z","receivedAt":"2019-12-30T04:15:49Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Dec 30, 2019 at 12:53:56AM +0100, Arnaud Bertrand wrote:\n> Hi Junio,\n> \n> It really depends how git is used. With big collaborative project\n> (like git or linux kernel), you are totally right.\n> for development limited to a company that has developments with team\n> of 10-20 developers and that uses\n> a correct SCM plan, the name of the branch is regulated and is\n> meaningful, mostly  linked to a bug tracking system\n> system. For audits and  traceability, the branch name is really\n> important... certainly more than the email of the developer ;-)\n> So the \"contamination\" is negligible compare to the bentefit in this context.\n> It will also helps the graphical tools to have a comprehensive\n> represeintation which can do git even better.\n\nWhy does it need to be the branch name?  You can add your own extra\nmetadata to the git description.  So for example, I might have a git\ncommit that looks like this:\n\n    ext4: avoid declaring fs inconsistent due to invalid file handles\n\n    If we receive a file handle, either from NFS or open_by_handle_at(2),\n    and it points at an inode which has not been initialized, and the file\n    system has metadata checksums enabled, we shouldn't try to get the\n    inode, discover the checksum is invalid, and then declare the file\n    system as being inconsistent.\n\n    ... <details of repro omitted to keep this email short>\n\n    Google-Bug-Id: 120690101\n    Upstream-5.0-SHA1: 8a363970d1dc38c4ec4ad575c862f776f468d057\n    Tested: used the repro to verify that open_by_handle_at(2)\n       will not declare the fs inconsistent\n    Effort: storage/ext4       \n    Signed-off-by: Theodore Ts'o <tytso@mit.edu>\n    Change-Id: Iafb6da7c360a4c34b882f7fd6a91e3bb\n\nThe tie-in to the bug tracking system is done via \"Google-Bug-Id:\".\nThe Effort tag is used to identify which subteam should be responsble\nfor rebasing the commit to a newer upstream kernel.  (E.g., how to\naccount for all of the patches made on top of 4.14.x when you are\nrebasing to the newer 4.19 long-term-stable kernel, to make sure all\nnot-yet-usptreamed commits are carried over during the rebase\nprocess.)\n\nThe Upstream-X.Y-SHA1: tag indicates that this is an upstream commit\nthat was backported to the internal kernel.  If the commit isn't an\nupstream backport, we have a policy (which is enforced via an\nautomated bot when the commit goes through Gerritt for review) that\nthere must be an \"Upstream-Plan: \" tag indicating how the committer\nplans to get the change upstream.\n\nThe automated review bot also enforces that a Tested: tag exists,\ndescribing how the developer tested the change, and Change-Id: is used\nto link the commit to Gerrit, which is how we enforce that all commits\nhave to be reviewed by a second engineer before it is allowed into the\nproduction kernel sources.  We also maintain all of the Gerrit\ncomments as history and so we can have accountability as to who\nreviewed a commit before it was submitted into the release repository.\n\nWe also have automated bots which will run checkpatch and note the\nwarnings from checkpatch as Gerrit commits; and if the kernel doesn't\nbuild on a variety of architetures and configurations (e.g., debug,\ninstaller, etc.) a bot can also report this and add -1 Gerrit review.\n\nSee?  You can do an awful lot without regulating and recording the\nbranch name used by the engineer.  We have full audit and traceability\nthrough the Gerrit reviews, and we can use the Google-Bug-Id to track\nwhich release versions of which kernels have which bugs fixed.\n\nThe bottom line is each company is going to have a different workflow\nfor doing reviews, linkage to bug tracking systems, auditability, etc.\nIf everybody were to demand their unique scheme was to be supported\ndirectly in Git, it would be a mess.  The scheme that I've described\nabove needs no special git features.  It just uses some git hooks as a\nconvenience to to developers to help them fill in these required\nfields, using Gerrit for commit review, and some bots which submit\nreviews to Gerrit.\n\nCheers,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"389057","messageId":"CAEW0o+gtya5tm6Wb474Srmb2j4E9ocm9p75=aZWjTASbApsb1A@mail.gmail.com","threadId":"52509","inReplyTo":"20191230041517.GA84036@mit.edu","subject":"Re: Feature request: add a metadata in the commit: the \"commited in branch\" information","fromName":"Arnaud Bertrand","fromEmail":"xda@abalgo.com","sentAt":"2019-12-30T11:59:52Z","receivedAt":"2019-12-30T12:23:18Z","isPatch":false,"sender":{"key":"xda@abalgo.com","avatar":null},"body":"Le lun. 30 déc. 2019 à 05:15, Theodore Y. Ts'o <tytso@mit.edu> a écrit :\n>\n> On Mon, Dec 30, 2019 at 12:53:56AM +0100, Arnaud Bertrand wrote:\n> > Hi Junio,\n> >\n> > It really depends how git is used. With big collaborative project\n> > (like git or linux kernel), you are totally right.\n> > for development limited to a company that has developments with team\n> > of 10-20 developers and that uses\n> > a correct SCM plan, the name of the branch is regulated and is\n> > meaningful, mostly  linked to a bug tracking system\n> > system. For audits and  traceability, the branch name is really\n> > important... certainly more than the email of the developer ;-)\n> > So the \"contamination\" is negligible compare to the bentefit in this context.\n> > It will also helps the graphical tools to have a comprehensive\n> > represeintation which can do git even better.\n>\n> Why does it need to be the branch name?  You can add your own extra\n> metadata to the git description.\n\nThat's typically my problem.  It is not possible \"by default\", I mean\n- It is only possible if the developer configure something\n- or if there is an upper layer that guarantee this\nBy default, there is no hook embedded with the clone. So, as far as I\nknow (and I hope I'm wrong!), you have to use upper layer tools or to\nchange environment variables to activate this feature. Furthermore, it\nwill never be used by third party tool to beautify the branch\nrepresentation.\n\nI think that the major problem is that git calls \"branches\" what it is\nnot. Git branches are, in reality, \"dynamic tags\". In other words,\nwhen you are on this tag and you perform a commit, this dynamic tag\nmoves with your commit. So it is not really a branch as clearcase,\nmercurial, svn or cvs defined it.\n\n\n> So for example, I might have a git\n> commit that looks like this:\n>\n>     ext4: avoid declaring fs inconsistent due to invalid file handles\n>\n>     If we receive a file handle, either from NFS or open_by_handle_at(2),\n>     and it points at an inode which has not been initialized, and the file\n>     system has metadata checksums enabled, we shouldn't try to get the\n>     inode, discover the checksum is invalid, and then declare the file\n>     system as being inconsistent.\n>\n>     ... <details of repro omitted to keep this email short>\n>\n>     Google-Bug-Id: 120690101\n>     Upstream-5.0-SHA1: 8a363970d1dc38c4ec4ad575c862f776f468d057\n>     Tested: used the repro to verify that open_by_handle_at(2)\n>        will not declare the fs inconsistent\n>     Effort: storage/ext4\n>     Signed-off-by: Theodore Ts'o <tytso@mit.edu>\n>     Change-Id: Iafb6da7c360a4c34b882f7fd6a91e3bb\n>\n> The tie-in to the bug tracking system is done via \"Google-Bug-Id:\".\n> The Effort tag is used to identify which subteam should be responsble\n> for rebasing the commit to a newer upstream kernel.  (E.g., how to\n> account for all of the patches made on top of 4.14.x when you are\n> rebasing to the newer 4.19 long-term-stable kernel, to make sure all\n> not-yet-usptreamed commits are carried over during the rebase\n> process.)\n>\n> The Upstream-X.Y-SHA1: tag indicates that this is an upstream commit\n> that was backported to the internal kernel.  If the commit isn't an\n> upstream backport, we have a policy (which is enforced via an\n> automated bot when the commit goes through Gerritt for review) that\n> there must be an \"Upstream-Plan: \" tag indicating how the committer\n> plans to get the change upstream.\n>\n> The automated review bot also enforces that a Tested: tag exists,\n> describing how the developer tested the change, and Change-Id: is used\n> to link the commit to Gerrit, which is how we enforce that all commits\n> have to be reviewed by a second engineer before it is allowed into the\n> production kernel sources.  We also maintain all of the Gerrit\n> comments as history and so we can have accountability as to who\n> reviewed a commit before it was submitted into the release repository.\n>\n> We also have automated bots which will run checkpatch and note the\n> warnings from checkpatch as Gerrit commits; and if the kernel doesn't\n> build on a variety of architetures and configurations (e.g., debug,\n> installer, etc.) a bot can also report this and add -1 Gerrit review.\n>\n> See?  You can do an awful lot without regulating and recording the\n> branch name used by the engineer.  We have full audit and traceability\n> through the Gerrit reviews, and we can use the Google-Bug-Id to track\n> which release versions of which kernels have which bugs fixed.\n>\n\nYou have convinced me that Gerrit is a very nice tool that complete git  ;-)\nHowever, one of the main point of git is that it is easy to setup\n(once the tool is installed, it is one second to setup a new\nrepository and track files)\nSo, I certainly don't want to reduce this strong point of git!\n\n> The bottom line is each company is going to have a different workflow\n> for doing reviews, linkage to bug tracking systems, auditability, etc.\n> If everybody were to demand their unique scheme was to be supported\n> directly in Git, it would be a mess.\n\nInclude the name of the branch is not harmless or fanciful, it is\nsomething important in most of the workflows.\nFor example:\nIf you check this article:\nhttps://nvie.com/posts/a-successful-git-branching-model/\nThe branchname is fundamental and the pictures he shows in its article\nwill never be achieved without keeping track of the branchname.\nGit lightened the notion of branch and therefore its use, which is a\ngood thing but on the other hand, why forget the branch history? And\ncertainly if it is only a light metadata in the commit ?\nDon't you agree that the branchname where modifications were done\ncould give really precious information ?\nDon't you agree that a representation like it is in the article above\nis more clear than the standard git representation ?\n\nI think that add the branchname as \"weak\" metadata, invisible except\nwith a dedicated request like specific placeholder in git log could be\na real big added value compared to the inconvenient of its absence.\n\nCheers,\n\nArnaud\n\n> The scheme that I've described\n> above needs no special git features.  It just uses some git hooks as a\n> convenience to to developers to help them fill in these required\n> fields, using Gerrit for commit review, and some bots which submit\n> reviews to Gerrit.\n>\n> Cheers,\n>\n>                                                 - Ted\n"},{"id":"389067","messageId":"85e17e0628e05279d1dbbe62b877caa90a43ec38.camel@mad-scientist.net","threadId":"52509","inReplyTo":"CAEW0o+gtya5tm6Wb474Srmb2j4E9ocm9p75=aZWjTASbApsb1A@mail.gmail.com","subject":"Re: Feature request: add a metadata in the commit: the \"commited in branch\" information","fromName":"Paul Smith","fromEmail":"paul@mad-scientist.net","sentAt":"2019-12-30T15:15:56Z","receivedAt":"2019-12-30T15:40:28Z","isPatch":false,"sender":{"key":"paul@mad-scientist.net","avatar":"https://avatars.githubusercontent.com/u/109636?v=4"},"body":"On Mon, 2019-12-30 at 12:59 +0100, Arnaud Bertrand wrote:\n> > Why does it need to be the branch name?  You can add your own extra\n> > metadata to the git description.\n> \n> That's typically my problem.  It is not possible \"by default\", I mean\n> - It is only possible if the developer configure something\n> - or if there is an upper layer that guarantee this\n> By default, there is no hook embedded with the clone. So, as far as I\n> know (and I hope I'm wrong!), you have to use upper layer tools or to\n> change environment variables to activate this feature.\n\nIn general I have found that trying to mandate what users do in their own\nrepositories on their own systems is a losing proposition.\n\nInstead, we put requirements on what content is pushed to the central\nrepository.  Because the central repository is managed by the SCM admin\nteam we always know only properly-constructed commits can appear there,\nwithout having to assume that every individual developer's local\nenvironment has been set up in a specific way.\n\nThis can be done with hooks in the central repository: there are Git hooks\nthat are run before any push is accepted, which can cause the push to be\nrejected, and hooks that are run after a push is accepted, which can be\nused for triggering other operations.\n\nSo if you have a requirement about contents of Git commit message format,\nfor example, you can enforce that via these hooks.  If someone attempts to\npush commits to the central repository and the commit message has an\nincorrect format then the push is rejected and they'll have to fix it\nbefore they can proceed to push.\n\nIn the environments I've been associated with we don't care about branch\nnames; instead everything is based on bug tracker identifiers.  Every\ncommit needs to be associated with a valid bug ID (added to the commit\nmessage) and the pre-push hook verifies this and rejects the commit if not.\nThen after the push is accepted, post-push hooks will update the bug\ntracker with information about the push (SHA, software version, etc.)  This\nensures that development and management can use the bug tracker as their\nprimary planning tool to know what has been accomplished and what is left\nto accomplish.  Since the commit message is persisted through cherry-picks, \netc. it allows us to know which bugs were fixed in which different patch\nrelease branches as well.\n\n"}]}