{"thread":{"id":"55505","subject":"RFC/Discussion - Submodule UX Improvements","startedAt":"2021-04-16T23:37:06Z","lastAt":"2021-04-22T15:33:06Z","messageCount":18,"participants":["Emily Shaffer","Christian Couder","Philippe Blain","Randall S. Becker","Aaron Schrab","Jacob Keller","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"422135","messageId":"YHofmWcIAidkvJiD@google.com","threadId":"55505","inReplyTo":null,"subject":"RFC/Discussion - Submodule UX Improvements","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2021-04-16T23:36:57Z","receivedAt":"2021-04-16T23:37:06Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Hi folks,\n\nAs hinted by a couple recent patches, I'm planning on some pretty big submodule\nwork over the next 6 months or so - and Ævar pointed out to me in\nhttps://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com that I probably\nshould share some of those plans ahead of time. :) So attached is a lightly\nmodified version of the doc that we've been working on internally at Google,\nfocusing on what we think would be an ideal submodule workflow.\n\nI'm hoping that folks will get a chance to read some or all of it and let us\nknow what sounds cool (or sounds extremely broken). The best spot to start is\nprobably the \"Overview\" section, which describes what the \"main path\" would look\nlike for a user working on a project with submodules. Most of the work that\nwe're planning on doing is under the \"What doesn't already work\" headings.\n\nThanks in advance for any time you spend reading/discussing :)\n\n - Emily\n\nBackground\n==========\n\nIt's worth mentioning that the main goal that's funding this work is to provide\nan alternative for users whose projects use repo\n(https://source.android.com/setup/develop#repo) today. That means that the main\nfocus is to try and reach feature parity to repo for an easier transition for\nthose who want to switch. As a result, some of the direction below is aimed\ntowards learning from what has worked well with repo (but hopefully more\nflexible for users who want to do more, or differently).\n\nThere are also a few things mentioned that are specifically targeted to ease use\nwith Gerrit, which is in wide use here at Google (and therefore also a\nconsideration we need to make to keep getting paid ;) ).\n\nOverview\n=======\n\nWhen the work is completed, users should be able to have a clean, obvious\nworkflow when using best practices:\n\nTo download the code, they should be able to run simply git clone\nhttps://example.com/superproject to download the project and all its submodules;\nif partial clone is configured, they should receive only the objects allowed by\nthe filter in their superproject as well as in each submodule.\n\nTo begin working on a feature, from the superproject they can 'git switch -c\nfeature', and since the new branch is being created, a new branch 'feature' will\nbe created for each submodule, pointing to the submodule's current 'HEAD'. They\ncan move to a submodule directory and begin to make changes, and when they\ncommit these changes normally with 'git commit' from the submodule directory,\nrunning git status in the superproject will reflect that a submodule has\nchanged. Next, they can switch to a second submodule, making and committing more\nchanges.\n\nWhen they are ready to send these changes which are ready for review but need to\nbe linked together, they can switch back to the superproject, where 'git status'\nindicates that there are changes in both submodules. They can commit these\nchanges to the superproject and use 'git push' to send a review; Git will\nrecurse into affected submodules and push those submodule commits appropriately\nas well.\n\nWhile the user is waiting for feedback on their review, to work on their next\ntask, they can 'git switch other-feature', which will checkout the branches\nspecified in the superproject commit at the tip of 'other-feature'; now the user\ncan continue working as before.\n\nWhen it's time to update their local repo, the user can do so as with a\nsingle-repo project. First they can 'git checkout main && git pull' (or 'git\npull -r'); Git will first checkout the branches associated with main in each\nsubmodule, then fetch and merge/rebase in each submodule appropriately. Finally,\nthey can 'git switch feature && git rebase', at which time Git will recursively\ncheckout the branches associated with 'feature' in each submodule and rebase\neach submodule appropriately.\n\nDetailed Design\n===============\n\nThe Well-Tread Path: Basic Contribution Workflow\n------------------------------------------------\n\n- git clone\n\n1. git clone initializes the directory indicated by the user\n2. git clone fetches the superproject\n3. git clone checks out the superproject at server's HEAD (or at another commit\n   as specified by the user, e.g. with --branch)\n4. git clone warns the user that a recommended hook/config setup exists and\n   provides a tip on how to install it\n5. For each submodule encountered in step 3, git clone is invoked for the\n   submodule, and steps 1-4 are repeated (but in directories indicated by the\n   superproject commit, not by the user).\n\nNote that this means options like '--branch' *don't* propagate directly to the\nsubmodules. If superproject branch \"foo\" points its submodule to branch \"main\",\nthen 'git clone --branch foo https://superproject.git' will clone\nsuperproject/submodule to branch 'main' instead. (It *may* be OK to take\n'--branch' to mean \"the branch specified by the parent *and* the branch named in\n--branch if applicable, but no other branches\".)\n\nWhat doesn't already work:\n\n  * --recurse-submodules should turn on submodule.recurse=true\n  * superproject gets top-level config inherited by submodules\n  * New --recurse-submodules --single-branch semantics\n  * Progress bar for clone (see work estimates)\n  * Recommended config from project owner\n\n\n-- Partial clone\n\n1. git clone initializes the directory indicated by the user\n2. git clone applies the appropriate configs for the partial clone filter\n   requested by the user\n  a) These configs go to the config file shared by superproject and submodules.\n3. git clone fetches the superproject\n4. git clone checks out the superproject at server's HEAD\n5. git clone warns the user that a recommended hook/config setup exists and\n   provides a tip on how to install it\n6. For each submodule encountered in step 4, git clone is invoked for the\n   submodule, and steps 1-4 are repeated (but in directories indicated by the\n   superproject commit, not by the user). The same filter supplied to the\n   superproject applies to the submodules.\n\n\nWhat doesn't already work:\n\n  * --filter=blob:none with submodules (it's using global variables)\n  * propagating --filter=blob:none to submodules (via submodules.config)\n  * Recommended config from project owner\n\n\n- git fetch\n\nBy default, git fetch looks for (1) the remote name(s) supplied at the command\nline, (2) the remote which the currently checked out branch is tracking, or (3)\nthe remote named origin, in that order. For submodules, there is no guarantee\nthat (1) has anything to do with the state of the submodule referenced by the\nsuperproject commit, so just start from (2).\n\nThis operation can be extremely long-running if the project contains many large\nsubmodules, so progress indicators should be displayed.\n\nCaveat: this will mean that we should be more careful about ensuring that\nsubmodule branches have tracking info set up correctly; that may be an issue for\nusers who want to branch within their submodule. This may be OK because users\nwill probably still have 'origin' as their submodule's remote, and if they want\nmore complicated behavior, they will be able to configure it.\n\nWhat doesn't already work:\n\n  * Make sure not to propagate (1) to submodules while recursing\n  * Fetching new submodules.\n  * Not having 0.95 success probability ** 100 = low success probability (that\n    is, we need more retries during submodule fetch)\n  * Progress indicators\n\n\n- git switch / git checkout\n\nSubmodules should continue to perform these operations the same way that they\nhave before, that is, the way that single-repo Git works. But superprojects\nshould behave as follows:\n\n\n-- Create mode (git switch -c / git checkout -b)\n\n1. The current worktree is checked for uncommitted changes to tracked files. The\n   current worktree of each submodule is also checked.\n2. A new branch is created on the superproject; that branch's ref is pointed to\n   the current HEAD.\n3. The new branch is checked out on the superproject.\n4. A new branch with the same name is created on each submodule.\n  a. If there is a naming conflict, we could prompt the user to resolve it, or\n     we could just check out the branch by that name and print a warning to the\n     user with advice on how to solve it (cd submodule && git switch -c\n     different-branch-name HEAD@{1}). Maybe we could skip the warning/advice if\n     the tree is identical to the tree we would have used as the start point\n     (that is, the user switched branches in the submodule, then said \"oh crap\"\n     and went back and switched branches in the superproject).\n  b. Tracking info is set appropriately on each new branch to the upstream of\n     the branch referenced by the parent of the new superproject commit, OR to\n     the default branch's upstream.\n5. The new branch is checked out on each of the submodules.\n\nWhat doesn't already work:\n\n  * Safety check when leaving uncommitted submodule changes\n  * Propagating branch names to submodules currently requires a custom hacky\n    repolike patch\n  * Error handling + graceful non-error handling if the branch already exists\n  * \"Knowing what branch to push to\": copying over which-branch-is-upstream info\n    ** Needs some UX help, push.default is a mess\n  * Tracking info setups\n\n\n-- Switching to an existing branch (git switch / git checkout)\n\n1. The current worktree is checked for uncommitted changes to tracked files. The\n   current worktree of each submodule is also checked.\n2. The requested branch is checked out on the superproject.\n3. The submodule commit or branch referenced by the newly-checked-out\n   superproject commit is checked out on each submodule.\n\nWhat doesn't already work:\n\n  * Same as in create mode\n\n\n- git status\n\n-- From superproject\nThe superproject is clean if:\n\n  * No tracked files in the superproject have been modified and not committed\n  * No tracked files in any submodules have been modified and not committed\n  * No commits in any submodules differ from the commits referenced by the tip\n    commit of the superproject\n\nAdvices should describe:\n\n  * How to commit or drop changes to files in the superproject\n  * How to commit or drop changes to files in the submodules\n  * How to commit changes to submodule references \n  * Which commit/branch to switch the submodule back to if the current work\n    should be dropped: \"Submodule \"foo\" no longer points to \"main\", 'git -C foo\n    switch main' to discard changes\"\n\nWhat doesn't already work:\n\n  * \"git status\" being super fast and actually possible to use.\n    ** (That is, we've seen it move very slowly on projects with many\n       submodules.)\n  * Advice updates to use the appropriate submodule-y commands.\n\n-- From submodule\n\ngit status's behavior for submodules does not change compared to\nsingle-repository Git, except that a red warning line will also display if the\nsuperproject commit does not point to the HEAD of the submodule. (This could\nlook similar to the detached-HEAD warning and tracking branch lines in git\nstatus today, e.g. \"HEAD is ahead of parent project by 2 commits\".)\n\nWhat doesn't already work:\n\n  * \"git status\" from a submodule being aware of the superproject.\n\n\n- git push\n\n-- From superproject\n\nIdeally, a push of the superproject commit results in a push of each submodule\nwhich changed, to the appropriate Gerrit upstream. Commits pushed this way\nacross submodules should somehow be associated in the Gerrit UI, similar to the\n\"submitted together\" display. This will need some work to make happen.\n\nWhat doesn't already work:\n\n  * Automatically setting Gerrit topic (with a hook)\n  * \"push --recurse-submodules\" knowing where to push to in submodules to\n    initiate a Gerrit review\n    ** From `branch` field in .gitmodules?\n    ** Gerrit accepting 'git push -o review origin main' pushes?\n    ** Review URL with a remote helper that rewrites refs/heads/main to\n       refs/for/main?\n    ** Need UX help\n\nFrom submodule\nNo change to client behavior is needed. With Gerrit submodule subscriptions, the\nserver knows how to generate superproject commits when merging submodule\ncommits.\n\n- git pull / git rebase\n\nNote: We're still thinking about this one :)\n\n1. Performs a fetch as described above\n2. For each superproject commit, replay the submodule commits against the newly\n   updated submodule base; then, make a new superproject commit containing those\n   changes\n\nWhat doesn't already work:\n\n  * Rewriting gitlinks in a superproject commit when 'rebase\n    --recurse-submodules'-ing\n  * Resuming after resolving a conflict during rebase\n\n- git merge\n\nThe story for merges is a little bit muddled... and for our goals we don't need\nit for quite a while, so we haven't thought much about it :) Any suggestions\nfolks have about reasonable ways to 'git merge --recurse-submodules' are totally\nwelcome. For now, though, we'll probably just stick in some error message saying\nthat merges with submodules isn't currently supported (maybe we will even add\nthat downstream).\n\nWhat doesn't already work:\n\n  * Erroring out for \"not supported\"\n\n\nAligning Teams\n--------------\n\nThere's two pieces of work that we are relying on a lot, and both have been\nmentioned upstream by now, so I'll just link out:\n\n1. Recommended Hook Configurations\n(https://lore.kernel.org/git/pull.908.v2.git.1616723016659.gitgitgadget@gmail.com)\n\n2. Shared Configuration Across Submodules\n(https://lore.kernel.org/git/20210408233936.533342-1-emilyshaffer@google.com)\n\nEdge Cases, Mess Recovery, & Power Users\n-------------------------------\n\n- Unstaged Changes in Submodules At Commit Time\n\n-- Related Changes (Single Branch)\n\nIf a user has unstaged changes in multiple submodules and runs 'git commit\n--all' from the superproject, they should be presented with an editor which\ncontains commit message drafts for each modified branch, including the\nsuperproject, separated by scissors or some other delineator. After providing a\ncommit message, Git should perform each submodule commit, then finally perform\nthe superproject commit based on the submodules' new commit IDs and apply the\nproposed superproject commit message.\n\n\nWhat doesn't already work:\n\n  * \"git commit --recurse-submodules\" that lets me write a commit message with\n    scissors dividing things in each repository\n\n-- Unrelated Changes (Separating Into Multiple Branches)\n\nIf a user has unstaged changes in multiple submodules and only wants to commit\nsome of them, and runs 'git add --patch' from the superproject, they should be\nwalked through 'git add --patch' for each submodule first. However, since this\ncould be a lengthy process, we need to think carefully about how the UX should\nlook compared to the existing `git add --patch` UX for single-repo projects.\n\nWhat doesn't already work:\n\n  * \"git add --patch\" that recurses through submodule hunks as well\n\n\n- Recovering from Exploratory Changes with 'git restore' and 'git reset'\n\nWhen a user has checked out some historical commit in at least one submodule for\nthe purpose of exploration/investigation, it should be easy to reset the entire\ntree back to the state defined by the superproject commit. Running git restore\n(or git reset) from the superproject should recurse by running git checkout on\neach submodule - and when there are no untracked changes in the submodule, it\ncan do this without asking for user intervention or approval.\n\nWhat doesn't already work:\n\n  * Add some tests for good restore/reset behavior and make them pass\n\n\n- Multiple Commits on a Superproject Branch\n\nGenerally, one superproject commit should represent one feature, where that one\nfeature may consist of multiple submodule commits. It could be thought of\nsimilarly to a merge commit, which brings a stack of related changes into the\nhistory and summarizes them a single commit, without squashing or losing\nhistory. So a user who has two commits in one superproject branch is working on\ntwo features, one of which depends on the other. Reordering those commits should\ninvolve replaying the commits in each submodule associated with each\nsuperproject commit:\n\n\n  superproject  submodule                    superproject  submodule\n\n       A ---------> a1                            B ----------> b1\n       |            |                             |             |\n       |            a2                            |             b2\n       |            |                             |             |\n       |            a3                            A ----------> a1\n       |            |          rebase             |             |\n       B ---------> b1         =====>             |             a2\n       |            |                             |             |\n       |            b2                            |             a3\n       |            |                             |             |\n       o            o                             o             o\n       |            |                             |             |\n       o            o                             o             o\n       |            |                             |             |\n\n\n- Branching in a Submodule\n\nIn addition to the 'git status' warning, users should also receive a warning\nlike detached-HEAD when switching branches in the submodule without a\nsuperproject commit - \"the branch you are leaving behind is not tracked by any\nsuperproject commit\". Users who are just working in and pushing from a single\nsubmodule may find this warning annoying, so it should be clear how to disable\nthat warning per-submodule.\n\n\n- Worktrees\n\nWhen a user runs 'git worktree add' from the superproject, each submodule in the\nnew worktree should also be created as a worktree of the corresponding submodule\nin the original project.\n\nWhat doesn't already work:\n\n  * worktrees and submodules getting along - submodules are now freshly cloned\n    when creating a superproject worktree\n\n- git clone --reference [--dissociate]\n\nWhen cloning with an alternate directory, submodules should also try to use\nobject stores associated with the referenced project instead of cloning from\ntheir remotes right away. It is unclear how much of this works today.\n\n\nWhat doesn't already work:\n\n  * Writing some tests and making them pass\n"},{"id":"422198","messageId":"CAP8UFD0Ct8NofMdds=w0k1-jjX638L6QJQEJWVxqJ6ZPSoJUjg@mail.gmail.com","threadId":"55505","inReplyTo":"YHofmWcIAidkvJiD@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2021-04-18T05:22:07Z","receivedAt":"2021-04-18T05:22:38Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi Emily,\n\nOn Sat, Apr 17, 2021 at 1:39 AM Emily Shaffer <emilyshaffer@google.com> wrote:\n>\n> Hi folks,\n>\n> As hinted by a couple recent patches, I'm planning on some pretty big submodule\n> work over the next 6 months or so - and Ævar pointed out to me in\n> https://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com that I probably\n> should share some of those plans ahead of time. :) So attached is a lightly\n> modified version of the doc that we've been working on internally at Google,\n> focusing on what we think would be an ideal submodule workflow.\n\nThanks for sharing this doc! My main concern with this is that we are\nlikely to have a GSoC student working soon on finishing to port `git\nsubmodule` to C code. And I wonder how that would interact with your\nwork.\n"},{"id":"422230","messageId":"0fc5c0f7-52f7-fb36-f654-ff5223a8809b@gmail.com","threadId":"55505","inReplyTo":"YHofmWcIAidkvJiD@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Philippe Blain","fromEmail":"levraiphilippeblain@gmail.com","sentAt":"2021-04-19T03:20:06Z","receivedAt":"2021-04-19T03:20:12Z","isPatch":false,"sender":{"key":"levraiphilippeblain@gmail.com","avatar":"https://avatars.githubusercontent.com/u/44212482?v=4"},"body":"Hi Emily,\n\nLe 2021-04-16 à 19:36, Emily Shaffer a écrit :\n> Hi folks,\n> \n> As hinted by a couple recent patches, I'm planning on some pretty big submodule\n> work over the next 6 months or so - and Ævar pointed out to me in\n> https://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com that I probably\n> should share some of those plans ahead of time. :) So attached is a lightly\n> modified version of the doc that we've been working on internally at Google,\n> focusing on what we think would be an ideal submodule workflow.\n> \n> I'm hoping that folks will get a chance to read some or all of it and let us\n> know what sounds cool (or sounds extremely broken). The best spot to start is\n> probably the \"Overview\" section, which describes what the \"main path\" would look\n> like for a user working on a project with submodules. Most of the work that\n> we're planning on doing is under the \"What doesn't already work\" headings.\n> \n> Thanks in advance for any time you spend reading/discussing :)\n\nThanks a lot for sharing this roadmap. There a lot of good ideas there, that would really\nimprove the situation for projects using subdmodules. I've added some toughts on specific\nitems below.\n\n> \n>   - Emily\n> \n> Background\n> ==========\n> \n> It's worth mentioning that the main goal that's funding this work is to provide\n> an alternative for users whose projects use repo\n> (https://source.android.com/setup/develop#repo) today. That means that the main\n> focus is to try and reach feature parity to repo for an easier transition for\n> those who want to switch. As a result, some of the direction below is aimed\n> towards learning from what has worked well with repo (but hopefully more\n> flexible for users who want to do more, or differently).\n> \n> There are also a few things mentioned that are specifically targeted to ease use\n> with Gerrit, which is in wide use here at Google (and therefore also a\n> consideration we need to make to keep getting paid ;) ).\n> \n> Overview\n> =======\n> \n> When the work is completed, users should be able to have a clean, obvious\n> workflow when using best practices:\n> \n> To download the code, they should be able to run simply git clone\n> https://example.com/superproject to download the project and all its submodules;\n\nPlaying the devil's advocate here, but some projects do not want / need all of\ntheir submodules in a \"regular\" checkout, so I guess that would have to be somehow\nconfigurable. I've always felt that since each project is different in that regard,\nit would be better if each project could declare if their submodules are non-optional\nand need to be also cloned when the superproject is cloned. Maybe an additional field\nin '.gitmodules', like a boolean 'submodule.<name>.optional', could be added,\nso that submodules that are optional are not cloned, but others are. If that setting\nis opt-in (meaning that it defaults to 'true', i.e., submodules are considered optional by default),\nthen it would be easier to argue for 'git clone' to mean 'git clone --recurse-submodules':\n'git clone' would clone the superproject and any non-optional submodule.\nThen eventually, when the usage of 'submodule.<name>.optional' becomes more widespread,\nwe can switch the default and then projects would need to explicitely declare their submodule\noptional if they don't want them cloned by a simple 'git clone'.\n\n\n> if partial clone is configured, they should receive only the objects allowed by\n> the filter in their superproject as well as in each submodule.\n> \n> To begin working on a feature, from the superproject they can 'git switch -c\n> feature', and since the new branch is being created, a new branch 'feature' will\n> be created for each submodule, pointing to the submodule's current 'HEAD'. They\n> can move to a submodule directory and begin to make changes, and when they\n> commit these changes normally with 'git commit' from the submodule directory,\n> running git status in the superproject will reflect that a submodule has\n> changed. Next, they can switch to a second submodule, making and committing more\n> changes.\n\nYes. Apart from recursive 'git checkout -b / git switch -c', the workflow you describe\nalready works well, from my experience.\n\n> \n> When they are ready to send these changes which are ready for review but need to\n> be linked together, they can switch back to the superproject, where 'git status'\n> indicates that there are changes in both submodules. They can commit these\n> changes to the superproject and use 'git push' to send a review; Git will\n> recurse into affected submodules and push those submodule commits appropriately\n> as well.\n> \n> While the user is waiting for feedback on their review, to work on their next\n> task, they can 'git switch other-feature', which will checkout the branches\n> specified in the superproject commit at the tip of 'other-feature'; now the user\n> can continue working as before.\n\nHere, I'm not sure what you mean by \"the branches (plural) specified in the superproject\ncommit at the tip of other-feature\". Today, with 'submodule.recurse = true', 'git checkout some-feature'\nalready checks out each submodule in detached HEAD at the commit recorded in the superproject commit\nat the tip of some-feature. It's unclear if you are proposing to instead record submodule branch\nnames in the superproject commit.. is that what's going on here ? (or is it just a typo ?)\n\n> \n> When it's time to update their local repo, the user can do so as with a\n> single-repo project. First they can 'git checkout main && git pull' (or 'git\n> pull -r'); Git will first checkout the branches associated with main in each\n> submodule, then fetch and merge/rebase in each submodule appropriately. \n\nWhat if some submodule does not use the same branch name for their primary integration branch?\nSometimes as a superproject using another project as a submodule, you do not\ncontrol that...\n\n> Finally,\n> they can 'git switch feature && git rebase', at which time Git will recursively\n> checkout the branches associated with 'feature' in each submodule and rebase\n> each submodule appropriately.\n> \n> Detailed Design\n> ===============\n> \n> The Well-Tread Path: Basic Contribution Workflow\n> ------------------------------------------------\n> \n> - git clone\n> \n> 1. git clone initializes the directory indicated by the user\n> 2. git clone fetches the superproject\n> 3. git clone checks out the superproject at server's HEAD (or at another commit\n>     as specified by the user, e.g. with --branch)\n> 4. git clone warns the user that a recommended hook/config setup exists and\n>     provides a tip on how to install it\n> 5. For each submodule encountered in step 3, git clone is invoked for the\n>     submodule, and steps 1-4 are repeated (but in directories indicated by the\n>     superproject commit, not by the user).\n> \n> Note that this means options like '--branch' *don't* propagate directly to the\n> submodules. If superproject branch \"foo\" points its submodule to branch \"main\",\n\nHere again, I'm not sure what you mean, because right now there is no concept of\nthe superproject having a submodule \"pointing to some branch\", only to a specific\ncommit. 'submodule.<name>.branch' is only ever used by the command 'git submodule update --remote'.\nIs there an implicit proposal to change that ?\n\n> then 'git clone --branch foo https://superproject.git' will clone\n> superproject/submodule to branch 'main' instead. (It *may* be OK to take\n> '--branch' to mean \"the branch specified by the parent *and* the branch named in\n> --branch if applicable, but no other branches\".)\n> \n> What doesn't already work:\n> \n>    * --recurse-submodules should turn on submodule.recurse=true\n\nThat's actually a good very idea, but maybe it should be explicitely mentioned, I think\n(in the output of the command I mean).\n\n>    * superproject gets top-level config inherited by submodules\n>    * New --recurse-submodules --single-branch semantics\n>    * Progress bar for clone (see work estimates)\n>    * Recommended config from project owner\n> \n> \n> -- Partial clone\n> \n> 1. git clone initializes the directory indicated by the user\n> 2. git clone applies the appropriate configs for the partial clone filter\n>     requested by the user\n>    a) These configs go to the config file shared by superproject and submodules.\n> 3. git clone fetches the superproject\n> 4. git clone checks out the superproject at server's HEAD\n> 5. git clone warns the user that a recommended hook/config setup exists and\n>     provides a tip on how to install it\n> 6. For each submodule encountered in step 4, git clone is invoked for the\n>     submodule, and steps 1-4 are repeated (but in directories indicated by the\n>     superproject commit, not by the user). The same filter supplied to the\n>     superproject applies to the submodules.\n> \n> \n> What doesn't already work:\n> \n>    * --filter=blob:none with submodules (it's using global variables)\n>    * propagating --filter=blob:none to submodules (via submodules.config)\n>    * Recommended config from project owner\n> \n> \n> - git fetch\n> \n> By default, git fetch looks for (1) the remote name(s) supplied at the command\n> line, (2) the remote which the currently checked out branch is tracking, or (3)\n> the remote named origin, in that order. For submodules, there is no guarantee\n> that (1) has anything to do with the state of the submodule referenced by the\n> superproject commit, so just start from (2).\n> \n> This operation can be extremely long-running if the project contains many large\n> submodules, so progress indicators should be displayed.\n> \n> Caveat: this will mean that we should be more careful about ensuring that\n> submodule branches have tracking info set up correctly; that may be an issue for\n> users who want to branch within their submodule. This may be OK because users\n> will probably still have 'origin' as their submodule's remote, and if they want\n> more complicated behavior, they will be able to configure it.\n> \n> What doesn't already work:\n> \n>    * Make sure not to propagate (1) to submodules while recursing\n>    * Fetching new submodules.\n>    * Not having 0.95 success probability ** 100 = low success probability (that\n>      is, we need more retries during submodule fetch)\n>    * Progress indicators\n\nI would add the following:\n\n- Fix 'git fetch upstream' when 'submodule.recurse' and 'fetch.recurseSubdmodules=on-demand'\nare both set  (the submodule is not fetched even if the superproject changed the submodule\ncommit).\n\n- Do not rely on 'origin' exising in the submodule (or being pushable to). Right now,\nrenaming the 'origin' remote to 'upstream' in a submodule, and using 'origin' for one's own\nfork of a submodule, (as is often done in the superproject), breaks 'git fetch --recurse-submodules'\n(or 'git fetch' if 'submodule.recurse' is set), in the sense that the fetch does not recurse\nto the submodule, as it should. I do not have a simple reproducer handy but\nI've seen it happen and there are a couple hard-coded \"origin\" in the submodule code [1], [2].\n\n> \n> \n> - git switch / git checkout\n> \n> Submodules should continue to perform these operations the same way that they\n> have before, that is, the way that single-repo Git works. But superprojects\n> should behave as follows:\n> \n> \n> -- Create mode (git switch -c / git checkout -b)\n> \n> 1. The current worktree is checked for uncommitted changes to tracked files. The\n>     current worktree of each submodule is also checked.\n> 2. A new branch is created on the superproject; that branch's ref is pointed to\n>     the current HEAD.\n> 3. The new branch is checked out on the superproject.\n> 4. A new branch with the same name is created on each submodule.\n\nThat might not be wanted by all, so I think it should be configurable.\n\n>    a. If there is a naming conflict, we could prompt the user to resolve it, or\n>       we could just check out the branch by that name and print a warning to the\n>       user with advice on how to solve it (cd submodule && git switch -c\n>       different-branch-name HEAD@{1}). Maybe we could skip the warning/advice if\n>       the tree is identical to the tree we would have used as the start point\n>       (that is, the user switched branches in the submodule, then said \"oh crap\"\n>       and went back and switched branches in the superproject).\n>    b. Tracking info is set appropriately on each new branch to the upstream of\n>       the branch referenced by the parent of the new superproject commit, OR to\n>       the default branch's upstream.\n\nThis last point is a little unclear: which \"new superproject commit\" ? (we are creating\na branch, so there is no new commit yet?). And again, you talk about a (submodule?) branch being referenced\nby a superproject commit, which is not a concept that actually exists today.\nAlso, usually tracking info is only set\nautomatically when using the form 'git checkout -b new-branch upstream/master' or\nthe like. Do you also propose that 'git checkout -b new-branch', by itself, should\nautomatically set tracking info ?\n\n\n> 5. The new branch is checked out on each of the submodules.\n> \n> What doesn't already work:\n> \n>    * Safety check when leaving uncommitted submodule changes\n\nYes, that has been reported several times ([3], [4], [5]). I have fixes for this,\nnot quite ready to send because I'm trying to write extensive tests (maybe too extensive)...\n\n>    * Propagating branch names to submodules currently requires a custom hacky\n>      repolike patch\n>    * Error handling + graceful non-error handling if the branch already exists\n>    * \"Knowing what branch to push to\": copying over which-branch-is-upstream info\n>      ** Needs some UX help, push.default is a mess\n>    * Tracking info setups\n> \n> -- Switching to an existing branch (git switch / git checkout)\n> \n> 1. The current worktree is checked for uncommitted changes to tracked files. The\n>     current worktree of each submodule is also checked.\n> 2. The requested branch is checked out on the superproject.\n> 3. The submodule commit or branch referenced by the newly-checked-out\n>     superproject commit is checked out on each submodule.\n> \n> What doesn't already work:\n> \n>    * Same as in create mode\n\nHere, I would add that 'git checkout --recurse-submodules', along with 'git clone --recurse-submodules',\nhave trouble with correctly checkout-ing an older commit that records a submodule that\nwas since removed from the project. The user experience around this use case is currently very very bad [6].\nThis is partly due to 'git clone --recurse-submodules' only cloning submodules that are recorded in\nthe tip commit of the default branch of the superproject, which could certainly be improved.\n\n> \n> \n> - git status\n> \n> -- From superproject\n> The superproject is clean if:\n> \n>    * No tracked files in the superproject have been modified and not committed\n>    * No tracked files in any submodules have been modified and not committed\n>    * No commits in any submodules differ from the commits referenced by the tip\n>      commit of the superproject\n> \n> Advices should describe:\n> \n>    * How to commit or drop changes to files in the superproject\n>    * How to commit or drop changes to files in the submodules\n>    * How to commit changes to submodule references\n>    * Which commit/branch to switch the submodule back to if the current work\n>      should be dropped: \"Submodule \"foo\" no longer points to \"main\", 'git -C foo\n>      switch main' to discard changes\"\n> \n> What doesn't already work:\n> \n>    * \"git status\" being super fast and actually possible to use.\n>      ** (That is, we've seen it move very slowly on projects with many\n>         submodules.)\n>    * Advice updates to use the appropriate submodule-y commands.\n\nI would add that 'git status' should show the submodule as \"rewind\" if the\ncurrently checked out submodule commit is *behind* what's recorded in the current superproject\ncommit. That is shown by 'git diff --submodule=<log | diff>' and 'git submodule summary'\nand is quite useful to prevent a following 'git commit -am' in the superproject to regress the submodule commit\nby mistake. It would be nice if 'git status' could also show this information (code in\nsubmodule.c::show_submodule_header).\n\n> \n> -- From submodule\n> \n> git status's behavior for submodules does not change compared to\n> single-repository Git, except that a red warning line will also display if the\n> superproject commit does not point to the HEAD of the submodule. (This could\n> look similar to the detached-HEAD warning and tracking branch lines in git\n> status today, e.g. \"HEAD is ahead of parent project by 2 commits\".)\n\nThat would be a nice addition :)\n\n> \n> What doesn't already work:\n> \n>    * \"git status\" from a submodule being aware of the superproject.\n> \n> \n> - git push\n> \n> -- From superproject\n> \n> Ideally, a push of the superproject commit results in a push of each submodule\n> which changed, to the appropriate Gerrit upstream. Commits pushed this way\n> across submodules should somehow be associated in the Gerrit UI, similar to the\n> \"submitted together\" display. This will need some work to make happen.\n> \n> What doesn't already work:\n> \n>    * Automatically setting Gerrit topic (with a hook)\n>    * \"push --recurse-submodules\" knowing where to push to in submodules to\n>      initiate a Gerrit review\n>      ** From `branch` field in .gitmodules?\n>      ** Gerrit accepting 'git push -o review origin main' pushes?\n>      ** Review URL with a remote helper that rewrites refs/heads/main to\n>         refs/for/main?\n>      ** Need UX help\n\nIt would be nice if 'git push' would not force users to use the same\nremote names and branch names in the superproject and the submodule.\nPrevious discussion around this that I had spotted are at [7] and [8].\n\n> \n>>From submodule\n> No change to client behavior is needed. With Gerrit submodule subscriptions, the\n> server knows how to generate superproject commits when merging submodule\n> commits.\n> \n> - git pull / git rebase\n> \n> Note: We're still thinking about this one :)\n> \n> 1. Performs a fetch as described above\n> 2. For each superproject commit, replay the submodule commits against the newly\n>     updated submodule base; then, make a new superproject commit containing those\n>     changes\n> \n> What doesn't already work:\n> \n>    * Rewriting gitlinks in a superproject commit when 'rebase\n>      --recurse-submodules'-ing\n>    * Resuming after resolving a conflict during rebase\n\nIn general, rebase is not well aware of 'submodule.recurse'. Even if you do not\nneed to rewrite superproject commits, there are a couple of use cases that are broken\nright now:\n\n- 'git rebase upstream/master' when upstream updated the submodule, will correctly\n(recursively) checkout upstream/master before starting the rebase, but upon\n'git rebase --abort', the submodule will stay checked out at the commit recorded in\n'upstream/master', which is confusing. This only happens when 'submodule.recurse' is true (!).\n- 'git rebase -i' which stops at a commit 'A' where the submodule commit is changed,\ndoes not correctly check out the submodule tree. It's checked out at the commit recorded in A~1\n(and this also only happens if submodule.recurse is true)\n- In some cases, like 'rebase -i'-ing across the addition of new submodules, at the end\nof the rebase the submodules are empty, and 'git submodule update' must be run to\nre-populate them.\n\n> \n> - git merge\n> \n> The story for merges is a little bit muddled... and for our goals we don't need\n> it for quite a while, so we haven't thought much about it :) Any suggestions\n> folks have about reasonable ways to 'git merge --recurse-submodules' are totally\n> welcome. For now, though, we'll probably just stick in some error message saying\n> that merges with submodules isn't currently supported (maybe we will even add\n> that downstream).\n\nWhat is \"downstream\" here ?\n\nAlso, there is quite a bit of a future plan in the commit message of\na6d7eb2c7a (pull: optionally rebase submodules (remote submodule changes only), 2017-06-23).\nIt would be nice to revisit this, I think (regarding both rebase and merge).\n\n> \n> What doesn't already work:\n> \n>    * Erroring out for \"not supported\"\n> \n> \n> Aligning Teams\n> --------------\n> \n> [ ... ] \n> \n> \n> - Worktrees\n> \n> When a user runs 'git worktree add' from the superproject, each submodule in the\n> new worktree should also be created as a worktree of the corresponding submodule\n> in the original project.\n> \n> What doesn't already work:\n> \n>    * worktrees and submodules getting along - submodules are now freshly cloned\n>      when creating a superproject worktree\n\nThat would certainly be nice. I've been using worktrees with submodule-containing\nprojects and everything has been working fine (there were 2 bugs but I fixed them).\nOnce we are not wasting disk space\nby re-cloning the submodules, we whould remove the 'not recommended' mention in the\ndocs aboout using worktrees with projects containing submodules.\n\n> \n> - git clone --reference [--dissociate]\n> \n> When cloning with an alternate directory, submodules should also try to use\n> object stores associated with the referenced project instead of cloning from\n> their remotes right away. It is unclear how much of this works today.\n> \n> \n> What doesn't already work:\n> \n>    * Writing some tests and making them pass\n> \n\n\nThanks again for providing these details,\n\nPhilippe.\n\n[1] https://github.com/git/git/blob/b0c09ab8796fb736efa432b8e817334f3e5ee75a/builtin/submodule--helper.c#L43-L51\n[2] https://github.com/git/git/blob/b0c09ab8796fb736efa432b8e817334f3e5ee75a/submodule.c#L1525\n[3] https://lore.kernel.org/git/CAHsG2VT4YB_nf8PrEmrHwK-iY-AQo0VDcvXGVsf8cEYXws4nig@mail.gmail.com/\n[4] https://lore.kernel.org/git/20200525094019.22padbzuk7ukr5uv@overdrive.tratt.net/T/#u\n[5] https://lore.kernel.org/git/05afbdeb-6c72-f14c-cdf0-e14894de05a3@gmail.com/T/#t\n[6] https://github.com/gitgitgadget/git/issues/752\n[7]https://lore.kernel.org/git/20170405174719.1297-6-bmwill@google.com/t/#m224c2475b1bad333e1118f68c80465b638ed87ee\n[8] https://public-inbox.org/git/20170627162307.GE161648@aiede.mtv.corp.google.com/\n"},{"id":"422311","messageId":"00dc01d7351b$6ffc6500$4ff52f00$@nexbridge.com","threadId":"55505","inReplyTo":"YHofmWcIAidkvJiD@google.com","subject":"RE: RFC/Discussion - Submodule UX Improvements","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-04-19T12:56:32Z","receivedAt":"2021-04-19T12:56:48Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"> -----Original Message-----\n> From: Emily Shaffer <emilyshaffer@google.com>\nOn April 16, 2021 7:37 PM, Emily Shaffer wrote:\n> As hinted by a couple recent patches, I'm planning on some pretty big\n> submodule work over the next 6 months or so - and Ævar pointed out to me\nin\n> https://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com that I\n> probably should share some of those plans ahead of time. :) So attached is\na\n> lightly modified version of the doc that we've been working on internally\nat\n> Google, focusing on what we think would be an ideal submodule workflow.\n> \n> I'm hoping that folks will get a chance to read some or all of it and let\nus know\n> what sounds cool (or sounds extremely broken). The best spot to start is\n> probably the \"Overview\" section, which describes what the \"main path\"\nwould\n> look like for a user working on a project with submodules. Most of the\nwork\n> that we're planning on doing is under the \"What doesn't already work\"\n> headings.\n> \n> Thanks in advance for any time you spend reading/discussing :)\n<big snip>\n\nJust adding my voice here, this is something my teams would be very happy to\nconsider.\n\n> - Worktrees\n> When a user runs 'git worktree add' from the superproject, each submodule\n>  in the new worktree should also be created as a worktree of the\ncorresponding\n>  submodule in the original project.\n> What doesn't already work:\n>   * worktrees and submodules getting along - submodules are now freshly\ncloned\n>     when creating a superproject worktree\n\nMy teams are currently debating the use of submodules (we have gone back and\nforth over the years on these) and worktrees (which seem to have some\npositive process implications for those more legacy-ish team members more\nused to a centralised workflows). I have not seen any worktree/submodule\ncombinations used but fear the worst - as in I'm pretty sure I know which of\nmy team members is going to try this. It is probably a separate matter to\nmake the two get along better.\n\nCheers,\nRandall\n\n-- Brief whoami:\nNonStop developer since approximately 211288444200000000\nUNIX developer since approximately 421664400\n-- In my real life, I talk too much.\n\n\n\n"},{"id":"422312","messageId":"YH1+C47AErrCUkHI@pug.qqx.org","threadId":"55505","inReplyTo":"YHofmWcIAidkvJiD@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Aaron Schrab","fromEmail":"aaron@schrab.com","sentAt":"2021-04-19T12:56:43Z","receivedAt":"2021-04-19T13:03:10Z","isPatch":false,"sender":{"key":"aaron@schrab.com","avatar":"https://avatars.githubusercontent.com/u/39620?v=4"},"body":"At 16:36 -0700 16 Apr 2021, Emily Shaffer <emilyshaffer@google.com> wrote:\n>- git switch / git checkout\n\n(snip)\n\n>4. A new branch with the same name is created on each submodule.\n>  a. If there is a naming conflict, we could prompt the user to resolve it, or\n>     we could just check out the branch by that name and print a warning to the\n>     user with advice on how to solve it (cd submodule && git switch -c\n>     different-branch-name HEAD@{1}). Maybe we could skip the warning/advice if\n>     the tree is identical to the tree we would have used as the start point\n>     (that is, the user switched branches in the submodule, then said \"oh crap\"\n>     and went back and switched branches in the superproject).\n>  b. Tracking info is set appropriately on each new branch to the upstream of\n>     the branch referenced by the parent of the new superproject commit, OR to\n>     the default branch's upstream.\n>5. The new branch is checked out on each of the submodules.\n\nIn many cases the branch name for the superproject isn't going to be \nappropriate for submodules.\n\nThis seems likely to create a LOT of junk branches. Do you also have a \nproposal for cleaning those up?\n"},{"id":"422322","messageId":"CA+P7+xqzsD+pU=-9YUYdGDAqT4uVk=XS4sdxA5WnAXL_7GwM5Q@mail.gmail.com","threadId":"55505","inReplyTo":"YHofmWcIAidkvJiD@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2021-04-19T19:14:48Z","receivedAt":"2021-04-19T19:15:01Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Fri, Apr 16, 2021 at 4:38 PM Emily Shaffer <emilyshaffer@google.com> wrote:\n>\n> Hi folks,\n>\n> As hinted by a couple recent patches, I'm planning on some pretty big submodule\n> work over the next 6 months or so - and Ævar pointed out to me in\n> https://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com that I probably\n> should share some of those plans ahead of time. :) So attached is a lightly\n> modified version of the doc that we've been working on internally at Google,\n> focusing on what we think would be an ideal submodule workflow.\n>\n> I'm hoping that folks will get a chance to read some or all of it and let us\n> know what sounds cool (or sounds extremely broken). The best spot to start is\n> probably the \"Overview\" section, which describes what the \"main path\" would look\n> like for a user working on a project with submodules. Most of the work that\n> we're planning on doing is under the \"What doesn't already work\" headings.\n>\n> Thanks in advance for any time you spend reading/discussing :)\n>\n>  - Emily\n>\n> Background\n> ==========\n>\n> It's worth mentioning that the main goal that's funding this work is to provide\n> an alternative for users whose projects use repo\n> (https://source.android.com/setup/develop#repo) today. That means that the main\n> focus is to try and reach feature parity to repo for an easier transition for\n> those who want to switch. As a result, some of the direction below is aimed\n> towards learning from what has worked well with repo (but hopefully more\n> flexible for users who want to do more, or differently).\n>\n> There are also a few things mentioned that are specifically targeted to ease use\n> with Gerrit, which is in wide use here at Google (and therefore also a\n> consideration we need to make to keep getting paid ;) ).\n>\n> Overview\n> =======\n>\n\nOne thing that I think I didn't see covered when I scanned this, that\nis something I find difficult or annoying to resolve is using \"blame\"\nwith submodules. I use blame a lot to do code history analysis to\nunderstand how something got to the way it is. (Often this helps\nresolve issues or bugs by using new context to understand why an old\nchange was broken).\n\nIt has bothered me in the past when I try to do \"git blame\n<path/to/submodule>\" and I get nothing. Obviously there are ways\naround this: you can for example just log the path and get the commit\nthat changed it most recently, or try to search for when the submodule\nwas set to a given commit.\n\nA sort of dream I had was a flow where I could do something from the\nparent like \"git blame <path/to/submodule>/submodule/file\" and have it\npresent a blame of that files contents keyed on the *parent* commit\nthat changed the submodule to have that line, as opposed to being\nforced to go into the submodule and figure out what commit introduced\nit and then go back to the parent and find out what commit changed the\nsubmodule to include that submodule commit.\n\n> When the work is completed, users should be able to have a clean, obvious\n> workflow when using best practices:\n>\n> To download the code, they should be able to run simply git clone\n> https://example.com/superproject to download the project and all its submodules;\n> if partial clone is configured, they should receive only the objects allowed by\n> the filter in their superproject as well as in each submodule.\n>\n> To begin working on a feature, from the superproject they can 'git switch -c\n> feature', and since the new branch is being created, a new branch 'feature' will\n> be created for each submodule, pointing to the submodule's current 'HEAD'. They\n> can move to a submodule directory and begin to make changes, and when they\n> commit these changes normally with 'git commit' from the submodule directory,\n> running git status in the superproject will reflect that a submodule has\n> changed. Next, they can switch to a second submodule, making and committing more\n> changes.\n>\n> When they are ready to send these changes which are ready for review but need to\n> be linked together, they can switch back to the superproject, where 'git status'\n> indicates that there are changes in both submodules. They can commit these\n> changes to the superproject and use 'git push' to send a review; Git will\n> recurse into affected submodules and push those submodule commits appropriately\n> as well.\n>\n> While the user is waiting for feedback on their review, to work on their next\n> task, they can 'git switch other-feature', which will checkout the branches\n> specified in the superproject commit at the tip of 'other-feature'; now the user\n> can continue working as before.\n>\n> When it's time to update their local repo, the user can do so as with a\n> single-repo project. First they can 'git checkout main && git pull' (or 'git\n> pull -r'); Git will first checkout the branches associated with main in each\n> submodule, then fetch and merge/rebase in each submodule appropriately. Finally,\n> they can 'git switch feature && git rebase', at which time Git will recursively\n> checkout the branches associated with 'feature' in each submodule and rebase\n> each submodule appropriately.\n>\n"},{"id":"422325","messageId":"013401d73552$287f49e0$797ddda0$@nexbridge.com","threadId":"55505","inReplyTo":"CA+P7+xqzsD+pU=-9YUYdGDAqT4uVk=XS4sdxA5WnAXL_7GwM5Q@mail.gmail.com","subject":"RE: RFC/Discussion - Submodule UX Improvements","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-04-19T19:28:15Z","receivedAt":"2021-04-19T19:28:25Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On April 19, 2021 3:15 PM, Jacob Keller wrote:\n> On Fri, Apr 16, 2021 at 4:38 PM Emily Shaffer <emilyshaffer@google.com>\n> wrote:\n> >\n> > Hi folks,\n> >\n> > As hinted by a couple recent patches, I'm planning on some pretty big\n> > submodule work over the next 6 months or so - and Ævar pointed out to\n> > me in https://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com\n> > that I probably should share some of those plans ahead of time. :) So\n> > attached is a lightly modified version of the doc that we've been\n> > working on internally at Google, focusing on what we think would be an ideal\n> submodule workflow.\n> >\n> > I'm hoping that folks will get a chance to read some or all of it and\n> > let us know what sounds cool (or sounds extremely broken). The best\n> > spot to start is probably the \"Overview\" section, which describes what\n> > the \"main path\" would look like for a user working on a project with\n> > submodules. Most of the work that we're planning on doing is under the\n> \"What doesn't already work\" headings.\n> >\n> > Thanks in advance for any time you spend reading/discussing :)\n> >\n> >  - Emily\n> >\n> > Background\n> > ==========\n> >\n> > It's worth mentioning that the main goal that's funding this work is\n> > to provide an alternative for users whose projects use repo\n> > (https://source.android.com/setup/develop#repo) today. That means that\n> > the main focus is to try and reach feature parity to repo for an\n> > easier transition for those who want to switch. As a result, some of\n> > the direction below is aimed towards learning from what has worked\n> > well with repo (but hopefully more flexible for users who want to do more, or\n> differently).\n> >\n> > There are also a few things mentioned that are specifically targeted\n> > to ease use with Gerrit, which is in wide use here at Google (and\n> > therefore also a consideration we need to make to keep getting paid ;) ).\n> >\n> > Overview\n> > =======\n> >\n> \n> One thing that I think I didn't see covered when I scanned this, that is\n> something I find difficult or annoying to resolve is using \"blame\"\n> with submodules. I use blame a lot to do code history analysis to understand\n> how something got to the way it is. (Often this helps resolve issues or bugs by\n> using new context to understand why an old change was broken).\n> \n> It has bothered me in the past when I try to do \"git blame\n> <path/to/submodule>\" and I get nothing. Obviously there are ways around this:\n> you can for example just log the path and get the commit that changed it most\n> recently, or try to search for when the submodule was set to a given commit.\n> \n> A sort of dream I had was a flow where I could do something from the parent\n> like \"git blame <path/to/submodule>/submodule/file\" and have it present a\n> blame of that files contents keyed on the *parent* commit that changed the\n> submodule to have that line, as opposed to being forced to go into the\n> submodule and figure out what commit introduced it and then go back to the\n> parent and find out what commit changed the submodule to include that\n> submodule commit.\n\nNot going to disagree, but are you looking for the blame on the submodule ref file itself or files in the submodule? It's hard to teach git to do a blame on a one-line file.\n\nOtherwise, and I think this is what you really are going for, teaching it to do a blame based on \"git blame <path/to/submodule>/submodule/file\" would be very nice and abstracts out the need for the user (or more importantly to me = scripts) to understand that a submodule is involved; however, it is opening up a very large door: \"should/could we teach git to abstract submodules out of every command\". This would potentially replace a significant part of the use cases for the \"git submodule foreach\" sub-command. In your ask, the current paradigm \"cd <path/to/submodule>/submodule && git blame file\" or pretty much every other command does work, but it requires the user/script to know you have a submodule in the path. So my question is: is this worth the effort? I don't have a good answer to that question. Half of my brain would like this very much/the other half is scared of the impact to the code.\n\nJust my musings.\n\nRandall\n\n"},{"id":"422447","messageId":"CA+P7+xrOuhG5ujQRYS0=o7S9=xD5zm6BGp5mBRt493Lme9xYcw@mail.gmail.com","threadId":"55505","inReplyTo":"013401d73552$287f49e0$797ddda0$@nexbridge.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2021-04-20T16:18:05Z","receivedAt":"2021-04-20T16:18:19Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Mon, Apr 19, 2021 at 12:28 PM Randall S. Becker\n<rsbecker@nexbridge.com> wrote:\n> On April 19, 2021 3:15 PM, Jacob Keller wrote:\n> > A sort of dream I had was a flow where I could do something from the parent\n> > like \"git blame <path/to/submodule>/submodule/file\" and have it present a\n> > blame of that files contents keyed on the *parent* commit that changed the\n> > submodule to have that line, as opposed to being forced to go into the\n> > submodule and figure out what commit introduced it and then go back to the\n> > parent and find out what commit changed the submodule to include that\n> > submodule commit.\n>\n> Not going to disagree, but are you looking for the blame on the submodule ref file itself or files in the submodule? It's hard to teach git to do a blame on a one-line file.\n>\n\nWell, I would like if \"git blame <path/to/submodule>\" did.. something\nother than just fail. Sometimes my brain is working in a \"blame where\nthis came from\" and I type that out and then get frustrated when it\nfails. Additionally...\n\n> Otherwise, and I think this is what you really are going for, teaching it to do a blame based on \"git blame <path/to/submodule>/submodule/file\" would be very nice and abstracts out the need for the user (or more importantly to me = scripts) to understand that a submodule is involved; however, it is opening up a very large door: \"should/could we teach git to abstract submodules out of every command\". This would potentially replace a significant part of the use cases for the \"git submodule foreach\" sub-command. In your ask, the current paradigm \"cd <path/to/submodule>/submodule && git blame file\" or pretty much every other command does work, but it requires the user/script to know you have a submodule in the path. So my question is: is this worth the effort? I don't have a good answer to that question. Half of my brain would like this very much/the other half is scared of the impact to the code.\n>\n> Just my musings.\n\nI'm not asking for \"git blame <path/to/submodule>/<file>\" to give the\nthe same outout as \"cd <path/to/submodule> && git blame <file>\"\n\nWhat i'm asking is: given this file, tell me which commit in the\nparent did the line get introduced. So basically I want to walk over\nthe changes to the submodule pointer and find out when it get\nintroduced into the parent, not when it got introduced into the\nsubmodule itself.\n\nThis is a related question, but it is actually not trivial to go\ninstantly from \"it was in xyz submodule commit\" to \"it was then pulled\nin by xyz parent commit\". It's something that is quite tedious to do\nmanually, especially since the submodule pointer could change\narbitrarily so knowing the submodule commit doesn't mean you can\nsimply grep for which commit set the submodule exactly to that commit.\nEssentially, I want a 'git blame' that ignores all changes which\naren't actually the submodule pointer, update.\n\nI think that's something that is much harder to do manually, but feels\nlike it should be relatively simple to implement within the blame\nalgorithm. I don't feel like this is something strictly replaceable by\n\"git submodule foreach\"\n\n>\n> Randall\n>\n"},{"id":"422459","messageId":"YH8hxQKarZW6sU+9@google.com","threadId":"55505","inReplyTo":"CA+P7+xrOuhG5ujQRYS0=o7S9=xD5zm6BGp5mBRt493Lme9xYcw@mail.gmail.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2021-04-20T18:47:33Z","receivedAt":"2021-04-20T18:47:41Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Tue, Apr 20, 2021 at 09:18:05AM -0700, Jacob Keller wrote:\n> \n> On Mon, Apr 19, 2021 at 12:28 PM Randall S. Becker\n> <rsbecker@nexbridge.com> wrote:\n> > On April 19, 2021 3:15 PM, Jacob Keller wrote:\n> > > A sort of dream I had was a flow where I could do something from the parent\n> > > like \"git blame <path/to/submodule>/submodule/file\" and have it present a\n> > > blame of that files contents keyed on the *parent* commit that changed the\n> > > submodule to have that line, as opposed to being forced to go into the\n> > > submodule and figure out what commit introduced it and then go back to the\n> > > parent and find out what commit changed the submodule to include that\n> > > submodule commit.\n> >\n> > Not going to disagree, but are you looking for the blame on the submodule ref file itself or files in the submodule? It's hard to teach git to do a blame on a one-line file.\n> >\n> \n> Well, I would like if \"git blame <path/to/submodule>\" did.. something\n> other than just fail. Sometimes my brain is working in a \"blame where\n> this came from\" and I type that out and then get frustrated when it\n> fails. Additionally...\n> \n> > Otherwise, and I think this is what you really are going for, teaching it to do a blame based on \"git blame <path/to/submodule>/submodule/file\" would be very nice and abstracts out the need for the user (or more importantly to me = scripts) to understand that a submodule is involved; however, it is opening up a very large door: \"should/could we teach git to abstract submodules out of every command\". This would potentially replace a significant part of the use cases for the \"git submodule foreach\" sub-command. In your ask, the current paradigm \"cd <path/to/submodule>/submodule && git blame file\" or pretty much every other command does work, but it requires the user/script to know you have a submodule in the path. So my question is: is this worth the effort? I don't have a good answer to that question. Half of my brain would like this very much/the other half is scared of the impact to the code.\n> >\n> > Just my musings.\n> \n> I'm not asking for \"git blame <path/to/submodule>/<file>\" to give the\n> the same outout as \"cd <path/to/submodule> && git blame <file>\"\n> \n> What i'm asking is: given this file, tell me which commit in the\n> parent did the line get introduced. So basically I want to walk over\n> the changes to the submodule pointer and find out when it get\n> introduced into the parent, not when it got introduced into the\n> submodule itself.\n> \n> This is a related question, but it is actually not trivial to go\n> instantly from \"it was in xyz submodule commit\" to \"it was then pulled\n> in by xyz parent commit\". It's something that is quite tedious to do\n> manually, especially since the submodule pointer could change\n> arbitrarily so knowing the submodule commit doesn't mean you can\n> simply grep for which commit set the submodule exactly to that commit.\n> Essentially, I want a 'git blame' that ignores all changes which\n> aren't actually the submodule pointer, update.\n> \n> I think that's something that is much harder to do manually, but feels\n> like it should be relatively simple to implement within the blame\n> algorithm. I don't feel like this is something strictly replaceable by\n> \"git submodule foreach\"\n\nI think I understand what you're saying. Something like the following\ntree:\n\nsuper   sub\nb------->4\n         3\n         2\na------->1\n\nproducing something like this:\n\n'git -C sub blame main.c'\n\n1 AU Thor\t2020-01-01\n2 CO Mitter\t2020-01-02\t\tint main() {\n4 AU Thor\t2020-01-04\t\t  printf(\"Hello world!\\n\");\n3 Dev E\t\t2020-01-03\t\t  return 0;\n2 CO Mitter\t2020-01-02\t\t}\n\nand\n'git blame sub/main.c'\n\na Mai N\t\t2020-01-01\nb Senior Dev\t2020-01-04\t\tint main() {\nb Senior Dev\t2020-01-04\t\t  printf(\"Hello world!\\n\");\nb Senior Dev\t2020-01-04\t\t  return 0;\nb Senior Dev\t2020-01-04\t\t}\n\nor to put it another way: if we are treating superproject commit as \"the\nwhole feature\", then it could be useful to see \"which feature added this\nchange\" instead of \"which atomic commit inside a feature added this\nchange\".\n\nTo me, it sounds expensive to compute... wouldn't you  need to say, for\neach blame line, \"is this commit an ancestor of the commit associated in\nTHIS superproject commit? ...how about the next superproject commit?\"\nBut I also don't have much experience with the blame implementation so\nmaybe I'm thinking naively :) :)\n\nAnd even if it is expensive, considering that Jacob and Randall both had\ndifferent ideas of what their ideal 'git blame' recursive behavior would\nbe, maybe it makes sense to use a flag to ask for the more expensive\nbehavior, e.g. 'git blame --show-superproject-commit sub/main.c'?\n\n - Emily\n"},{"id":"422460","messageId":"YH8iTDNZpsoCu+lx@google.com","threadId":"55505","inReplyTo":"YH1+C47AErrCUkHI@pug.qqx.org","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2021-04-20T18:49:48Z","receivedAt":"2021-04-20T18:49:58Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Mon, Apr 19, 2021 at 08:56:43AM -0400, Aaron Schrab wrote:\n> \n> At 16:36 -0700 16 Apr 2021, Emily Shaffer <emilyshaffer@google.com> wrote:\n> > - git switch / git checkout\n> \n> (snip)\n> \n> > 4. A new branch with the same name is created on each submodule.\n> >  a. If there is a naming conflict, we could prompt the user to resolve it, or\n> >     we could just check out the branch by that name and print a warning to the\n> >     user with advice on how to solve it (cd submodule && git switch -c\n> >     different-branch-name HEAD@{1}). Maybe we could skip the warning/advice if\n> >     the tree is identical to the tree we would have used as the start point\n> >     (that is, the user switched branches in the submodule, then said \"oh crap\"\n> >     and went back and switched branches in the superproject).\n> >  b. Tracking info is set appropriately on each new branch to the upstream of\n> >     the branch referenced by the parent of the new superproject commit, OR to\n> >     the default branch's upstream.\n> > 5. The new branch is checked out on each of the submodules.\n> \n> In many cases the branch name for the superproject isn't going to be\n> appropriate for submodules.\n> \n> This seems likely to create a LOT of junk branches. Do you also have a\n> proposal for cleaning those up?\n\nYeah, I think we have a point internally for \"clean up alllll the\nsubmodule branches that are unreferenced/already merged\". You're right\nthat in a workflow where I have a superproject with eight submodules,\nbecause I need them to build, but only do active development on one\nsubmodule out of the eight, I'll have a ton of junk refs in the other\nseven submodules. Yuck :)\n\n - Emily\n"},{"id":"422463","messageId":"019c01d7361b$82344700$869cd500$@nexbridge.com","threadId":"55505","inReplyTo":"YH8iTDNZpsoCu+lx@google.com","subject":"RE: RFC/Discussion - Submodule UX Improvements","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-04-20T19:29:34Z","receivedAt":"2021-04-20T19:29:51Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On April 20, 2021 2:50 PM, Emily Shaffer wrote:\n> On Mon, Apr 19, 2021 at 08:56:43AM -0400, Aaron Schrab wrote:\n> >\n> > At 16:36 -0700 16 Apr 2021, Emily Shaffer <emilyshaffer@google.com>\n> wrote:\n> > > - git switch / git checkout\n> >\n> > (snip)\n> >\n> > > 4. A new branch with the same name is created on each submodule.\n> > >  a. If there is a naming conflict, we could prompt the user to resolve\nit, or\n> > >     we could just check out the branch by that name and print a\nwarning to\n> the\n> > >     user with advice on how to solve it (cd submodule && git switch -c\n> > >     different-branch-name HEAD@{1}). Maybe we could skip the\n> warning/advice if\n> > >     the tree is identical to the tree we would have used as the start\npoint\n> > >     (that is, the user switched branches in the submodule, then said\n\"oh crap\"\n> > >     and went back and switched branches in the superproject).\n> > >  b. Tracking info is set appropriately on each new branch to the\nupstream of\n> > >     the branch referenced by the parent of the new superproject\ncommit, OR\n> to\n> > >     the default branch's upstream.\n> > > 5. The new branch is checked out on each of the submodules.\n> >\n> > In many cases the branch name for the superproject isn't going to be\n> > appropriate for submodules.\n> >\n> > This seems likely to create a LOT of junk branches. Do you also have a\n> > proposal for cleaning those up?\n> \n> Yeah, I think we have a point internally for \"clean up alllll the\nsubmodule\n> branches that are unreferenced/already merged\". You're right that in a\n> workflow where I have a superproject with eight submodules, because I need\n> them to build, but only do active development on one submodule out of the\n> eight, I'll have a ton of junk refs in the other seven submodules. Yuck :)\n\nIn fact, this yuck is a reason why many organizations have gone to\nmonolithic repositories instead of multiple smaller ones - because of the\ntouch points. However, the argument for using multiple smaller repos mirrors\nthis particular use case, so while \"yuck\", it might have value when\nmirroring what happens in the issue tracking systems that have massive touch\npoints. We were there and moved to monolithic per product release group, but\nwhen we had the other approach, this particular feature actually would have\nhelped a whole lot. I wonder whether this mess might have more value than we\nthink.\n\nRegards,\nRandall\n\n-- Brief whoami:\nNonStop developer since approximately 211288444200000000\nUNIX developer since approximately 421664400\nMVS not admitting to anything\n-- In my real life, I talk too much.\n\n\n\n"},{"id":"422465","messageId":"01a001d7361c$b73c8560$25b59020$@nexbridge.com","threadId":"55505","inReplyTo":"YH8hxQKarZW6sU+9@google.com","subject":"RE: RFC/Discussion - Submodule UX Improvements","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2021-04-20T19:38:13Z","receivedAt":"2021-04-20T19:38:24Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On April 20, 2021 2:48 PM, Emily Shaffer wrote:\n> On Tue, Apr 20, 2021 at 09:18:05AM -0700, Jacob Keller wrote:\n> >\n> > On Mon, Apr 19, 2021 at 12:28 PM Randall S. Becker\n> > <rsbecker@nexbridge.com> wrote:\n> > > On April 19, 2021 3:15 PM, Jacob Keller wrote:\n> > > > A sort of dream I had was a flow where I could do something from\n> > > > the parent like \"git blame <path/to/submodule>/submodule/file\" and\n> > > > have it present a blame of that files contents keyed on the\n> > > > *parent* commit that changed the submodule to have that line, as\n> > > > opposed to being forced to go into the submodule and figure out\n> > > > what commit introduced it and then go back to the parent and find\n> > > > out what commit changed the submodule to include that submodule\n> commit.\n> > >\n> > > Not going to disagree, but are you looking for the blame on the\nsubmodule\n> ref file itself or files in the submodule? It's hard to teach git to do a\nblame on a\n> one-line file.\n> > >\n> >\n> > Well, I would like if \"git blame <path/to/submodule>\" did.. something\n> > other than just fail. Sometimes my brain is working in a \"blame where\n> > this came from\" and I type that out and then get frustrated when it\n> > fails. Additionally...\n> >\n> > > Otherwise, and I think this is what you really are going for, teaching\nit to do\n> a blame based on \"git blame <path/to/submodule>/submodule/file\" would be\n> very nice and abstracts out the need for the user (or more importantly to\nme =\n> scripts) to understand that a submodule is involved; however, it is\nopening up a\n> very large door: \"should/could we teach git to abstract submodules out of\nevery\n> command\". This would potentially replace a significant part of the use\ncases for\n> the \"git submodule foreach\" sub-command. In your ask, the current paradigm\n> \"cd <path/to/submodule>/submodule && git blame file\" or pretty much every\n> other command does work, but it requires the user/script to know you have\na\n> submodule in the path. So my question is: is this worth the effort? I\ndon't have a\n> good answer to that question. Half of my brain would like this very\nmuch/the\n> other half is scared of the impact to the code.\n> > >\n> > > Just my musings.\n> >\n> > I'm not asking for \"git blame <path/to/submodule>/<file>\" to give the\n> > the same outout as \"cd <path/to/submodule> && git blame <file>\"\n> >\n> > What i'm asking is: given this file, tell me which commit in the\n> > parent did the line get introduced. So basically I want to walk over\n> > the changes to the submodule pointer and find out when it get\n> > introduced into the parent, not when it got introduced into the\n> > submodule itself.\n> >\n> > This is a related question, but it is actually not trivial to go\n> > instantly from \"it was in xyz submodule commit\" to \"it was then pulled\n> > in by xyz parent commit\". It's something that is quite tedious to do\n> > manually, especially since the submodule pointer could change\n> > arbitrarily so knowing the submodule commit doesn't mean you can\n> > simply grep for which commit set the submodule exactly to that commit.\n> > Essentially, I want a 'git blame' that ignores all changes which\n> > aren't actually the submodule pointer, update.\n> >\n> > I think that's something that is much harder to do manually, but feels\n> > like it should be relatively simple to implement within the blame\n> > algorithm. I don't feel like this is something strictly replaceable by\n> > \"git submodule foreach\"\n> \n> I think I understand what you're saying. Something like the following\n> tree:\n> \n> super   sub\n> b------->4\n>          3\n>          2\n> a------->1\n> \n> producing something like this:\n> \n> 'git -C sub blame main.c'\n> \n> 1 AU Thor\t2020-01-01\n> 2 CO Mitter\t2020-01-02\t\tint main() {\n> 4 AU Thor\t2020-01-04\t\t  printf(\"Hello world!\\n\");\n> 3 Dev E\t\t2020-01-03\t\t  return 0;\n> 2 CO Mitter\t2020-01-02\t\t}\n> \n> and\n> 'git blame sub/main.c'\n> \n> a Mai N\t\t2020-01-01\n> b Senior Dev\t2020-01-04\t\tint main() {\n> b Senior Dev\t2020-01-04\t\t  printf(\"Hello world!\\n\");\n> b Senior Dev\t2020-01-04\t\t  return 0;\n> b Senior Dev\t2020-01-04\t\t}\n> \n> or to put it another way: if we are treating superproject commit as \"the\nwhole\n> feature\", then it could be useful to see \"which feature added this change\"\n> instead of \"which atomic commit inside a feature added this change\".\n> \n> To me, it sounds expensive to compute... wouldn't you  need to say, for\neach\n> blame line, \"is this commit an ancestor of the commit associated in THIS\n> superproject commit? ...how about the next superproject commit?\"\n> But I also don't have much experience with the blame implementation so\n> maybe I'm thinking naively :) :)\n> \n> And even if it is expensive, considering that Jacob and Randall both had\n> different ideas of what their ideal 'git blame' recursive behavior would\nbe,\n> maybe it makes sense to use a flag to ask for the more expensive behavior,\ne.g.\n> 'git blame --show-superproject-commit sub/main.c'?\n\nI was partly trying to figure out which path Jacob was requesting and \"both\"\nseem useful to me. Looking at our own super-repo history and comparing what\nis in one specific submodule, we have the commit on the submodule ref file\nflipping repeatedly between two commits during a period of time (in\ntimestamp order) where an there were multiple submodule topic branches in\nparallel with super-repo topic branches. I'm not saying that was a good\nthing, just the reality of one particular icky submodule. Once back on the\nmain branch, things moved to something rational, but a blame of changing\nsubmodule contents of sub/main.c --show-superproject-commit would lead to\nsomething inherently non-deterministic until the topic branches are all\npruned and/or merged, at least in this degenerative situation. After the\nmerge, everything was back to deterministic and simple to compute.\n\nRandall\n\n"},{"id":"422493","messageId":"YH9drebF84mx2t5r@google.com","threadId":"55505","inReplyTo":"0fc5c0f7-52f7-fb36-f654-ff5223a8809b@gmail.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2021-04-20T23:03:09Z","receivedAt":"2021-04-20T23:03:17Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Sun, Apr 18, 2021 at 11:20:06PM -0400, Philippe Blain wrote:\n> > To download the code, they should be able to run simply git clone\n> > https://example.com/superproject to download the project and all its submodules;\n> \n> Playing the devil's advocate here, but some projects do not want / need all of\n> their submodules in a \"regular\" checkout, so I guess that would have to be somehow\n> configurable. I've always felt that since each project is different in that regard,\n> it would be better if each project could declare if their submodules are non-optional\n> and need to be also cloned when the superproject is cloned. Maybe an additional field\n> in '.gitmodules', like a boolean 'submodule.<name>.optional', could be added,\n> so that submodules that are optional are not cloned, but others are. If that setting\n> is opt-in (meaning that it defaults to 'true', i.e., submodules are considered optional by default),\n> then it would be easier to argue for 'git clone' to mean 'git clone --recurse-submodules':\n> 'git clone' would clone the superproject and any non-optional submodule.\n> Then eventually, when the usage of 'submodule.<name>.optional' becomes more widespread,\n> we can switch the default and then projects would need to explicitely declare their submodule\n> optional if they don't want them cloned by a simple 'git clone'.\n\nThis is actually a point we discussed internally and I cut out of the\ndoc before sharing, because it is very far down our roadmap (not\nexpecting to address until probably the second half of the year). As I\nunderstand it, this can also be achieved today by setting\n'submodule.path/to/module.active = false' in the superproject's\n.git/config.\n\nHowever, it seems to me like this would be a really cool application of\nsparse-checkout, especially if you could distribute the sparse-checkout\nconfig (for example, during clone* ;) ) before a user has the chance to\ntry and clone all necessary repos for the initial checkout.\n\n* https://lore.kernel.org/git/pull.908.git.1616105016055.gitgitgadget@gmail.com\n> > While the user is waiting for feedback on their review, to work on their next\n> > task, they can 'git switch other-feature', which will checkout the branches\n> > specified in the superproject commit at the tip of 'other-feature'; now the user\n> > can continue working as before.\n> \n> Here, I'm not sure what you mean by \"the branches (plural) specified in the superproject\n> commit at the tip of other-feature\". Today, with 'submodule.recurse = true', 'git checkout some-feature'\n> already checks out each submodule in detached HEAD at the commit recorded in the superproject commit\n> at the tip of some-feature. It's unclear if you are proposing to instead record submodule branch\n> names in the superproject commit.. is that what's going on here ? (or is it just a typo ?)\n\nYeah, I'm not sure that it makes sense to record branch name in the\ncommit itself - but I could see it being useful to do some inference on\nthe client side and put the user in some state besides detached-HEAD on\ncheckout. Hmmm.\n\n> \n> > \n> > When it's time to update their local repo, the user can do so as with a\n> > single-repo project. First they can 'git checkout main && git pull' (or 'git\n> > pull -r'); Git will first checkout the branches associated with main in each\n> > submodule, then fetch and merge/rebase in each submodule appropriately.\n> \n> What if some submodule does not use the same branch name for their primary integration branch?\n> Sometimes as a superproject using another project as a submodule, you do not\n> control that...\n\nYeah, you're right that that's an important consideration - \"how can I\nteach my superproject to default to a different branch than the name of\nthe superproject's current branch?\" I wonder whether the branch config\nin .gitmodules (or an equivalent in superproject's .git/config) would\nmake sense to try and use here?\n\n> > Note that this means options like '--branch' *don't* propagate directly to the\n> > submodules. If superproject branch \"foo\" points its submodule to branch \"main\",\n> \n> Here again, I'm not sure what you mean, because right now there is no concept of\n> the superproject having a submodule \"pointing to some branch\", only to a specific\n> commit. 'submodule.<name>.branch' is only ever used by the command 'git submodule update --remote'.\n> Is there an implicit proposal to change that ?\n\nThere is no proposal to change the concept of superproject commits\nreferencing submodule commits. I do not think it is a good idea to try\nto have superproject commits reference submodule branch names. No. :) I\nthink the answer here is the same as above - detached-HEAD is an\ninconvenient state for the user unless they specifically ask for it, and\nit would be better to check out some branch predictably if possible.\n\nThat means that it should be hard for users to end up in a state where\nthe submodule commit 123 referenced by the superproject commit abc\ndoesn't have some ref pointing to it in the submodule; I think this is\nwhat I was trying to get at with the 'git status' improvements and 'git\ncheckout'/'git switch' warnings.\n\n> \n> > then 'git clone --branch foo https://superproject.git' will clone\n> > superproject/submodule to branch 'main' instead. (It *may* be OK to take\n> > '--branch' to mean \"the branch specified by the parent *and* the branch named in\n> > --branch if applicable, but no other branches\".)\n> > \n> > What doesn't already work:\n> > \n> >    * --recurse-submodules should turn on submodule.recurse=true\n> \n> That's actually a good very idea, but maybe it should be explicitely mentioned, I think\n> (in the output of the command I mean).\n> \n> >    * superproject gets top-level config inherited by submodules\n> >    * New --recurse-submodules --single-branch semantics\n> >    * Progress bar for clone (see work estimates)\n> >    * Recommended config from project owner\n> > \n> > \n> > -- Partial clone\n> > \n> > 1. git clone initializes the directory indicated by the user\n> > 2. git clone applies the appropriate configs for the partial clone filter\n> >     requested by the user\n> >    a) These configs go to the config file shared by superproject and submodules.\n> > 3. git clone fetches the superproject\n> > 4. git clone checks out the superproject at server's HEAD\n> > 5. git clone warns the user that a recommended hook/config setup exists and\n> >     provides a tip on how to install it\n> > 6. For each submodule encountered in step 4, git clone is invoked for the\n> >     submodule, and steps 1-4 are repeated (but in directories indicated by the\n> >     superproject commit, not by the user). The same filter supplied to the\n> >     superproject applies to the submodules.\n> > \n> > \n> > What doesn't already work:\n> > \n> >    * --filter=blob:none with submodules (it's using global variables)\n> >    * propagating --filter=blob:none to submodules (via submodules.config)\n> >    * Recommended config from project owner\n> > \n> > \n> > - git fetch\n> > \n> > By default, git fetch looks for (1) the remote name(s) supplied at the command\n> > line, (2) the remote which the currently checked out branch is tracking, or (3)\n> > the remote named origin, in that order. For submodules, there is no guarantee\n> > that (1) has anything to do with the state of the submodule referenced by the\n> > superproject commit, so just start from (2).\n> > \n> > This operation can be extremely long-running if the project contains many large\n> > submodules, so progress indicators should be displayed.\n> > \n> > Caveat: this will mean that we should be more careful about ensuring that\n> > submodule branches have tracking info set up correctly; that may be an issue for\n> > users who want to branch within their submodule. This may be OK because users\n> > will probably still have 'origin' as their submodule's remote, and if they want\n> > more complicated behavior, they will be able to configure it.\n> > \n> > What doesn't already work:\n> > \n> >    * Make sure not to propagate (1) to submodules while recursing\n> >    * Fetching new submodules.\n> >    * Not having 0.95 success probability ** 100 = low success probability (that\n> >      is, we need more retries during submodule fetch)\n> >    * Progress indicators\n> \n> I would add the following:\n> \n> - Fix 'git fetch upstream' when 'submodule.recurse' and 'fetch.recurseSubdmodules=on-demand'\n> are both set  (the submodule is not fetched even if the superproject changed the submodule\n> commit).\n\nInteresting. Sounds like it's worth writing a test case to see what does\nhappen/what should happen and make it work :)\n\n> \n> - Do not rely on 'origin' exising in the submodule (or being pushable to). Right now,\n> renaming the 'origin' remote to 'upstream' in a submodule, and using 'origin' for one's own\n> fork of a submodule, (as is often done in the superproject), breaks 'git fetch --recurse-submodules'\n> (or 'git fetch' if 'submodule.recurse' is set), in the sense that the fetch does not recurse\n> to the submodule, as it should. I do not have a simple reproducer handy but\n> I've seen it happen and there are a couple hard-coded \"origin\" in the submodule code [1], [2].\n\nThis sounds to me like a specific example of a more generalized goal,\nwhich may or may not have ended up in this doc(?) to appropriately\nchoose the right remote for fetching and pushing. So, definitely :)\n\n> > \n> > \n> > - git switch / git checkout\n> > \n> > Submodules should continue to perform these operations the same way that they\n> > have before, that is, the way that single-repo Git works. But superprojects\n> > should behave as follows:\n> > \n> > \n> > -- Create mode (git switch -c / git checkout -b)\n> > \n> > 1. The current worktree is checked for uncommitted changes to tracked files. The\n> >     current worktree of each submodule is also checked.\n> > 2. A new branch is created on the superproject; that branch's ref is pointed to\n> >     the current HEAD.\n> > 3. The new branch is checked out on the superproject.\n> > 4. A new branch with the same name is created on each submodule.\n> \n> That might not be wanted by all, so I think it should be configurable.\n> \n> >    a. If there is a naming conflict, we could prompt the user to resolve it, or\n> >       we could just check out the branch by that name and print a warning to the\n> >       user with advice on how to solve it (cd submodule && git switch -c\n> >       different-branch-name HEAD@{1}). Maybe we could skip the warning/advice if\n> >       the tree is identical to the tree we would have used as the start point\n> >       (that is, the user switched branches in the submodule, then said \"oh crap\"\n> >       and went back and switched branches in the superproject).\n> >    b. Tracking info is set appropriately on each new branch to the upstream of\n> >       the branch referenced by the parent of the new superproject commit, OR to\n> >       the default branch's upstream.\n> \n> This last point is a little unclear: which \"new superproject commit\" ? (we are creating\n> a branch, so there is no new commit yet?). And again, you talk about a (submodule?) branch being referenced\n> by a superproject commit, which is not a concept that actually exists today.\n\nYeah, I can clean up the wording here, thanks for pointing it out.\n\n> Also, usually tracking info is only set\n> automatically when using the form 'git checkout -b new-branch upstream/master' or\n> the like. Do you also propose that 'git checkout -b new-branch', by itself, should\n> automatically set tracking info ?\n\nYes - that is an approach that we want to explore, to solve the general\npush/fetch remote+branch problem.\n\n> \n> \n> > 5. The new branch is checked out on each of the submodules.\n> > \n> > What doesn't already work:\n> > \n> >    * Safety check when leaving uncommitted submodule changes\n> \n> Yes, that has been reported several times ([3], [4], [5]). I have fixes for this,\n> not quite ready to send because I'm trying to write extensive tests (maybe too extensive)...\n> \n> >    * Propagating branch names to submodules currently requires a custom hacky\n> >      repolike patch\n> >    * Error handling + graceful non-error handling if the branch already exists\n> >    * \"Knowing what branch to push to\": copying over which-branch-is-upstream info\n> >      ** Needs some UX help, push.default is a mess\n> >    * Tracking info setups\n> > \n> > -- Switching to an existing branch (git switch / git checkout)\n> > \n> > 1. The current worktree is checked for uncommitted changes to tracked files. The\n> >     current worktree of each submodule is also checked.\n> > 2. The requested branch is checked out on the superproject.\n> > 3. The submodule commit or branch referenced by the newly-checked-out\n> >     superproject commit is checked out on each submodule.\n> > \n> > What doesn't already work:\n> > \n> >    * Same as in create mode\n> \n> Here, I would add that 'git checkout --recurse-submodules', along with 'git clone --recurse-submodules',\n> have trouble with correctly checkout-ing an older commit that records a submodule that\n> was since removed from the project. The user experience around this use case is currently very very bad [6].\n> This is partly due to 'git clone --recurse-submodules' only cloning submodules that are recorded in\n> the tip commit of the default branch of the superproject, which could certainly be improved.\n\nYeah, we are aware of this pain internally too, thanks for pointing it\nout.\n\n> \n> > \n> > \n> > - git status\n> > \n> > -- From superproject\n> > The superproject is clean if:\n> > \n> >    * No tracked files in the superproject have been modified and not committed\n> >    * No tracked files in any submodules have been modified and not committed\n> >    * No commits in any submodules differ from the commits referenced by the tip\n> >      commit of the superproject\n> > \n> > Advices should describe:\n> > \n> >    * How to commit or drop changes to files in the superproject\n> >    * How to commit or drop changes to files in the submodules\n> >    * How to commit changes to submodule references\n> >    * Which commit/branch to switch the submodule back to if the current work\n> >      should be dropped: \"Submodule \"foo\" no longer points to \"main\", 'git -C foo\n> >      switch main' to discard changes\"\n> > \n> > What doesn't already work:\n> > \n> >    * \"git status\" being super fast and actually possible to use.\n> >      ** (That is, we've seen it move very slowly on projects with many\n> >         submodules.)\n> >    * Advice updates to use the appropriate submodule-y commands.\n> \n> I would add that 'git status' should show the submodule as \"rewind\" if the\n> currently checked out submodule commit is *behind* what's recorded in the current superproject\n> commit. That is shown by 'git diff --submodule=<log | diff>' and 'git submodule summary'\n> and is quite useful to prevent a following 'git commit -am' in the superproject to regress the submodule commit\n> by mistake. It would be nice if 'git status' could also show this information (code in\n> submodule.c::show_submodule_header).\n\nOh interesting, that's a good point. Thanks.\n\n> \n> > \n> > -- From submodule\n> > \n> > git status's behavior for submodules does not change compared to\n> > single-repository Git, except that a red warning line will also display if the\n> > superproject commit does not point to the HEAD of the submodule. (This could\n> > look similar to the detached-HEAD warning and tracking branch lines in git\n> > status today, e.g. \"HEAD is ahead of parent project by 2 commits\".)\n> \n> That would be a nice addition :)\n> \n> > \n> > What doesn't already work:\n> > \n> >    * \"git status\" from a submodule being aware of the superproject.\n> > \n> > \n> > - git push\n> > \n> > -- From superproject\n> > \n> > Ideally, a push of the superproject commit results in a push of each submodule\n> > which changed, to the appropriate Gerrit upstream. Commits pushed this way\n> > across submodules should somehow be associated in the Gerrit UI, similar to the\n> > \"submitted together\" display. This will need some work to make happen.\n> > \n> > What doesn't already work:\n> > \n> >    * Automatically setting Gerrit topic (with a hook)\n> >    * \"push --recurse-submodules\" knowing where to push to in submodules to\n> >      initiate a Gerrit review\n> >      ** From `branch` field in .gitmodules?\n> >      ** Gerrit accepting 'git push -o review origin main' pushes?\n> >      ** Review URL with a remote helper that rewrites refs/heads/main to\n> >         refs/for/main?\n> >      ** Need UX help\n> \n> It would be nice if 'git push' would not force users to use the same\n> remote names and branch names in the superproject and the submodule.\n> Previous discussion around this that I had spotted are at [7] and [8].\n> \n> > \n> > > From submodule\n> > No change to client behavior is needed. With Gerrit submodule subscriptions, the\n> > server knows how to generate superproject commits when merging submodule\n> > commits.\n> > \n> > - git pull / git rebase\n> > \n> > Note: We're still thinking about this one :)\n> > \n> > 1. Performs a fetch as described above\n> > 2. For each superproject commit, replay the submodule commits against the newly\n> >     updated submodule base; then, make a new superproject commit containing those\n> >     changes\n> > \n> > What doesn't already work:\n> > \n> >    * Rewriting gitlinks in a superproject commit when 'rebase\n> >      --recurse-submodules'-ing\n> >    * Resuming after resolving a conflict during rebase\n> \n> In general, rebase is not well aware of 'submodule.recurse'. Even if you do not\n> need to rewrite superproject commits, there are a couple of use cases that are broken\n> right now:\n> \n> - 'git rebase upstream/master' when upstream updated the submodule, will correctly\n> (recursively) checkout upstream/master before starting the rebase, but upon\n> 'git rebase --abort', the submodule will stay checked out at the commit recorded in\n> 'upstream/master', which is confusing. This only happens when 'submodule.recurse' is true (!).\n> - 'git rebase -i' which stops at a commit 'A' where the submodule commit is changed,\n> does not correctly check out the submodule tree. It's checked out at the commit recorded in A~1\n> (and this also only happens if submodule.recurse is true)\n> - In some cases, like 'rebase -i'-ing across the addition of new submodules, at the end\n> of the rebase the submodules are empty, and 'git submodule update' must be run to\n> re-populate them.\n\nInteresting, thanks for pointing these out.\n\n> \n> > \n> > - git merge\n> > \n> > The story for merges is a little bit muddled... and for our goals we don't need\n> > it for quite a while, so we haven't thought much about it :) Any suggestions\n> > folks have about reasonable ways to 'git merge --recurse-submodules' are totally\n> > welcome. For now, though, we'll probably just stick in some error message saying\n> > that merges with submodules isn't currently supported (maybe we will even add\n> > that downstream).\n> \n> What is \"downstream\" here ?\n\n\"Downstream\" meaning the version (fork? ehh) of Git that we build and\nship to developers at Google. We carry a handful of patches - mostly for\nstuff that only makes sense internally, like certain transports or\nauthentication helpers - and occasionally experimental stuff (for\nexample, we ship config-based hooks to Googlers this way right now). If\nwe're expecting \"No, you can't merge with submodules!\" to be a temporary\nerror message, then it might not make sense to try and upstream that\nerror string at all.\n\n\n\nThanks for the thorough read and all the pointers, I really appreciate\nit.\n\n - Emily\n"},{"id":"422495","messageId":"YH9feTykPhimIA13@google.com","threadId":"55505","inReplyTo":"CAP8UFD0Ct8NofMdds=w0k1-jjX638L6QJQEJWVxqJ6ZPSoJUjg@mail.gmail.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2021-04-20T23:10:49Z","receivedAt":"2021-04-20T23:10:59Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Sun, Apr 18, 2021 at 07:22:07AM +0200, Christian Couder wrote:\n> \n> Hi Emily,\n> \n> On Sat, Apr 17, 2021 at 1:39 AM Emily Shaffer <emilyshaffer@google.com> wrote:\n> >\n> > Hi folks,\n> >\n> > As hinted by a couple recent patches, I'm planning on some pretty big submodule\n> > work over the next 6 months or so - and Ævar pointed out to me in\n> > https://lore.kernel.org/git/87v98p17im.fsf@evledraar.gmail.com that I probably\n> > should share some of those plans ahead of time. :) So attached is a lightly\n> > modified version of the doc that we've been working on internally at Google,\n> > focusing on what we think would be an ideal submodule workflow.\n> \n> Thanks for sharing this doc! My main concern with this is that we are\n> likely to have a GSoC student working soon on finishing to port `git\n> submodule` to C code. And I wonder how that would interact with your\n> work.\n\nI discussed this a little with Jonathan N and Albert and we think it\nprobably won't matter too much. If anything, I expect mostly we would\ntouch the submodule--helper, and not the 'git submodule' builtin. But\njust in case - it would be useful if any GSoC student were publishing\ntheir code to a feature branch (on a fork, maybe) so that I could keep\nan eye out for possible conflicts that way. Or, at very least, CCing me\nand Jonathan N on patches :)\n\n - Emily\n"},{"id":"422499","messageId":"xmqq1rb4355h.fsf@gitster.g","threadId":"55505","inReplyTo":"YH9drebF84mx2t5r@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-04-20T23:30:50Z","receivedAt":"2021-04-20T23:30:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n> This is actually a point we discussed internally and I cut out of the\n> doc before sharing, because it is very far down our roadmap (not\n> expecting to address until probably the second half of the year). As I\n> understand it, this can also be achieved today by setting\n> 'submodule.path/to/module.active = false' in the superproject's\n> .git/config.\n\nYeah, I think we also added support to choose which submodules can\nbe \"active\" based on the attributes system.\n\nThree are many ways to apply band-aid to a tree that should have\nbeen a monolithic single repository but has been split into many\nsubmodules only because we historically did not scale well.  As you\nmeantioned, sparse-checkout and lazy/partial cloning may change the\npicture drastically, not just \"sparse\" may allow such an \"a set of\nartificially split out submodules\" to be selectively populated, but\nmore directly clone and work with only the parts you are interested\nin a monolithic repository.\n"},{"id":"422520","messageId":"a8ac4042-a0dd-47eb-8419-7b7d19da7cec@gmail.com","threadId":"55505","inReplyTo":"YH9drebF84mx2t5r@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Philippe Blain","fromEmail":"levraiphilippeblain@gmail.com","sentAt":"2021-04-21T02:27:19Z","receivedAt":"2021-04-21T02:27:24Z","isPatch":false,"sender":{"key":"levraiphilippeblain@gmail.com","avatar":"https://avatars.githubusercontent.com/u/44212482?v=4"},"body":"\n\nLe 2021-04-20 à 19:03, Emily Shaffer a écrit :\n\n>>> When it's time to update their local repo, the user can do so as with a\n>>> single-repo project. First they can 'git checkout main && git pull' (or 'git\n>>> pull -r'); Git will first checkout the branches associated with main in each\n>>> submodule, then fetch and merge/rebase in each submodule appropriately.\n>>\n>> What if some submodule does not use the same branch name for their primary integration branch?\n>> Sometimes as a superproject using another project as a submodule, you do not\n>> control that...\n> \n> Yeah, you're right that that's an important consideration - \"how can I\n> teach my superproject to default to a different branch than the name of\n> the superproject's current branch?\" I wonder whether the branch config\n> in .gitmodules (or an equivalent in superproject's .git/config) would\n> make sense to try and use here?\n\nI think it depends on the workflow. Re-reading the above, I would definitely *not* want\n'git pull --recurse-submodules' in the superproject to go into each submodule\nand do 'git pull' there ! Because maybe some submodule introduced breaking changes\nin its API or something and I do not want to deal with that now; I just want to update my tree\nwith the latest changes *to the superproject* (and maybe to the submodules *if* they\nwere updated by the superproject, but not if they were updated in the submodule upstream project).\nFor me, 'git pull --recurse-submodules'\nhas mostly the right behaviour today, except what does not work (doing something useful\nwhen both sides record changes to the submodule pointer).\n\n\n> \n>> Also, usually tracking info is only set\n>> automatically when using the form 'git checkout -b new-branch upstream/master' or\n>> the like. Do you also propose that 'git checkout -b new-branch', by itself, should\n>> automatically set tracking info ?\n> \n> Yes - that is an approach that we want to explore, to solve the general\n> push/fetch remote+branch problem.\n\nYeah, it would be nice if the triangular workflow capabilities of Git would be expanded\n(if I understand correctly that's what you are hinting at here). My personal TODO list\nfor that has the following items (just dumping that here in case it's useful to someone):\n\n# improve UI/UX around 'branch.pushRemote' and 'remote.pushDefault'\n- git branch --verbose could show difference with @{push} in addition to / instead of @{upstream}\n- git status \"\n- git prompt \"\n- add config branch.<name>.pushBranch (or pushRef)\n- add 'git branch --set-push-to remote/name' to set branch.name.pushRemote and branch.name.pushRef\n- add 'git push -p <remote> <branch>' to set 'branch.name.pushRemote' and 'branch.name.pushRef' (and warn if push.default is not 'current') OR:\n- allow 'branch.pushRemote' and 'remote.pushDefault' to work if push.default=simple\n- reword push.default section in git-config  (very unclear)\n\nhttps://lore.kernel.org/git/87d0q72du2.fsf@javad.com/t/#u\nhttps://lore.kernel.org/git/20130607124146.GF28668@sociomantic.com/t/#u\n\n\nCheers,\n\nPhilippe.\n"},{"id":"422535","messageId":"CA+P7+xprW6tBH8Yo=-qxD0J8ZeeFspOnAp1RShscNcy7v_4btQ@mail.gmail.com","threadId":"55505","inReplyTo":"YH8hxQKarZW6sU+9@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2021-04-21T06:57:17Z","receivedAt":"2021-04-21T06:57:31Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Tue, Apr 20, 2021 at 11:47 AM Emily Shaffer <emilyshaffer@google.com> wrote:\n>\n> On Tue, Apr 20, 2021 at 09:18:05AM -0700, Jacob Keller wrote:\n> >\n> > On Mon, Apr 19, 2021 at 12:28 PM Randall S. Becker\n> > <rsbecker@nexbridge.com> wrote:\n> > > On April 19, 2021 3:15 PM, Jacob Keller wrote:\n> > > > A sort of dream I had was a flow where I could do something from the parent\n> > > > like \"git blame <path/to/submodule>/submodule/file\" and have it present a\n> > > > blame of that files contents keyed on the *parent* commit that changed the\n> > > > submodule to have that line, as opposed to being forced to go into the\n> > > > submodule and figure out what commit introduced it and then go back to the\n> > > > parent and find out what commit changed the submodule to include that\n> > > > submodule commit.\n> > >\n> > > Not going to disagree, but are you looking for the blame on the submodule ref file itself or files in the submodule? It's hard to teach git to do a blame on a one-line file.\n> > >\n> >\n> > Well, I would like if \"git blame <path/to/submodule>\" did.. something\n> > other than just fail. Sometimes my brain is working in a \"blame where\n> > this came from\" and I type that out and then get frustrated when it\n> > fails. Additionally...\n> >\n> > > Otherwise, and I think this is what you really are going for, teaching it to do a blame based on \"git blame <path/to/submodule>/submodule/file\" would be very nice and abstracts out the need for the user (or more importantly to me = scripts) to understand that a submodule is involved; however, it is opening up a very large door: \"should/could we teach git to abstract submodules out of every command\". This would potentially replace a significant part of the use cases for the \"git submodule foreach\" sub-command. In your ask, the current paradigm \"cd <path/to/submodule>/submodule && git blame file\" or pretty much every other command does work, but it requires the user/script to know you have a submodule in the path. So my question is: is this worth the effort? I don't have a good answer to that question. Half of my brain would like this very much/the other half is scared of the impact to the code.\n> > >\n> > > Just my musings.\n> >\n> > I'm not asking for \"git blame <path/to/submodule>/<file>\" to give the\n> > the same outout as \"cd <path/to/submodule> && git blame <file>\"\n> >\n> > What i'm asking is: given this file, tell me which commit in the\n> > parent did the line get introduced. So basically I want to walk over\n> > the changes to the submodule pointer and find out when it get\n> > introduced into the parent, not when it got introduced into the\n> > submodule itself.\n> >\n> > This is a related question, but it is actually not trivial to go\n> > instantly from \"it was in xyz submodule commit\" to \"it was then pulled\n> > in by xyz parent commit\". It's something that is quite tedious to do\n> > manually, especially since the submodule pointer could change\n> > arbitrarily so knowing the submodule commit doesn't mean you can\n> > simply grep for which commit set the submodule exactly to that commit.\n> > Essentially, I want a 'git blame' that ignores all changes which\n> > aren't actually the submodule pointer, update.\n> >\n> > I think that's something that is much harder to do manually, but feels\n> > like it should be relatively simple to implement within the blame\n> > algorithm. I don't feel like this is something strictly replaceable by\n> > \"git submodule foreach\"\n>\n> I think I understand what you're saying. Something like the following\n> tree:\n>\n> super   sub\n> b------->4\n>          3\n>          2\n> a------->1\n>\n> producing something like this:\n>\n> 'git -C sub blame main.c'\n>\n> 1 AU Thor       2020-01-01\n> 2 CO Mitter     2020-01-02              int main() {\n> 4 AU Thor       2020-01-04                printf(\"Hello world!\\n\");\n> 3 Dev E         2020-01-03                return 0;\n> 2 CO Mitter     2020-01-02              }\n>\n> and\n> 'git blame sub/main.c'\n>\n> a Mai N         2020-01-01\n> b Senior Dev    2020-01-04              int main() {\n> b Senior Dev    2020-01-04                printf(\"Hello world!\\n\");\n> b Senior Dev    2020-01-04                return 0;\n> b Senior Dev    2020-01-04              }\n>\n> or to put it another way: if we are treating superproject commit as \"the\n> whole feature\", then it could be useful to see \"which feature added this\n> change\" instead of \"which atomic commit inside a feature added this\n> change\".\n>\n\nRight. I often want to find out when some change actually made it into\nthe super project.\n\n> To me, it sounds expensive to compute... wouldn't you  need to say, for\n> each blame line, \"is this commit an ancestor of the commit associated in\n> THIS superproject commit? ...how about the next superproject commit?\"\n> But I also don't have much experience with the blame implementation so\n> maybe I'm thinking naively :) :)\n\nWell I imagine it has to be similar to how we compute the blame for a\nregular file? I imagine we start at some commit and walk backwards up\nthe tree, no?\n\nI imagine the current blame algorithm starts from the current commit\nand walks backwards through the commit history, determining which\ncommit was last to have a given line.\n\nIn the submodule case I highlighted, we would be doing the same thing:\nFollow the super project history. When you find a submodule file, pull\nits contents from the matching submodule commit that the parent\nhistory saw. No need to dig any further into the submodule commit\nhistory, just give me that contents and then I can treat it as if that\ncontents was what was in the super project for this commit, and use\nthe normal blame algorithm.\n\nIt's much more difficult to do that manually (hence why we invented\nblame/annotate in the first place), and trying to go from \"git -C\n<submodule> blame file\" to then figure out which super project commit\nintroduced the change is also tedious and non-trivial considering you\nmight now have intermediate or unrelated changes (i.e. it's actually\npossible that that particular commit *never* made it into the super\nproject at all, because it got skipped over, and it might even be\nafter the file got re-written)\n\nMy idea for how blame of submodujles work is to essentially pretend as\nif you had subtree merged the contents of the submodule into regular\nparent project files with those paths, and then do blame on that using\njust the parent project history.... If that makes sense?\n\n>\n> And even if it is expensive, considering that Jacob and Randall both had\n> different ideas of what their ideal 'git blame' recursive behavior would\n> be, maybe it makes sense to use a flag to ask for the more expensive\n> behavior, e.g. 'git blame --show-superproject-commit sub/main.c'?\n>\n\nRight I imagine that in some ways both are useful, and it depends on\nthe context of what you're looking for.\n\nThe reason I bring up the blame example is because the idea for what I\nwant is quite tedious to mimic by hand, and requires more than just a\nsimple git submodule foreach or a cd into the submodule to operate on\nit as a standalone repository.\n\n>  - Emily\n"},{"id":"422692","messageId":"CA+P7+xrTSBVWcX2uZmHAErKWXAFHesGNv5cizeTfmX5yuJCD-g@mail.gmail.com","threadId":"55505","inReplyTo":"YHofmWcIAidkvJiD@google.com","subject":"Re: RFC/Discussion - Submodule UX Improvements","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2021-04-22T15:32:52Z","receivedAt":"2021-04-22T15:33:06Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Fri, Apr 16, 2021 at 4:38 PM Emily Shaffer <emilyshaffer@google.com> wrote:\n> -- Create mode (git switch -c / git checkout -b)\n>\n> 1. The current worktree is checked for uncommitted changes to tracked files. The\n>    current worktree of each submodule is also checked.\n> 2. A new branch is created on the superproject; that branch's ref is pointed to\n>    the current HEAD.\n> 3. The new branch is checked out on the superproject.\n> 4. A new branch with the same name is created on each submodule.\n>   a. If there is a naming conflict, we could prompt the user to resolve it, or\n>      we could just check out the branch by that name and print a warning to the\n>      user with advice on how to solve it (cd submodule && git switch -c\n>      different-branch-name HEAD@{1}). Maybe we could skip the warning/advice if\n>      the tree is identical to the tree we would have used as the start point\n>      (that is, the user switched branches in the submodule, then said \"oh crap\"\n>      and went back and switched branches in the superproject).\n>   b. Tracking info is set appropriately on each new branch to the upstream of\n>      the branch referenced by the parent of the new superproject commit, OR to\n>      the default branch's upstream.\n> 5. The new branch is checked out on each of the submodules.\n>\n> What doesn't already work:\n>\n>   * Safety check when leaving uncommitted submodule changes\n>   * Propagating branch names to submodules currently requires a custom hacky\n>     repolike patch\n>   * Error handling + graceful non-error handling if the branch already exists\n>   * \"Knowing what branch to push to\": copying over which-branch-is-upstream info\n>     ** Needs some UX help, push.default is a mess\n>   * Tracking info setups\n>\n\n\nAs someone who uses submodules extensively for various projects, I'm\nnot sure about propagating branches into the submodules.\n\nI think i'd only want this behavior if/when I intend to work on a\nsubmodule. Because of the nature of submodules being distinct, we tend\ntowards doing submodule work separately, merging it, and then pulling\nthat change into the super project.\n\n>\n> -- Switching to an existing branch (git switch / git checkout)\n>\n> 1. The current worktree is checked for uncommitted changes to tracked files. The\n>    current worktree of each submodule is also checked.\n> 2. The requested branch is checked out on the superproject.\n> 3. The submodule commit or branch referenced by the newly-checked-out\n>    superproject commit is checked out on each submodule.\n>\n> What doesn't already work:\n>\n>   * Same as in create mode\n\nI'd imagine there are multiple cases here. For cases where you're not\nactively developing submodule, you want to just checkout the right\ncontents (i.e. what is tracked by the super project). But if you're\ndeveloping the submodule in conjunction with the super project you\nmight want to instead checkout the matching work (as in above where\nyou create a branch within the submodule?)\n\n> - git status\n>\n> -- From superproject\n> The superproject is clean if:\n>\n>   * No tracked files in the superproject have been modified and not committed\n>   * No tracked files in any submodules have been modified and not committed\n>   * No commits in any submodules differ from the commits referenced by the tip\n>     commit of the superproject\n>\n> Advices should describe:\n>\n>   * How to commit or drop changes to files in the superproject\n>   * How to commit or drop changes to files in the submodules\n>   * How to commit changes to submodule references\n>   * Which commit/branch to switch the submodule back to if the current work\n>     should be dropped: \"Submodule \"foo\" no longer points to \"main\", 'git -C foo\n>     switch main' to discard changes\"\n>\n> What doesn't already work:\n>\n>   * \"git status\" being super fast and actually possible to use.\n>     ** (That is, we've seen it move very slowly on projects with many\n>        submodules.)\n>   * Advice updates to use the appropriate submodule-y commands.\n>\n\nYea, a slow status means people tend to not use it!\n\n> -- From submodule\n>\n> git status's behavior for submodules does not change compared to\n> single-repository Git, except that a red warning line will also display if the\n> superproject commit does not point to the HEAD of the submodule. (This could\n> look similar to the detached-HEAD warning and tracking branch lines in git\n> status today, e.g. \"HEAD is ahead of parent project by 2 commits\".)\n>\n> What doesn't already work:\n>\n>   * \"git status\" from a submodule being aware of the superproject.\n>\n\nThis seems like a very good improvement. One of the biggest complaints\nabout submodules I've had to deal with when helping coworkers is the\nfact that submodules weren't moved forward automatically, and that\nthey had no real idea that the submodule was different. This tended to\nlead towards commits including submodule rewinds on accident.\n\n\n> - Worktrees\n>\n> When a user runs 'git worktree add' from the superproject, each submodule in the\n> new worktree should also be created as a worktree of the corresponding submodule\n> in the original project.\n>\n> What doesn't already work:\n>\n>   * worktrees and submodules getting along - submodules are now freshly cloned\n>     when creating a superproject worktree\n\nThis is something I would love to see fixed!  Right now using work\ntrees on a project with submodules is problematic. Especially a\nproject with many submodules, as this ends up making many extra\nclones, taking disk space and network time to setup.\n\n>\n> - git clone --reference [--dissociate]\n>\n> When cloning with an alternate directory, submodules should also try to use\n> object stores associated with the referenced project instead of cloning from\n> their remotes right away. It is unclear how much of this works today.\n>\n>\n> What doesn't already work:\n>\n>   * Writing some tests and making them pass\n"}]}