{"thread":{"id":"61795","subject":"git subtree bugs (mishandled merges, recursion depth)","startedAt":"2024-07-17T16:55:11Z","lastAt":"2026-04-17T04:14:59Z","messageCount":4,"participants":["Ian Jackson","Colin Stagner"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"498842","messageId":"26263.63341.878041.155047@chiark.greenend.org.uk","threadId":"61795","inReplyTo":null,"subject":"git subtree bugs (mishandled merges, recursion depth)","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2024-07-17T16:55:09Z","receivedAt":"2024-07-17T16:55:11Z","isPatch":false,"body":"I have what ought to be a fairly straightforward situation that\ngit-subtree seems to be mishandling.\n\nSteps to reproduce:\n\n git clone https://gitlab.torproject.org/tpo/core/arti.git\n cd arti\n git checkout 01d02118cdda30636e606fc1a89b3e04f28b8ad1\n git subtree split -P maint/rust-maint-common\n\nExpected behaviour:\n\n git subtree (hopefully fairly rapidly) prints a the commitid of the\n tip of a branch suitable for merging back to the upstream repo, which\n is at https://gitlab.torproject.org/tpo/core//rust-maint-common\n\n The resulting history ought to have a few dozen commits,\n most of which are the upstream history of the subtree.\n\nActual behaviour (git 2.45.2, Debian amd64 1:2.45.2-1 .deb):\n\n $ git subtree split -P maint/rust-maint-common\n /usr/lib/git-core/git-subtree: 318: Maximum function recursion depth (1000) reached\n $\n\nActual behaviour (git 2.20.1, Debian ancient 1:2.20.1-2+deb10u9):\n\n Takes a very long time.  Everntually produces an output commit\n which has most of arti.git#main in its history.\n\nNotes about the source repository:\n\n The state of arti.git:maint/rust-maint-common is the result of the\n following:\n   (i) create a new rust-maint-common.git, and add and edit files\n     (many of these changes came via gitlab MRs, there are merges)\n   (ii) in arti.git, `git subtree add`, and make further changes,\n     to files both within and without the subtree\n   (iii) Make a gitlab MR from (ii) and merge it into arti.git#main.\n     (resulting in a fairly merge-rich history)\n     https://gitlab.torproject.org/tpo/core/arti/-/merge_requests/2267\n\nA workaround:\n\n If I check out main^2 (01d02118cdda30636e606fc1a89b3e04f28b8ad1^2)\n and run git-subtree split using the ancient version of git, it still\n takes ages, but the output is correct.  So the old version of git has\n a bug meaning it can produce higly excessive output, when merges are\n present.\n\n This workaround is only available because right now the history of\n the subtree's files, within arti.git, is fairly simple.\n\n With the new version of git, I get the \"recursion depth\" error,\n regardless.\n\nThanks for your attention.\n\nIan.\n\n-- \nIan Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.  \n\nPronouns: they/he.  If I emailed you from @fyvzl.net or @evade.org.uk,\nthat is a private address which bypasses my fierce spamfilter.\n"},{"id":"541705","messageId":"e9611b58-3886-4f04-8f49-16d140ebfc15@howdoi.land","threadId":"61795","inReplyTo":"26263.63341.878041.155047@chiark.greenend.org.uk","subject":"Re: git subtree bugs (mishandled merges, recursion depth)","fromName":"Colin Stagner","fromEmail":"ask+git@howdoi.land","sentAt":"2026-04-16T01:26:02Z","receivedAt":"2026-04-16T01:26:22Z","isPatch":false,"body":"Hello Ian, does this git-subtree issue still affect you?\n\nOn 7/17/24 11:55, Ian Jackson wrote:\n> Steps to reproduce:\n> \n>   git clone https://gitlab.torproject.org/tpo/core/arti.git\n>   cd arti\n>   git checkout 01d02118cdda30636e606fc1a89b3e04f28b8ad1\n>   git subtree split -P maint/rust-maint-common\n>\n> Actual behaviour (git 2.45.2, Debian amd64 1:2.45.2-1 .deb):\n> \n>   $ git subtree split -P maint/rust-maint-common\n>   /usr/lib/git-core/git-subtree: 318: Maximum function recursion depth (1000) reached\n>   $\nOn Debian's POSIX sh, shell recursion is artificially limited to 1000 \ncalls. This is not typical behavior; most distros I've tested do not cap \nit. bash has a configurable recursion depth limit, but sh ignores it.\n\nI've proposed a fix for the recursion depth issue in:\n\n<https://lore.kernel.org/git/20260305-cs-subtree-split-recursion-v2-0-7266be870ba9@howdoi.land>\n\nIf you have the time, I'd appreciate some testing and/or a code review.\n\n> Expected behaviour:\n>\n>   The resulting history ought to have a few dozen commits,\n>   most of which are the upstream history of the subtree.\n\n\n> Actual behaviour (git 2.20.1, Debian ancient 1:2.20.1-2+deb10u9):\n> \n>  Takes a very long time.  Everntually produces an output commit\n>  which has most of arti.git#main in its history.\n\nEven with my patch series applied, there are many more than a \"few dozen \ncommits\" in the history. For me this splits as\n\n     9a2422685e6cc05625f47a1fe709f1908f31fc87\n\nwith 12307 commits in the history graph.\n\nThe reason for this is likely e7b07376e5 (Merge branch \n'rs/subtree-fixes', 2018-10-26), which was merged around that time. \nPrevious versions discarded too much history, and that patch series \nadded more merge-base ancestry checks.\n\nWhen merges come into play, the task of choosing which history is \n\"important\" and which history is \"not important\" is not always clear-cut.\n\nColin\n\n"},{"id":"541744","messageId":"27104.62121.658449.222834@chiark.greenend.org.uk","threadId":"61795","inReplyTo":"e9611b58-3886-4f04-8f49-16d140ebfc15@howdoi.land","subject":"Re: git subtree bugs (mishandled merges, recursion depth)","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2026-04-16T14:31:05Z","receivedAt":"2026-04-16T14:31:07Z","isPatch":false,"body":"Colin Stagner writes (\"Re: git subtree bugs (mishandled merges, recursion depth)\"):\n> On 7/17/24 11:55, Ian Jackson wrote:\n> > Actual behaviour (git 2.20.1, Debian ancient 1:2.20.1-2+deb10u9):\n> > \n> >  Takes a very long time.  Everntually produces an output commit\n> >  which has most of arti.git#main in its history.\n> \n> Even with my patch series applied, there are many more than a \"few dozen \n> commits\" in the history. For me this splits as\n\nHi.  (For future reference, that patch series is\n  [PATCH v2 0/3] contrib/subtree: reduce recursion during split\nin the other thread.)\n\n>      9a2422685e6cc05625f47a1fe709f1908f31fc87\n> \n> with 12307 commits in the history graph.\n> \n> The reason for this is likely e7b07376e5 (Merge branch \n> 'rs/subtree-fixes', 2018-10-26), which was merged around that time. \n> Previous versions discarded too much history, and that patch series \n> added more merge-base ancestry checks.\n> \n> When merges come into play, the task of choosing which history is \n> \"important\" and which history is \"not important\" is not always clear-cut.\n\nI have some thoughts about this.\n\nI didn't find a formal description of git-subtree's data model, or how\ngit subtree split works, precisely.  So I'm going to make some\nsuppositions.\n\nI observe that git-subtree split doesn't record any metadata in the\nsplit versions of the commits (for example, the downstream project\ncommitid they were split from).\n\nRepeated splits ought ideally not to constantly generate additional\nmaterial.  So the algorithm ought to be deterministic.  An easy way to\ndo that is to make splitting a pure function from downstream commits\nto subtree commits.\n\nIf one can run git subtree split on every commit in the downstream\nthat has a git subtree merge as an ancestor, then one might think that\nmeans the split must produce as many commits as there are in the\ndowntream.\n\nBut we can map multiple downstream commits to the same subtree\ncommit.  Consider the cases, for some downstream commit D.\n\n 0. D is a single parent commit that *does* change the subtree.\n    This becomes a new commit with parent split(D~).\n\n 1. D is a single parent commit that doesn't change the subtree:\n    We reuse the parent's split: split(D) = split(D~)\n\n 2. D is a multi-parent commit.  Determine \\forall{i} split(D^i).\n    Discard all split(D^i) which are ancestors of any split(D^j).\n    If any remaining split(D^i) is not subtree-treesame D,\n    or there is more than one remaining split(D^i),\n    construct a new commit with those remaining split(D^i) as parents.\n    Otherwise all remaining split(D^i) are the same,\n    and they are treesame to D, so discard: split(D) = split(D^i).\n\n 3. D is a subtree merge commit.  split(D^1) is explicitly stated\n    in the git-subtree metadata.  Calculate split(D^0) as above.\n    Then calculate split(D) according to point 2.\n\nIn fact, 0 and 1 are special cases of 2.\n\nDo you think it would be worth me prototyping this?  I think at least\nfor my case it would produce considerably fewer commits, but until I\ntry it that's just guesswork.\n\nIan.\n\n-- \nIan Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.  \n\nPronouns: they/he.  If I emailed you from @fyvzl.net or @evade.org.uk,\nthat is a private address which bypasses my fierce spamfilter.\n"},{"id":"541806","messageId":"d1b51a09-acc8-4354-b5bc-f50a13e7c3fd@howdoi.land","threadId":"61795","inReplyTo":"27104.62121.658449.222834@chiark.greenend.org.uk","subject":"Re: git subtree bugs (mishandled merges, recursion depth)","fromName":"Colin Stagner","fromEmail":"ask+git@howdoi.land","sentAt":"2026-04-17T04:14:43Z","receivedAt":"2026-04-17T04:14:59Z","isPatch":false,"body":"On 4/16/26 09:31, Ian Jackson wrote:\n> Colin Stagner writes (\"Re: git subtree bugs (mishandled merges, recursion depth)\"):\n>> When merges come into play, the task of choosing which history is\n>> \"important\" and which history is \"not important\" is not always clear-cut.\n> \n> I have some thoughts about this.\n> \n> I didn't find a formal description of git-subtree's data model, or how\n> git subtree split works, precisely.  So I'm going to make some\n> suppositions.\n> \n> I observe that git-subtree split doesn't record any metadata in the\n> split versions of the commits (for example, the downstream project\n> commitid they were split from).\n\nThis would be helpful information to have, but git-subtree does not \nrecord it in general.\n\n`split --rejoin` mode *does* record the parent→split commitid mapping. \nBut this is only recorded in the input history, and only for a single \ncommit. --rejoin actually isn't as helpful as it may first appear.\n\nKeep in mind that no part of history is fixed. Both the input history \nand/or the split history might get rebased. If that happens, the \noriginal input→split commitid mapping is either unhelpful or misleading.\n\n\n> Repeated splits ought ideally not to constantly generate additional\n> material.  So the algorithm ought to be deterministic.\n\nThe algorithm is deterministic. It does need to reconstruct the complete \nsplit history every time, which is inefficient.\n\nThe real issue is this line from the git-subtree manual page:\n\n     Repeated splits of exactly the same history are guaranteed\n     to be identical (i.e. to produce the same commit IDs) as\n     long as the settings passed to split (such as --annotate)\n     are the same.\n\nwhich means that git-subtree SHOULDN'T change how splits work… but it \nhas anyway. There have been many fixes to subtree-split over the years \nthat have changed its behavior in history-incompatible ways. This \nguarantee really hasn't held up.\n\nWe really need something committed, like a config file, that records:\n\n1. how the user wants splits to happen *now*; and\n2. how the old split history was split before\n\nThat way, we can introduce new functionality or approaches without \nintroducing breakage.\n\n\n> An easy way to do that is to make splitting a pure\n> function from downstream commits to subtree commits.\n> \n> If one can run git subtree split on every commit in the downstream\n> that has a git subtree merge as an ancestor, then one might think that\n> means the split must produce as many commits as there are in the\n> downtream.\n\nPart of the problem here may be that `subtree split` is not particularly \naware of `subtree merge`. All split sees is that the tree has a \ndifferent prefix. Changing this would introduce history-breakage, and it \nwould be good to make it opt-in.\n\nYou can also perform subtree merges without `git-subtree`. Merges done \nthis way don't record any information in the trailers.\n\n\n> But we can map multiple downstream commits to the same subtree\n> commit.  Consider the cases, for some downstream commit D.\n> \n>   0. D is a single parent commit that *does* change the subtree.\n>      This becomes a new commit with parent split(D~).\n> \n>   1. D is a single parent commit that doesn't change the subtree:\n>      We reuse the parent's split: split(D) = split(D~)\n\nI believe both (0) and (1) are current behavior.\n\n>   2. D is a multi-parent commit.  Determine \\forall{i} split(D^i).\n>      Discard all split(D^i) which are ancestors of any split(D^j).\n>      If any remaining split(D^i) is not subtree-treesame D,\n>      or there is more than one remaining split(D^i),\n>      construct a new commit with those remaining split(D^i) as parents.\n>      Otherwise all remaining split(D^i) are the same,\n>      and they are treesame to D, so discard: split(D) = split(D^i).\n\nThe merge processing is where a lot of history-breaking changes have \noccurred. There are probably lots of edge-cases to discover along the \nway. I recommend drawing lots and lots of pictures.\n\n>   3. D is a subtree merge commit.  split(D^1) is explicitly stated\n>      in the git-subtree metadata.  Calculate split(D^0) as above.\n>      Then calculate split(D) according to point 2.\n\nIn a git-subtree-merge, the merge always has two parents. The split \nhistory is always on the second parent. The trailers are mostly useful \nto tell you what the prefix was… but you should still check the commitid \nto verify that no rebase has happened.\n\n\n> Do you think it would be worth me prototyping this?\n\nIf it's worth it to you, then it's worth it. If it's good, others will \nwant it too.\n\nI've tried lots of git workflows, including submodules and subtrees. But \nI've found that a regular, plain `git merge`—with a \"merge upwards\" \nworkflow—is by far the fastest, easiest, and most flexible. Whenever \npossible, I arrange projects so they can work this way.\n\nI do use plenty of subtree merges to bring in dependencies. subtree \nmerges work great for this.\n\nAt least for now, a `subtree split` does not undo a `subtree merge`. If \nI'm looking to submit changes I've made on top of a subtree merge, I use \nformat-patch\n\n     git format-patch --relative=my/subtree/prefix ...\n\nto make patches which are layout-compatible with the original repo. Then \nI `git am` the patches in the original repo as a topic branch. This \nworks well for my purposes.\n\nI only use `split` to produce repos with read-only views of \nsubdirectories. This is mostly due to a project design quirk that is \nbeyond my control.\n\nIf I fork a project, I'll just fork the entire thing. No need to split it.\n\nThis is how I work, but you'll find plenty of other opinions out there.\n\nBest of luck,\n\nColin\n\n"}]}