{"thread":{"id":"50902","subject":"[GSoC] [RFC] Proposal: Teach git stash to handle unmerged index entries.","startedAt":"2019-04-09T15:39:40Z","lastAt":"2019-04-10T05:09:43Z","messageCount":3,"participants":["Kapil Jain","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"373498","messageId":"CAMknYEMh=CN6GGRPD_fkafHy84e49JY5dK2dAgX6Z7542dJ-Uw@mail.gmail.com","threadId":"50902","inReplyTo":null,"subject":"[GSoC] [RFC] Proposal: Teach git stash to handle unmerged index entries.","fromName":"Kapil Jain","fromEmail":"jkapil.cs@gmail.com","sentAt":"2019-04-09T15:39:25Z","receivedAt":"2019-04-09T15:39:40Z","isPatch":false,"sender":{"key":"jkapil.cs@gmail.com","avatar":null},"body":"Plan to implement the project.\n\nObjective:\nTeach git stash to handle unmerged index entries.\n\nDescription:\nWhen the index is unmerged, git stash refuses to do anything. That is\nunnecessary, though, as it could easily craft e.g. an octopus merge of\nthe various stages. A subsequent git stash apply can detect that\noctopus and re-generate the unmerged index.\n\n\nImplementation Idea:\nPerforming an octopus merge of all `stage n` (n>0) unmerged index\nentries, could solve the problem, but\n\nWhat if there are conflicts in merging ?\nIn this case, we would store(commit) the conflicted state, so they can\nbe regenerated when git stash is applied.\n\nHow to store the conflicted files ?\ncreate a tree from the merge using `git-write-tree`\nand then commit that tree using `git-commit-tree`.\n\n\nRelevant Discussions:\nhttps://colabti.org/irclogger/irclogger_log/git-devel?date=2019-04-05#l92\nhttps://colabti.org/irclogger/irclogger_log/git-devel?date=2019-04-09#l47\n\n\nIdea Execution Plan: Divided into 2 parts.\n\nPart 1: Store the unmerged index entries this part will work with `git\nstash push`\n\nstash.sh: file would be changed to accommodate the below implementation.\n\nStep 1:\nExtract all the unmerged entries from index file and store them in a\ntemporary index file.\n\nread-cache.c: this file is responsible for reading index file,\nprobably this implementation will end up in this file.\n\nStep 2:\ncache-tree.c: study and implement a slightly modified version of the\nfunction `write_index_as_tree()`\n\nint write_index_as_tree(struct object_id *oid, struct index_state\n*index_state, const char *index_path, int flags, const char *prefix);\n\nthis function is responsible for writing tree from index file.\nCurrently in this function, the index must be in a fully merged state,\nand we are dealing with its exact opposite. So a version to write tree\nfor unmerged index entries will be implemented.\n\nStep 3:\nwrite-tree.c: some possible changes will go here, so as to use the\nmodified version of write_index_as_tree() function.\n\nStep 4:\nuse git-commit-tree to commit the written tree and store the hash in\nsome file say `stash_conflicting_merge`\n\nStep 5:\nWrite tests for all implementation till this point.\n\nPart 2: Retrieve the tree hash and regenerate the state of repository\nas it was earlier.\n\nStep 6:\nModify implementation of `git stash apply` for regenerating the committed tree.\n\nStep 7:\nWrite tests.\n"},{"id":"373568","messageId":"xmqqo95e1nik.fsf@gitster-ct.c.googlers.com","threadId":"50902","inReplyTo":"CAMknYEMh=CN6GGRPD_fkafHy84e49JY5dK2dAgX6Z7542dJ-Uw@mail.gmail.com","subject":"Re: [GSoC] [RFC] Proposal: Teach git stash to handle unmerged index entries.","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-04-10T04:40:19Z","receivedAt":"2019-04-10T04:40:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kapil Jain <jkapil.cs@gmail.com> writes:\n\n> Plan to implement the project.\n>\n> Objective:\n>\n> Description:\n>\n> Implementation Idea:\n>\n> Relevant Discussions:\n>\n> Idea Execution Plan: Divided into 2 parts.\n\nTwo things missing before implementation idea are design, and more\nimportantly, the success criteria.  What lets you and your mentor\ndeclare victory?\n\nAs to the design, it does not quite matter if you add four or more\nseparate trees to represent stage #[0123] entries in the index to\nthe already octopus merge commit that represents a stash entry\n(i.e. when keeping the untracked ones, I think the stash entry's\n\"result of the merge\" tree records the state of the tracked files in\nthe working tree, and the \"result of the merge\" commit records the\nthe-current HEAD, a commit that records the state of the index and\nanothre commit that records the state of the untracked files, as its\nparents---that's already a 3-parent octopus).\n\nThe fact that a stash entry is represented as a merge commit is a\nmere implementation detail, and there is *NO* need to worry about\nresolving merge conflicts while recording a stash.  If the result of\nthis GSoC task is to be any usable together with the current version\nin a backward compatible way, you must record these extra states as\nextra parents of the merge, so it is sort of given already that\nyou'd be using some form of an octopus merge.\n\nThe real challenge would be how the unstashing part of such a stash\nentry that records unmerged state should work.  Personally I do not\nthink it will be very useful to allow unstashing such a stash entry\non top of any arbitrary commit---rather, I suspect that the user\nwould want to come back to the exact HEAD the user had trouble\nresolving conflicts at, without having to first checking it out.\nIOW, a usual way to use \"git stash\" is\n\n\t$ git checkout topic\n\t$ edit edit edit\n\t... I am happily hacking away ...\n\t... the boss appears with an ultra-urgent task ...\n\t$ git stash save -m WIP\n\t$ git checkout master\n\t$ edit-and-build-and-test\n\t$ git commit\n\t... now the emergency is over ...\n\t$ git checkout topic\n\t... sync with the work others may have done on topic\n\t... while I was dealing with the boss\n\t$ git pull --rebase origin topic\n\t$ git stash pop\n\nIOW, it is expected to be applied on top of an updated commit.\n\nBut I have a moderately strong suspicion that a stash that holds\nunmerged state (i.e. a conflicted merge in progress) is created with\na use case, which is very different from the normal use case, in\nmind.  When creating such a stash entry, the above sequence would go\nmore like this:\n\n\t$ git checkout topic\n\t$ git merge ...\n\t... oops, conflicted, and it takes time to resolve ...\n\t$ edit edit inspect edit\n\t... the boss appears\n\t$ git stash save -m \"Merge in progress\"\n\t$ git checkout master\n\t... deal with the emergency the same way ...\n\t$ git checkout topic\n\t... go back to the conflict resolution first without\n\t... touching what may have happened on the branch in\n\t... the meantime---a human brain cannot afford to deal\n\t... with two or more parallel conflicts at the same\n\t... time.\n\t$ git stash pop\n\t... now deal with the conflict we were looking at\n\t... before the boss interrupted us.\n\t$ edit inspect edit\n\t... be satisfied with the result\n\t$ git commit\n\t... now let's see if others have something else that\n\t... is interesting\n\t$ git pull --rebase origin topic\n\nAnd if we assume that the primary use of a stash for a conflicted\nstate is to bring us back to the exact state (rather than allowing\nus to pretend as if we started form a different HEAD), it might even\nmake sense to teach \"git stash pop\" step to barf if HEAD does not\nmatch the first parent of the merge commit that represents the stash\nentry being applied (again, stash^{tree} is the working tree,\nstash^1 is then-current HEAD).  That would make the application side\na lot simpler and manageable by developers who are not intimately\nfamiliar with the code.\n\nOthers may disagree with the above assumption (i.e. \"a stash for a\nconflicted state does not have to be applicable), though, making\nyour task a lot harder ;-).\n\nQuite honestly, I do not think you can design a system that attempts\nto \"stash apply/pop\" a recorded unmerged state on top of any\narbitrary HEAD and leave a state useful for the end user to deal with\nwhen the \"stash apply/pop\" step itself introduces _new_ conflicts\ndue to the differences between the then-current HEAD the stash entry\nis based on and the HEAD the \"stash apply\" is attempted on top of.\nEven the current \"stash apply/pop with the change between the HEAD\nand the index\" does punt when it cannot make a clean application,\nand that is without any unmerged entries in the recorded index\nstate.\n\nThe key point is \"a state useful for the end user\"---it is easy to\nbuild a system that claims to leave a state created from the updated\nHEAD and what's recorded in a stash entry that the end users cannot\nuse as a stating point to make progress, but that is not something\nour users would want.\n\nHave fun.\n"},{"id":"373570","messageId":"xmqqk1g21m5p.fsf@gitster-ct.c.googlers.com","threadId":"50902","inReplyTo":"xmqqo95e1nik.fsf@gitster-ct.c.googlers.com","subject":"Re: [GSoC] [RFC] Proposal: Teach git stash to handle unmerged index entries.","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-04-10T05:09:38Z","receivedAt":"2019-04-10T05:09:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> As to the design, it does not quite matter if you add four or more\n> separate trees to represent stage #[0123] entries in the index to\n> the already octopus merge commit that represents a stash entry ...\n\nI forgot that I was planning to expand on this part while writing\nthe message I am following up.\n\nThere are a few things you must take into account while designing a\nnew format for a stash entry:\n\n - Your new feature will *NOT* be the last extension to the stash\n   subsystem.  Always leave room to other developers to extend it\n   further, without breaking backward compatiblity when your new\n   feature is int in use.\n\n - Even though you may have never encountered in your projects,\n   higher stage entries can have duplicates.  When merging two\n   branches into your current branch, and there are three merge\n   bases for such an octopus merge, the system (and the index\n   format) is designed to allow a merge backend to store 3 stage #1\n   entries (because there are that many common ancestor versions in\n   the example), 1 stage #2 entry (because there is only one\n   \"current brahch\" a merge is made into) and 2 stage #3 entries\n   (because there are that many other branches you are merging into\n   the current branch), all for the same path.\n\nSo, a design that says:\n\n   A stash entry in the current system is recorded as a merge\n   commit, whose tree represents the state of the tracked working\n   tree files, whose first parent records the HEAD commit the stash\n   entry was created on, and whose second parent records the tree\n   that would have been created if \"git write-tree\" were done on the\n   index when the stash entry was created.  Optionally, it can have\n   the third parent whose tree records the state of untracked files.\n\n   Let's add three more parents.  IOW, the fourth parent's tree\n   records the result of \"git write-tree\" of the index after\n   removing all the entries other than those at stage #1 and moving\n   the remainder from stage #1 down to stage #0, and similarly the\n   fifth is for stage #2 and the sixth is for stage #3.\n\nis bad at multiple counts.\n\n - It does not say what should happen to the third parent when this\n   new \"record unmerged state\" feature is used without using the\n   \"record untracked paths\" feature.\n\n - It does not allow multiple stage #1 and/or stage #3 entries.\n\nFor the first point, I think a trick to record the same commit as\nthe first parent may be a good hack to say \"this is not used\"; we\nmight need to allow commit-tree not to complain about duplicate\nparents if we go that route.\n\nFOr the second one, there may be multiple solutions.  A\nquick-and-dirty and obvious way may be to add only one new parent to\nthe merge commit that represents a stash entry (i.e. the fourth\nparent).  Make that new parent a merge of three commits, each of\nwhich represents what was in stage #1, stage #2 and stage #3 (we can\nreuse the second parent of the stash entry that usually records the\nindex state to store stage #0 entries).\n\nAs we allow multiple stage #1 or stage #3 entries in the index, and\nthere is no fundamental reason why we should not allow multiple\nstage #2 entries, make each of these three commits able to represent\nmultiple entries at the same stage, perhaps by\n\n - iterate over the index and count the maximum occurrence of the\n   same path at the same stage #$n;\n - make that stage #$n commit a merge of that many parent commits.\n   The tree recorded in that stage #$n commit can be an empty tree.\n\nI am not saying this is a good design.  I am merely showing the\nexpected level of detail when your design gets in a presentable\nshape and shared with the list.\n\nHave fun.\n\n\n"}]}