{"thread":{"id":"31179","subject":"GSOC remote-svn: branch detection","startedAt":"2012-08-03T09:43:30Z","lastAt":"2012-08-07T21:26:26Z","messageCount":5,"participants":["Florian Achleitner","Jonathan Nieder","Dmitry Ivankov","Ramkumar Ramachandra"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"196404","messageId":"12682331.q6WHVv9EKU@flomedio","threadId":"31179","inReplyTo":null,"subject":"GSOC remote-svn: branch detection","fromName":"Florian Achleitner","fromEmail":"florian.achleitner.2.6.31@gmail.com","sentAt":"2012-08-03T09:43:30Z","receivedAt":"2012-08-03T09:43:30Z","isPatch":false,"sender":{"key":"florian.achleitner.2.6.31@gmail.com","avatar":"https://avatars.githubusercontent.com/u/880777?v=4"},"body":"Hi!\n\nI'm playing around in vcs-svn/ to start a framework for detecting and \nprocessing branches  in svndumps. So I wanted to let you know about my ideas.\n\nTwo approaches:\n1. Import linearly and split later:\nOne idea is to import from svn linearly, i.e. one revision on top of it's \npredecessor, like now, and detect and split branches afterwards. The svn \nmetadata is stored in git notes, so the required information would be \navailable.\n+ allows recovery, because the linear history is always here.\n+ it's easier to peek around in the git history than in the svn dump during \nimport to do the branch detection.\n- requires creation of new commits in the branch detection stage.\n- this results in double commits and awkward history, linear vs. branched.\n\n2. Split during import:\nDetect branches as they are created while reading the svn dump and identify to \nwhich branch a following node belongs.\nFirst step is to restructure svndump.c to be able to buffer one complete \nrevision for inspection before starting to write a commit to fast import.\nProbably it's possible to feed the blobs to fast import directly and only \nbuffer node data and defer commit creation, but not the data.\nCurrently, at the beginning of a new revision on the svn side, a new commit is \ncreated on top of a constant ref. When we support branches, we don't know the \nref, i.e. the branch(es), the revision changes, before reading all the 'Node-\n*' lines.\n+ feels more 'right'\n- requires revision buffering\n\nGenerally:\nDetect branches as they are created by 'Node-copyfrom*' to some commonly used \nbranch directories, like branches/. More complex branch detection can be \nimplemented later, of course.\nStore detected branches permanently (necessary for incremental fetches), and \nassign every file modification to one of those branches, if possible. Else \nassign them to, hm .. \nIf a revision modifies more than one branch, create multiple commits.\n\nThanks for your comments and ideas! \n\n--\nFlorian\n"},{"id":"196425","messageId":"20120803181728.GA21745@copier","threadId":"31179","inReplyTo":"12682331.q6WHVv9EKU@flomedio","subject":"Re: GSOC remote-svn: branch detection","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-08-03T18:17:28Z","receivedAt":"2012-08-03T18:17:28Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nFlorian Achleitner wrote:\n\n> Two approaches:\n> 1. Import linearly and split later:\n> One idea is to import from svn linearly, i.e. one revision on top of it's \n> predecessor, like now, and detect and split branches afterwards. The svn \n> metadata is stored in git notes, so the required information would be \n> available.\n> + allows recovery, because the linear history is always here.\n> + it's easier to peek around in the git history than in the svn dump during \n> import to do the branch detection.\n> - requires creation of new commits in the branch detection stage.\n> - this results in double commits and awkward history, linear vs. branched.\n\nI don't think you've captured the real pros and cons here.\n\n+ Divides responsibility between a component that fetches and a component\nthat splits branches, making for easier debugging, independent refactoring\nof components, reuse in other contexts (e.g., splitting out branches in\nother similar VCSen, etc)\n\n- Divides responsibility between a component that fetches and a component\nthat splits branches, which is tricky because it involves designing an\ninterface between them and documenting it.  And maybe a different\ninterface would be better.\n\nThere are also performance and history-clarity ramifications as you've\nmentioned, but they do not seem as important.\n\nHope that helps,\nJonathan\n\n> 2. Split during import:\n"},{"id":"196451","messageId":"CA+gfSn__=b1JfL+6LMCqYvJo5mJL_=p7JdRxFspK27OW=0oFLA@mail.gmail.com","threadId":"31179","inReplyTo":"20120803181728.GA21745@copier","subject":"Re: GSOC remote-svn: branch detection","fromName":"Dmitry Ivankov","fromEmail":"divanorama@gmail.com","sentAt":"2012-08-04T06:40:18Z","receivedAt":"2012-08-04T06:40:18Z","isPatch":false,"sender":{"key":"divanorama@gmail.com","avatar":"https://avatars.githubusercontent.com/u/158999?v=4"},"body":"Hi,\n\nOn Sat, Aug 4, 2012 at 12:17 AM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Hi,\n>\n> Florian Achleitner wrote:\n>\n>> Two approaches:\n>> 1. Import linearly and split later:\n>> One idea is to import from svn linearly, i.e. one revision on top of it's\n>> predecessor, like now, and detect and split branches afterwards. The svn\n>> metadata is stored in git notes, so the required information would be\n>> available.\n>> + allows recovery, because the linear history is always here.\nThis is a good one, but I'd put questions another way:\n- do we want to query svn server only for newer revisions even if our\nsettings changed (branch layout ones for example), maybe we don't mind\nsome queries in settings change case (like git-svn.perl)?\n- do we want to be able to filter svn history early (like take\ntrunk,branches,tags, skip tests_data as it's huge but sometimes there\nare svn cp to/from it, or maybe the repo has weird permissions or even\nis corrupted)?\n- do we just want a completely separate (fast) (local) storage like\nsvn dump file to use it for imports and settings changes?\n\nI personally still haven't decided on those. My set of pros/cons:\n+ should be the simplest thing for simple small repos\n+ keeps all the original data details and looks quite robust\n- becomes complicated if we don't want or can't import some parts of\nthe history. While git-svn.perl somehow handles is.\n- looks like a thing to store and access svn dump information, do we\nreally want it to be in a form of git objects (almost sure), how\nstable, flexible, independent from svn helper should it be (that's\nwhat Jonathan talks about).\n\nWeird idea: what if we keep everything in one huge git tree like\nrXX/{data,props,copy-from,..}/path/path/path/file. It should represent\nall the known svn info so far. Ok, I know it's a late stage now and\nthis thing is completely raw, just posting to have it written out\nsomewhere :)\n\n>> + it's easier to peek around in the git history than in the svn dump during\n>> import to do the branch detection.\n>> - requires creation of new commits in the branch detection stage.\n>> - this results in double commits and awkward history, linear vs. branched.\n>\n> I don't think you've captured the real pros and cons here.\n>\n> + Divides responsibility between a component that fetches and a component\n> that splits branches, making for easier debugging, independent refactoring\n> of components, reuse in other contexts (e.g., splitting out branches in\n> other similar VCSen, etc)\n>\n> - Divides responsibility between a component that fetches and a component\n> that splits branches, which is tricky because it involves designing an\n> interface between them and documenting it.  And maybe a different\n> interface would be better.\n>\n> There are also performance and history-clarity ramifications as you've\n> mentioned, but they do not seem as important.\n>\n> Hope that helps,\n> Jonathan\n>\n>> 2. Split during import:\n"},{"id":"196469","messageId":"CALkWK0mu1=NEUZzB1VPAf0DU_nguuq_nJ-9Rn7Pj6zeNfoZGtA@mail.gmail.com","threadId":"31179","inReplyTo":"20120803181728.GA21745@copier","subject":"Re: GSOC remote-svn: branch detection","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2012-08-04T18:23:58Z","receivedAt":"2012-08-04T18:23:58Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi,\n\nFlorian Achleitner wrote:\n> 1. Import linearly and split later:\n\nI think this approach will be a lot less messy if you can cleanly\nseparate the fetching component from the mapper.  Currently, svndump\nre-creates the layout of the SVN repository.  And the series you\nposted last week contains a patch that attaches a note with SVN\nmetadata to each commit.  Do you have thoughts on how the mapping will\ntake place?\n\nRam\n"},{"id":"196624","messageId":"3476983.FSv5Fk2g49@flobuntu","threadId":"31179","inReplyTo":"CALkWK0mu1=NEUZzB1VPAf0DU_nguuq_nJ-9Rn7Pj6zeNfoZGtA@mail.gmail.com","subject":"Re: GSOC remote-svn: branch detection","fromName":"Florian Achleitner","fromEmail":"florian.achleitner.2.6.31@gmail.com","sentAt":"2012-08-07T21:26:26Z","receivedAt":"2012-08-07T21:26:26Z","isPatch":false,"sender":{"key":"florian.achleitner.2.6.31@gmail.com","avatar":"https://avatars.githubusercontent.com/u/880777?v=4"},"body":"On Saturday 04 August 2012 23:53:58 Ramkumar Ramachandra wrote:\n> Hi,\n> \n> Florian Achleitner wrote:\n> > 1. Import linearly and split later:\n> I think this approach will be a lot less messy if you can cleanly\n> separate the fetching component from the mapper.  Currently, svndump\n> re-creates the layout of the SVN repository.  And the series you\n> posted last week contains a patch that attaches a note with SVN\n> metadata to each commit.  Do you have thoughts on how the mapping will\n> take place?\n\nThe mapping itself is currently a black box for me, it's internals could be \nrather complex. It could get a function like is_branch_start, that is called \nwith a node ctx and tells if this is likely to be the start of branch. The \ndetected branches are stored and upcoming changes in the associated \ndirectories are mapped to a commit on a branch.\nThe detection of branch starts and the list of existing branches can be taken \nfrom whatever logic we want. So that's approx. the idea.\n\nCurrently I'm working on more basic preparations. I want to split the creation \nof commits and the creation of blobs in svndump.c.\nThis is necessary because fast import requires a branch name as an argument to \nthe 'commit' command, and\ncurrently a 'commit' command is started when a new revision is encountered in \nthe svndump.\nBut to decide on which branch the commit should go, or even if it will be more \nthan one commit, it is necessary to read all the nodes first.\nTo prevent buffering the node content, I want to replace the inline data format \n(currently used) by 'blob' commands.\nWhile parsing the dump, every node change creates a blob command to feed the \ndata immediately into fast-import while the node metadata (struct node_ctx) is \nstored at least until the revision ends. Then the blobs can be put on a linear \nmaster tree and other branch trees. The node metadata could also be read from \nnotes, if remapping branches.\nThat's not so easy to do, because the current implementation mixes tree-\noperations and blob-operations heavily, and relies on only one global \nnode_ctx.\n\n> \n> Ram\n\nFlo\n"}]}