{"thread":{"id":"24292","subject":"help moving boost.org to git","startedAt":"2010-07-05T14:16:36Z","lastAt":"2010-07-06T18:29:33Z","messageCount":19,"participants":["Eric Niebler","Erik Faye-Lund","Johannes Sixt","Sverre Rabbelier","Finn Arne Gangstad","Avery Pennarun","Greg Troxel","Dave Abrahams","Jakub Narebski","David Abrahams","Raja R Harinath"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"144823","messageId":"4C31E944.30801@boostpro.com","threadId":"24292","inReplyTo":null,"subject":"help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-05T14:16:36Z","receivedAt":"2010-07-05T14:16:36Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"I have a question about the best approach to take for refactoring a\nlarge svn project into git. The project, boost.org, is a collection of\nC++ libraries (>100) that are mostly independent. (There may be\ncross-library dependencies, but we plan to handle that at a higher\nlevel.) After the move to git, we'd like each library to be in its own\ngit repository. Boost can then be a stitching-together of these, using\nsubmodules or something (opinions welcome). It's an old project with\nlots of history that we don't want to lose. The naive approach of simply\nforking into N repositories for the N libraries and deleting the\nunwanted files in each is unworkable because we'll end up with all the\nhistory duplicated everywhere ... >100 repositories, each larger than 100Mb.\n\nSo, what are the options? Can I somehow delete from each repository the\nhistory that is irrelevant? Is these some feature of git I don't know\nabout that can solve this problem for us?\n\n(Caveat: I'm new to git and still getting up to speed. An acceptable\nanswer is: go off an learn about feature X and come back to us.)\n\nAt boost, We've already discussed a few possible approaches. Feel free\nto comment and/or criticize any of the solutions suggested here:\n\n  http://github.com/ryppl/ryppl/issues#issue/4\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144824","messageId":"AANLkTinWnyB0C7HlyLz93lszfkWd81Y7reeWRpMNwHuv@mail.gmail.com","threadId":"24292","inReplyTo":"4C31E944.30801@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@googlemail.com","sentAt":"2010-07-05T14:48:47Z","receivedAt":"2010-07-05T14:48:47Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Mon, Jul 5, 2010 at 4:16 PM, Eric Niebler <eric@boostpro.com> wrote:\n> I have a question about the best approach to take for refactoring a\n> large svn project into git. The project, boost.org, is a collection of\n> C++ libraries (>100) that are mostly independent. (There may be\n> cross-library dependencies, but we plan to handle that at a higher\n> level.) After the move to git, we'd like each library to be in its own\n> git repository. Boost can then be a stitching-together of these, using\n> submodules or something (opinions welcome). It's an old project with\n> lots of history that we don't want to lose. The naive approach of simply\n> forking into N repositories for the N libraries and deleting the\n> unwanted files in each is unworkable because we'll end up with all the\n> history duplicated everywhere ... >100 repositories, each larger than 100Mb.\n>\n> So, what are the options? Can I somehow delete from each repository the\n> history that is irrelevant? Is these some feature of git I don't know\n> about that can solve this problem for us?\n>\n\nYou're probably looking for git-filter-branch. This tool can be used\nwith the --subdirectory-filter option to filter out a specific\nsubdirectory to it's own branch. Or if the project isn't split into\nsubdirectories, you can use the --tree-filter option to filter\nspecific files if you want.\n\nSee http://www.kernel.org/pub/software/scm/git/docs/git-filter-branch.html\nfor details\n\n-- \nErik \"kusma\" Faye-Lund\n"},{"id":"144825","messageId":"4C31F0D4.1040207@viscovery.net","threadId":"24292","inReplyTo":"4C31E944.30801@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2010-07-05T14:48:52Z","receivedAt":"2010-07-05T14:48:52Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 7/5/2010 16:16, schrieb Eric Niebler:\n> I have a question about the best approach to take for refactoring a\n> large svn project into git. The project, boost.org, is a collection of\n> C++ libraries (>100) that are mostly independent. (There may be\n> cross-library dependencies, but we plan to handle that at a higher\n> level.) After the move to git, we'd like each library to be in its own\n> git repository.\n\nYou could use svn2git: http://gitorious.org/svn2git\nKDE uses it to split its SVN repository into pieces. The tool is driven by\na \"ruleset\" that specifies SVN subdirectories and revision numbers that\nmake up a module.\n\n-- \n\"Atomic objects are neither active nor radioactive.\" --\nProgramming Languages -- C++, Final Committee Draft (Doc.N3092)\n"},{"id":"144833","messageId":"4C321BAF.8020303@boostpro.com","threadId":"24292","inReplyTo":"4C31F0D4.1040207@viscovery.net","subject":"Re: help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-05T17:51:43Z","receivedAt":"2010-07-05T17:51:43Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"On 7/5/2010 10:48 AM, Johannes Sixt wrote:\n> Am 7/5/2010 16:16, schrieb Eric Niebler:\n>> I have a question about the best approach to take for refactoring a\n>> large svn project into git. The project, boost.org, is a collection of\n>> C++ libraries (>100) that are mostly independent. (There may be\n>> cross-library dependencies, but we plan to handle that at a higher\n>> level.) After the move to git, we'd like each library to be in its own\n>> git repository.\n> \n> You could use svn2git: http://gitorious.org/svn2git\n> KDE uses it to split its SVN repository into pieces. The tool is driven by\n> a \"ruleset\" that specifies SVN subdirectories and revision numbers that\n> make up a module.\n\nI'm off to learn about filter-branch, tree-filter and svn2git. Thanks\nfor the suggestions. More questions to come, I'm sure.\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144837","messageId":"AANLkTilwCsMLy7jyKFoRLu07HW6E3C76Kmkm4UVg6OBU@mail.gmail.com","threadId":"24292","inReplyTo":"4C321BAF.8020303@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2010-07-05T18:43:48Z","receivedAt":"2010-07-05T18:43:48Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Mon, Jul 5, 2010 at 19:51, Eric Niebler <eric@boostpro.com> wrote:\n\n> I'm off to learn about filter-branch, tree-filter and svn2git. Thanks\n> for the suggestions. More questions to come, I'm sure.\n\nAlso have a look at git-subtree.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"144852","messageId":"20100705220443.GA23727@pvv.org","threadId":"24292","inReplyTo":"4C31E944.30801@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Finn Arne Gangstad","fromEmail":"finnag@pvv.org","sentAt":"2010-07-05T22:04:43Z","receivedAt":"2010-07-05T22:04:43Z","isPatch":false,"sender":{"key":"finnag@pvv.org","avatar":"https://gravatar.com/avatar/b421ddd58c3f0f93aa473e17b98bb8d53c221fef741746bc8cb59fae4ec6d95e?d=mp&s=160"},"body":"On Mon, Jul 05, 2010 at 10:16:36AM -0400, Eric Niebler wrote:\n> I have a question about the best approach to take for refactoring a\n> large svn project into git. The project, boost.org, is a collection of\n> C++ libraries (>100) that are mostly independent. (There may be\n> cross-library dependencies, but we plan to handle that at a higher\n> level.) After the move to git, we'd like each library to be in its own\n> git repository. Boost can then be a stitching-together of these, using\n> submodules or something (opinions welcome). It's an old project with\n> lots of history that we don't want to lose. The naive approach of simply\n> forking into N repositories for the N libraries and deleting the\n> unwanted files in each is unworkable because we'll end up with all the\n> history duplicated everywhere ... >100 repositories, each larger than 100Mb.\n\nIf the libraries are not independent (i.e. some commits are across\nmultiple libraries), submodules will give you some interesting\nchallenges to put it mildly.\n\nThe current boost 1.43 is 29344 files, is this all there is? This\nshould fit eaily into a single repository. The Linux kernel is much\nlarger, and that is sort of the canonical single repo git project. I\n_strongly_ recommend that you go for a single repo if you can make it\nwork.\n\nIf you manage to create a single git repo with the history you want,\nit is trivial to split out separate repositories of subdirectories\nlater (and those repos will then be comparatively small). git subtree\nallegedly automates this process more or less (I have not used it, but\nhave heard good things about it). What about having a single \"master\nrepository\", and then using subtree to create single-library repos for\nthe library developers if they want a smaller repo to play around in?\n\n> So,, what are the options? Can I somehow delete from each repository the\n> history that is irrelevant? Is these some feature of git I don't know\n> about that can solve this problem for us?\n\nHow do you define \"irrelevant\"? Do you only require enough history for\ngit annotate/blame to give correct results?  Or does this only refer\nto multiple repositories sharing the same ancient history?\n\n> At boost, We've already discussed a few possible approaches. Feel free\n> to comment and/or criticize any of the solutions suggested here:\n> \n>   http://github.com/ryppl/ryppl/issues#issue/4\n\nIt is unclear from the discussion if you will change to git, or use\ngit in addition to svn? This will have some impact on how to go about\nthis.\n\n- Finn Arne\n"},{"id":"144853","messageId":"4C32668E.9040000@boostpro.com","threadId":"24292","inReplyTo":"20100705220443.GA23727@pvv.org","subject":"Re: help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-05T23:11:10Z","receivedAt":"2010-07-05T23:11:10Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"On 7/5/2010 6:04 PM, Finn Arne Gangstad wrote:\n> On Mon, Jul 05, 2010 at 10:16:36AM -0400, Eric Niebler wrote:\n>> I have a question about the best approach to take for refactoring a\n>> large svn project into git. The project, boost.org, is a collection of\n>> C++ libraries (>100) that are mostly independent. (There may be\n>> cross-library dependencies, but we plan to handle that at a higher\n>> level.) After the move to git, we'd like each library to be in its own\n>> git repository. Boost can then be a stitching-together of these, using\n>> submodules or something (opinions welcome). It's an old project with\n>> lots of history that we don't want to lose. The naive approach of simply\n>> forking into N repositories for the N libraries and deleting the\n>> unwanted files in each is unworkable because we'll end up with all the\n>> history duplicated everywhere ... >100 repositories, each larger than 100Mb.\n> \n> If the libraries are not independent (i.e. some commits are across\n> multiple libraries), submodules will give you some interesting\n> challenges to put it mildly.\n\nYou have correctly assessed the situation. There *are* cross-library\ncommits in our history. What are the implications of this for\nmodularlization?\n\n> The current boost 1.43 is 29344 files, is this all there is? \n\nYes.\n\n> This\n> should fit eaily into a single repository. The Linux kernel is much\n> larger, and that is sort of the canonical single repo git project. I\n> _strongly_ recommend that you go for a single repo if you can make it\n> work.\n\nIt does fit into one repo, but that doesn't meet our needs for the\nfuture. Users want to install and build library X and its dependencies,\nnot all of boost. This is increasingly becoming a problem as boost\ngrows. Imagine if a perl programmer had to download all of CPAN to use\nor hack on any one perl module. Or if contributing to CPAN meant getting\nthe whole shebang, history and all. I'm sure even in the Linux kernel,\nnot *every* third-party driver is maintained in the master git repo.\n\nWe are aiming to make boost a clearing-house for C++ libraries (like\nCPAN, or PyPi for python), turning the official boost distribution into\nlittle more than a well-tested collection of the libraries that have\npassed our peer-review and regression test process.\n\nIn fact, the modularization has already been done, and work is well\nunderway on the infrastructure to support dependency tracking. But the\nmodularization is not history-preserving and needs to be redone.\n\n> If you manage to create a single git repo with the history you want,\n> it is trivial to split out separate repositories of subdirectories\n> later (and those repos will then be comparatively small). git subtree\n> allegedly automates this process more or less (I have not used it, but\n> have heard good things about it). What about having a single \"master\n> repository\", and then using subtree to create single-library repos for\n> the library developers if they want a smaller repo to play around in?\n\nThis sounds like it might be ok, but I need to research it.\n\n>> So,, what are the options? Can I somehow delete from each repository the\n>> history that is irrelevant? Is these some feature of git I don't know\n>> about that can solve this problem for us?\n> \n> How do you define \"irrelevant\"? Do you only require enough history for\n> git annotate/blame to give correct results?  Or does this only refer\n> to multiple repositories sharing the same ancient history?\n\nIf multiple repositories share the same ancient history, wouldn't that\ngive git annotate/blame enough information? Sorry, git newbie here.\n\n>> At boost, We've already discussed a few possible approaches. Feel free\n>> to comment and/or criticize any of the solutions suggested here:\n>>\n>>   http://github.com/ryppl/ryppl/issues#issue/4\n> \n> It is unclear from the discussion if you will change to git, or use\n> git in addition to svn? This will have some impact on how to go about\n> this.\n\nThe plan is to move to git. However, we don't expect this to happen\novernight, so a way to continue to pull changes from a svn mirror while\nthe new git repositories are being set up would be ideal.\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144855","messageId":"AANLkTimAqL8gvgIisLpWE6xj2p0jEZD5wetdGYJnOpdr@mail.gmail.com","threadId":"24292","inReplyTo":"4C32668E.9040000@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-05T23:32:17Z","receivedAt":"2010-07-05T23:32:17Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"(note: on this mailing list, you shouldn't drop names from the cc:\nline when replying to a thread)\n\nOn Mon, Jul 5, 2010 at 7:11 PM, Eric Niebler <eric@boostpro.com> wrote:\n> On 7/5/2010 6:04 PM, Finn Arne Gangstad wrote:\n>> This\n>> should fit eaily into a single repository. The Linux kernel is much\n>> larger, and that is sort of the canonical single repo git project. I\n>> _strongly_ recommend that you go for a single repo if you can make it\n>> work.\n>\n> It does fit into one repo, but that doesn't meet our needs for the\n> future. Users want to install and build library X and its dependencies,\n> not all of boost. This is increasingly becoming a problem as boost\n> grows. Imagine if a perl programmer had to download all of CPAN to use\n> or hack on any one perl module. Or if contributing to CPAN meant getting\n> the whole shebang, history and all. I'm sure even in the Linux kernel,\n> not *every* third-party driver is maintained in the master git repo.\n\nActually, that's mostly not true; there are a few third-party drivers\nthat don't make it into the core Linux repo, but that's mostly because\nthey haven't been accepted by the kernel maintainers for whatever\nreason (often quality or duplication, I guess).  The goal for the vast\nmajority of Linux drivers is indeed to get merged into the Linux core.\n\n...and it works pretty well, all things considered.  It's certainly\nnot the only way to do it for every project, but it's actually a\npretty good way.  The kernel repo history runs to hundreds of megs\nnowadays, but on a modern Internet connection that's not a big deal.\nAnd then you never have to worry about downloading more modules later.\n You also never have versioning problems.\n\n> We are aiming to make boost a clearing-house for C++ libraries (like\n> CPAN, or PyPi for python), turning the official boost distribution into\n> little more than a well-tested collection of the libraries that have\n> passed our peer-review and regression test process.\n\nOf course you will want to have some kind of really excellent\nversioned dependency fetching system (exactly like CPAN or PyPi or\nruby gems) if you want this to be nice.  git's submodules stuff is\nalmost certainly not going to add any features you need/want.  On the\nother hand, cloning a separate git repo is pretty easy to write your\nCPAN-like script around.\n\n> In fact, the modularization has already been done, and work is well\n> underway on the infrastructure to support dependency tracking. But the\n> modularization is not history-preserving and needs to be redone.\n\nIf your code doesn't move too many files around, then splitting out\nthe history is pretty easy with git-subtree (a tool I wrote that's not\npart of git):\n\n   git subtree split --prefix=/path/to/subdir\n\nAnd you get a new history for just that subdir.  That might do exactly\nwhat you want.  It also works iteratively, so you can export your\nhistory from svn, then re-export the changes as they occur over time.\n\n>>> So,, what are the options? Can I somehow delete from each repository the\n>>> history that is irrelevant? Is these some feature of git I don't know\n>>> about that can solve this problem for us?\n>>\n>> How do you define \"irrelevant\"? Do you only require enough history for\n>> git annotate/blame to give correct results?  Or does this only refer\n>> to multiple repositories sharing the same ancient history?\n>\n> If multiple repositories share the same ancient history, wouldn't that\n> give git annotate/blame enough information? Sorry, git newbie here.\n\nYes, it would.  But how much of the ancient history do you want?  If\nyou want all of it, you don't save any space in your repo.\n\n> The plan is to move to git. However, we don't expect this to happen\n> overnight, so a way to continue to pull changes from a svn mirror while\n> the new git repositories are being set up would be ideal.\n\nThis isn't too hard to do; you just need some scripts around git-svn\nand git-subtree (or whatever tool you use to do the splitting).  We've\ndone this at work for a couple of years now and it's working fine.\n\nThe confusing part is taking *submissions* back through both channels.\n If you value your sanity, you probably want to only allow submissions\nback via svn while you're running the two in parallel; but that makes\ngit's added features a lot less useful, so you probably want to run in\nparallel for only a short time.\n\nHave fun,\n\nAvery\n"},{"id":"144856","messageId":"4C3275C0.8000406@boostpro.com","threadId":"24292","inReplyTo":"AANLkTimAqL8gvgIisLpWE6xj2p0jEZD5wetdGYJnOpdr@mail.gmail.com","subject":"Re: help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-06T00:16:00Z","receivedAt":"2010-07-06T00:16:00Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"On 7/5/2010 7:32 PM, Avery Pennarun wrote:\n> (note: on this mailing list, you shouldn't drop names from the cc:\n> line when replying to a thread)\n\nNoted, thanks.\n\n> On Mon, Jul 5, 2010 at 7:11 PM, Eric Niebler <eric@boostpro.com> wrote:\n>> On 7/5/2010 6:04 PM, Finn Arne Gangstad wrote:\n>>> This\n>>> should fit eaily into a single repository. The Linux kernel is much\n>>> larger, and that is sort of the canonical single repo git project. I\n>>> _strongly_ recommend that you go for a single repo if you can make it\n>>> work.\n>>\n>> It does fit into one repo, but that doesn't meet our needs for the\n>> future. Users want to install and build library X and its dependencies,\n>> not all of boost. This is increasingly becoming a problem as boost\n>> grows. Imagine if a perl programmer had to download all of CPAN to use\n>> or hack on any one perl module. Or if contributing to CPAN meant getting\n>> the whole shebang, history and all. I'm sure even in the Linux kernel,\n>> not *every* third-party driver is maintained in the master git repo.\n> \n> Actually, that's mostly not true; there are a few third-party drivers\n> that don't make it into the core Linux repo\n<snip discussion showing my ignorance of Linux's repository structure>\n\nThanks for the correction. The CPAN/PyPi analogy is still apt.\n\n>> We are aiming to make boost a clearing-house for C++ libraries (like\n>> CPAN, or PyPi for python), turning the official boost distribution into\n>> little more than a well-tested collection of the libraries that have\n>> passed our peer-review and regression test process.\n> \n> Of course you will want to have some kind of really excellent\n> versioned dependency fetching system (exactly like CPAN or PyPi or\n> ruby gems) if you want this to be nice.  git's submodules stuff is\n> almost certainly not going to add any features you need/want.  On the\n> other hand, cloning a separate git repo is pretty easy to write your\n> CPAN-like script around.\n\nIndeed, we are stealing the work of the python guys. Pip does most of\nwhat we want. They've graciously been accepting our patches so it\nhappily clones git repos in order to satisfy dependencies now. It is\nsome kind of really excellent! :-)\n\n>> In fact, the modularization has already been done, and work is well\n>> underway on the infrastructure to support dependency tracking. But the\n>> modularization is not history-preserving and needs to be redone.\n> \n> If your code doesn't move too many files around, then splitting out\n> the history is pretty easy with git-subtree (a tool I wrote that's not\n> part of git):\n> \n>    git subtree split --prefix=/path/to/subdir\n> \n> And you get a new history for just that subdir.  That might do exactly\n> what you want.  It also works iteratively, so you can export your\n> history from svn, then re-export the changes as they occur over time.\n\nThis looks like it here:\n\n  http://github.com/apenwarr/git-subtree\n\nI'll have to read the docs. Thanks for the tip.\n\n>>>> So,, what are the options? Can I somehow delete from each repository the\n>>>> history that is irrelevant? Is these some feature of git I don't know\n>>>> about that can solve this problem for us?\n>>>\n>>> How do you define \"irrelevant\"? Do you only require enough history for\n>>> git annotate/blame to give correct results?  Or does this only refer\n>>> to multiple repositories sharing the same ancient history?\n>>\n>> If multiple repositories share the same ancient history, wouldn't that\n>> give git annotate/blame enough information? Sorry, git newbie here.\n> \n> Yes, it would.  But how much of the ancient history do you want?  If\n> you want all of it, you don't save any space in your repo.\n\nRepos, plural. We'd save space because the history wouldn't be\nduplicated in each one. Right? Or else I'm confused and this something\nthat will become clear after I understand what git subtree does.\n\nRight now, the other boost developers are pushing for a solution that\nuses grafts. I'm fuzzy on what they are exactly, but it seems that we'd\nfreeze a svn mirror and have anybody interested in history put grafts in\ntheir local repository pointing back at the mirror. I don't know enough\nyet to say what the pros/cons of this approach might be wrt git subtree.\n\n>> The plan is to move to git. However, we don't expect this to happen\n>> overnight, so a way to continue to pull changes from a svn mirror while\n>> the new git repositories are being set up would be ideal.\n> \n> This isn't too hard to do; you just need some scripts around git-svn\n> and git-subtree (or whatever tool you use to do the splitting).  We've\n> done this at work for a couple of years now and it's working fine.\n\nCool.\n\n> The confusing part is taking *submissions* back through both channels.\n> If you value your sanity, you probably want to only allow submissions\n> back via svn while you're running the two in parallel; but that makes\n> git's added features a lot less useful, so you probably want to run in\n> parallel for only a short time.\n\nOh my! I don't think we'd open the git repositories for changes until\nafter we close down svn. This problem is hard enough.\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144857","messageId":"rmimxu5bh2s.fsf@fnord.ir.bbn.com","threadId":"24292","inReplyTo":"4C31E944.30801@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Greg Troxel","fromEmail":"gdt@ir.bbn.com","sentAt":"2010-07-06T00:16:11Z","receivedAt":"2010-07-06T00:16:11Z","isPatch":false,"sender":{"key":"gdt@ir.bbn.com","avatar":null},"body":"\nYou have found the core issue with svn/git: svn allows you to have a\nlarge repo with everything (and atomic commits across it) and to have\nusers check out parts of the repo separately.  git does not, because the\nsvn separate checkouts model only works with a remote repository that\nyou don't keep a copy of.  With git, cloning the repo gets you the whole\nthing.\n\nOne thought is that you may want to separate how you organize boost\nsources in git and how you release them.  It's possible to have a single\ngit repo for all libraries and have atomic commits but then create\ndistfiles for each library separately.\n\ngit becomes a bit slow when repositories get really large (although\nother tools are not any better - the problem is in the sheer number of\nvnode ops necessary for the semantics).  I have a repo with all of\nNetBSD's \"src\", \"xsrc\" and \"pkgsrc\" in it, and \"git status\" can take\nseveral seconds because it is calling stat on 230K files.  With only 23K\nfiles, things should be ok.\n\nMy advice (which is not really about git) is to figure out whether you\nwant:\n\n  A) a set of interrelated libraries on which you will allow atomic\n  commits that change interfaces/usage in multiple libraries\n\nor\n\n  B) a set of independent libaries which have commits to separate\n  libraries, and for which you insist that each library have an API and\n  ABI compatiblity story, so that even when upgraded other libraries can\n  continue to use it.\n\n\nFor A, you probably want one git repo, much as you have one svn repo\nnow.  For B, multiple git repos are the right answer.\n"},{"id":"144858","messageId":"4C3277E9.3060800@boostpro.com","threadId":"24292","inReplyTo":"rmimxu5bh2s.fsf@fnord.ir.bbn.com","subject":"Re: help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-06T00:25:13Z","receivedAt":"2010-07-06T00:25:13Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"On 7/5/2010 8:16 PM, Greg Troxel wrote:\n> \n> You have found the core issue with svn/git: svn allows you to have a\n> large repo with everything (and atomic commits across it) and to have\n> users check out parts of the repo separately.  git does not, because the\n> svn separate checkouts model only works with a remote repository that\n> you don't keep a copy of.  With git, cloning the repo gets you the whole\n> thing.\n\nMakes sense.\n\n> One thought is that you may want to separate how you organize boost\n> sources in git and how you release them.  It's possible to have a single\n> git repo for all libraries and have atomic commits but then create\n> distfiles for each library separately.\n> \n> git becomes a bit slow when ...\n<snip>\n\nIt can't get any worse than svn. We haven't run into any perf problems\nwith git yet. That's not our primary concern.\n\n> My advice (which is not really about git) is to figure out whether you\n> want:\n> \n>   A) a set of interrelated libraries on which you will allow atomic\n>   commits that change interfaces/usage in multiple libraries\n> \n> or\n> \n>   B) a set of independent libaries which have commits to separate\n>   libraries, and for which you insist that each library have an API and\n>   ABI compatiblity story, so that even when upgraded other libraries can\n>   continue to use it.\n> \n> \n> For A, you probably want one git repo, much as you have one svn repo\n> now.  For B, multiple git repos are the right answer.\n\nI'll take B FTW! :-) The idea is to open up and distribute C++ library\ndevelopment. Versioned dependency tracking will be handled at a higher\nlevel with per-project metadata and a tool (pip) that resolves\ndependencies. API compatibility is handled with peer review and\nregression testing. ABI compatibility is not an issue because we're\ndistributing source code.\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144859","messageId":"loom.20100706T034440-190@post.gmane.org","threadId":"24292","inReplyTo":"4C32668E.9040000@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Dave Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2010-07-06T01:46:56Z","receivedAt":"2010-07-06T01:46:56Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"Eric Niebler <eric <at> boostpro.com> writes:\n\n> We are aiming to make boost a clearing-house for C++ libraries (like\n> CPAN, or PyPi for python), \n\nClarification: that's our goal for Ryppl, not Boost.\n\n> turning the official boost distribution into\n> little more than a well-tested collection of the libraries that have\n> passed our peer-review and regression test process.\n\nExactly right.\n\n--\nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144888","messageId":"m3ocelm1tj.fsf@localhost.localdomain","threadId":"24292","inReplyTo":"loom.20100706T034440-190@post.gmane.org","subject":"Re: help moving boost.org to git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-07-06T08:51:00Z","receivedAt":"2010-07-06T08:51:00Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dave Abrahams <dave@boostpro.com> writes:\n\n> Eric Niebler <eric <at> boostpro.com> writes:\n> \n> > We are aiming to make boost a clearing-house for C++ libraries (like\n> > CPAN, or PyPi for python), \n> \n> Clarification: that's our goal for Ryppl, not Boost.\n\nBy the way, could you please add information about Ryppl to Git Wiki?\nhttps://git.wiki.kernel.org/index.php/InterfacesFrontendsAndTools\n\nThanks in advance\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"144894","messageId":"m2iq4sj3us.wl%dave@boostpro.com","threadId":"24292","inReplyTo":"m3ocelm1tj.fsf@localhost.localdomain","subject":"Re: help moving boost.org to git","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2010-07-06T10:34:35Z","receivedAt":"2010-07-06T10:34:35Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"At Tue, 06 Jul 2010 01:51:00 -0700 (PDT),\nJakub Narebski wrote:\n> \n> Dave Abrahams <dave@boostpro.com> writes:\n> \n> > Eric Niebler <eric <at> boostpro.com> writes:\n> > \n> > > We are aiming to make boost a clearing-house for C++ libraries (like\n> > > CPAN, or PyPi for python), \n> > \n> > Clarification: that's our goal for Ryppl, not Boost.\n> \n> By the way, could you please add information about Ryppl to Git Wiki?\n> https://git.wiki.kernel.org/index.php/InterfacesFrontendsAndTools\n\nWe'd be happy to, but: have you read the document at ryppl.org, and do\nyou really think it's appropriate?  Even though ryppl is still very\nalpha?\n\n[BTW, I can't reach that page right now; it's just forever “waiting\nfor reply...”]\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144938","messageId":"878w5o7iq7.fsf@hariville.hurrynot.org","threadId":"24292","inReplyTo":"4C31F0D4.1040207@viscovery.net","subject":"Re: help moving boost.org to git","fromName":"Raja R Harinath","fromEmail":"harinath@hurrynot.org","sentAt":"2010-07-06T15:06:24Z","receivedAt":"2010-07-06T15:06:24Z","isPatch":false,"sender":{"key":"harinath@hurrynot.org","avatar":"https://avatars.githubusercontent.com/u/4610?v=4"},"body":"Hi,\n\nJohannes Sixt <j.sixt@viscovery.net> writes:\n\n> Am 7/5/2010 16:16, schrieb Eric Niebler:\n>> I have a question about the best approach to take for refactoring a\n>> large svn project into git. The project, boost.org, is a collection of\n>> C++ libraries (>100) that are mostly independent. (There may be\n>> cross-library dependencies, but we plan to handle that at a higher\n>> level.) After the move to git, we'd like each library to be in its own\n>> git repository.\n>\n> You could use svn2git: http://gitorious.org/svn2git\n> KDE uses it to split its SVN repository into pieces. The tool is driven by\n> a \"ruleset\" that specifies SVN subdirectories and revision numbers that\n> make up a module.\n\nI'm also involved in moving a large SVN project to git (the mono\nproject).  I have found and fixed several issues with svn2git\n\n  git://gitorious.org/~harinath/svn2git/rrh-svn2git.git\n\n- Hari\n"},{"id":"144923","messageId":"AANLkTikkKhvzsczKJwjsc0kmCmWQGAIUzc__Wr20Dbwd@mail.gmail.com","threadId":"24292","inReplyTo":"4C3275C0.8000406@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-06T17:27:56Z","receivedAt":"2010-07-06T17:27:56Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Mon, Jul 5, 2010 at 8:16 PM, Eric Niebler <eric@boostpro.com> wrote:\n> On 7/5/2010 7:32 PM, Avery Pennarun wrote:\n>> Eric Niebler wrote:\n>>> If multiple repositories share the same ancient history, wouldn't that\n>>> give git annotate/blame enough information? Sorry, git newbie here.\n>>\n>> Yes, it would.  But how much of the ancient history do you want?  If\n>> you want all of it, you don't save any space in your repo.\n>\n> Repos, plural. We'd save space because the history wouldn't be\n> duplicated in each one. Right? Or else I'm confused and this something\n> that will become clear after I understand what git subtree does.\n\nThe statement \"multiple repositories share the same ancient history\"\nabove is the part that's confusing.  If you use a tool like\ngit-subtree or git-filter-branch, you're actually generating a \"new\nhistory\" based on the original history.  The \"new history\" obviously\ncontains fewer files than the original, which would take less space.\nBut if you want multiple repositories to \"share the same ancient\nhistory\" you can't rewrite it, and thus you aren't saving any space in\nany one repo.\n\nI'm assuming you want to rewrite history to save space (since that's\nwhat this thread is about).  And git annotate/blame will work as long\nas your rewritten history contains all the files you care about in\nthat repo.\n\n> Right now, the other boost developers are pushing for a solution that\n> uses grafts. I'm fuzzy on what they are exactly, but it seems that we'd\n> freeze a svn mirror and have anybody interested in history put grafts in\n> their local repository pointing back at the mirror. I don't know enough\n> yet to say what the pros/cons of this approach might be wrt git subtree.\n\nThe primary advantage of grafts is that you can do something easy\n*right now* and then fix it all up later.  eg. if you screw up your\nhistory extraction and do it better later, you can just re-graft it\nand you're done.\n\nA secondary advantage of grafts is that cloning the \"primary\"\nrepository will be tiny since it doesn't have much ancient history.\n\nA disadvantage of grafts is that each user has to deal with grafts in\nhis cloned repo, and unless he does, things like 'git log' and 'git\nblame' won't show anything from the grafted history.  Supposedly 'git\nreplace' was designed to help with this issue, but I've never used it\nso I don't know for sure.\n\nAnd of course, grafts don't actually do any history rewriting for you.\n You could split out a subtree's history and then graft it on, but the\nsplitting process is still the same as it would be without grafts.\nThe alternative would be to *not* rewrite history, just keep the\nentire history of the whole project in one place, and graft it on if\nyou really need it.  That's actually pretty clean (and accurately\nreflects exactly what *really happened*, which is a nice feature to\nhave in a vcs history), but you'll then never have a single repo of\njust one subproject with the entire history of that subproject.  That\nlatter turns out to not actually be very important in practice, so you\nmight want to do it.\n\n>> The confusing part is taking *submissions* back through both channels.\n>> If you value your sanity, you probably want to only allow submissions\n>> back via svn while you're running the two in parallel; but that makes\n>> git's added features a lot less useful, so you probably want to run in\n>> parallel for only a short time.\n>\n> Oh my! I don't think we'd open the git repositories for changes until\n> after we close down svn. This problem is hard enough.\n\nIt can be done, and I've done it :)  But you're wise to avoid that situation.\n\nHave fun,\n\nAvery\n"},{"id":"144933","messageId":"4C336F3D.1010906@boostpro.com","threadId":"24292","inReplyTo":"AANLkTikkKhvzsczKJwjsc0kmCmWQGAIUzc__Wr20Dbwd@mail.gmail.com","subject":"Re: help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-06T18:00:29Z","receivedAt":"2010-07-06T18:00:29Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"On 7/6/2010 1:27 PM, Avery Pennarun wrote:\n> On Mon, Jul 5, 2010 at 8:16 PM, Eric Niebler <eric@boostpro.com> wrote:\n>> On 7/5/2010 7:32 PM, Avery Pennarun wrote:\n>>> Eric Niebler wrote:\n>>>> If multiple repositories share the same ancient history, wouldn't that\n>>>> give git annotate/blame enough information? Sorry, git newbie here.\n>>>\n>>> Yes, it would.  But how much of the ancient history do you want?  If\n>>> you want all of it, you don't save any space in your repo.\n>>\n>> Repos, plural. We'd save space because the history wouldn't be\n>> duplicated in each one. Right? Or else I'm confused and this something\n>> that will become clear after I understand what git subtree does.\n> \n> The statement \"multiple repositories share the same ancient history\"\n> above is the part that's confusing.  If you use a tool like\n> git-subtree or git-filter-branch, you're actually generating a \"new\n> history\" based on the original history.  The \"new history\" obviously\n> contains fewer files than the original, which would take less space.\n> But if you want multiple repositories to \"share the same ancient\n> history\" you can't rewrite it, and thus you aren't saving any space in\n> any one repo.\n\nI think I have reached understanding! Thank you. It *would* save if I\npull down, say, 100 of these new repos+ancient history because git would\njust store the ancient history locally once. I'm also guessing git is\nsmart enough to avoid /downloading/ the ancient history 100x.\n\n> I'm assuming you want to rewrite history to save space (since that's\n> what this thread is about).  And git annotate/blame will work as long\n> as your rewritten history contains all the files you care about in\n> that repo.\n\nRight. I now understand that, too.\n\n>> Right now, the other boost developers are pushing for a solution that\n>> uses grafts. I'm fuzzy on what they are exactly, but it seems that we'd\n>> freeze a svn mirror and have anybody interested in history put grafts in\n>> their local repository pointing back at the mirror. I don't know enough\n>> yet to say what the pros/cons of this approach might be wrt git subtree.\n> \n> The primary advantage of grafts is that you can do something easy\n> *right now* and then fix it all up later.  eg. if you screw up your\n> history extraction and do it better later, you can just re-graft it\n> and you're done.\n\nHow does one screw up the history extraction, if one is not doing any\nfancy history rewriting (in this scenario)? Be there dragons?\n\n> A secondary advantage of grafts is that cloning the \"primary\"\n> repository will be tiny since it doesn't have much ancient history.\n\nRight. Only those who ask for it will pay for it. And only developers\nwill have need of it, and not all developers at that.\n\n> A disadvantage of grafts is that each user has to deal with grafts in\n> his cloned repo, and unless he does, things like 'git log' and 'git\n> blame' won't show anything from the grafted history.  Supposedly 'git\n> replace' was designed to help with this issue, but I've never used it\n> so I don't know for sure.\n\nI'll add it to the list of things to learn about.\n\n> And of course, grafts don't actually do any history rewriting for you.\n> You could split out a subtree's history and then graft it on, but the\n> splitting process is still the same as it would be without grafts.\n> The alternative would be to *not* rewrite history, just keep the\n> entire history of the whole project in one place, and graft it on if\n> you really need it.  That's actually pretty clean (and accurately\n> reflects exactly what *really happened*, which is a nice feature to\n> have in a vcs history), but you'll then never have a single repo of\n> just one subproject with the entire history of that subproject.  That\n> latter turns out to not actually be very important in practice, so you\n> might want to do it.\n\nThat's starting to sound pretty good.\n\nThanks,\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"144935","messageId":"AANLkTikfTFw_UdV1ia58MbWxH4h8TJAr-Y5WPvlXCjeJ@mail.gmail.com","threadId":"24292","inReplyTo":"4C336F3D.1010906@boostpro.com","subject":"Re: help moving boost.org to git","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-07-06T18:13:47Z","receivedAt":"2010-07-06T18:13:47Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Tue, Jul 6, 2010 at 2:00 PM, Eric Niebler <eric@boostpro.com> wrote:\n> On 7/6/2010 1:27 PM, Avery Pennarun wrote:\n>> The primary advantage of grafts is that you can do something easy\n>> *right now* and then fix it all up later.  eg. if you screw up your\n>> history extraction and do it better later, you can just re-graft it\n>> and you're done.\n>\n> How does one screw up the history extraction, if one is not doing any\n> fancy history rewriting (in this scenario)? Be there dragons?\n\nWell, \"rewriting history\" necessarily involves changing things about\nthe permanent record.  Every time you change things, you have a\npotential to change them incorrectly.  So in general, not rewriting is\nless error-prone than rewriting :)\n\nSpecifically, with a tool like git-subtree, it only really works if a\nparticular subproject has always existed in the same subdir of your\nrepo since it started.  If the subdir was ever renamed, or if some of\nthe files were previously part of one subdir but then moved around,\ngit-subtree doesn't (currently) know how to deal with that.\ngit-filter-branch can do anything you want, but you have to teach it\nhow, which is obviously even *more* error prone.\n\nThings are also a little messy if you have some kind of top-level\ndirectory with build infrastructure shared by all the subdirs.  Does\nthe top-level Makefile have a list of the subdirs it needs to build?\nIf so, there's no way to extract only a subset of true history that\nwill still build correctly - it'll be looking for directories that you\nexplicitly removed.  You could update the Makefiles programmatically\nin every single revision, but that's starting to get extremely\nmessy... and your history stops representing what *real life* really\nlooked like at the time.\n\nIf your subdirs haven't been moving around (which sounds like that\nmight be the case for you), and you don't have any top-level files\nthat you care about, rewriting might turn out to be straightforward.\nYou could also make the decision on a subdir-by-subdir basis, I guess.\n\nHave fun,\n\nAvery\n"},{"id":"144937","messageId":"4C33760D.9000404@boostpro.com","threadId":"24292","inReplyTo":"AANLkTikfTFw_UdV1ia58MbWxH4h8TJAr-Y5WPvlXCjeJ@mail.gmail.com","subject":"Re: help moving boost.org to git","fromName":"Eric Niebler","fromEmail":"eric@boostpro.com","sentAt":"2010-07-06T18:29:33Z","receivedAt":"2010-07-06T18:29:33Z","isPatch":false,"sender":{"key":"eric@boostpro.com","avatar":"https://gravatar.com/avatar/046c71b2ddb07b6d6e18992a12406654941dcf003b2ba762e6cfb1309243f8cc?d=mp&s=160"},"body":"On 7/6/2010 2:13 PM, Avery Pennarun wrote:\n<snip>\n> Specifically, with a tool like git-subtree, it only really works if a\n> particular subproject has always existed in the same subdir of your\n> repo since it started.  If the subdir was ever renamed, or if some of\n> the files were previously part of one subdir but then moved around,\n> git-subtree doesn't (currently) know how to deal with that.\n\nBah! Yes, directories have moved around in our svn repro. :-( In\nparticular, we've had cases where libraries in boost began life as\nsub-projects of a different library and then got spun off.\n\n> git-filter-branch can do anything you want, but you have to teach it\n> how, which is obviously even *more* error prone.\n\nI can only imagine.\n\n> Things are also a little messy if you have some kind of top-level\n> directory with build infrastructure shared by all the subdirs.  Does\n> the top-level Makefile have a list of the subdirs it needs to build?\n\nBah! Yes, the build, the docs and the test infrastructure all currently\nshare files across our submodules-to-be. Surely other projects have\nencountered this problem before, right? (KDE, I'm looking in your\ndirection.)\n\n> If so, there's no way to extract only a subset of true history that\n> will still build correctly - it'll be looking for directories that you\n> explicitly removed.  You could update the Makefiles programmatically\n> in every single revision, but that's starting to get extremely\n> messy... and your history stops representing what *real life* really\n> looked like at the time.\n\nI see what you mean.\n\n> If your subdirs haven't been moving around (which sounds like that\n> might be the case for you), and you don't have any top-level files\n> that you care about, rewriting might turn out to be straightforward.\n> You could also make the decision on a subdir-by-subdir basis, I guess.\n\nMore evidence that the fancy filter/branch/subtree/svn2git/whatever\nutilities aren't going to get us where we'd like to be. A simple\nconversion and grafts look like the only workable approach.\n\n> Have fun,\n\nHaving heaps! Thanks,\n\n-- \nEric Niebler\nBoostPro Computing\nhttp://www.boostpro.com\n"}]}