git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: help moving boost.org to git

From
Eric Niebler <eric@boostpro.com>
Date
Jul 5, 2010, 23:11 UTC
Message-ID
<4C32668E.9040000@boostpro.com>
In-Reply-To
<20100705220443.GA23727@pvv.org>
On 7/5/2010 6:04 PM, Finn Arne Gangstad wrote:
Show 16 quoted lines
> On Mon, Jul 05, 2010 at 10:16:36AM -0400, Eric Niebler wrote:
>> I have a question about the best approach to take for refactoring a
>> large svn project into git. The project, boost.org, is a collection of
>> C++ libraries (>100) that are mostly independent. (There may be
>> cross-library dependencies, but we plan to handle that at a higher
>> level.) After the move to git, we'd like each library to be in its own
>> git repository. Boost can then be a stitching-together of these, using
>> submodules or something (opinions welcome). It's an old project with
>> lots of history that we don't want to lose. The naive approach of simply
>> forking into N repositories for the N libraries and deleting the
>> unwanted files in each is unworkable because we'll end up with all the
>> history duplicated everywhere ... >100 repositories, each larger than 100Mb.
> 
> If the libraries are not independent (i.e. some commits are across
> multiple libraries), submodules will give you some interesting
> challenges to put it mildly.

You have correctly assessed the situation. There *are* cross-library commits in our history. What are the implications of this for modularlization?

> The current boost 1.43 is 29344 files, is this all there is? 
Yes.
Show 5 quoted lines
> This
> should fit eaily into a single repository. The Linux kernel is much
> larger, and that is sort of the canonical single repo git project. I
> _strongly_ recommend that you go for a single repo if you can make it
> work.

It does fit into one repo, but that doesn't meet our needs for the future. Users want to install and build library X and its dependencies, not all of boost. This is increasingly becoming a problem as boost grows. Imagine if a perl programmer had to download all of CPAN to use or hack on any one perl module. Or if contributing to CPAN meant getting the whole shebang, history and all. I'm sure even in the Linux kernel, not *every* third-party driver is maintained in the master git repo.

We are aiming to make boost a clearing-house for C++ libraries (like CPAN, or PyPi for python), turning the official boost distribution into little more than a well-tested collection of the libraries that have passed our peer-review and regression test process.

In fact, the modularization has already been done, and work is well underway on the infrastructure to support dependency tracking. But the modularization is not history-preserving and needs to be redone.

Show 7 quoted lines
> If you manage to create a single git repo with the history you want,
> it is trivial to split out separate repositories of subdirectories
> later (and those repos will then be comparatively small). git subtree
> allegedly automates this process more or less (I have not used it, but
> have heard good things about it). What about having a single "master
> repository", and then using subtree to create single-library repos for
> the library developers if they want a smaller repo to play around in?
This sounds like it might be ok, but I need to research it.
Show 7 quoted lines
>> So,, what are the options? Can I somehow delete from each repository the
>> history that is irrelevant? Is these some feature of git I don't know
>> about that can solve this problem for us?
> 
> How do you define "irrelevant"? Do you only require enough history for
> git annotate/blame to give correct results?  Or does this only refer
> to multiple repositories sharing the same ancient history?

If multiple repositories share the same ancient history, wouldn't that give git annotate/blame enough information? Sorry, git newbie here.

Show 8 quoted lines
>> At boost, We've already discussed a few possible approaches. Feel free
>> to comment and/or criticize any of the solutions suggested here:
>>
>>   http://github.com/ryppl/ryppl/issues#issue/4
> 
> It is unclear from the discussion if you will change to git, or use
> git in addition to svn? This will have some impact on how to go about
> this.

The plan is to move to git. However, we don't expect this to happen overnight, so a way to continue to pull changes from a svn mirror while the new git repositories are being set up would be ideal.

-- 
Eric Niebler
BoostPro Computing
http://www.boostpro.com
Previous: Finn Arne GangstadNext: Avery Pennarun
Message 8 of 19 in “help moving boost.org to git”
  1. Eric NieblerJul 5, 2010
  2. Erik Faye-LundJul 5, 2010
  3. Johannes SixtJul 5, 2010
  4. Eric NieblerJul 5, 2010
  5. Sverre RabbelierJul 5, 2010
  6. Raja R HarinathJul 6, 2010
  7. Finn Arne GangstadJul 5, 2010
  8. Eric NieblerJul 5, 2010
  9. Avery PennarunJul 5, 2010
  10. Eric NieblerJul 6, 2010
  11. Avery PennarunJul 6, 2010
  12. Eric NieblerJul 6, 2010
  13. Avery PennarunJul 6, 2010
  14. Eric NieblerJul 6, 2010
  15. Dave AbrahamsJul 6, 2010
  16. Jakub NarebskiJul 6, 2010
  17. David AbrahamsJul 6, 2010
  18. Greg TroxelJul 6, 2010
  19. Eric NieblerJul 6, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.