# Joining historical repository using grafts or replace

8 messages from 2014-10-30 to 2014-11-01. Participants: Dmitry Oksenchuk, W. Trevor King, Christian Couder.
Thread: https://gitlist.dev/t/37845

## Dmitry Oksenchuk, 2014-10-30 15:39

Subject: Joining historical repository using grafts or replace
Message-ID: <CA+POfmvCiNBF=P-OvQBTROVhaLtOdgNTDgPNyS=97bupSGk=4g@mail.gmail.com>
URL: https://gitlist.dev/e/CA%2BPOfmvCiNBF%3DP-OvQBTROVhaLtOdgNTDgPNyS%3D97bupSGk%3D4g%40mail.gmail.com

```
Hello,

We're in the middle of conversion of a large CVS repository (20 years,
70K commits, 1K branches, 10K tags) to Git and considering two
separate Git repositories: "historical" with CVS history and "working"
created without history from heads of active branches (10 active
branches). This allows us to have small fast "working" repository for
developers who don't want to have full history locally and ability to
rewrite history in "historical" repository (for example, to add
parents to merge commits or to fix conversion mistakes) without
affecting commit hashes in "working" repository (the hashes can be
stored in bug tracker or in the code).

The first idea was to use grafs to join branch roots in "working"
repository with branches in "historical" repository like in linux
repository but it seems that grafts are known as a "horrible hack" (
http://marc.info/?l=git&m=131127600030310&w=2
http://permalink.gmane.org/gmane.comp.version-control.git/177153 )

Since Git 1.6.5 "replace" can also be used to join the histories by
replacing branch roots in "working" repository with branch heads in
"historical" repository.

Both grafts and replace will be used locally. Grafts is a bit easier
to distribute (simple copying, replaces should be created via bash
script).

Are there any disadvantages of using grafts and replace? Will both of
them be supported in future versions of Git?

Thank you,
Dmitry

```

## W. Trevor King, 2014-10-30 15:44

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <20141030154426.GU15443@odin.tremily.us>
URL: https://gitlist.dev/e/20141030154426.GU15443%40odin.tremily.us
In-Reply-To: <CA+POfmvCiNBF=P-OvQBTROVhaLtOdgNTDgPNyS=97bupSGk=4g@mail.gmail.com>

```
On Thu, Oct 30, 2014 at 06:39:56PM +0300, Dmitry Oksenchuk wrote:
> We're in the middle of conversion of a large CVS repository (20
> years, 70K commits, 1K branches, 10K tags) to Git and considering
> two separate Git repositories: "historical" with CVS history and
> "working" created without history from heads of active branches (10
> active branches). This allows us to have small fast "working"
> repository for developers who don't want to have full history
> locally and ability to rewrite history in "historical" repository
> (for example, to add parents to merge commits or to fix conversion
> mistakes) without affecting commit hashes in "working" repository
> (the hashes can be stored in bug tracker or in the code).

A number of projects have done something like this (e.g. Linux).
Modern Gits have good support for shallow repositories though, so I'd
just make one full repository and leave it to developers to decide how
deep they want their local copy to be.

Cheers,
Trevor

-- 
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy

```

## Christian Couder, 2014-10-30 16:54

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <CAP8UFD3_fAWRdxQgAbfxYZSzrmy1Aza=nuZh-uSJsKOdRj+LVA@mail.gmail.com>
URL: https://gitlist.dev/e/CAP8UFD3_fAWRdxQgAbfxYZSzrmy1Aza%3DnuZh-uSJsKOdRj%2BLVA%40mail.gmail.com
In-Reply-To: <CA+POfmvCiNBF=P-OvQBTROVhaLtOdgNTDgPNyS=97bupSGk=4g@mail.gmail.com>

```
Hi,

On Thu, Oct 30, 2014 at 4:39 PM, Dmitry Oksenchuk <oksenchuk89@gmail.com> wrote:
> Hello,
>
> We're in the middle of conversion of a large CVS repository (20 years,
> 70K commits, 1K branches, 10K tags) to Git and considering two
> separate Git repositories: "historical" with CVS history and "working"
> created without history from heads of active branches (10 active
> branches). This allows us to have small fast "working" repository for
> developers who don't want to have full history locally and ability to
> rewrite history in "historical" repository (for example, to add
> parents to merge commits or to fix conversion mistakes) without
> affecting commit hashes in "working" repository (the hashes can be
> stored in bug tracker or in the code).

This might be a good idea. Did you already test that the small
repository is really faster than the full repository?

> The first idea was to use grafs to join branch roots in "working"
> repository with branches in "historical" repository like in linux
> repository but it seems that grafts are known as a "horrible hack" (
> http://marc.info/?l=git&m=131127600030310&w=2
> http://permalink.gmane.org/gmane.comp.version-control.git/177153 )
>
> Since Git 1.6.5 "replace" can also be used to join the histories by
> replacing branch roots in "working" repository with branch heads in
> "historical" repository.
>
> Both grafts and replace will be used locally. Grafts is a bit easier
> to distribute (simple copying, replaces should be created via bash
> script).

First, you might want to have a look at:

http://git-scm.com/book/en/v2/Git-Tools-Replace

as it looks like it describes your use case very well.

> Are there any disadvantages of using grafts and replace? Will both of
> them be supported in future versions of Git?

My opinion is that grafts have no advantage compared to replace refs.

Once you have created your replace refs, they can be managed like
other git refs, so they are easier to distribute.

Basically if you want to get the full history on a computer you just need to do:

git fetch 'refs/replace/*:refs/replace/*'

Best,
Christian.

```

## Dmitry Oksenchuk, 2014-10-30 17:41

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <CA+POfmvXEjDV9Vap6NDX7HvOMjEVG4mVe1uWFSTQy5g_c+vJnw@mail.gmail.com>
URL: https://gitlist.dev/e/CA%2BPOfmvXEjDV9Vap6NDX7HvOMjEVG4mVe1uWFSTQy5g_c%2BvJnw%40mail.gmail.com
In-Reply-To: <CAP8UFD3_fAWRdxQgAbfxYZSzrmy1Aza=nuZh-uSJsKOdRj+LVA@mail.gmail.com>

```
Hi Christian,

Thanks for your reply.

2014-10-30 19:54 GMT+03:00 Christian Couder <christian.couder@gmail.com>:
> On Thu, Oct 30, 2014 at 4:39 PM, Dmitry Oksenchuk <oksenchuk89@gmail.com> wrote:
>> We're in the middle of conversion of a large CVS repository (20 years,
>> 70K commits, 1K branches, 10K tags) to Git and considering two
>> separate Git repositories: "historical" with CVS history and "working"
>> created without history from heads of active branches (10 active
>> branches). This allows us to have small fast "working" repository for
>> developers who don't want to have full history locally and ability to
>> rewrite history in "historical" repository (for example, to add
>> parents to merge commits or to fix conversion mistakes) without
>> affecting commit hashes in "working" repository (the hashes can be
>> stored in bug tracker or in the code).
>
> This might be a good idea. Did you already test that the small
> repository is really faster than the full repository?

Yes, because of such amount of refs, push in "historical" repository
takes 12 sec, push in "working" repository takes 0.4 sec, push in
"joined" repository takes 2 sec. Local operations with history like
log and blame work with the same speed in "joined" repository as in
"historical" repository.

>> Are there any disadvantages of using grafts and replace? Will both of
>> them be supported in future versions of Git?
>
> My opinion is that grafts have no advantage compared to replace refs.
>
> Once you have created your replace refs, they can be managed like
> other git refs, so they are easier to distribute.
>
> Basically if you want to get the full history on a computer you just need to do:
>
> git fetch 'refs/replace/*:refs/replace/*'

That's true but you still need to have another remote with full
history because it has lots of tags and branches that will be cloned
by initial clone.

Regards,
Dmitry

```

## Dmitry Oksenchuk, 2014-10-30 17:56

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <CA+POfms4+fPnEv2nBPJx+hPRSMZyR66wpvfDf+SFy3X_8t2H-Q@mail.gmail.com>
URL: https://gitlist.dev/e/CA%2BPOfms4%2BfPnEv2nBPJx%2BhPRSMZyR66wpvfDf%2BSFy3X_8t2H-Q%40mail.gmail.com
In-Reply-To: <20141030154426.GU15443@odin.tremily.us>

```
2014-10-30 18:44 GMT+03:00 W. Trevor King <wking@tremily.us>:
> On Thu, Oct 30, 2014 at 06:39:56PM +0300, Dmitry Oksenchuk wrote:
>> We're in the middle of conversion of a large CVS repository (20
>> years, 70K commits, 1K branches, 10K tags) to Git and considering
>> two separate Git repositories: "historical" with CVS history and
>> "working" created without history from heads of active branches (10
>> active branches). This allows us to have small fast "working"
>> repository for developers who don't want to have full history
>> locally and ability to rewrite history in "historical" repository
>> (for example, to add parents to merge commits or to fix conversion
>> mistakes) without affecting commit hashes in "working" repository
>> (the hashes can be stored in bug tracker or in the code).
>
> A number of projects have done something like this (e.g. Linux).
> Modern Gits have good support for shallow repositories though, so I'd
> just make one full repository and leave it to developers to decide how
> deep they want their local copy to be.

Good point. Shallow clone allows a developer to have a small fast
repository if history is not needed.
But having new history in one repository with CVS history prevents us
from rewriting it in case of conversion mistakes or desire to restore
parents in merge commits.

Thanks,
Dmitry

```

## Christian Couder, 2014-10-31 08:45

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <CAP8UFD0ZN1Dq3kAs=nQE3fY8RydqUhSDvYksy7J2igTpt9p5WQ@mail.gmail.com>
URL: https://gitlist.dev/e/CAP8UFD0ZN1Dq3kAs%3DnQE3fY8RydqUhSDvYksy7J2igTpt9p5WQ%40mail.gmail.com
In-Reply-To: <CA+POfmvXEjDV9Vap6NDX7HvOMjEVG4mVe1uWFSTQy5g_c+vJnw@mail.gmail.com>

```
Hi Dmitry,

On Thu, Oct 30, 2014 at 6:41 PM, Dmitry Oksenchuk <oksenchuk89@gmail.com> wrote:
> 2014-10-30 19:54 GMT+03:00 Christian Couder <christian.couder@gmail.com>:
>>
>> This might be a good idea. Did you already test that the small
>> repository is really faster than the full repository?
>
> Yes, because of such amount of refs, push in "historical" repository
> takes 12 sec, push in "working" repository takes 0.4 sec, push in
> "joined" repository takes 2 sec. Local operations with history like
> log and blame work with the same speed in "joined" repository as in
> "historical" repository.

What does "joined" mean? Does it mean joined using grafts? Or joined
using replace refs? Or just the unsplit full repository?

Also what is interesting is if local operations work with the same
speed in the small "working" repository as in the unsplit full
repository.

>>> Are there any disadvantages of using grafts and replace? Will both of
>>> them be supported in future versions of Git?
>>
>> My opinion is that grafts have no advantage compared to replace refs.
>>
>> Once you have created your replace refs, they can be managed like
>> other git refs, so they are easier to distribute.
>>
>> Basically if you want to get the full history on a computer you just need to do:
>>
>> git fetch 'refs/replace/*:refs/replace/*'

By the way the above should be:

git fetch origin 'refs/replace/*:refs/replace/*'

> That's true but you still need to have another remote with full
> history because it has lots of tags and branches that will be cloned
> by initial clone.

Yeah, you might want to have another remote for that reason, but this
is true with both grafts and replace refs.

Best,
Christian.

```

## Dmitry Oksenchuk, 2014-10-31 15:47

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <CA+POfmt9cQcKshc3LgGpBqHdbsgjoN7Ou7eN3sGqJcDeXwgPyA@mail.gmail.com>
URL: https://gitlist.dev/e/CA%2BPOfmt9cQcKshc3LgGpBqHdbsgjoN7Ou7eN3sGqJcDeXwgPyA%40mail.gmail.com
In-Reply-To: <CAP8UFD0ZN1Dq3kAs=nQE3fY8RydqUhSDvYksy7J2igTpt9p5WQ@mail.gmail.com>

```
Hi Christian,

> On Thu, Oct 30, 2014 at 6:41 PM, Dmitry Oksenchuk <oksenchuk89@gmail.com> wrote:
>> 2014-10-30 19:54 GMT+03:00 Christian Couder <christian.couder@gmail.com>:
>>>
>>> This might be a good idea. Did you already test that the small
>>> repository is really faster than the full repository?
>>
>> Yes, because of such amount of refs, push in "historical" repository
>> takes 12 sec, push in "working" repository takes 0.4 sec, push in
>> "joined" repository takes 2 sec. Local operations with history like
>> log and blame work with the same speed in "joined" repository as in
>> "historical" repository.
>
> What does "joined" mean? Does it mean joined using grafts? Or joined
> using replace refs? Or just the unsplit full repository?

It's joined using grafts or replace. In both cases performance is the same.

> Also what is interesting is if local operations work with the same
> speed in the small "working" repository as in the unsplit full
> repository.

Speed of operations like git diff, git add, git commit is exactly the
same in both repositories.
Operations like git log and git blame work much faster in repository
without history (not surprisingly :)
For example, git log in small repository takes 0.2 sec, in full
repository - 0.8 sec. git blame in full repository can take up to 9
sec for large files with long history.

Regards,
Dmitry

```

## Christian Couder, 2014-11-01 15:03

Subject: Re: Joining historical repository using grafts or replace
Message-ID: <CAP8UFD17DOajgFWTTDC12qz3m4wKJTn2XLn+MWUE8Fc_XyTBTQ@mail.gmail.com>
URL: https://gitlist.dev/e/CAP8UFD17DOajgFWTTDC12qz3m4wKJTn2XLn%2BMWUE8Fc_XyTBTQ%40mail.gmail.com
In-Reply-To: <CA+POfmt9cQcKshc3LgGpBqHdbsgjoN7Ou7eN3sGqJcDeXwgPyA@mail.gmail.com>

```
Hi Dmitry,

On Fri, Oct 31, 2014 at 4:47 PM, Dmitry Oksenchuk <oksenchuk89@gmail.com> wrote:
> Hi Christian,
>
>> On Thu, Oct 30, 2014 at 6:41 PM, Dmitry Oksenchuk <oksenchuk89@gmail.com> wrote:
>>>
>>> Yes, because of such amount of refs, push in "historical" repository
>>> takes 12 sec, push in "working" repository takes 0.4 sec, push in
>>> "joined" repository takes 2 sec. Local operations with history like
>>> log and blame work with the same speed in "joined" repository as in
>>> "historical" repository.
>>
>> What does "joined" mean? Does it mean joined using grafts? Or joined
>> using replace refs? Or just the unsplit full repository?
>
> It's joined using grafts or replace. In both cases performance is the same.
>
>> Also what is interesting is if local operations work with the same
>> speed in the small "working" repository as in the unsplit full
>> repository.
>
> Speed of operations like git diff, git add, git commit is exactly the
> same in both repositories.
> Operations like git log and git blame work much faster in repository
> without history (not surprisingly :)
> For example, git log in small repository takes 0.2 sec, in full
> repository - 0.8 sec. git blame in full repository can take up to 9
> sec for large files with long history.

Ok, thanks for the information. I think it shows that indeed it makes
sense to split your repo.

Best,
Christian.

```
