threads / discuss / 63260

Way to "impersonate" remote or sync remotes without fetching everything?

Subject: Way to "impersonate" remote or sync remotes without fetching everything?

## tl;dr

6 messages between Apr 5, 2025 and Apr 14, 2025.

replies: 5people: 2as markdown or json

Klaus Frank· Apr 5, 2025, 14:01 UTC · lore
Hi,

what is the best way to sync multiple remotes with each other without having to pull everything into a local checkout first?

Now as many projects have moved off of github onto their own gitlab instances it is kinda hard to keep track of all of the contributions as you basically have to create a fork on each of these instances to open PRs and because I'd like to have my contributions also on my gitlab server too I've now the issue of having to keep multiple remote repositories in sync. And then sometimes there are additional remotes from others that you'd like to interact with as well, so now these need to be mirrored as well. Then there is also the issue of archiving, e.g. freeing oneself from a dependency getting deleted upstream that I see at more and more companies I work with. They want to have an internal fork of everything (someties also wanting to review all of the commits before pulling them into their internal git server; but for most it was enough to just fail on rewritten/force-pushed history so far). And just for my personal projects I'd like to be able to have them on github, gitlab, and my own gitlab server, while being able to make edits (and accept PRs) at any of these and keep the repos in sync (or at least automatically sync them as long as nothing conflicts).

Is there a way to use some of the more advanced features of git to accomplis this? Like e.g. the alternative object database mechanism in a local temp project and pushing to the other remotes or something?

If any of you ran into similar issues in the past, how did you solve them?
What I tried so far in order:
My initial thought was to look for a way to do:
1. "impersonate" a remote (something like "git clone --bare" but
    without actually cloning any of the objects and querying the remote
    when needed)
2. Any git command within that repo will now be treated as if it was
    ran within/by the impersonated remote. E.g. It'll consider all of
    the objects the remote has as its own.
3. Add all of the other remotes as usual (the impersonated one wouldn't
    be reported as remote but as the local one the others were added to)
4. Doing a "git push --all" to any of these remotes will cause git to
    download thouse objects from the impersonated one that it needs to
    push towards that remote (and hasn't cached locally already).

As I wasn't able to find such a feature I tried to workaround this. At first I tried to use a shallow clone but that didn't really work as I didn't know where I'd have to start in order to be able to give all of the remotes a common commit to be detected as belonging together (esp. if one of the remotes didn't exist and had to be created by pushing to it...).

Then I managed to find these two suboptimal ones so far:
1. pull-push-style:
     1. "Pull --mirror --bare" the first one
     2. "Push --all" to all of the other(s)
2. Without always pulling the same server for the entire repo,
    regardless of it having changes or not:
     1. Create new and empty repo locally
     2. Add all of the remotes
     3. Fetch from the nearest server,
     4. "lfs fetch --all" from the nearest server
     5. fetch from all others
     6. "lfs fetch --all" from all others
     7. Hackishly update all of the local refs to basically be the ones
        of the remote that should be used as source (aka "rm -rf
        .git/refs/heads/*; cp -r .git/refs/remotes/origin/*
        .git/refs/heads/")
     8. "Push --all" to all of the remotes

The 1st one is the most simple, it almost always works (only fails in some very rare cases, like when the remote contains "zero-padded file modes" and such) but it'll cause unecessary load on the origin server as it has to always first make an entire clone of the repo. Even if the remotes are already in sync. Having local state would partially would solve that but then I'd need a lot of disk space and I wouldn't be able to just run this within a CI job. => Therefore undesirable.

The 2nd one is more sophisticated, it may look kinda righit at first, but it keeps breaking all the time for countless different reasons. Even though it allows to only pull new/changed commits from the origin server (aka it doesn't overload a single server unecessarily), it'll still clone all of the repos from the local git server onto the CI worker, esp. considering larger repos using git-lfs (or trying to do it in parallel with multiple/all of my repos) it causes the CI worker to run out of disk space as well as a lot of unecessary network and disk IO => Therefore also undesirable.

Also not to mention that none of my current approaches can do a real sync, they are all relying upon having one of the remotes designated as the authoritative source and if any of the others changed instead it'll just fail.

Sincerely, Klaus Frank

D. Ben Knoble· Apr 11, 2025, 18:43 UTC · re: Klaus Frank · lore

Re: Way to "impersonate" remote or sync remotes without fetching everything?

On Sat, Apr 5, 2025 at 10:01 AM Klaus Frank <vger.kernel.org@frank.fyi> wrote:
> Also not to mention that none of my current approaches can do a real
> sync, they are all relying upon having one of the remotes designated as
> the authoritative source and if any of the others changed instead it'll
> just fail.

Maybe I haven't totally understood your use-case, but what if the authoritative source is your local repository, and then you push to all your remote mirrors to publish your trees? That is, I don't think it's a good idea to have a remote worker automatically pushing changes across all the mirrors; rather, you get to be in control of when you push to those mirrors.

(I thought there was a push-equivalent of remotes.<group>, which I was going to suggest as being helpful for this kind of mode where some remotes are your mirrors and others are collaborators, but I can't find it.) Configuring a new remote with pushurls that point to the other remotes [1] seems to be the way to make pushing to multiple remotes easy.

Of course, this doesn't help CI, but then it can just pick any mirror it wants to fetch a commit to build?

[1]: https://stackoverflow.com/a/14290145/4400820
-- 
D. Ben Knoble
Klaus Frank· Apr 11, 2025, 19:02 UTC · re: D. Ben Knoble · lore

Re: Way to "impersonate" remote or sync remotes without fetching everything?

On 2025-04-11 20:43:24, D. Ben Knoble wrote:
> Maybe I haven't totally understood your use-case, but what if the
> authoritative source is your local repository, and then you push to

There is no local repository, that's kinda the source of all of this. The sync script runs in a CI/CD. I'm kinda abusing CI/CD here to run a kind of cron job, in a separate repository that does the sync, maybe it is easier to just call it scheduled pipeline/action or just stateless cron job?

Lets make a more quick example:

gdm is being developed here: https://gitlab.gnome.org/GNOME/gdm so in order to make a PR I'll have to create a fork in that GitLab instance so now we're at 2 repositories. Then I want to have my own independent archive mirror in my own gitlab instance. Then I also want to mirror it onto gitlab.com and github.com just for the sake of this example. Now we're at 5 remotes.

Now I'd like to have a script in CI/CD (that runs server side) to sync all of them. In example the gnome.org one could probably mostly be the autoritative source (except for the branches that contain my changes).

D. Ben Knoble· Apr 13, 2025, 21:52 UTC · re: Klaus Frank · lore

Re: Way to "impersonate" remote or sync remotes without fetching everything?

On Fri, Apr 11, 2025 at 3:02 PM Klaus Frank <vger.kernel.org@frank.fyi> wrote:
Show 23 quoted lines
>
> On 2025-04-11 20:43:24, D. Ben Knoble wrote:
> > Maybe I haven't totally understood your use-case, but what if the
> > authoritative source is your local repository, and then you push to
>
> There is no local repository, that's kinda the source of all of this.
> The sync script runs in a CI/CD. I'm kinda abusing CI/CD here to run
> a kind of cron job, in a separate repository that does the sync, maybe it
> is easier to just call it scheduled pipeline/action or just stateless
> cron job?
>
> Lets make a more quick example:
>
> gdm is being developed here: https://gitlab.gnome.org/GNOME/gdm
> so in order to make a PR I'll have to create a fork in that GitLab
> instance so now we're at 2 repositories. Then I want to have my own
> independent archive mirror in my own gitlab instance. Then I also
> want to mirror it onto gitlab.com and github.com just for the sake of
> this example. Now we're at 5 remotes.
>
> Now I'd like to have a script in CI/CD (that runs server side) to sync
> all of them. In example the gnome.org one could probably mostly be the
> autoritative source (except for the branches that contain my changes).

That all makes sense, except: why need the sync (cron) job? Treat a local copy as authoritative for you and push to all your remotes. This puts you in control at the cost of not happening automatically. (You could conceivably have a local cron job that did this.)

-- 
D. Ben Knoble
Klaus Frank· Apr 14, 2025, 00:34 UTC · re: D. Ben Knoble · lore

Re: Way to "impersonate" remote or sync remotes without fetching everything?

On 2025-04-13 23:52:14, D. Ben Knoble wrote:
> That all makes sense, except: why need the sync (cron) job? Treat a
> local copy as authoritative for you and push to all your remotes. This
> puts you in control at the cost of not happening automatically. (You
> could conceivably have a local cron job that did this.)

Cause I don't want to have it locally on e.g. a notebook that can break mainly :D (I know myself I won't be making enough backups if that is their primary location)

But I have been thinking about something similar earlier today. Maybe I should just trash my software forge (gitlab) and just use git from the cli via ssh. Then writing a cron job to do the syncing would also be easier as I'd have a local copy to work with and maybe add some Stagit sparkles to replace the web-ui (aka some static page generator for git repos) https://codemadness.org/stagit.html

(Still not quite satisfies with Stagit nor cgit, maybe I'll find something that needs, less files generated on the server side and some "git in javascript" to display the more exotic views instead of doing it server side like cgit does)

-- Klaus Frank

D. Ben Knoble· Apr 14, 2025, 19:40 UTC · re: Klaus Frank · lore

Re: Way to "impersonate" remote or sync remotes without fetching everything?

On Mon, Apr 14, 2025 at 11:44 AM Klaus Frank <vger.kernel.org@frank.fyi> wrote:
Show 11 quoted lines
>
> On 2025-04-13 23:52:14, D. Ben Knoble wrote:
> > That all makes sense, except: why need the sync (cron) job? Treat a
> > local copy as authoritative for you and push to all your remotes. This
> > puts you in control at the cost of not happening automatically. (You
> > could conceivably have a local cron job that did this.)
>
> Cause I don't want to have it locally on e.g. a notebook that can break
> mainly :D
> (I know myself I won't be making enough backups if that is their primary
> location)

If you push to other remotes (or cron, as below), then you'll have "backups for free"?

Show 6 quoted lines
>
> But I have been thinking about something similar earlier today. Maybe I
> should just
> trash my software forge (gitlab) and just use git from the cli via ssh.
> Then writing a cron job to do the syncing would also be easier as I'd have
> a local copy to work with
[…]
-- 
D. Ben Knoble

← back to recent threads