git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: multiple working directories for long-running builds (was: "git merge" merges too much!)

From
GWGreg A. Woods <woods@planix.com>
Date
Dec 1, 2009, 22:44 UTC
Message-ID
<m1NFbSE-000kn2C@most.weird.com>
In-Reply-To
<20091201211830.GE11235@dpotapov.dyndns.org>
At Wed, 2 Dec 2009 00:18:30 +0300, Dmitry Potapov <dpotapov@gmail.com> wrote:
Subject: Re: multiple working directories for long-running builds (was:	"git merge" merges too much!)
> 
> AFAIK, "git archive" is cheaper than git clone.

It depends on what you mean by "cheaper" It's clearly going to require less disk space. However it's also clearly going to require more disk bandwidth, potentially a _LOT_ more disk bandwidth.

> I do not say it is fast
> for huge project, but if you want to run a process such as clean build
> and test that takes a long time anyway, it does not add much to the
> total time.

I think you need to try throwing around an archive of, say, 50,000 small files a few times simultaneously on your system to appreciate the issue.

(i.e. consider the load on a storage subsystem, say a SAN or NAS, where with your suggestion there might be a dozen or more developers running "git archive" frequently enough that even three or four might be doing it at the same time, and this on top of all the i/o bandwidth required for the builds all of the other developers are also running at the same time.)

> > Disk bandwidth is almost always more expensive than disk space.
> 
> Disk bandwidth is certainly more expensive than disk space, and the
> whole point was to avoid a lot of disk bandwidth by using hot cache.

Huh? Throwing around the archive has nothing to do with the build system in this case.

Please let me worry about optimizing the builds -- that's well under control already and is not really yet an issue for the VCS, at least not yet, and maybe never in many cases.

I'm just not willing to even consider using what would really be the most simplistic and most expensive form of updating a working directory as could ever be imagined. "Git archive" is truly unintelligent, as-is.

Perhaps if "git archive" could talk intelligently to an rsync process and be smart about updating an existing working directory it would be the ideal answer, but _NEVER_ with the current method of just unpacking an archive over an existing directory! (Now there's a good Google SoC, or masters, project for someone eager to learn about rsync & git internals!)

Local filesystem "git clone" is usable in many scenarios, but it just won't work nearly so efficiently in a scenario where users have local repos on their workstations and use an NFS NAS to feed the build servers. As I understand it this 'git-new-workdir' script will work though since it uses symlinks that can be pointed across the mount back to the local disk on the user's workstation. They can just mount the build directory and go into it and run a "git checkout" and start another build on the build server(s).

A major further advantage of multiple working directories is that this eliminates one more point of failure -- i.e. you don't end up with multiple copies of the repo that _should_ be effectively read-only for everything but "push", and perhaps then only to one branch. I don't like giving developers too much rope, especially in all the wrong places. "git archive" does achieve the same even better I suppose, but without something like a "--format=rsync" option it's completely out of the question.

> Another thing to consider is that if you put a really huge project in one
> Git repo than Git may not be as fast as you may want, because Git tracks
> the whole project as the whole. So, you may want to split your project in
> a few relatively independent modules (See git submodule).
Indeed -- but sometimes I think this is not feasible either.

I know of at least three very real-world projects where there are tens of thousands of small files that really must be managed as one unit, and where running a build in that tree could take a whole day or two on even the fastest currently available dedicated build server. Eg. pkgsrc.

-- 
						Greg A. Woods
						Planix, Inc.

<woods@planix.com>       +1 416 218 0099        http://www.planix.com/
Previous: Jeff EplerNext: Dmitry Potapov
Message 30 of 33 in “"git merge" merges too much!”
  1. Greg A. WoodsNov 29, 2009
  2. Jeff KingNov 29, 2009
  3. Greg A. WoodsNov 30, 2009
  4. Dmitry PotapovNov 30, 2009
  5. Greg A. WoodsDec 1, 2009
  6. Dmitry PotapovDec 1, 2009
  7. Greg A. WoodsDec 1, 2009
  8. Dmitry PotapovDec 2, 2009
  9. Nanako ShiraishiDec 2, 2009
  10. Jeff KingDec 2, 2009
  11. Greg A. WoodsDec 3, 2009
  12. Junio C HamanoDec 3, 2009
  13. Greg A. WoodsDec 3, 2009
  14. Jeff KingDec 3, 2009
  15. Uri OkrentDec 3, 2009
  16. Marko KreenDec 3, 2009
  17. Greg A. WoodsDec 9, 2009
  18. Jeff KingDec 3, 2009
  19. Junio C HamanoNov 29, 2009
  20. Greg A. WoodsNov 30, 2009
  21. Junio C HamanoNov 30, 2009
  22. Dmitry PotapovNov 30, 2009
  23. Greg A. WoodsDec 1, 2009
  24. Dmitry PotapovDec 1, 2009
  25. Greg A. WoodsDec 1, 2009
  26. Dmitry PotapovDec 1, 2009
  27. Greg A. WoodsDec 1, 2009
  28. Dmitry PotapovDec 1, 2009
  29. Jeff EplerDec 1, 2009
  30. Greg A. WoodsDec 1, 2009
  31. Dmitry PotapovDec 2, 2009
  32. Greg A. WoodsDec 3, 2009
  33. Junio C HamanoDec 2, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.