threads / discuss / 18881

integrating make and git

Subject: integrating make and git

## tl;dr

14 messages between Apr 15, 2009 and Apr 18, 2009.

replies: 13people: 11as markdown or json

E R· Apr 15, 2009, 15:19 UTC · lore

I have an idea about integrating make with git, and I'm wondering if it is a reasonable thing to do.

First of all, I am under the impression that git can quickly compute a hash of a directory and its contents. Is that correct?

If so, suppose you using git to manage revision control of a project which has some components like 'lib1', 'lib2', etc. Typically you would perform something like: make clean; make all and 'make all' would perform 'make lib1' and 'make lib2'. When checking out a different revision of the project you would have to perform another 'make clean' before 'make all' since you aren't sure of what's changed and the timestamps of the derived files will be more recent than the timestamps of the source files.

Now suppose that making 'lib1' only depends on the source code in a certain directory. The idea is to associate the hash of the source directory for lib1 with its the derived files. Make can check this to determine if the component really needs to be rebuilt. Then as you move around in the repository you can avoid rebuilding components that haven't changed.

Good, bad, ugly?
Matthieu Moy· Apr 15, 2009, 15:41 UTC · re: E R · lore

Re: integrating make and git

E R <pc88mxer@gmail.com> writes:
Show 5 quoted lines
> When checking out a
> different revision of the project you would have to perform another
> 'make clean' before 'make all' since you aren't sure of what's changed
> and the timestamps of the derived files will be more recent than the
> timestamps of the source files.

The last assumption is incorrect. git checkout will touch the files it modifies, and won't play with timestamp precisely to save you from having to do "make clean" each time you use git.

-- 
Matthieu
Daniel Barkalow· Apr 15, 2009, 16:20 UTC · re: E R · lore

Re: integrating make and git

On Wed, 15 Apr 2009, E R wrote:
Show 14 quoted lines
> I have an idea about integrating make with git, and I'm wondering if
> it is a reasonable thing to do.
> 
> First of all, I am under the impression that git can quickly compute a
> hash of a directory and its contents. Is that correct?
> 
> If so, suppose you using git to manage revision control of a project
> which has some components like 'lib1', 'lib2', etc. Typically you
> would perform something like: make clean; make all and 'make all'
> would perform 'make lib1' and 'make lib2'. When checking out a
> different revision of the project you would have to perform another
> 'make clean' before 'make all' since you aren't sure of what's changed
> and the timestamps of the derived files will be more recent than the
> timestamps of the source files.

No, the timestamps of the changed source files will be newer than the timestamps of the derived files. Git doesn't backdate files in working directories, in order to avoid causing the problem you're trying to fix. (And because getting the history is so quick and easy with git that looking at dates on files in the filesystem is kind of pointless.)

	-Daniel
*This .sig left intentionally blank*
E R· Apr 15, 2009, 16:47 UTC · re: Daniel Barkalow · lore

Re: integrating make and git

Ok - I was wrong about the timestamps not getting updated. Thanks for that correction.

However, what about the idea of associating the result of a build with the hash of the source files used by the build, and using git to compute the hash?

On Wed, Apr 15, 2009 at 11:20 AM, Daniel Barkalow <barkalow@iabervon.org> wrote:
Show 26 quoted lines
> On Wed, 15 Apr 2009, E R wrote:
>
>> I have an idea about integrating make with git, and I'm wondering if
>> it is a reasonable thing to do.
>>
>> First of all, I am under the impression that git can quickly compute a
>> hash of a directory and its contents. Is that correct?
>>
>> If so, suppose you using git to manage revision control of a project
>> which has some components like 'lib1', 'lib2', etc. Typically you
>> would perform something like: make clean; make all and 'make all'
>> would perform 'make lib1' and 'make lib2'. When checking out a
>> different revision of the project you would have to perform another
>> 'make clean' before 'make all' since you aren't sure of what's changed
>> and the timestamps of the derived files will be more recent than the
>> timestamps of the source files.
>
> No, the timestamps of the changed source files will be newer than the
> timestamps of the derived files. Git doesn't backdate files in working
> directories, in order to avoid causing the problem you're trying to fix.
> (And because getting the history is so quick and easy with git that
> looking at dates on files in the filesystem is kind of pointless.)
>
>        -Daniel
> *This .sig left intentionally blank*
>
Robin Rosenberg· Apr 15, 2009, 17:30 UTC · re: E R · lore

Re: integrating make and git

onsdag 15 april 2009 18:47:52 skrev E R <pc88mxer@gmail.com>:
Show 6 quoted lines
> Ok - I was wrong about the timestamps not getting updated. Thanks for
> that correction.
> 
> However, what about the idea of associating the result of a build with
> the hash of the source files used by the build, and using git to
> compute the hash?

Take a look at ccache. It doesn't use Git, but it uses hashes of source, and compiler flags and associates that with the resulting object files, so it can avoid compiling. If you are building largs C/C++ (especially C++) projects you want it.

-- robin
Jeff King· Apr 16, 2009, 08:26 UTC · re: Robin Rosenberg · lore

Re: integrating make and git

On Wed, Apr 15, 2009 at 07:30:32PM +0200, Robin Rosenberg wrote:
> Take a look at ccache. It doesn't use Git, but it uses hashes of source, and
> compiler flags and associates that with the resulting object files, so it
> can avoid compiling. If you are building largs C/C++ (especially C++)
> projects you want it. 

In theory, one could improve something like ccache by asking git the sha-1 of the file. Since git maintains a cache based on stat info, you can get away with not looking at the file contents at all (which saves CPU time in hashing, but also helps a lot when building from a cold cache).

In practice, this doesn't help because:
  1. ccache looks at more than just the file itself. I believe it
     actually runs it through cpp and hashes that.
  2. People combine ccache with make; if the stat data hasn't changed,
     in most cases, you will skip building before you even get to
     ccache.

But one could probably design a system to replace both ccache and make that relies on git's fast sha-1 reporting to avoid duplicate work. I suspect nobody has bothered because make+ccache is "fast enough" that the added complexity would not be worth it.

-Peff
Matthieu Moy· Apr 16, 2009, 09:55 UTC · re: Jeff King · lore

Re: integrating make and git

Jeff King <peff@peff.net> writes:
> But one could probably design a system to replace both ccache and make
> that relies on git's fast sha-1 reporting to avoid duplicate work. I
> suspect nobody has bothered because make+ccache is "fast enough" that
> the added complexity would not be worth it.

AIUI, ClearCase does something similar to that. Call that an immense bloatware where everything has to come together, or a nice integration of different tools, I never used it, so I don't know (I heard the first option more than the second ...).

-- 
Matthieu
Nguyen Thai Ngoc Duy· Apr 16, 2009, 12:50 UTC · re: Matthieu Moy · lore

Re: integrating make and git

On Thu, Apr 16, 2009 at 7:55 PM, Matthieu Moy <Matthieu.Moy@imag.fr> wrote:
Show 11 quoted lines
> Jeff King <peff@peff.net> writes:
>
>> But one could probably design a system to replace both ccache and make
>> that relies on git's fast sha-1 reporting to avoid duplicate work. I
>> suspect nobody has bothered because make+ccache is "fast enough" that
>> the added complexity would not be worth it.
>
> AIUI, ClearCase does something similar to that. Call that an immense
> bloatware where everything has to come together, or a nice integration
> of different tools, I never used it, so I don't know (I heard the
> first option more than the second ...).

It does help on really big projects (complete rebuild may take one day, complete recollect built objects and link them with clearmake take about one hour). C++-based projects may like it due to long time compilation.

-- 
Duy
Daniel Barkalow· Apr 15, 2009, 21:01 UTC · re: E R · lore

Re: integrating make and git

On Wed, 15 Apr 2009, E R wrote:
Show 6 quoted lines
> Ok - I was wrong about the timestamps not getting updated. Thanks for
> that correction.
> 
> However, what about the idea of associating the result of a build with
> the hash of the source files used by the build, and using git to
> compute the hash?

It's a reasonable idea, in general, but may or may not be useful for any particular problem. In general, the objects in your subdirectories are also doing to depend on some but not most things from an include directory, and so there's not much benefit you can get on a per-directory granularity. On the other hand, I've gotten good results by embedding the commit sha1 in generated object files, which allowed me to exactly identify different builds much later, and even figure out what the source that went into them was.

	-Daniel
*This .sig left intentionally blank*
John Bito· Apr 15, 2009, 21:34 UTC · re: Daniel Barkalow · lore

Re: integrating make and git

If you're not already using make for a project, think before you start. If you're building something that will go into distribution in source form, you probably should need to use it (via automake & autoconf). For the stuff that I'm doing in a more focused environment, I use boost-build/bjam (http://www.boost.org/users/download/boost_jam_3_1_17). This provides very clean organization of release/debug builds as well as a much more expressive language than make.

It's easy to create makefiles that are quite brittle and the plumbing to establish portable builds is really complex. I was away from make working on Java & Ruby for almost ten years. Though I have more than 10 years of experience with make, I'm very happy to have replaced gobs of makefile code with the boost-build package and a few, short Jamfiles.

YMMV John

Ben Jackson· Apr 16, 2009, 03:50 UTC · re: E R · lore

Re: integrating make and git

E R <pc88mxer <at> gmail.com> writes:
> Now suppose that making 'lib1' only depends on the source code in a
> certain directory. The idea is to associate the hash of the source
> directory for lib1 with its the derived files. Make can check this to
> determine if the component really needs to be rebuilt.

ClearCase has "wink-ins" which are very much like this. It knows that a given object was produced from a certain set of sources with a particular command. When someone wants to recreate that object (not even necessarily the original builder) it can "wink in" the result. Typically a brand new "view" (a ClearCase working directory) build will consist of winking in a ton of objects rather than building anything. I'm not sure how much of this is due to cleverness in clearmake and how much is due to the view being implemented as a virtual filesystem (which can see every repository file being read as part of a build).

--Ben
David Kågedal· Apr 16, 2009, 08:05 UTC · re: Ben Jackson · lore

Re: integrating make and git

Ben Jackson <ben@ben.com> writes:
Show 15 quoted lines
> E R <pc88mxer <at> gmail.com> writes:
>
>> Now suppose that making 'lib1' only depends on the source code in a
>> certain directory. The idea is to associate the hash of the source
>> directory for lib1 with its the derived files. Make can check this to
>> determine if the component really needs to be rebuilt.
>
> ClearCase has "wink-ins" which are very much like this.  It knows that a given
> object was produced from a certain set of sources with a particular command. 
> When someone wants to recreate that object (not even necessarily the original
> builder) it can "wink in" the result.  Typically a brand new "view" (a ClearCase
> working directory) build will consist of winking in a ton of objects rather than
> building anything.  I'm not sure how much of this is due to cleverness in
> clearmake and how much is due to the view being implemented as a virtual
> filesystem (which can see every repository file being read as part of a build).

It very much depends on implementing its own file system, since it otherwise would have no idea what the *real* build dependencies are.

That one of the nice things about clearmake, by the way. You don't have to worry much about describing the dependencies, since it will figure it out all by itself when you first build the project.

But I don't think there is much in CC for git to copy... :-)
-- 
David Kågedal
Dmitry Potapov· Apr 17, 2009, 17:24 UTC · re: David Kågedal · lore

Re: integrating make and git

On Thu, Apr 16, 2009 at 12:05 PM, David Kågedal <davidk@lysator.liu.se> wrote:
Show 20 quoted lines
> Ben Jackson <ben@ben.com> writes:
>
>> E R <pc88mxer <at> gmail.com> writes:
>>
>>> Now suppose that making 'lib1' only depends on the source code in a
>>> certain directory. The idea is to associate the hash of the source
>>> directory for lib1 with its the derived files. Make can check this to
>>> determine if the component really needs to be rebuilt.
>>
>> ClearCase has "wink-ins" which are very much like this.  It knows that a given
>> object was produced from a certain set of sources with a particular command.
>> When someone wants to recreate that object (not even necessarily the original
>> builder) it can "wink in" the result.  Typically a brand new "view" (a ClearCase
>> working directory) build will consist of winking in a ton of objects rather than
>> building anything.  I'm not sure how much of this is due to cleverness in
>> clearmake and how much is due to the view being implemented as a virtual
>> filesystem (which can see every repository file being read as part of a build).
>
> It very much depends on implementing its own file system, since it
> otherwise would have no idea what the *real* build dependencies are.

Not necessary... You can use LD_PRELOAD to intercept 'open' (and all needed syscalls), but it works only on those platforms where LD_PRELOAD is supported. IIRC, there was some tool that did this, but I have never used it. I am pretty happy with ccache :)

Dmitry
Ealdwulf Wuffinga· Apr 18, 2009, 07:03 UTC · re: E R · lore

Re: integrating make and git

This idea sounds very much like vesta (http://www.vestasys.org) - except that vesta has its own version control system, rather than using git. It has its own filesystem (actually a user-mode nfs server). It's GPL, so in theory you could swipe its filesystem code, or replace its vcs with git. Don't know how hard that would be, though - doesn't sound trivial. It's worth looking at vesta for ideas, though, even if you don't use the code. The owners are fairly responsive on their IRC channel, too.

Ealdwulf

← back to recent threads