threads / discuss / 43121

Re: Adding a new file as if it had existed

Subject: Re: Adding a new file as if it had existed

## tl;dr

11 messages between Dec 12, 2006 and Dec 13, 2006.

replies: 10people: 6as markdown or json

Bahadir Balban· Dec 12, 2006, 10:05 UTC · lore

Adding a new file as if it had existed

Hi,

When I initialise a git repository, I use a subset of files in the project and leave out irrelevant files for performance reasons. Then when I need to make changes to a file not yet in the repository, the file is treated as new, and if I reset the change or change branches the file is gone.

Is there a good way of adding new files to git as if they had existed from the initial commit (or even better, since a particular commit)? This way I would only track the new changes I made to an existing file.

Thanks,
Junio C Hamano· Dec 12, 2006, 10:13 UTC · re: Bahadir Balban · lore
"Bahadir Balban" <bahadir.balban@gmail.com> writes:
> Is there a good way of adding new files to git as if they had existed
> from the initial commit (or even better, since a particular commit)?
> This way I would only track the new changes I made to an existing
> file.
No.

I do not understand why not adding all the files you care about eventually anyway in the initial commit is needed for "performance reasons", if you do not touch majority of them for a long time. Care to explain?

Bahadir Balban· Dec 12, 2006, 11:32 UTC · re: Junio C Hamano · lore
On 12/12/06, Junio C Hamano <junkio@cox.net> wrote:
Show 6 quoted lines
> No.
>
> I do not understand why not adding all the files you care about
> eventually anyway in the initial commit is needed for
> "performance reasons", if you do not touch majority of them for
> a long time.  Care to explain?

If I don't know which files I may be touching in the future for implementing some feature, then I am obliged to add all the files even if they are irrelevant. I said "performance reasons" assuming all the file hashes need checked for every commit -a to see if they're changed, but I just tried on a PIII and it seems not so slow.

Johannes Schindelin· Dec 12, 2006, 12:07 UTC · re: Bahadir Balban · lore
Hi,
On Tue, 12 Dec 2006, Bahadir Balban wrote:
Show 10 quoted lines
> On 12/12/06, Junio C Hamano <junkio@cox.net> wrote:
> > No.
> > 
> > I do not understand why not adding all the files you care about
> > eventually anyway in the initial commit is needed for
> > "performance reasons", if you do not touch majority of them for
> > a long time.  Care to explain?
> 
> If I don't know which files I may be touching in the future for
> implementing some feature,

When I use an SCM, it is to track the revisions of a project. It seems you are content to have only parts of a revision? That does not make sense to me.

> I said "performance reasons" assuming all the file hashes need checked 
> for every commit -a to see if they're changed, but I just tried on a 
> PIII and it seems not so slow.
Bingo!
You just felt the consequences of the "index".

Ciao, Dscho

Andy Parkins· Dec 12, 2006, 12:26 UTC · re: Bahadir Balban · lore
On Tuesday 2006 December 12 11:32, Bahadir Balban wrote:
Show 5 quoted lines
> If I don't know which files I may be touching in the future for
> implementing some feature, then I am obliged to add all the files even
> if they are irrelevant. I said "performance reasons" assuming all the
> file hashes need checked for every commit -a to see if they're
> changed, but I just tried on a PIII and it seems not so slow.
Here's a handy rule of thumb I've learned in my use of git:
 "git is fast.  Really fast."

That'll hold you in good stead. In my experience there is no operation in git that is slow. I've got some trees that are for embedded work and hold the whole linux kernel, often more than once. Subversion, which I used previously, took literally hours to import the whole tree. Git takes minutes.

As to your direct concern: git doesn't hash every file at every commit. There is no need. git has an "index" that is used to prepare a commit; at the time you do the actual commit, git already knows which files are being checked in. Obviously, Linus uses git for managing the linux kernel, he's said before that he wanted a version control system that can do multiple commits /per second/. git can do that.

In short - don't worry about making life easy for git - it's a workhorse and does a grand job.

Andy
-- 
Dr Andy Parkins, M Eng (hons), MIEE
Andreas Ericsson· Dec 12, 2006, 13:20 UTC · re: Andy Parkins · lore
Andy Parkins wrote:
Show 12 quoted lines
> On Tuesday 2006 December 12 11:32, Bahadir Balban wrote:
> 
>> If I don't know which files I may be touching in the future for
>> implementing some feature, then I am obliged to add all the files even
>> if they are irrelevant. I said "performance reasons" assuming all the
>> file hashes need checked for every commit -a to see if they're
>> changed, but I just tried on a PIII and it seems not so slow.
> 
> Here's a handy rule of thumb I've learned in my use of git:
> 
>  "git is fast.  Really fast."
> 

Almost alarmingly so. When I started using git (back in May/June last year, when git was 2 - 3 months old), I was worried at first because it didn't seem to actually *do* anything, but just returned me to the prompt immediately.

Show 8 quoted lines
> 
> As to your direct concern: git doesn't hash every file at every commit.  There 
> is no need.  git has an "index" that is used to prepare a commit; at the time 
> you do the actual commit, git already knows which files are being checked in.  
> 
> In short - don't worry about making life easy for git - it's a workhorse and 
> does a grand job.
> 

Yup. Now I've gone the other way around and think other scm's are broken when they chew disk for 10 seconds whenever I try to do anything with them. I usually end up importing the other repo into git and do my work there.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Junio C Hamano· Dec 12, 2006, 18:31 UTC · re: Bahadir Balban · lore
"Bahadir Balban" <bahadir.balban@gmail.com> writes:
> ... I said "performance reasons" assuming all the
> file hashes need checked for every commit -a to see if they're
> changed, but I just tried on a PIII and it seems not so slow.
Ok.

Other people have already cleared the fear for 'commit' case, so I hope you are happier.

There is one thing we could further optimize, though.

Switching branches with 100k blobs in a commit even when there are a handful paths different between the branches would still need to populate the index by reading two trees and collapsing them into a single stage. In theory, we should be able to do a lot better if two-tree case of read-tree took advanrage of cache-tree information. If ce_match_stat() says Ok for all paths in a subdirectory and the cached tree object name for that subdirectory in the index match what we are reading from the new tree, we should be able to skip reading that subdirectory (and its subdirectories) from the new tree object at all.

Anybody interested to give it a try?
Andreas Ericsson· Dec 13, 2006, 09:40 UTC · re: Junio C Hamano · lore
Junio C Hamano wrote:
Show 17 quoted lines
> "Bahadir Balban" <bahadir.balban@gmail.com> writes:
> 
> There is one thing we could further optimize, though.
> 
> Switching branches with 100k blobs in a commit even when there
> are a handful paths different between the branches would still
> need to populate the index by reading two trees and collapsing
> them into a single stage.  In theory, we should be able to do a
> lot better if two-tree case of read-tree took advanrage of
> cache-tree information.  If ce_match_stat() says Ok for all
> paths in a subdirectory and the cached tree object name for that
> subdirectory in the index match what we are reading from the new
> tree, we should be able to skip reading that subdirectory (and
> its subdirectories) from the new tree object at all.
> 
> Anybody interested to give it a try?
> 

I'm not vell-versed enough in git internals to have my hopes high of making something useful of it, but if you give me a pointer of where to start I'd be happy to try, and perhaps learn something in the process.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Johannes Schindelin· Dec 13, 2006, 15:46 UTC · re: Andreas Ericsson · lore
Hi,
On Wed, 13 Dec 2006, Andreas Ericsson wrote:
Show 22 quoted lines
> Junio C Hamano wrote:
> > "Bahadir Balban" <bahadir.balban@gmail.com> writes:
> > 
> > There is one thing we could further optimize, though.
> > 
> > Switching branches with 100k blobs in a commit even when there
> > are a handful paths different between the branches would still
> > need to populate the index by reading two trees and collapsing
> > them into a single stage.  In theory, we should be able to do a
> > lot better if two-tree case of read-tree took advanrage of
> > cache-tree information.  If ce_match_stat() says Ok for all
> > paths in a subdirectory and the cached tree object name for that
> > subdirectory in the index match what we are reading from the new
> > tree, we should be able to skip reading that subdirectory (and
> > its subdirectories) from the new tree object at all.
> > 
> > Anybody interested to give it a try?
> > 
> 
> I'm not vell-versed enough in git internals to have my hopes high of 
> making something useful of it, but if you give me a pointer of where to 
> start I'd be happy to try, and perhaps learn something in the process.
Okay, I'll have a stab at explaining it.

For huge working directories, you usually have a huge number of trees. The idea of cache_tree is to remember not only the stat information of the blobs in the index, but to cache the hashes of the trees also (until they are invalidated, e.g. by an update-index). This avoids recalculation of the hashes when committing.

This cache is accessible by the global variable active_cache_tree. It is best accessed by the function cache_tree_find(), which you call like that:

	struct cache_tree *ct = cache_tree_find(active_cache_tree, path);

where the variable "path" may contain slashes. The SHA1 of the corresponding tree is in ct->sha1, and you can check if the hash is still valid by asking

	if (cache_tree_fully_valid(ct))
		/* still valid */

AFAIU Junio would like to take the shortcut of doing nothing at all when (twoway) reading a tree whose hash is identical to the hash stored in the corresponding cache_tree _and_ when the cache is still fully valid.

Ciao, Dscho

Andreas Ericsson· Dec 13, 2006, 15:52 UTC · re: Johannes Schindelin · lore
Johannes Schindelin wrote:
Show 50 quoted lines
> Hi,
> 
> On Wed, 13 Dec 2006, Andreas Ericsson wrote:
> 
>> Junio C Hamano wrote:
>>> "Bahadir Balban" <bahadir.balban@gmail.com> writes:
>>>
>>> There is one thing we could further optimize, though.
>>>
>>> Switching branches with 100k blobs in a commit even when there
>>> are a handful paths different between the branches would still
>>> need to populate the index by reading two trees and collapsing
>>> them into a single stage.  In theory, we should be able to do a
>>> lot better if two-tree case of read-tree took advanrage of
>>> cache-tree information.  If ce_match_stat() says Ok for all
>>> paths in a subdirectory and the cached tree object name for that
>>> subdirectory in the index match what we are reading from the new
>>> tree, we should be able to skip reading that subdirectory (and
>>> its subdirectories) from the new tree object at all.
>>>
>>> Anybody interested to give it a try?
>>>
>> I'm not vell-versed enough in git internals to have my hopes high of 
>> making something useful of it, but if you give me a pointer of where to 
>> start I'd be happy to try, and perhaps learn something in the process.
> 
> Okay, I'll have a stab at explaining it.
> 
> For huge working directories, you usually have a huge number of trees. The 
> idea of cache_tree is to remember not only the stat information of the 
> blobs in the index, but to cache the hashes of the trees also (until they 
> are invalidated, e.g. by an update-index). This avoids recalculation of 
> the hashes when committing.
> 
> This cache is accessible by the global variable active_cache_tree. It is 
> best accessed by the function cache_tree_find(), which you call like that:
> 
> 	struct cache_tree *ct = cache_tree_find(active_cache_tree, path);
> 
> where the variable "path" may contain slashes. The SHA1 of the 
> corresponding tree is in ct->sha1, and you can check if the hash is still 
> valid by asking
> 
> 	if (cache_tree_fully_valid(ct))
> 		/* still valid */
> 
> AFAIU Junio would like to take the shortcut of doing nothing at all when 
> (twoway) reading a tree whose hash is identical to the hash stored in the 
> corresponding cache_tree _and_ when the cache is still fully valid.
> 
Seems you wrote half the code for me already. :)

Thanks for the excellent explanation. I'll see if I can grok it further tonight.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Jakub Narebski· Dec 12, 2006, 12:36 UTC · re: Bahadir Balban · lore
Bahadir Balban wrote:
Show 10 quoted lines
> When I initialise a git repository, I use a subset of files in the
> project and leave out irrelevant files for performance reasons. Then
> when I need to make changes to a file not yet in the repository, the
> file is treated as new, and if I reset the change or change branches
> the file is gone.
> 
> Is there a good way of adding new files to git as if they had existed
> from the initial commit (or even better, since a particular commit)?
> This way I would only track the new changes I made to an existing
> file.

Generally, it is not possible without rewriting history. In git (in any sane SCM) commits are atomic; there is no CVS-like bunch of per-file histories. You can use cg-admin-rewritehist from Cogito (alternate UI for git)... but as it was said somewhere else git is fast. And the rule of thumb: check first, then optimize.

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git

← back to recent threads