{"thread":{"id":"43121","subject":"Re: Adding a new file as if it had existed","startedAt":"2006-12-12T10:05:08Z","lastAt":"2006-12-13T15:52:46Z","messageCount":11,"participants":["Johannes Schindelin","Andreas Ericsson","Junio C Hamano","Andy Parkins","Bahadir Balban","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"295971","messageId":"7ac1e90c0612120205k38b2fc14jbfd8ea682406efb2@mail.gmail.com","threadId":"43121","inReplyTo":null,"subject":"Adding a new file as if it had existed","fromName":"Bahadir Balban","fromEmail":"bahadir.balban@gmail.com","sentAt":"2006-12-12T10:05:08Z","receivedAt":"2006-12-12T10:05:08Z","isPatch":false,"sender":{"key":"bahadir.balban@gmail.com","avatar":null},"body":"Hi,\n\nWhen I initialise a git repository, I use a subset of files in the\nproject and leave out irrelevant files for performance reasons. Then\nwhen I need to make changes to a file not yet in the repository, the\nfile is treated as new, and if I reset the change or change branches\nthe file is gone.\n\nIs there a good way of adding new files to git as if they had existed\nfrom the initial commit (or even better, since a particular commit)?\nThis way I would only track the new changes I made to an existing\nfile.\n\nThanks,\n"},{"id":"296820","messageId":"7vhcw1whfx.fsf@assigned-by-dhcp.cox.net","threadId":"43121","inReplyTo":"7ac1e90c0612120205k38b2fc14jbfd8ea682406efb2@mail.gmail.com","subject":"Re: Adding a new file as if it had existed","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-12T10:13:06Z","receivedAt":"2006-12-12T10:13:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Bahadir Balban\" <bahadir.balban@gmail.com> writes:\n\n> Is there a good way of adding new files to git as if they had existed\n> from the initial commit (or even better, since a particular commit)?\n> This way I would only track the new changes I made to an existing\n> file.\n\nNo.\n\nI do not understand why not adding all the files you care about\neventually anyway in the initial commit is needed for\n\"performance reasons\", if you do not touch majority of them for\na long time.  Care to explain?\n\n"},{"id":"295999","messageId":"7ac1e90c0612120332o20d6778bsa16a788fdc04a3a1@mail.gmail.com","threadId":"43121","inReplyTo":"7vhcw1whfx.fsf@assigned-by-dhcp.cox.net","subject":"Re: Adding a new file as if it had existed","fromName":"Bahadir Balban","fromEmail":"bahadir.balban@gmail.com","sentAt":"2006-12-12T11:32:28Z","receivedAt":"2006-12-12T11:32:28Z","isPatch":false,"sender":{"key":"bahadir.balban@gmail.com","avatar":null},"body":"On 12/12/06, Junio C Hamano <junkio@cox.net> wrote:\n> No.\n>\n> I do not understand why not adding all the files you care about\n> eventually anyway in the initial commit is needed for\n> \"performance reasons\", if you do not touch majority of them for\n> a long time.  Care to explain?\n\nIf I don't know which files I may be touching in the future for\nimplementing some feature, then I am obliged to add all the files even\nif they are irrelevant. I said \"performance reasons\" assuming all the\nfile hashes need checked for every commit -a to see if they're\nchanged, but I just tried on a PIII and it seems not so slow.\n\n"},{"id":"294409","messageId":"Pine.LNX.4.63.0612121305430.2807@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43121","inReplyTo":"7ac1e90c0612120332o20d6778bsa16a788fdc04a3a1@mail.gmail.com","subject":"Re: Adding a new file as if it had existed","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-12T12:07:22Z","receivedAt":"2006-12-12T12:07:22Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 12 Dec 2006, Bahadir Balban wrote:\n\n> On 12/12/06, Junio C Hamano <junkio@cox.net> wrote:\n> > No.\n> > \n> > I do not understand why not adding all the files you care about\n> > eventually anyway in the initial commit is needed for\n> > \"performance reasons\", if you do not touch majority of them for\n> > a long time.  Care to explain?\n> \n> If I don't know which files I may be touching in the future for\n> implementing some feature,\n\nWhen I use an SCM, it is to track the revisions of a project. It seems you \nare content to have only parts of a revision? That does not make sense to \nme.\n\n> I said \"performance reasons\" assuming all the file hashes need checked \n> for every commit -a to see if they're changed, but I just tried on a \n> PIII and it seems not so slow.\n\nBingo!\n\nYou just felt the consequences of the \"index\".\n\nCiao,\nDscho\n"},{"id":"295108","messageId":"200612121226.32772.andyparkins@gmail.com","threadId":"43121","inReplyTo":"7ac1e90c0612120332o20d6778bsa16a788fdc04a3a1@mail.gmail.com","subject":"Re: Adding a new file as if it had existed","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2006-12-12T12:26:29Z","receivedAt":"2006-12-12T12:26:29Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2006 December 12 11:32, Bahadir Balban wrote:\n\n> If I don't know which files I may be touching in the future for\n> implementing some feature, then I am obliged to add all the files even\n> if they are irrelevant. I said \"performance reasons\" assuming all the\n> file hashes need checked for every commit -a to see if they're\n> changed, but I just tried on a PIII and it seems not so slow.\n\nHere's a handy rule of thumb I've learned in my use of git:\n\n \"git is fast.  Really fast.\"\n\nThat'll hold you in good stead.  In my experience there is no operation in git \nthat is slow.  I've got some trees that are for embedded work and hold the \nwhole linux kernel, often more than once.  Subversion, which I used \npreviously, took literally hours to import the whole tree.  Git takes \nminutes.\n\nAs to your direct concern: git doesn't hash every file at every commit.  There \nis no need.  git has an \"index\" that is used to prepare a commit; at the time \nyou do the actual commit, git already knows which files are being checked in.  \nObviously, Linus uses git for managing the linux kernel, he's said before \nthat he wanted a version control system that can do multiple commits /per \nsecond/.  git can do that.\n\nIn short - don't worry about making life easy for git - it's a workhorse and \ndoes a grand job.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIEE\n"},{"id":"297135","messageId":"elm7ji$m6g$1@sea.gmane.org","threadId":"43121","inReplyTo":"7ac1e90c0612120205k38b2fc14jbfd8ea682406efb2@mail.gmail.com","subject":"Re: Adding a new file as if it had existed","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-12-12T12:36:16Z","receivedAt":"2006-12-12T12:36:16Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Bahadir Balban wrote:\n\n> When I initialise a git repository, I use a subset of files in the\n> project and leave out irrelevant files for performance reasons. Then\n> when I need to make changes to a file not yet in the repository, the\n> file is treated as new, and if I reset the change or change branches\n> the file is gone.\n> \n> Is there a good way of adding new files to git as if they had existed\n> from the initial commit (or even better, since a particular commit)?\n> This way I would only track the new changes I made to an existing\n> file.\n\nGenerally, it is not possible without rewriting history. In git (in any\nsane SCM) commits are atomic; there is no CVS-like bunch of per-file\nhistories. You can use cg-admin-rewritehist from Cogito (alternate UI\nfor git)... but as it was said somewhere else git is fast. And the rule\nof thumb: check first, then optimize.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n\n"},{"id":"296623","messageId":"457EACA0.7050208@op5.se","threadId":"43121","inReplyTo":"200612121226.32772.andyparkins@gmail.com","subject":"Re: Adding a new file as if it had existed","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-12T13:20:32Z","receivedAt":"2006-12-12T13:20:32Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Andy Parkins wrote:\n> On Tuesday 2006 December 12 11:32, Bahadir Balban wrote:\n> \n>> If I don't know which files I may be touching in the future for\n>> implementing some feature, then I am obliged to add all the files even\n>> if they are irrelevant. I said \"performance reasons\" assuming all the\n>> file hashes need checked for every commit -a to see if they're\n>> changed, but I just tried on a PIII and it seems not so slow.\n> \n> Here's a handy rule of thumb I've learned in my use of git:\n> \n>  \"git is fast.  Really fast.\"\n> \n\nAlmost alarmingly so. When I started using git (back in May/June last \nyear, when git was 2 - 3 months old), I was worried at first because it \ndidn't seem to actually *do* anything, but just returned me to the \nprompt immediately.\n\n> \n> As to your direct concern: git doesn't hash every file at every commit.  There \n> is no need.  git has an \"index\" that is used to prepare a commit; at the time \n> you do the actual commit, git already knows which files are being checked in.  \n> \n> In short - don't worry about making life easy for git - it's a workhorse and \n> does a grand job.\n> \n\nYup. Now I've gone the other way around and think other scm's are broken \nwhen they chew disk for 10 seconds whenever I try to do anything with \nthem. I usually end up importing the other repo into git and do my work \nthere.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"294763","messageId":"7vzm9tuft7.fsf@assigned-by-dhcp.cox.net","threadId":"43121","inReplyTo":"7ac1e90c0612120332o20d6778bsa16a788fdc04a3a1@mail.gmail.com","subject":"Re: Adding a new file as if it had existed","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-12T18:31:16Z","receivedAt":"2006-12-12T18:31:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Bahadir Balban\" <bahadir.balban@gmail.com> writes:\n\n> ... I said \"performance reasons\" assuming all the\n> file hashes need checked for every commit -a to see if they're\n> changed, but I just tried on a PIII and it seems not so slow.\n\nOk.\n\nOther people have already cleared the fear for 'commit' case, so\nI hope you are happier.\n\nThere is one thing we could further optimize, though.\n\nSwitching branches with 100k blobs in a commit even when there\nare a handful paths different between the branches would still\nneed to populate the index by reading two trees and collapsing\nthem into a single stage.  In theory, we should be able to do a\nlot better if two-tree case of read-tree took advanrage of\ncache-tree information.  If ce_match_stat() says Ok for all\npaths in a subdirectory and the cached tree object name for that\nsubdirectory in the index match what we are reading from the new\ntree, we should be able to skip reading that subdirectory (and\nits subdirectories) from the new tree object at all.\n\nAnybody interested to give it a try?\n\n"},{"id":"294124","messageId":"457FCA8C.6000300@op5.se","threadId":"43121","inReplyTo":"7vzm9tuft7.fsf@assigned-by-dhcp.cox.net","subject":"Re: Adding a new file as if it had existed","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-13T09:40:28Z","receivedAt":"2006-12-13T09:40:28Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> \"Bahadir Balban\" <bahadir.balban@gmail.com> writes:\n> \n> There is one thing we could further optimize, though.\n> \n> Switching branches with 100k blobs in a commit even when there\n> are a handful paths different between the branches would still\n> need to populate the index by reading two trees and collapsing\n> them into a single stage.  In theory, we should be able to do a\n> lot better if two-tree case of read-tree took advanrage of\n> cache-tree information.  If ce_match_stat() says Ok for all\n> paths in a subdirectory and the cached tree object name for that\n> subdirectory in the index match what we are reading from the new\n> tree, we should be able to skip reading that subdirectory (and\n> its subdirectories) from the new tree object at all.\n> \n> Anybody interested to give it a try?\n> \n\nI'm not vell-versed enough in git internals to have my hopes high of \nmaking something useful of it, but if you give me a pointer of where to \nstart I'd be happy to try, and perhaps learn something in the process.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"293877","messageId":"Pine.LNX.4.63.0612131611050.3635@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43121","inReplyTo":"457FCA8C.6000300@op5.se","subject":"Re: Adding a new file as if it had existed","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-13T15:46:32Z","receivedAt":"2006-12-13T15:46:32Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 13 Dec 2006, Andreas Ericsson wrote:\n\n> Junio C Hamano wrote:\n> > \"Bahadir Balban\" <bahadir.balban@gmail.com> writes:\n> > \n> > There is one thing we could further optimize, though.\n> > \n> > Switching branches with 100k blobs in a commit even when there\n> > are a handful paths different between the branches would still\n> > need to populate the index by reading two trees and collapsing\n> > them into a single stage.  In theory, we should be able to do a\n> > lot better if two-tree case of read-tree took advanrage of\n> > cache-tree information.  If ce_match_stat() says Ok for all\n> > paths in a subdirectory and the cached tree object name for that\n> > subdirectory in the index match what we are reading from the new\n> > tree, we should be able to skip reading that subdirectory (and\n> > its subdirectories) from the new tree object at all.\n> > \n> > Anybody interested to give it a try?\n> > \n> \n> I'm not vell-versed enough in git internals to have my hopes high of \n> making something useful of it, but if you give me a pointer of where to \n> start I'd be happy to try, and perhaps learn something in the process.\n\nOkay, I'll have a stab at explaining it.\n\nFor huge working directories, you usually have a huge number of trees. The \nidea of cache_tree is to remember not only the stat information of the \nblobs in the index, but to cache the hashes of the trees also (until they \nare invalidated, e.g. by an update-index). This avoids recalculation of \nthe hashes when committing.\n\nThis cache is accessible by the global variable active_cache_tree. It is \nbest accessed by the function cache_tree_find(), which you call like that:\n\n\tstruct cache_tree *ct = cache_tree_find(active_cache_tree, path);\n\nwhere the variable \"path\" may contain slashes. The SHA1 of the \ncorresponding tree is in ct->sha1, and you can check if the hash is still \nvalid by asking\n\n\tif (cache_tree_fully_valid(ct))\n\t\t/* still valid */\n\nAFAIU Junio would like to take the shortcut of doing nothing at all when \n(twoway) reading a tree whose hash is identical to the hash stored in the \ncorresponding cache_tree _and_ when the cache is still fully valid.\n\nCiao,\nDscho\n"},{"id":"295219","messageId":"458021CE.1000407@op5.se","threadId":"43121","inReplyTo":"Pine.LNX.4.63.0612131611050.3635@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Adding a new file as if it had existed","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-13T15:52:46Z","receivedAt":"2006-12-13T15:52:46Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Wed, 13 Dec 2006, Andreas Ericsson wrote:\n> \n>> Junio C Hamano wrote:\n>>> \"Bahadir Balban\" <bahadir.balban@gmail.com> writes:\n>>>\n>>> There is one thing we could further optimize, though.\n>>>\n>>> Switching branches with 100k blobs in a commit even when there\n>>> are a handful paths different between the branches would still\n>>> need to populate the index by reading two trees and collapsing\n>>> them into a single stage.  In theory, we should be able to do a\n>>> lot better if two-tree case of read-tree took advanrage of\n>>> cache-tree information.  If ce_match_stat() says Ok for all\n>>> paths in a subdirectory and the cached tree object name for that\n>>> subdirectory in the index match what we are reading from the new\n>>> tree, we should be able to skip reading that subdirectory (and\n>>> its subdirectories) from the new tree object at all.\n>>>\n>>> Anybody interested to give it a try?\n>>>\n>> I'm not vell-versed enough in git internals to have my hopes high of \n>> making something useful of it, but if you give me a pointer of where to \n>> start I'd be happy to try, and perhaps learn something in the process.\n> \n> Okay, I'll have a stab at explaining it.\n> \n> For huge working directories, you usually have a huge number of trees. The \n> idea of cache_tree is to remember not only the stat information of the \n> blobs in the index, but to cache the hashes of the trees also (until they \n> are invalidated, e.g. by an update-index). This avoids recalculation of \n> the hashes when committing.\n> \n> This cache is accessible by the global variable active_cache_tree. It is \n> best accessed by the function cache_tree_find(), which you call like that:\n> \n> \tstruct cache_tree *ct = cache_tree_find(active_cache_tree, path);\n> \n> where the variable \"path\" may contain slashes. The SHA1 of the \n> corresponding tree is in ct->sha1, and you can check if the hash is still \n> valid by asking\n> \n> \tif (cache_tree_fully_valid(ct))\n> \t\t/* still valid */\n> \n> AFAIU Junio would like to take the shortcut of doing nothing at all when \n> (twoway) reading a tree whose hash is identical to the hash stored in the \n> corresponding cache_tree _and_ when the cache is still fully valid.\n> \n\nSeems you wrote half the code for me already. :)\n\nThanks for the excellent explanation. I'll see if I can grok it further \ntonight.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"}]}