{"thread":{"id":"43125","subject":"Re: cygwin, 44k files: how to commit only index?","startedAt":"2006-12-07T14:27:36Z","lastAt":"2006-12-09T08:27:46Z","messageCount":18,"participants":["Alex Riesen","Junio C Hamano","Christian MICHON","Shawn Pearce","Torgil Svensson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"296703","messageId":"81b0412b0612070627r3ff0b394s124d95fbf8084f16@mail.gmail.com","threadId":"43125","inReplyTo":null,"subject":"cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-12-07T14:27:36Z","receivedAt":"2006-12-07T14:27:36Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"I have a kind of awkward project to work with (~44k files, many binaries).\n\nThe normal \"git commit\", which seem to be more than enough\nfor anything and anyone else, is a really annoying procedure\nin my context. It spend too much time refreshing index and\ngenerating list of the files for the commit message.\n\nAt first I stopped using git commit -a (doing only update-index),\nnow I'm about to start using write-tree/commit-tree/update-ref\ndirectly. It helps, but sometimes I really miss -F/-C. It's also\nugly: I can (and almost did) commit an unchanged tree.\n\nIs there any simple way to modify git commit for such a workflow?\nFailing that, any simple and _fast_ way to find out if the index\n"},{"id":"293919","messageId":"7vd56vtt2g.fsf@assigned-by-dhcp.cox.net","threadId":"43125","inReplyTo":"81b0412b0612070627r3ff0b394s124d95fbf8084f16@mail.gmail.com","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-07T19:16:39Z","receivedAt":"2006-12-07T19:16:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Alex Riesen\" <raa.lkml@gmail.com> writes:\n\n> I have a kind of awkward project to work with (~44k files, many binaries).\n>\n> The normal \"git commit\", which seem to be more than enough\n> for anything and anyone else, is a really annoying procedure\n> in my context. It spend too much time refreshing index and\n> generating list of the files for the commit message.\n>\n> At first I stopped using git commit -a (doing only update-index),\n\nI am not sure what you are trying.  Do you mean stat() is slow\non your filesystem?\n\n> Is there any simple way to modify git commit for such a workflow?\n> Failing that, any simple and _fast_ way to find out if the index\n> is any different from HEAD? (so that I don't produce empty commits).\n\nMaybe you want \"assume unchanged\"?\n"},{"id":"296935","messageId":"20061207192632.GC12143@spearce.org","threadId":"43125","inReplyTo":"7vd56vtt2g.fsf@assigned-by-dhcp.cox.net","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-07T19:26:32Z","receivedAt":"2006-12-07T19:26:32Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> \"Alex Riesen\" <raa.lkml@gmail.com> writes:\n> \n> > I have a kind of awkward project to work with (~44k files, many binaries).\n> >\n> > The normal \"git commit\", which seem to be more than enough\n> > for anything and anyone else, is a really annoying procedure\n> > in my context. It spend too much time refreshing index and\n> > generating list of the files for the commit message.\n> >\n> > At first I stopped using git commit -a (doing only update-index),\n> \n> I am not sure what you are trying.  Do you mean stat() is slow\n> on your filesystem?\n\nIts Cygwin/NTFS.  lstat() is slow.  readdir() is slow.  I have the\nsame problem on my Cygwin systems.\n \n> > Is there any simple way to modify git commit for such a workflow?\n> > Failing that, any simple and _fast_ way to find out if the index\n> > is any different from HEAD? (so that I don't produce empty commits).\n> \n> Maybe you want \"assume unchanged\"?\n\nYes, basically.  The Cygwin/NTFS issues Alex is pointing out are\nexactly why git-gui has a \"Trust File Modification Timestamp\" option\non both a per-repository and global level.  My larger repositories\n(~10k files) are difficult to work with without that option enabled.\n\n-- \n"},{"id":"298763","messageId":"20061207193555.GD12143@spearce.org","threadId":"43125","inReplyTo":"20061207192632.GC12143@spearce.org","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-07T19:35:55Z","receivedAt":"2006-12-07T19:35:55Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Shawn Pearce <spearce@spearce.org> wrote:\n> Its Cygwin/NTFS.  lstat() is slow.  readdir() is slow.  I have the\n> same problem on my Cygwin systems.\n\nJust to be clear, I'm not trying to blame Cygwin here.\n\nWindows' dir command is slow.  Windows Explorer is slow while\nbrowsing directories.  Eclipse chugs hard while doing any directory\nscans (it normally runs very fast if its not rescanning the entire\ndirectory structure).  The drive is just plain slow.\n\nYea, I know, get a faster disk... but some bean counters don't\nbelieve that a $50 more expensive disk could ever save enough time\nto warrant the extra $50 captial expenditure...\n\nI spend at least an hour a week waiting for enough IO to finish so\nthat the mouse pointer will move again.  *sigh*\n\n-- \n"},{"id":"296333","messageId":"7vhcw7scln.fsf@assigned-by-dhcp.cox.net","threadId":"43125","inReplyTo":"20061207192632.GC12143@spearce.org","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-07T19:57:40Z","receivedAt":"2006-12-07T19:57:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n>> I am not sure what you are trying.  Do you mean stat() is slow\n>> on your filesystem?\n>\n> Its Cygwin/NTFS.  lstat() is slow.  readdir() is slow.  I have the\n> same problem on my Cygwin systems.\n>  \n>> > Is there any simple way to modify git commit for such a workflow?\n>> > Failing that, any simple and _fast_ way to find out if the index\n>> > is any different from HEAD? (so that I don't produce empty commits).\n>> \n>> Maybe you want \"assume unchanged\"?\n>\n> Yes, basically.\n\nThen maybe \"git grep assume.unchanged\" would help?\n"},{"id":"298369","messageId":"20061207202931.GB12502@spearce.org","threadId":"43125","inReplyTo":"7vhcw7scln.fsf@assigned-by-dhcp.cox.net","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-07T20:29:31Z","receivedAt":"2006-12-07T20:29:31Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n> \n> >> I am not sure what you are trying.  Do you mean stat() is slow\n> >> on your filesystem?\n> >> \n> >> Maybe you want \"assume unchanged\"?\n> >\n> > Yes, basically.\n> \n> Then maybe \"git grep assume.unchanged\" would help?\n\nHmm.  OK, maybe I should have answered \"No\"\" to your first question.\nI keep looking at the assume unchanaged feature of update-index,\nbut refuse to use it because I'm a lazy guy who will forget to tell\nthe index a file has been modified.  Consequently I'm going to miss\na change during a commit.\n\nWhat may help (and without using assume unchanged) is:\n\n * skip the `update-index --refresh` part of git-status/git-commit\n * skip the status template in COMMIT_MSG when using the editor\n\nAs Git will still at least make sure a `commit -a` includes\neverything that is dirty.\n\nFiles whose modification dates may have been messed with (but\nwhose content are unchanged) will just go through expensive SHA1\ncomputation to arrive at the same value, which is fine.\n\nUsers skipping the first part are doing so under the assumption that\ntheir modification dates are usually always correct, and that then\nthey aren't the SHA1 computation of a handful of files is cheap\ncompared to stat'ing the entire set of files.\n\nUsers skipping the second part are doing so under the assumption\nthat knowing the names of the files they are committing doesn't\nreally improve their odds of writing a good commit message.\n\n-- \n"},{"id":"295872","messageId":"46d6db660612071326m4817165l992e8d6e7bd673c5@mail.gmail.com","threadId":"43125","inReplyTo":"20061207193555.GD12143@spearce.org","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Christian MICHON","fromEmail":"christian.michon@gmail.com","sentAt":"2006-12-07T21:26:30Z","receivedAt":"2006-12-07T21:26:30Z","isPatch":false,"sender":{"key":"christian.michon@gmail.com","avatar":"https://gravatar.com/avatar/8a7c327b21187fbcab5c27640a49450eec72e0355dc292501197f27a5a744ec4?d=mp&s=160"},"body":"On 12/7/06, Shawn Pearce <spearce@spearce.org> wrote:\n> Shawn Pearce <spearce@spearce.org> wrote:\n> > Its Cygwin/NTFS.  lstat() is slow.  readdir() is slow.  I have the\n> > same problem on my Cygwin systems.\n>\n> Just to be clear, I'm not trying to blame Cygwin here.\n>\n> Windows' dir command is slow.  Windows Explorer is slow while\n> browsing directories.  Eclipse chugs hard while doing any directory\n> scans (it normally runs very fast if its not rescanning the entire\n> directory structure).  The drive is just plain slow.\n> (...)\n\nbefore buying any new hardware, you could easily imagine the\nfollowing scenario (I'm also \"stuck\" with windows, so it's an idea\nI've been toying around for a week or so).\n\nThere're virtualizers around, on which networking capabilities can\nbe activated. And we could easily create a vm with linux+git\ninside, using ext2/ext3/ext4 fs virtual disks (you'd benefit from\nwindows cache actually...)\n\nexample: YTech_Subversion_Appliance_v1.1 (ubuntu + subversion).\n\nI've no prototype yet, but I've 2 scenario possible:\n1) use vmplayer and a minimal uclibc initramfs with git onboard\n2) use qemu+kqemu and a similar mini-distro (but right now networking\nis an issue on windows hosts: I'm exploring tunneling)\n\nThe 1st scenario is \"easy\". And I start to prefer this idea over\neven mingw porting of git (I tried and it's hard, really).\n\nBut again, maybe jgit would be a better universal solution.\n\n-- \n"},{"id":"297639","messageId":"7vzm9zqsnj.fsf@assigned-by-dhcp.cox.net","threadId":"43125","inReplyTo":"20061207202931.GB12502@spearce.org","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-07T21:53:52Z","receivedAt":"2006-12-07T21:53:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> What may help (and without using assume unchanged) is:\n>\n>  * skip the `update-index --refresh` part of git-status/git-commit\n>  * skip the status template in COMMIT_MSG when using the editor\n>\n> As Git will still at least make sure a `commit -a` includes\n> everything that is dirty.\n>\n> Files whose modification dates may have been messed with (but\n> whose content are unchanged) will just go through expensive SHA1\n> computation to arrive at the same value, which is fine.\n>\n> Users skipping the first part are doing so under the assumption that\n> their modification dates are usually always correct, and that then\n> they aren't the SHA1 computation of a handful of files is cheap\n> compared to stat'ing the entire set of files.\n>\n> Users skipping the second part are doing so under the assumption\n> that knowing the names of the files they are committing doesn't\n> really improve their odds of writing a good commit message.\n\nThe second part is not about a good commit message but more\nabout a path that should have been updated but forgotten (the\nsame mistake you would be likely to make and that is the reason\nassume-unchanged is not good for you).\n\nI do not mind too much if you added a new --quick option to \"git\ncommit\" for this rather specialized need.\n"},{"id":"293883","messageId":"20061207221503.GA4990@steel.home","threadId":"43125","inReplyTo":"7vd56vtt2g.fsf@assigned-by-dhcp.cox.net","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"fork0@t-online.de","sentAt":"2006-12-07T22:15:03Z","receivedAt":"2006-12-07T22:15:03Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Junio C Hamano, Thu, Dec 07, 2006 20:16:39 +0100:\n> > I have a kind of awkward project to work with (~44k files, many binaries).\n> >\n> > The normal \"git commit\", which seem to be more than enough\n> > for anything and anyone else, is a really annoying procedure\n> > in my context. It spend too much time refreshing index and\n> > generating list of the files for the commit message.\n> >\n> > At first I stopped using git commit -a (doing only update-index),\n> \n> I am not sure what you are trying.  Do you mean stat() is slow\n> on your filesystem?\n\nincredibly slow. That and the matter of having 44000 files to process\nwith that slow stat().\n\n> > Is there any simple way to modify git commit for such a workflow?\n> > Failing that, any simple and _fast_ way to find out if the index\n> > is any different from HEAD? (so that I don't produce empty commits).\n> \n> Maybe you want \"assume unchanged\"?\n> \n\nIf that is core.ignoreState you mean, than maybe this is what I mean.\nI haven't tried it yet (now I wonder myself why I haven't tried it).\nBut (I'm repeating myself, in <81b0412b0612060235l5d5f93d0hd1aaf34924f7783@mail.gmail.com>)\nI do not really understand how it _can_ help: \"I ask because it does\nnot ignore stat info, as the name implies. Because if it would,\nthere'd be no point of calling lstat at all, wouldn't it?\" That last\nquestion was about refresh_cache_entry - it calls lstat\nunconditionally.\n\nStill, I guess I'll have to try it.\n\nBut aside from me trying ignoreState, can anyone help me with that\nquestion regarding checking if the index is any different from HEAD?\nBecause even on a very brocken filesystem and 40k files in a repo you\nsometimes do want to call git-update-index --refresh just to be sure\nyou haven't missed anything. And than it'll quickly become annoying\nflicking ignoreState back and forth.\n"},{"id":"298015","messageId":"7vr6vbqqzh.fsf@assigned-by-dhcp.cox.net","threadId":"43125","inReplyTo":"20061207221503.GA4990@steel.home","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-07T22:29:54Z","receivedAt":"2006-12-07T22:29:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"fork0@t-online.de (Alex Riesen) writes:\n\n> But aside from me trying ignoreState, can anyone help me with that\n> question regarding checking if the index is any different from HEAD?\n\nComparing index and HEAD should be cheap on a system with slow\nlstat(), I think, as \"git-diff-index --cached HEAD\" should just\nignore the working tree altogether.  Is that what you want?\n"},{"id":"297455","messageId":"20061208052705.GA4318@steel.home","threadId":"43125","inReplyTo":"7vr6vbqqzh.fsf@assigned-by-dhcp.cox.net","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"fork0@t-online.de","sentAt":"2006-12-08T05:27:05Z","receivedAt":"2006-12-08T05:27:05Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Junio C Hamano, Thu, Dec 07, 2006 23:29:54 +0100:\n> > But aside from me trying ignoreState, can anyone help me with that\n> > question regarding checking if the index is any different from HEAD?\n> \n> Comparing index and HEAD should be cheap on a system with slow\n> lstat(), I think, as \"git-diff-index --cached HEAD\" should just\n> ignore the working tree altogether.  Is that what you want?\n> \n\nyes, except that it'll compare the whole trees. Could I make it stop\nat first mismatch? \"-q|--quiet\" for git-diff-index perhaps?\nIt's just not only stat, but also, open, read, mmap (yes, I try to use\nit for packs) and close are really slow here as well.\n"},{"id":"295808","messageId":"7vzm9ynahc.fsf@assigned-by-dhcp.cox.net","threadId":"43125","inReplyTo":"20061208052705.GA4318@steel.home","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-08T06:54:39Z","receivedAt":"2006-12-08T06:54:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"fork0@t-online.de (Alex Riesen) writes:\n\n> yes, except that it'll compare the whole trees. Could I make it stop\n> at first mismatch? \"-q|--quiet\" for git-diff-index perhaps?\n> It's just not only stat, but also, open, read, mmap (yes, I try to use\n> it for packs) and close are really slow here as well.\n\nThat sounds like optimizing for a wrong case -- you expect the\nindex to match HEAD and trying to catch mistakes by detecting\na mismatch, right?\n\nHaving said that, I should point out that it is a low hanging\nfruit to optimize \"diff-index --cached\" for cases where index\nis expected to mostly match HEAD.\n\nThe current code for \"diff-index --cached\" reads the whole tree\ninto the index as stage #1 entries (diff-lib.c::run_diff_index),\nand then compares stage #0 (from the original index contents)\nand stage #1 (the tree parameter from the command line).  Even\nif you stop at the first mismatch, you would already have paid\nthe overhead to open and read all tree objects before even\nstarting the comparison.\n\nHowever, this code is from the ancient time before cache-tree\nwas introduced in the index.  If the index is expected to mostly\nmatch HEAD, most of the cache-tree nodes are up-to-date, and\nwhole subtree can be skipped with a single comparison between\ntwo tree SHA-1s at a shallower level of the directory tree.\n\nIn 'pu' (jc/diff topic), I have a very generic code to walk the\nindex, working tree and zero or more trees in parallel, taking\nadvantage of cache-tree.  If somebody is interested to learn the\ninternals of git, some of the code could be lifted from there\nand simplified to walk just the index and a single tree, and I\nthink that would optimize \"diff-index --cached\" quite a bit.\n\nA very unscientific test of running in the kernel repository I\njust pulled (hot cache) on my box is:\n\n$ /usr/bin/time git diff-index -r --cached --abbrev v2.6.19 >/tmp/1\n0.91user 0.20system 0:01.12elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+10949minor)pagefaults 0swaps\n\nwhile the para-walk to produce the moral equivalent is:\n\n$ /usr/bin/time test-para --no-work v2.6.19 >/tmp/2\n0.11user 0.02system 0:00.13elapsed 98%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+4524minor)pagefaults 0swaps\n"},{"id":"296866","messageId":"81b0412b0612072327x77477584jb9131b26b0854f2@mail.gmail.com","threadId":"43125","inReplyTo":"7vzm9ynahc.fsf@assigned-by-dhcp.cox.net","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-12-08T07:27:08Z","receivedAt":"2006-12-08T07:27:08Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 12/8/06, Junio C Hamano <junkio@cox.net> wrote:\n> > yes, except that it'll compare the whole trees. Could I make it stop\n> > at first mismatch? \"-q|--quiet\" for git-diff-index perhaps?\n> > It's just not only stat, but also, open, read, mmap (yes, I try to use\n> > it for packs) and close are really slow here as well.\n>\n> That sounds like optimizing for a wrong case -- you expect the\n> index to match HEAD and trying to catch mistakes by detecting\n> a mismatch, right?\n\nI expect the index to differ from HEAD. The test is to avoid the mistake\nof doing an empty commit.\n\n> Having said that, I should point out that it is a low hanging\n> fruit to optimize \"diff-index --cached\" for cases where index\n> is expected to mostly match HEAD.\n>\n> The current code for \"diff-index --cached\" reads the whole tree\n> into the index as stage #1 entries (diff-lib.c::run_diff_index),\n> and then compares stage #0 (from the original index contents)\n> and stage #1 (the tree parameter from the command line).  Even\n> if you stop at the first mismatch, you would already have paid\n> the overhead to open and read all tree objects before even\n> starting the comparison.\n\nBut I don't have to pay for the overhead of comparing all\nentries, if I can stop at first mismatch and exit with non-0.\nI think it'd make a difference (at least some difference).\nBut, if we could avoid loading of the entries which\nwill be never compared anyway, the speedup will be\nof course more substantial...\n\n> In 'pu' (jc/diff topic), I have a very generic code to walk the\n> index, working tree and zero or more trees in parallel, taking\n> advantage of cache-tree.  If somebody is interested to learn the\n> internals of git, some of the code could be lifted from there\n> and simplified to walk just the index and a single tree, and I\n> think that would optimize \"diff-index --cached\" quite a bit.\n\n"},{"id":"294542","messageId":"7vhcw6n8iz.fsf@assigned-by-dhcp.cox.net","threadId":"43125","inReplyTo":"81b0412b0612072327x77477584jb9131b26b0854f2@mail.gmail.com","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-12-08T07:36:52Z","receivedAt":"2006-12-08T07:36:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Alex Riesen\" <raa.lkml@gmail.com> writes:\n\n>> The current code for \"diff-index --cached\" reads the whole tree\n>> into the index as stage #1 entries (diff-lib.c::run_diff_index),\n>> and then compares stage #0 (from the original index contents)\n>> and stage #1 (the tree parameter from the command line).  Even\n>> if you stop at the first mismatch, you would already have paid\n>> the overhead to open and read all tree objects before even\n>> starting the comparison.\n>\n> But I don't have to pay for the overhead of comparing all\n> entries, if I can stop at first mismatch and exit with non-0.\n\nBench it if you doubt me.\n\nI'd bet that the time spent in comparison between stages inside\nindex (and remember, you are not generating textual diff, only\ncomparing the SHA-1) is dwarfed by the overhead of populating\nthe stage #1 of the index with what is read from all the tree\nobjects.\n\n"},{"id":"295937","messageId":"81b0412b0612072348q13faaab0s9ba7a235f7fd64dc@mail.gmail.com","threadId":"43125","inReplyTo":"7vhcw6n8iz.fsf@assigned-by-dhcp.cox.net","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-12-08T07:48:09Z","receivedAt":"2006-12-08T07:48:09Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 12/8/06, Junio C Hamano <junkio@cox.net> wrote:\n> >> The current code for \"diff-index --cached\" reads the whole tree\n> >> into the index as stage #1 entries (diff-lib.c::run_diff_index),\n> >> and then compares stage #0 (from the original index contents)\n> >> and stage #1 (the tree parameter from the command line).  Even\n> >> if you stop at the first mismatch, you would already have paid\n> >> the overhead to open and read all tree objects before even\n> >> starting the comparison.\n> >\n> > But I don't have to pay for the overhead of comparing all\n> > entries, if I can stop at first mismatch and exit with non-0.\n>\n> Bench it if you doubt me.\n\nI don't question that the overhead of comparing is very much\nunnoticable. It just that it surely isn't zero, and it will grow with\nthe size of repo (linearly, right?)... And I am sure that this\nrepo will _only_ grow (typical corporate project).\n\n> I'd bet that the time spent in comparison between stages inside\n> index (and remember, you are not generating textual diff, only\n> comparing the SHA-1) is dwarfed by the overhead of populating\n> the stage #1 of the index with what is read from all the tree\n> objects.\n\nI already understood that. I just haven't found yet what can I do\n"},{"id":"294491","messageId":"81b0412b0612080043y6ff3462ev5f8b4c4cf40182f5@mail.gmail.com","threadId":"43125","inReplyTo":"81b0412b0612072327x77477584jb9131b26b0854f2@mail.gmail.com","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-12-08T08:43:49Z","receivedAt":"2006-12-08T08:43:49Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 12/8/06, Alex Riesen <raa.lkml@gmail.com> wrote:\n> On 12/8/06, Junio C Hamano <junkio@cox.net> wrote:\n> > In 'pu' (jc/diff topic), I have a very generic code to walk the\n> > index, working tree and zero or more trees in parallel, taking\n> > advantage of cache-tree.  If somebody is interested to learn the\n> > internals of git, some of the code could be lifted from there\n> > and simplified to walk just the index and a single tree, and I\n> > think that would optimize \"diff-index --cached\" quite a bit.\n>\n> Will try to look at it.\n>\n\nAnd now I'm playing with that (against test-para.c from pu).\nI expect it to be broken by that webGmail, so it may not\napply to anything, but you'll get the idea. More clearly than\nfrom me trying to explain.\n\ncommit 83642cdaca6dc1a2f94aa41923bc9e8f02d0e12f\nAuthor: Alex Riesen <raa.lkml@gmail.com>\nDate:   Fri Dec 8 09:38:18 2006 +0100\n\n    add --quiet to test-para: stop at the first difference\n\ndiff --git a/test-para.c b/test-para.c\nindex bce5f0c..99d3792 100644\n--- a/test-para.c\n+++ b/test-para.c\n@@ -21,6 +21,7 @@ int main(int ac, const char **av)\n \tunsigned char trees[64][20];\n \tint num_tree = 0, i, using_head, show_all = 0;\n \tint index_wanted = 1, work_wanted = 1, tree_wanted = 0;\n+\tint quiet = 0;\n \tconst char *prefix;\n \tconst char **pathspec;\n\n@@ -43,6 +44,8 @@ int main(int ac, const char **av)\n \t\t\twork_wanted = 0;\n \t\telse if (!strcmp(av[1] + 2, \"no-index\"))\n \t\t\tindex_wanted = 0;\n+\t\telse if (!strcmp(av[1] + 2, \"quiet\"))\n+\t\t\tquiet = 1;\n \t\telse if (!av[1][2])\n \t\t\tbreak;\n \t\telse\n@@ -118,7 +121,9 @@ int main(int ac, const char **av)\n \t\t\t\tshow_one(z, e->name, e->namelen,\n \t\t\t\t\t e->hash, e->mode);\n \t\t\t}\n-\t\t\telse\n+\t\t\telse {\n+\t\t\t\tif (quiet)\n+\t\t\t\t\texit(1);\n \t\t\t\tfor (i = 0; i < w.num_trees + 2; i++) {\n \t\t\t\t\tchar numbuf[10];\n\n@@ -149,6 +154,7 @@ int main(int ac, const char **av)\n \t\t\t\t\tshow_one(z, e->name, e->namelen,\n \t\t\t\t\t\t e->hash, e->mode);\n \t\t\t\t}\n+\t\t\t}\n \t\t}\n\n"},{"id":"294917","messageId":"81b0412b0612080616t3d43f739gc201879fddcc20a3@mail.gmail.com","threadId":"43125","inReplyTo":"20061207221503.GA4990@steel.home","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-12-08T14:16:41Z","receivedAt":"2006-12-08T14:16:41Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 12/7/06, Alex Riesen <fork0@t-online.de> wrote:\n> > Maybe you want \"assume unchanged\"?\n>\n> If that is core.ignoreState you mean, than maybe this is what I mean.\n> I haven't tried it yet (now I wonder myself why I haven't tried it).\n> But (I'm repeating myself, in\n> <81b0412b0612060235l5d5f93d0hd1aaf34924f7783@mail.gmail.com>)\n> I do not really understand how it _can_ help: \"I ask because it does\n> not ignore stat info, as the name implies. Because if it would,\n> there'd be no point of calling lstat at all, wouldn't it?\" That last\n> question was about refresh_cache_entry - it calls lstat\n> unconditionally.\n>\n> Still, I guess I'll have to try it.\n>\n\nTried. No noticeable difference:\n\n$ git repo-config core.ignorestat true; time gup --refresh\nreal    0m8.004s\nuser    0m1.936s\nsys     0m5.702s\n$ git repo-config core.ignorestat false; time gup --refresh\nreal    0m7.787s\nuser    0m1.890s\nsys     0m5.703s\n$\n"},{"id":"297530","messageId":"e7bda7770612090027x22a06ca5i6d9b768f0ad3c4ad@mail.gmail.com","threadId":"43125","inReplyTo":"46d6db660612071326m4817165l992e8d6e7bd673c5@mail.gmail.com","subject":"Re: cygwin, 44k files: how to commit only index?","fromName":"Torgil Svensson","fromEmail":"torgil.svensson@gmail.com","sentAt":"2006-12-09T08:27:46Z","receivedAt":"2006-12-09T08:27:46Z","isPatch":false,"sender":{"key":"torgil.svensson@gmail.com","avatar":null},"body":"On 12/7/06, Christian MICHON <christian.michon@gmail.com> wrote:\n> On 12/7/06, Shawn Pearce <spearce@spearce.org> wrote:\n> > Shawn Pearce <spearce@spearce.org> wrote:\n> > > Its Cygwin/NTFS.  lstat() is slow.  readdir() is slow.  I have the\n> > > same problem on my Cygwin systems.\n> >\n> > Just to be clear, I'm not trying to blame Cygwin here.\n> >\n> > Windows' dir command is slow.  Windows Explorer is slow while\n> > browsing directories.\n\nI think this is a very common scenario costing hideous amounts of\nmoney around the globe.\n\nIf you have lot's of files in a folder, don't even think of\naccidentally touching those folders in Windows Explorer, if you do -\nkeep Process Explorer or similar ready. I've ended up using (even w/o\nCygwin) scripts, automatic compressing and even a database functioning\nas directory cache - basically creating accessibility layers for a\ndisabled file-system.\n\n\n>\n> before buying any new hardware, you could easily imagine the\n> following scenario (I'm also \"stuck\" with windows, so it's an idea\n> I've been toying around for a week or so).\n>\n> There're virtualizers around, on which networking capabilities can\n> be activated. And we could easily create a vm with linux+git\n> inside, using ext2/ext3/ext4 fs virtual disks (you'd benefit from\n> windows cache actually...)\n>\n> example: YTech_Subversion_Appliance_v1.1 (ubuntu + subversion).\n>\n> I've no prototype yet, but I've 2 scenario possible:\n> 1) use vmplayer and a minimal uclibc initramfs with git onboard\n> 2) use qemu+kqemu and a similar mini-distro (but right now networking\n> is an issue on windows hosts: I'm exploring tunneling)\n>\n> The 1st scenario is \"easy\". And I start to prefer this idea over\n> even mingw porting of git (I tried and it's hard, really).\n>\n> But again, maybe jgit would be a better universal solution.\n>\n> --\n> Christian\n> -\n\nVery interesting!  Have you a time-frame for this?  Maybe even\n"}]}