{"thread":{"id":"32862","subject":"inotify to minimize stat() calls","startedAt":"2013-02-08T21:10:28Z","lastAt":"2013-04-30T00:27:12Z","messageCount":88,"participants":["Ramkumar Ramachandra","Junio C Hamano","Duy Nguyen","Robert Zeh","demerphq","Erik Faye-Lund","Martin Fick","Karsten Blees","Jeff King","Magnus Bäck","Ævar Arnfjörð Bjarmason","Drew Northup","Torsten Bögershausen","Nguyễn Thái Ngọc Duy","Thomas Rast"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"209040","messageId":"CALkWK0=EP0Lv1F_BArub7SpL9rgFhmPtpMOCgwFqfJmVE=oa=A@mail.gmail.com","threadId":"32862","inReplyTo":null,"subject":"inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-08T21:10:28Z","receivedAt":"2013-02-08T21:10:28Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi,\n\nFor large repositories, many simple git commands like `git status`\ntake a while to respond.  I understand that this is because of large\nnumber of stat() calls to figure out which files were changed.  I\noverheard that Mercurial wants to solve this problem using itnotify,\nbut the idea bothers me because it's not portable.  Will Git ever\nconsider using inotify on Linux?  What is the downside?\n\nRam\n"},{"id":"209042","messageId":"7vehgqzc2p.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"CALkWK0=EP0Lv1F_BArub7SpL9rgFhmPtpMOCgwFqfJmVE=oa=A@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-08T22:15:58Z","receivedAt":"2013-02-08T22:15:58Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ramkumar Ramachandra <artagnon@gmail.com> writes:\n\n> ...  Will Git ever\n> consider using inotify on Linux?  What is the downside?\n\nI think this has come up from time to time, but my understanding is\nthat nobody thought things through to find a good layer in the\ncodebase to interface to an external daemon that listens to inotify\nevents yet.  It is not something like \"somebody decreed that we\nwould never consider because of such and such downsides.\"  We are\nnot there yet.\n"},{"id":"209043","messageId":"7va9rezaoy.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"7vehgqzc2p.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-08T22:45:49Z","receivedAt":"2013-02-08T22:45:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>\n>> ...  Will Git ever\n>> consider using inotify on Linux?  What is the downside?\n>\n> I think this has come up from time to time, but my understanding is\n> that nobody thought things through to find a good layer in the\n> codebase to interface to an external daemon that listens to inotify\n> events yet.  It is not something like \"somebody decreed that we\n> would never consider because of such and such downsides.\"  We are\n> not there yet.\n\nI checked read-cache.c and preload-index.c code.  To get the\ndiscussion rolling, I think something like the outline below may be\na good starting point and a feasible weekend hack for somebody\ncompetent:\n\n * At the beginning of preload_index(), instead of spawning the\n   worker thread and doing the lstat() check ourselves, we open a\n   socket to our daemon (see below) that watches this repository and\n   make a request for lstat update.  The request will contain:\n\n    - The SHA1 checksum of the index file we just read (to ensure\n      that we and our daemon share the same baseline to\n      communicate); and\n\n    - the pathspec data.\n\n   Our daemon, if it already has a fresh data available, will give\n   us a list of <path, lstat result>.  Our main process runs a loop\n   that is equivalent to what preload_thread() runs but uses the\n   lstat() data we obtained from the daemon.  If our daemon says it\n   does not have a fresh data (or somehow our daemon is dead), we do\n   the work ourselves.\n\n * Our daemon watches the index file and the working tree, and\n   waits for the above consumer.  First it reads the index (and\n   remembers what it read), and whenever an inotify event comes,\n   does the lstat() and remembers the result.  It never writes\n   to the index, and does not hold the index lock.  Whenever the\n   index file changes, it needs to reload the index, and discard\n   lstat() data it already has for paths that are lost from the\n   updated index.\n"},{"id":"209065","messageId":"CACsJy8DW=tkEy2iOAZxQ+ZyVQ+L11JsPcSxrES5YY7gECmX7UQ@mail.gmail.com","threadId":"32862","inReplyTo":"7va9rezaoy.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-09T02:10:25Z","receivedAt":"2013-02-09T02:10:25Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Feb 9, 2013 at 5:45 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>>\n>>> ...  Will Git ever\n>>> consider using inotify on Linux?  What is the downside?\n>>\n>> I think this has come up from time to time, but my understanding is\n>> that nobody thought things through to find a good layer in the\n>> codebase to interface to an external daemon that listens to inotify\n>> events yet.  It is not something like \"somebody decreed that we\n>> would never consider because of such and such downsides.\"  We are\n>> not there yet.\n>\n> I checked read-cache.c and preload-index.c code.  To get the\n> discussion rolling, I think something like the outline below may be\n> a good starting point and a feasible weekend hack for somebody\n> competent:\n>\n>  * At the beginning of preload_index(), instead of spawning the\n>    worker thread and doing the lstat() check ourselves, we open a\n>    socket to our daemon (see below) that watches this repository and\n\nCan we replace \"open a socket to our daemon\" with \"open a special file\nin .git to get stat data written by our daemon\"? TCP/IP socket means\nsystem-wide daemon, not attractive. UNIX socket is not available on\nWindows (although there may be named pipe, I don't know).\n\n>    make a request for lstat update.  The request will contain:\n>\n>     - The SHA1 checksum of the index file we just read (to ensure\n>       that we and our daemon share the same baseline to\n>       communicate); and\n>\n>     - the pathspec data.\n>\n>    Our daemon, if it already has a fresh data available, will give\n>    us a list of <path, lstat result>.  Our main process runs a loop\n>    that is equivalent to what preload_thread() runs but uses the\n>    lstat() data we obtained from the daemon.  If our daemon says it\n>    does not have a fresh data (or somehow our daemon is dead), we do\n>    the work ourselves.\n>\n>  * Our daemon watches the index file and the working tree, and\n>    waits for the above consumer.  First it reads the index (and\n>    remembers what it read), and whenever an inotify event comes,\n>    does the lstat() and remembers the result.  It never writes\n>    to the index, and does not hold the index lock.  Whenever the\n>    index file changes, it needs to reload the index, and discard\n>    lstat() data it already has for paths that are lost from the\n>    updated index.\n\n\n-- \nDuy\n"},{"id":"209067","messageId":"7vwquiw6u3.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"CACsJy8DW=tkEy2iOAZxQ+ZyVQ+L11JsPcSxrES5YY7gECmX7UQ@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-09T02:37:24Z","receivedAt":"2013-02-09T02:37:24Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> Can we replace \"open a socket to our daemon\" with \"open a special file\n> in .git to get stat data written by our daemon\"? TCP/IP socket means\n> system-wide daemon, not attractive. UNIX socket is not available on\n> Windows (although there may be named pipe, I don't know).\n\nI do not think TCP/IP socket is too bad (you have to be able to read\nthe index file to be able to ask questions to the daemon to begin\nwith, so you must have list of paths already; the answer from the\ndaemon would not leak anything more sensitive than you can already\nknow), and UNIX domain socket is not too bad either.\n\nJust like the implementation detail of the daemon itself may differ\non platforms (does Windows have the identical inotify interface?  I\ndoubt it), I expect the RPC mechanism between the daemon and the\nclient would be platform dependent.  So take that \"open a socket\" as\na generic way to say \"have these two communicate with some magic\",\nnothing more.\n"},{"id":"209068","messageId":"7vsj56w5y9.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"7va9rezaoy.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-09T02:56:30Z","receivedAt":"2013-02-09T02:56:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I checked read-cache.c and preload-index.c code.  To get the\n> discussion rolling, I think something like the outline below may be\n> a good starting point and a feasible weekend hack for somebody\n> competent:\n>\n>  * At the beginning of preload_index(), instead of spawning the\n>    worker thread and doing the lstat() check ourselves, we open a\n>    socket to our daemon (see below) that watches this repository and\n>    make a request for lstat update.  The request will contain:\n>\n>     - The SHA1 checksum of the index file we just read (to ensure\n>       that we and our daemon share the same baseline to\n>       communicate); and\n>\n>     - the pathspec data.\n>\n>    Our daemon, if it already has a fresh data available, will give\n>    us a list of <path, lstat result>.  Our main process runs a loop\n>    that is equivalent to what preload_thread() runs but uses the\n>    lstat() data we obtained from the daemon.  If our daemon says it\n>    does not have a fresh data (or somehow our daemon is dead), we do\n>    the work ourselves.\n>\n>  * Our daemon watches the index file and the working tree, and\n>    waits for the above consumer.  First it reads the index (and\n>    remembers what it read), and whenever an inotify event comes,\n>    does the lstat() and remembers the result.  It never writes\n>    to the index, and does not hold the index lock.  Whenever the\n>    index file changes, it needs to reload the index, and discard\n>    lstat() data it already has for paths that are lost from the\n>    updated index.\n\nI left the details unsaid in thee above because I thought it was\nfairly obvious from the nature of the \"outline\", but let me spend a\nfew more lines to avoid confusion.\n\n - The way the daemon \"watches\" the changes to the working tree and\n   the index may well be very platform dependent.  I said \"inotify\"\n   above, but the mechanism does not have to be inotify.\n\n - The channel the daemon and the client communicates would also be\n   system dependent.  UNIX domain socket in $GIT_DIR/ with a\n   well-known name would be one possibility but it does not have to\n   be the only option.\n\n - The data given from the daemon to the client does not have to\n   include full lstat() information.  They start from the same index\n   info, and the only thing preload_index() wants to know is for\n   which paths it should call ce_mark_uptodate(ce), so the answer\n   given by our daemon can be a list of paths.\n"},{"id":"209069","messageId":"9AF8A28B-71FE-4BBC-AD55-1DD3FDE8FFC3@gmail.com","threadId":"32862","inReplyTo":"7vsj56w5y9.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-02-09T03:36:15Z","receivedAt":"2013-02-09T03:36:15Z","isPatch":false,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"The delay for commands like git status is much worse on Windows than Linux; for my workflow I would be happy with a Windows only implementation. \n\n>From the description so far, I have some question: how does the daemon get started and stopped?  Is there one per repository --- this seems to be implied by putting the unix domain socket in $GIT_DIR. Could we automatically reject connections from anything other than localhost when using TCP?\n\nRobert Zeh\n\nOn Feb 8, 2013, at 8:56 PM, Junio C Hamano <gitster@pobox.com> wrote:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n>> I checked read-cache.c and preload-index.c code.  To get the\n>> discussion rolling, I think something like the outline below may be\n>> a good starting point and a feasible weekend hack for somebody\n>> competent:\n>> \n>> * At the beginning of preload_index(), instead of spawning the\n>>   worker thread and doing the lstat() check ourselves, we open a\n>>   socket to our daemon (see below) that watches this repository and\n>>   make a request for lstat update.  The request will contain:\n>> \n>>    - The SHA1 checksum of the index file we just read (to ensure\n>>      that we and our daemon share the same baseline to\n>>      communicate); and\n>> \n>>    - the pathspec data.\n>> \n>>   Our daemon, if it already has a fresh data available, will give\n>>   us a list of <path, lstat result>.  Our main process runs a loop\n>>   that is equivalent to what preload_thread() runs but uses the\n>>   lstat() data we obtained from the daemon.  If our daemon says it\n>>   does not have a fresh data (or somehow our daemon is dead), we do\n>>   the work ourselves.\n>> \n>> * Our daemon watches the index file and the working tree, and\n>>   waits for the above consumer.  First it reads the index (and\n>>   remembers what it read), and whenever an inotify event comes,\n>>   does the lstat() and remembers the result.  It never writes\n>>   to the index, and does not hold the index lock.  Whenever the\n>>   index file changes, it needs to reload the index, and discard\n>>   lstat() data it already has for paths that are lost from the\n>>   updated index.\n> \n> I left the details unsaid in thee above because I thought it was\n> fairly obvious from the nature of the \"outline\", but let me spend a\n> few more lines to avoid confusion.\n> \n> - The way the daemon \"watches\" the changes to the working tree and\n>   the index may well be very platform dependent.  I said \"inotify\"\n>   above, but the mechanism does not have to be inotify.\n> \n> - The channel the daemon and the client communicates would also be\n>   system dependent.  UNIX domain socket in $GIT_DIR/ with a\n>   well-known name would be one possibility but it does not have to\n>   be the only option.\n> \n> - The data given from the daemon to the client does not have to\n>   include full lstat() information.  They start from the same index\n>   info, and the only thing preload_index() wants to know is for\n>   which paths it should call ce_mark_uptodate(ce), so the answer\n>   given by our daemon can be a list of paths.\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"209077","messageId":"CALkWK0ms76MYa7ddp4F2uyntkFFoBO7gA5JH5Om=OrHxh-encQ@mail.gmail.com","threadId":"32862","inReplyTo":"7vsj56w5y9.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-09T11:32:04Z","receivedAt":"2013-02-09T11:32:04Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> I checked read-cache.c and preload-index.c code.  To get the\n>> discussion rolling, I think something like the outline below may be\n>> a good starting point and a feasible weekend hack for somebody\n>> competent:\n>>\n>>  * At the beginning of preload_index(), instead of spawning the\n>>    worker thread and doing the lstat() check ourselves, we open a\n>>    socket to our daemon (see below) that watches this repository and\n>>    make a request for lstat update.  The request will contain:\n>>\n>>     - The SHA1 checksum of the index file we just read (to ensure\n>>       that we and our daemon share the same baseline to\n>>       communicate); and\n>>\n>>     - the pathspec data.\n>>\n>>    Our daemon, if it already has a fresh data available, will give\n>>    us a list of <path, lstat result>.  Our main process runs a loop\n>>    that is equivalent to what preload_thread() runs but uses the\n>>    lstat() data we obtained from the daemon.  If our daemon says it\n>>    does not have a fresh data (or somehow our daemon is dead), we do\n>>    the work ourselves.\n>>\n>>  * Our daemon watches the index file and the working tree, and\n>>    waits for the above consumer.  First it reads the index (and\n>>    remembers what it read), and whenever an inotify event comes,\n>>    does the lstat() and remembers the result.  It never writes\n>>    to the index, and does not hold the index lock.  Whenever the\n>>    index file changes, it needs to reload the index, and discard\n>>    lstat() data it already has for paths that are lost from the\n>>    updated index.\n>\n> I left the details unsaid in thee above because I thought it was\n> fairly obvious from the nature of the \"outline\", but let me spend a\n> few more lines to avoid confusion.\n>\n>  - The way the daemon \"watches\" the changes to the working tree and\n>    the index may well be very platform dependent.  I said \"inotify\"\n>    above, but the mechanism does not have to be inotify.\n\nIs the BSD kernel's inotify the same as the one on Linux?  Must we\ndesign something that's generic enough from the start?\n\nMore importantly, do you know of a platform-independent inotify\nimplementation in C?  A quick Googling turned up QFileSystemWatcher\n[1], a part of QT.\n\n[1]: http://qt-project.org/doc/qt-4.8/qfilesystemwatcher.html\n\n>  - The channel the daemon and the client communicates would also be\n>    system dependent.  UNIX domain socket in $GIT_DIR/ with a\n>    well-known name would be one possibility but it does not have to\n>    be the only option.\n\nUNIX domain sockets are also preferred because we'd never want to\nconnect to a watch daemon over the network?\n\nThen the communication channel code also has to be generic enough.\n\n>  - The data given from the daemon to the client does not have to\n>    include full lstat() information.  They start from the same index\n>    info, and the only thing preload_index() wants to know is for\n>    which paths it should call ce_mark_uptodate(ce), so the answer\n>    given by our daemon can be a list of paths.\n\nRight.\n"},{"id":"209078","messageId":"CALkWK0mttn6E+D-22UBbvDCuNEy_jNOtBaKPS-a8mTbO2uAF3g@mail.gmail.com","threadId":"32862","inReplyTo":"9AF8A28B-71FE-4BBC-AD55-1DD3FDE8FFC3@gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-09T12:05:55Z","receivedAt":"2013-02-09T12:05:55Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Robert Zeh wrote:\n> From the description so far, I have some question: how does the daemon get started and stopped?  Is there one per repository ...\n\nWhat about getting systemd to watch everything for us?  Then we can\njust have one daemon reporting filesystem changes over one global\nsocket.  It's API should be the inotify subset:\n\n   systemd_add_watch\n   systemd_remove_watch\n\nExcept systemd_add_watch also accepts a UNIX socket to send lstat\nevents to.  Our preload_index() is just reduced to making one\nsystemd_add_watch() call the very first time and updating the index as\nnecessary.  Now, what about desktops with huge uptimes (like mine)?\nWon't they get polluted with too many useless watches over time?\nSimple: timeout.  If nobody reads from the UNIX socket for two hours\nafter a systemd_add_watch, execute systemd_remove_watch automatically.\n\nSomeone must implement a similar daemon on other platforms reporting\ninformation in exactly the same way (although with different\ninternals).  IP sockets are system-wide and all platforms have them,\nso the communication channel is also standardized.\n\nThis is much better than Junio's suggestion to study possible\nimplementations on all platforms and designing a generic daemon/\ncommunication channel.  That's no weekend project.\n"},{"id":"209079","messageId":"CALkWK0k9q62kBjrx5JwDpMOQD7wi4timpT0r-jEN2QEzdx8b_Q@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0mttn6E+D-22UBbvDCuNEy_jNOtBaKPS-a8mTbO2uAF3g@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-09T12:11:54Z","receivedAt":"2013-02-09T12:11:54Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Ramkumar Ramachandra wrote:\n> What about getting systemd to watch everything for us?  Then we can\n> just have one daemon reporting filesystem changes over one global\n> socket.  It's API should be the inotify subset:\n\nEr, not one global socket: many little sockets as described later.\n(The idea was just forming while I was writing this paragraph)\n"},{"id":"209082","messageId":"CALkWK0nQVjKpyef8MDYMs0D9HJGCL8egypT3YWSdU8EYTO7Y+w@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0mttn6E+D-22UBbvDCuNEy_jNOtBaKPS-a8mTbO2uAF3g@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-09T12:53:49Z","receivedAt":"2013-02-09T12:53:49Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Ramkumar Ramachandra wrote:\n> What about getting systemd to watch everything for us?\n\nsystemd is the perfect candidate!  It already has an inotify watcher:\nsee systemd.path(5).  Can't be used as-is because it spawns processes\non events, which is a non-scalable design.  Secondly, it uses static\n.path files to define the rules which is no good for us.  So, we need\nto add an API to it, and ask it to report events over IP sockets.  The\nAPI part is simple too, because it already has a DBUS API for many\nthings [1]; it's just a matter of extending it.\n\nYes, I know.  This introduces dbus as an additional optional\nnon-portable dependency.  Do you have suggestions for alternatives\nthat aren't complicated?\n\n[1]: http://www.freedesktop.org/wiki/Software/systemd/dbus\n"},{"id":"209083","messageId":"CACsJy8CEHzqH1X=v4yau0SyZwrZp1r6hNp=yXD+eZh1q_BS-0g@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0nQVjKpyef8MDYMs0D9HJGCL8egypT3YWSdU8EYTO7Y+w@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-09T12:59:41Z","receivedAt":"2013-02-09T12:59:41Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Feb 9, 2013 at 7:53 PM, Ramkumar Ramachandra <artagnon@gmail.com> wrote:\n> Ramkumar Ramachandra wrote:\n>> What about getting systemd to watch everything for us?\n>\n> systemd is the perfect candidate!\n\nHow about this as a start? I did not really check what it does, but it\ndoes not look complicate enough to pull systemd in.\n\nhttp://article.gmane.org/gmane.comp.version-control.git/151934\n\nYouo may want to search the mail archive. This topic has come up a few\ntimes before, there may be other similar patches.\n-- \nDuy\n"},{"id":"209085","messageId":"CALkWK0=6_n4rf6AWci6J+uhGHpjTUmK7YFdVHuSJedN2zLWtMA@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8CEHzqH1X=v4yau0SyZwrZp1r6hNp=yXD+eZh1q_BS-0g@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-09T17:10:20Z","receivedAt":"2013-02-09T17:10:20Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Duy Nguyen wrote:\n> How about this as a start? I did not really check what it does, but it\n> does not look complicate enough to pull systemd in.\n>\n> http://article.gmane.org/gmane.comp.version-control.git/151934\n\nClever hack.  I didn't know that there was a switch called\ncore.ignoreStat which will disable automatic lstat() calls altogether.\n So, Finn advises that we set this switch and run igit instead of git.\n There's a git-inotify-daemon which runs inotifywait with -m forever,\nupdating a modified_files hash.  When it is sent a TERM from igit\n(which is what happens immediately upon execution), it writes all this\ncollected information about modified files to a named pipe that igit\npasses to it.  igit then does a git update-index --assume-unchained\n--stdin to read the data from the pipe.  Towards the end of its life,\nigit starts up a fresh git-inotify-daemon for future invocations.\n\nFinn notes in the commit message that it offers no speedup, because\n.gitignore files in every directory still have to be read.  I think\nthis is silly: we really should be caching .gitignore, and touching it\nonly when lstat() reports that the file has changed.\n\nAs far as a real implementation that we'd want to merge into git.git\nis concerned, I have a few comments:\nRunning multiple daemons on-the-fly for monitoring filesystem changes\nis not elegant at all.  Keeping track of the state of so many loose\ndaemons is a hard problem: how do we ensure any semblance of\nreliability without that?  Systemd is a very big improvement over the\nlegacy of a hundred loose shell scripts that SysVInit demanded.  It\nmonitors and babysits daemons; it uses cgroups to even kill\nmisbehaving daemons.  I can inspect running daemons at any time, and\nhave a uniform way to start/ stop/ restart them.\n\nOkay, now you're asking me to consider a system-wide daemon\nindependent of systemd.  It has to run with root privileges so it has\naccess to everyone's repositories, which means that people have to\ntrust it beyond doubt.  What does it do?  It has a generic API to\nwatch filesystem paths and report events over an IP socket.  Do you\nthink that this will only be useful to git?  Every other version\ncontrol system (and presumably many other pieces of software) will\nwant to use it.  One huge downside I see of making this part of\nsystemd is Ubuntu.  They've decided not to use systemd for some\nunfathomable reason.\n\nReally, the elephant in the room right now seems to be .gitignore.\nUntil that is fixed, there is really no use of writing this inotify\ndaemon, no?  Can someone enlighten me on how exactly .gitignore files\nare processed?\n\n> Youo may want to search the mail archive. This topic has come up a few\n> times before, there may be other similar patches.\n\nThe thread you linked me to is a 2010 email, and now it's 2013.  We've\nbeen silent about inotify for three years?\n\nThanks for your inputs, Duy.\n"},{"id":"209087","messageId":"CALkWK0=5s0WoA5Y-2wmJRsthtckaMAcRK=JhqhduMty1Pr=Lqw@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0=6_n4rf6AWci6J+uhGHpjTUmK7YFdVHuSJedN2zLWtMA@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-09T18:56:44Z","receivedAt":"2013-02-09T18:56:44Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Ramkumar Ramachandra wrote:\n> Okay, now you're asking me to consider a system-wide daemon\n> independent of systemd.  It has to run with root privileges so it has\n> access to everyone's repositories, which means that people have to\n> trust it beyond doubt.  What does it do?  It has a generic API to\n> watch filesystem paths and report events over an IP socket.  Do you\n> think that this will only be useful to git?  Every other version\n> control system (and presumably many other pieces of software) will\n> want to use it.  One huge downside I see of making this part of\n> systemd is Ubuntu.  They've decided not to use systemd for some\n> unfathomable reason.\n\nAfter some thought, I've decided that extending systemd is not the way\nto go.  And the dbus API is really an overkill.  Writing a simple\nsystem-wide daemon shouldn't be a challenge; the hard part is getting\ngit to use it properly.\n"},{"id":"209088","messageId":"7vliaxwa9p.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"CALkWK0mttn6E+D-22UBbvDCuNEy_jNOtBaKPS-a8mTbO2uAF3g@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-09T19:35:30Z","receivedAt":"2013-02-09T19:35:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ramkumar Ramachandra <artagnon@gmail.com> writes:\n\n> This is much better than Junio's suggestion to study possible\n> implementations on all platforms and designing a generic daemon/\n> communication channel.  That's no weekend project.\n\nIt appears that you misunderstood what I wrote.  That was not \"here\nis a design; I want it in my system.  Go implemment it\".\n\nIt was \"If somebody wants to discuss it but does not know where to\nbegin, doing a small experiment like this and reporting how well it\nworked here may be one way to do so.\", nothing more.\n"},{"id":"209115","messageId":"CACsJy8DeM5--WVXg3b65RxLBS7Jho-7KmcGwWk7B5uAx77yOEw@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0=6_n4rf6AWci6J+uhGHpjTUmK7YFdVHuSJedN2zLWtMA@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-10T05:24:58Z","receivedAt":"2013-02-10T05:24:58Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 10, 2013 at 12:10 AM, Ramkumar Ramachandra\n<artagnon@gmail.com> wrote:\n> Finn notes in the commit message that it offers no speedup, because\n> .gitignore files in every directory still have to be read.  I think\n> this is silly: we really should be caching .gitignore, and touching it\n> only when lstat() reports that the file has changed.\n>\n> ...\n>\n> Really, the elephant in the room right now seems to be .gitignore.\n> Until that is fixed, there is really no use of writing this inotify\n> daemon, no?  Can someone enlighten me on how exactly .gitignore files\n> are processed?\n\n.gitignore is a different issue. I think it's mainly used with\nread_directory/fill_directory to collect ignored files (or not-ignored\nfiles). And it's not always used (well, status and add does, but diff\nshould not). I think wee need to measure how much mass lstat\nelimination gains us (especially on big repos) and how much\n.gitignore/.gitattributes caching does. I don't think .gitignore has\nsuch a big impact though. strace on git.git tells me \"git status\"\nissues about 2500 lstat calls, and just 740 open+getdents calls (on\ntotal 3800 syscalls). I will think if we can do something about\n.gitignore/.gitattributes.\n-- \nDuy\n"},{"id":"209117","messageId":"20130210111732.GA24377@lanh","threadId":"32862","inReplyTo":"CACsJy8DeM5--WVXg3b65RxLBS7Jho-7KmcGwWk7B5uAx77yOEw@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-10T11:17:32Z","receivedAt":"2013-02-10T11:17:32Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 10, 2013 at 12:24:58PM +0700, Duy Nguyen wrote:\n> On Sun, Feb 10, 2013 at 12:10 AM, Ramkumar Ramachandra\n> <artagnon@gmail.com> wrote:\n> > Finn notes in the commit message that it offers no speedup, because\n> > .gitignore files in every directory still have to be read.  I think\n> > this is silly: we really should be caching .gitignore, and touching it\n> > only when lstat() reports that the file has changed.\n> >\n> > ...\n> >\n> > Really, the elephant in the room right now seems to be .gitignore.\n> > Until that is fixed, there is really no use of writing this inotify\n> > daemon, no?  Can someone enlighten me on how exactly .gitignore files\n> > are processed?\n>\n> .gitignore is a different issue. I think it's mainly used with\n> read_directory/fill_directory to collect ignored files (or not-ignored\n> files). And it's not always used (well, status and add does, but diff\n> should not). I think wee need to measure how much mass lstat\n> elimination gains us (especially on big repos) and how much\n> .gitignore/.gitattributes caching does.\n\nOK let's count. I start with a \"standard\" repository, linux-2.6. This\nis the number from strace -T on \"git status\" (*). The first column is\naccumulated time, the second the number of syscalls.\n\ntop syscalls sorted     top syscalls sorted\nby acc. time            by number\n----------------------------------------------\n0.401906 40950 lstat    0.401906 40950 lstat\n0.190484 5343 getdents\t0.150055 5374 open\n0.150055 5374 open\t0.190484 5343 getdents\n0.074843 2806 close\t0.074843 2806 close\n0.003216 157 read\t0.003216 157 read\n\nThe following patch pretends every entry is uptodate without\nlstat. With the patch, we can see refresh code is the cause of mass\nlstat, as lstat disappears:\n\n0.185347 5343 getdents  0.144173 5374 open\n0.144173 5374 open\t0.185347 5343 getdents\n0.071844 2806 close\t0.071844 2806 close\n0.004918 135 brk\t0.003378 157 read\n0.003378 157 read\t0.004918 135 brk\n\n-- 8< --\ndiff --git a/read-cache.c b/read-cache.c\nindex 827ae55..94d8ed8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1018,6 +1018,10 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n \tif (ce_uptodate(ce))\n \t\treturn ce;\n\n+#if 1\n+\tce_mark_uptodate(ce);\n+\treturn ce;\n+#endif\n \t/*\n \t * CE_VALID or CE_SKIP_WORKTREE means the user promised us\n \t * that the change to the work tree does not matter and told\n-- 8< --\n\nThe following patch eliminates untracked search code. As we can see,\nopen+getdents also disappears with this patch:\n\n0.462909 40950 lstat   0.462909 40950 lstat\n0.003417 129 brk       0.003417 129 brk\n0.000762 53 read       0.000762 53 read\n0.000720 36 open       0.000720 36 open\n0.000544 12 munmap     0.000454 33 close\n\nSo from syscalls point of view, we know what code issues most of\nthem. Let's see how much time we gain be these patches, which is an\napproximate of the gain by inotify support. This time I measure on\ngentoo-x86.git [1] because this one has really big worktree (100k\nfiles)\n\n        unmodified  read-cache.c  dir.c     both\nreal    0m0.550s    0m0.479s      0m0.287s  0m0.213s\nuser    0m0.305s    0m0.315s\t  0m0.201s  0m0.182s\nsys     0m0.240s    0m0.157s\t  0m0.084s  0m0.030s\n\nand the syscall picture on gentoo-x86.git:\n\n1.106615 101942 lstat    1.106615 101942 lstat\n0.667235 47083 getdents\t 0.641604 47114 open\n0.641604 47114 open\t 0.667235 47083 getdents\n0.286711 23573 close\t 0.286711 23573 close\n0.005842 350 brk\t 0.005842 350 brk\n\nWe can see that shortcuting untracked code gives bigger gain than\nindex refresh code. So I have to agree that .gitignore may be the big\nelephant in this particular case.\n\nBear in mind though this is Linux, where lstat is fast. On systems\nwith slow lstat, these timings could look very different due to the\nlarge number of lstat calls compared to open+getdents. I really like\nto see similar numbers on Windows.\n\nread_directory/fill_directory code is mostly used by \"git add\" (not\nwith -u) and \"git status\", while refresh code is executed in add,\ncheckout, commit/status, diff, merge. So while smaller gain, reducing\nlstat calls could benefit in more cases.\n\nA relatively slow \"git add\" is acceptable. \"git status\" should be\nfast. Although in my workflow, I do \"git diff [--stat] [--cached]\"\nmuch more often than \"git status\" so relatively slow \"git status\" does\nnot hurt me much. But people may do it differently.\n\nOn speeding up read_directory with inotify support. I haven't thought\nit through, but I think we could save (or get it via socket) a list of\nuntracked files in .git, regardless ignore status, with the help from\ninotify. When this list is verified valid, read_directory could be\nmodified to traverse the tree using this list (plus the index) instead\nof opendir+readdir. Not sure how the change might look though.\n\n\n[1] http://git-exp.overlays.gentoo.org/gitweb/?p=exp/gentoo-x86.git;a=summary\n\n(*) the script to produce those numbers is\n\n-- 8< --\n#!/bin/sh\n\nexport LANG=C\nstrace -T \"$@\" 2>&1 >/dev/null |\n\tsed 's/\\(^[^(]*\\)(.*<\\([0-9.]*\\)>$/\\1 \\2/' |\n\tawk '{\n\t  sec[$1]+=$2;\n\t  count[$1]++;\n\t}\n\tEND {\n\t  for (i in sec)\n\t    printf(\"%f %d %s\\n\", sec[i], count[i], i);\n\t  }' >/tmp/s\n\nsort -nr /tmp/s | head -n5\nsort -nrk2 /tmp/s | head -n5\n-- 8< --\n"},{"id":"209118","messageId":"20130210112205.GA28434@lanh","threadId":"32862","inReplyTo":"20130210111732.GA24377@lanh","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-10T11:22:05Z","receivedAt":"2013-02-10T11:22:05Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 10, 2013 at 06:17:32PM +0700, Duy Nguyen wrote:\n> The following patch eliminates untracked search code. As we can see,\n> open+getdents also disappears with this patch:\n> \n> 0.462909 40950 lstat   0.462909 40950 lstat\n> 0.003417 129 brk       0.003417 129 brk\n> 0.000762 53 read       0.000762 53 read\n> 0.000720 36 open       0.000720 36 open\n> 0.000544 12 munmap     0.000454 33 close\n\n.. and the patch is missing:\n\n-- 8< --\ndiff --git a/dir.c b/dir.c\nindex 57394e4..1963c6f 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -1439,8 +1439,10 @@ int read_directory(struct dir_struct *dir, const char *path, int len, const char\n \t\treturn dir->nr;\n \n \tsimplify = create_simplify(pathspec);\n+#if 0\n \tif (!len || treat_leading_path(dir, path, len, simplify))\n \t\tread_directory_recursive(dir, path, len, 0, simplify);\n+#endif\n \tfree_simplify(simplify);\n \tqsort(dir->entries, dir->nr, sizeof(struct dir_entry *), cmp_name);\n \tqsort(dir->ignored, dir->ignored_nr, sizeof(struct dir_entry *), cmp_name);\n-- 8< --\n"},{"id":"209121","messageId":"CANgJU+WYSD8RHb19EP0M89=Y_XskfjDtFWf51qjg4ur+rDb3ug@mail.gmail.com","threadId":"32862","inReplyTo":"20130210111732.GA24377@lanh","subject":"Re: inotify to minimize stat() calls","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2013-02-10T13:26:55Z","receivedAt":"2013-02-10T13:26:55Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On 10 February 2013 12:17, Duy Nguyen <pclouds@gmail.com> wrote:\n> Bear in mind though this is Linux, where lstat is fast. On systems\n> with slow lstat, these timings could look very different due to the\n> large number of lstat calls compared to open+getdents. I really like\n> to see similar numbers on Windows.\n\nIs windows stat really so slow? I encountered this perception in\nwindows Perl in the past, and I know that on windows Perl stat\n*appears* slow compared to *nix, because in order to satisfy the full\n*nix stat interface, specifically the nlink field, it must open and\nclose the file*. As of 5.10 this can be disabled by setting a magic\nvar ${^WIN32_SLOPPY_STAT} to a true value, which makes a significant\nimprovement to the performance of the Perl level stat implementation.\nI would not be surprised if the cygwin implementation of stat() has\nthe same issue as Perl did, and that stat appears much slower than it\nactually need be if you don't care about the nlink field.\n\nYves\n* http://perl5.git.perl.org/perl.git/blob/HEAD:/win32/win32.c#l1492\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"209139","messageId":"CACsJy8CsV9JtgQ60k4gV_P5qi+Uw9zP=By7cpCFxNSLVqrBcog@mail.gmail.com","threadId":"32862","inReplyTo":"CANgJU+WYSD8RHb19EP0M89=Y_XskfjDtFWf51qjg4ur+rDb3ug@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-10T15:35:46Z","receivedAt":"2013-02-10T15:35:46Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 10, 2013 at 8:26 PM, demerphq <demerphq@gmail.com> wrote:\n> On 10 February 2013 12:17, Duy Nguyen <pclouds@gmail.com> wrote:\n>> Bear in mind though this is Linux, where lstat is fast. On systems\n>> with slow lstat, these timings could look very different due to the\n>> large number of lstat calls compared to open+getdents. I really like\n>> to see similar numbers on Windows.\n>\n> Is windows stat really so slow?\n\nI can't say. I haven't used Windows for months (and git on Windows for years)..\n\n> I encountered this perception in\n> windows Perl in the past, and I know that on windows Perl stat\n> *appears* slow compared to *nix, because in order to satisfy the full\n> *nix stat interface, specifically the nlink field, it must open and\n> close the file*. As of 5.10 this can be disabled by setting a magic\n> var ${^WIN32_SLOPPY_STAT} to a true value, which makes a significant\n> improvement to the performance of the Perl level stat implementation.\n> I would not be surprised if the cygwin implementation of stat() has\n> the same issue as Perl did, and that stat appears much slower than it\n> actually need be if you don't care about the nlink field.\n\nThe native port of git uses get_file_attr (in\ncompat/mingw.c:do_lstat()) to simulate lstat and always sets nlink to\n1. I assume this means git does not care about nlink field. I don't\nknow about cygwin though.\n\n> Yves\n> * http://perl5.git.perl.org/perl.git/blob/HEAD:/win32/win32.c#l1492\n-- \nDuy\n"},{"id":"209141","messageId":"CALkWK0kLieAfPihX3j=BzD+ndo-g-2210Za2xN=HbHcRVwgMtA@mail.gmail.com","threadId":"32862","inReplyTo":"20130210111732.GA24377@lanh","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-10T16:45:41Z","receivedAt":"2013-02-10T16:45:41Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"On Sun, Feb 10, 2013 at 4:47 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Sun, Feb 10, 2013 at 12:24:58PM +0700, Duy Nguyen wrote:\n>> On Sun, Feb 10, 2013 at 12:10 AM, Ramkumar Ramachandra\n>> <artagnon@gmail.com> wrote:\n>> > Finn notes in the commit message that it offers no speedup, because\n>> > .gitignore files in every directory still have to be read.  I think\n>> > this is silly: we really should be caching .gitignore, and touching it\n>> > only when lstat() reports that the file has changed.\n>> >\n>> > ...\n>> >\n>> > Really, the elephant in the room right now seems to be .gitignore.\n>> > Until that is fixed, there is really no use of writing this inotify\n>> > daemon, no?  Can someone enlighten me on how exactly .gitignore files\n>> > are processed?\n>>\n>> .gitignore is a different issue. I think it's mainly used with\n>> read_directory/fill_directory to collect ignored files (or not-ignored\n>> files). And it's not always used (well, status and add does, but diff\n>> should not). I think wee need to measure how much mass lstat\n>> elimination gains us (especially on big repos) and how much\n>> .gitignore/.gitattributes caching does.\n>\n> OK let's count. I start with a \"standard\" repository, linux-2.6. This\n> is the number from strace -T on \"git status\" (*). The first column is\n> accumulated time, the second the number of syscalls.\n>\n> top syscalls sorted     top syscalls sorted\n> by acc. time            by number\n> ----------------------------------------------\n> 0.401906 40950 lstat    0.401906 40950 lstat\n> 0.190484 5343 getdents  0.150055 5374 open\n> 0.150055 5374 open      0.190484 5343 getdents\n> 0.074843 2806 close     0.074843 2806 close\n> 0.003216 157 read       0.003216 157 read\n>\n> The following patch pretends every entry is uptodate without\n> lstat. With the patch, we can see refresh code is the cause of mass\n> lstat, as lstat disappears:\n>\n> 0.185347 5343 getdents  0.144173 5374 open\n> 0.144173 5374 open      0.185347 5343 getdents\n> 0.071844 2806 close     0.071844 2806 close\n> 0.004918 135 brk        0.003378 157 read\n> 0.003378 157 read       0.004918 135 brk\n\nOkay, we're saving 40k lstat() calls.\n\n> -- 8< --\n> diff --git a/read-cache.c b/read-cache.c\n> index 827ae55..94d8ed8 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1018,6 +1018,10 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n>         if (ce_uptodate(ce))\n>                 return ce;\n>\n> +#if 1\n> +       ce_mark_uptodate(ce);\n> +       return ce;\n> +#endif\n>         /*\n>          * CE_VALID or CE_SKIP_WORKTREE means the user promised us\n>          * that the change to the work tree does not matter and told\n> -- 8< --\n\nSo you're skipping the rest of refresh_cache_ent(), which contains our\nlstat() and returning immediately.  Instead of marking paths with the\n\"assume unchanged\" bit, as core.ignoreStat does, you're directly\nattacking the function that refreshes the index and bypassing the\nlstat() call.  How are they different?  read-cache.c:1030 checks\nce->flags & CE_VALID (which is set in read-cache.c:88 if\nassume_unchanged) and bypasses the lstat() call anyway.  So why didn't\nyou just set core.ignoreStat for your test?\n\n> The following patch eliminates untracked search code. As we can see,\n> open+getdents also disappears with this patch:\n>\n> 0.462909 40950 lstat   0.462909 40950 lstat\n> 0.003417 129 brk       0.003417 129 brk\n> 0.000762 53 read       0.000762 53 read\n> 0.000720 36 open       0.000720 36 open\n> 0.000544 12 munmap     0.000454 33 close\n\nOkay, 5k open and 5k getdents calls are gone, but what does this mean?\n\nPulling the patch from your next email to figure this out:\n> -- 8< --\n> diff --git a/dir.c b/dir.c\n> index 57394e4..1963c6f 100644\n> --- a/dir.c\n> +++ b/dir.c\n> @@ -1439,8 +1439,10 @@ int read_directory(struct dir_struct *dir, const char *path, int len, const char\n>                 return dir->nr;\n>\n>         simplify = create_simplify(pathspec);\n> +#if 0\n>         if (!len || treat_leading_path(dir, path, len, simplify))\n>                 read_directory_recursive(dir, path, len, 0, simplify);\n> +#endif\n>         free_simplify(simplify);\n>         qsort(dir->entries, dir->nr, sizeof(struct dir_entry *), cmp_name);\n>         qsort(dir->ignored, dir->ignored_nr, sizeof(struct dir_entry *), cmp_name);\n> -- 8< --\n\nAh, read_directory(), from the .gitignore/ exclude angle.  Yes,\nread_directory() seems to be the main culprit there, from my reading\nof Documentation/technical/api-directory-listing.txt.\n\nSo, what did you do?  You short-circuited the function into never\nexecuting read_directory_recursive(), so the opendir() and readdir()\nare gone.  I'm confused about what this means: will new directories\nfail to appear as \"untracked\" now?  Either way, I understand that\nyou've factored out the .gitignore/ excludes.  Let's look at the\ntimings now.\n\n> So from syscalls point of view, we know what code issues most of\n> them. Let's see how much time we gain be these patches, which is an\n> approximate of the gain by inotify support. This time I measure on\n> gentoo-x86.git [1] because this one has really big worktree (100k\n> files)\n>\n>         unmodified  read-cache.c  dir.c     both\n> real    0m0.550s    0m0.479s      0m0.287s  0m0.213s\n> user    0m0.305s    0m0.315s      0m0.201s  0m0.182s\n> sys     0m0.240s    0m0.157s      0m0.084s  0m0.030s\n\nSo, the .gitignore/ exclude does seem to be the elephant in the room,\nafter all!  There are only minor gains from not updating the index.\n\n> and the syscall picture on gentoo-x86.git:\n>\n> 1.106615 101942 lstat    1.106615 101942 lstat\n> 0.667235 47083 getdents  0.641604 47114 open\n> 0.641604 47114 open      0.667235 47083 getdents\n> 0.286711 23573 close     0.286711 23573 close\n> 0.005842 350 brk         0.005842 350 brk\n\nThe lstat to getdents/ open is higher than in linux-2.6.git here, but\nthe profit in eliminating lstat is still very small.\n\n> We can see that shortcuting untracked code gives bigger gain than\n> index refresh code. So I have to agree that .gitignore may be the big\n> elephant in this particular case.\n\nYes :)\n\n> Bear in mind though this is Linux, where lstat is fast. On systems\n> with slow lstat, these timings could look very different due to the\n> large number of lstat calls compared to open+getdents. I really like\n> to see similar numbers on Windows.\n\nI see.\n\n> read_directory/fill_directory code is mostly used by \"git add\" (not\n> with -u) and \"git status\", while refresh code is executed in add,\n> checkout, commit/status, diff, merge. So while smaller gain, reducing\n> lstat calls could benefit in more cases.\n\nGood point, although my major complaint with big repositories is\nstatus and diff.  Checkout isn't too bad if the branches don't diverge\nmuch, add is fast enough, but diff is a big problem.\n\n> A relatively slow \"git add\" is acceptable. \"git status\" should be\n> fast. Although in my workflow, I do \"git diff [--stat] [--cached]\"\n> much more often than \"git status\" so relatively slow \"git status\" does\n> not hurt me much. But people may do it differently.\n\nHm.\n\n> On speeding up read_directory with inotify support. I haven't thought\n> it through, but I think we could save (or get it via socket) a list of\n> untracked files in .git, regardless ignore status, with the help from\n> inotify. When this list is verified valid, read_directory could be\n> modified to traverse the tree using this list (plus the index) instead\n> of opendir+readdir. Not sure how the change might look though.\n\nWe could even tell if .gitignore has changed with inotify support, and\ntell exactly when we need to update our path treatment.  And yes, we\ncan have inotify to tell us about new files and directories directly,\nso we can traverse them.  I think we should go ahead with the\nsystem-wide inotify daemon for now: its design needs to be discussed,\nso that it's generic enough for all our usecases.  I'm thinking there\nshould be atleast two distinct calls to:\n1. Report all changed paths, for use with read-cache.c.\n2. Report only new paths, for use with dir.c.\nOr can we figure out which of the changed paths are new ourselves?\n\nKudos to your great work on getting getting these numbers, Duy!\n"},{"id":"209142","messageId":"CABPQNSZ282Lre=sy-+ZQdJA9JnGqQguq2bQDOwvjb0fP+1-w8Q@mail.gmail.com","threadId":"32862","inReplyTo":"20130210111732.GA24377@lanh","subject":"Re: inotify to minimize stat() calls","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2013-02-10T16:58:11Z","receivedAt":"2013-02-10T16:58:11Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Sun, Feb 10, 2013 at 12:17 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Sun, Feb 10, 2013 at 12:24:58PM +0700, Duy Nguyen wrote:\n>> On Sun, Feb 10, 2013 at 12:10 AM, Ramkumar Ramachandra\n>> <artagnon@gmail.com> wrote:\n>> > Finn notes in the commit message that it offers no speedup, because\n>> > .gitignore files in every directory still have to be read.  I think\n>> > this is silly: we really should be caching .gitignore, and touching it\n>> > only when lstat() reports that the file has changed.\n>> >\n>> > ...\n>> >\n>> > Really, the elephant in the room right now seems to be .gitignore.\n>> > Until that is fixed, there is really no use of writing this inotify\n>> > daemon, no?  Can someone enlighten me on how exactly .gitignore files\n>> > are processed?\n>>\n>> .gitignore is a different issue. I think it's mainly used with\n>> read_directory/fill_directory to collect ignored files (or not-ignored\n>> files). And it's not always used (well, status and add does, but diff\n>> should not). I think wee need to measure how much mass lstat\n>> elimination gains us (especially on big repos) and how much\n>> .gitignore/.gitattributes caching does.\n>\n> OK let's count. I start with a \"standard\" repository, linux-2.6. This\n> is the number from strace -T on \"git status\" (*). The first column is\n> accumulated time, the second the number of syscalls.\n>\n> top syscalls sorted     top syscalls sorted\n> by acc. time            by number\n> ----------------------------------------------\n> 0.401906 40950 lstat    0.401906 40950 lstat\n> 0.190484 5343 getdents  0.150055 5374 open\n> 0.150055 5374 open      0.190484 5343 getdents\n> 0.074843 2806 close     0.074843 2806 close\n> 0.003216 157 read       0.003216 157 read\n>\n> The following patch pretends every entry is uptodate without\n> lstat. With the patch, we can see refresh code is the cause of mass\n> lstat, as lstat disappears:\n>\n> 0.185347 5343 getdents  0.144173 5374 open\n> 0.144173 5374 open      0.185347 5343 getdents\n> 0.071844 2806 close     0.071844 2806 close\n> 0.004918 135 brk        0.003378 157 read\n> 0.003378 157 read       0.004918 135 brk\n>\n> -- 8< --\n> diff --git a/read-cache.c b/read-cache.c\n> index 827ae55..94d8ed8 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1018,6 +1018,10 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n>         if (ce_uptodate(ce))\n>                 return ce;\n>\n> +#if 1\n> +       ce_mark_uptodate(ce);\n> +       return ce;\n> +#endif\n>         /*\n>          * CE_VALID or CE_SKIP_WORKTREE means the user promised us\n>          * that the change to the work tree does not matter and told\n> -- 8< --\n>\n> The following patch eliminates untracked search code. As we can see,\n> open+getdents also disappears with this patch:\n>\n> 0.462909 40950 lstat   0.462909 40950 lstat\n> 0.003417 129 brk       0.003417 129 brk\n> 0.000762 53 read       0.000762 53 read\n> 0.000720 36 open       0.000720 36 open\n> 0.000544 12 munmap     0.000454 33 close\n>\n> So from syscalls point of view, we know what code issues most of\n> them. Let's see how much time we gain be these patches, which is an\n> approximate of the gain by inotify support. This time I measure on\n> gentoo-x86.git [1] because this one has really big worktree (100k\n> files)\n>\n>         unmodified  read-cache.c  dir.c     both\n> real    0m0.550s    0m0.479s      0m0.287s  0m0.213s\n> user    0m0.305s    0m0.315s      0m0.201s  0m0.182s\n> sys     0m0.240s    0m0.157s      0m0.084s  0m0.030s\n>\n> and the syscall picture on gentoo-x86.git:\n>\n> 1.106615 101942 lstat    1.106615 101942 lstat\n> 0.667235 47083 getdents  0.641604 47114 open\n> 0.641604 47114 open      0.667235 47083 getdents\n> 0.286711 23573 close     0.286711 23573 close\n> 0.005842 350 brk         0.005842 350 brk\n>\n> We can see that shortcuting untracked code gives bigger gain than\n> index refresh code. So I have to agree that .gitignore may be the big\n> elephant in this particular case.\n>\n> Bear in mind though this is Linux, where lstat is fast. On systems\n> with slow lstat, these timings could look very different due to the\n> large number of lstat calls compared to open+getdents. I really like\n> to see similar numbers on Windows.\n\nKarsten Blees has done something similar-ish on Windows, and he posted\nthe results here:\n\nhttps://groups.google.com/forum/#!topic/msysgit/fL_jykUmUNE/discussion\n\nI also seem to remember he doing a ReadDirectoryChangesW version, but\nI don't remember what happened with that.\n"},{"id":"209143","messageId":"CAKXa9=qQwJqxZLxhAS35QeF1+dwH+ukod0NfFggVCuUZHz-USg@mail.gmail.com","threadId":"32862","inReplyTo":"7vliaxwa9p.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-02-10T19:03:00Z","receivedAt":"2013-02-10T19:03:00Z","isPatch":false,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"On Sat, Feb 9, 2013 at 1:35 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>\n>> This is much better than Junio's suggestion to study possible\n>> implementations on all platforms and designing a generic daemon/\n>> communication channel.  That's no weekend project.\n>\n> It appears that you misunderstood what I wrote.  That was not \"here\n> is a design; I want it in my system.  Go implemment it\".\n>\n> It was \"If somebody wants to discuss it but does not know where to\n> begin, doing a small experiment like this and reporting how well it\n> worked here may be one way to do so.\", nothing more.\n\nWhat if instead of communicating over a socket, the daemon\ndumped a file containing all of the lstat information after git\nwrote a file? By definition the daemon should know about file writes.\n\nThere would be no network communication, which I think would make\nthings more secure. It would simplify the rendezvous by insisting on\nwell known locations in $GIT_DIR.\n\nRobert Zeh\n"},{"id":"209144","messageId":"201302101226.12646.mfick@codeaurora.org","threadId":"32862","inReplyTo":"CAKXa9=qQwJqxZLxhAS35QeF1+dwH+ukod0NfFggVCuUZHz-USg@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2013-02-10T19:26:12Z","receivedAt":"2013-02-10T19:26:12Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Sunday, February 10, 2013 12:03:00 pm Robert Zeh wrote:\n> On Sat, Feb 9, 2013 at 1:35 PM, Junio C Hamano \n<gitster@pobox.com> wrote:\n> > Ramkumar Ramachandra <artagnon@gmail.com> writes:\n> >> This is much better than Junio's suggestion to study\n> >> possible implementations on all platforms and\n> >> designing a generic daemon/ communication channel. \n> >> That's no weekend project.\n> > \n> > It appears that you misunderstood what I wrote.  That\n> > was not \"here is a design; I want it in my system.  Go\n> > implemment it\".\n> > \n> > It was \"If somebody wants to discuss it but does not\n> > know where to begin, doing a small experiment like\n> > this and reporting how well it worked here may be one\n> > way to do so.\", nothing more.\n> \n> What if instead of communicating over a socket, the\n> daemon dumped a file containing all of the lstat\n> information after git wrote a file? By definition the\n> daemon should know about file writes.\n\nBut git doesn't, how will it know when the file is written?\nWill it use inotify, or poll (kind of defeats the point)?\n\n-Martin\n"},{"id":"209147","messageId":"7vhaljudos.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"20130210112205.GA28434@lanh","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-10T20:16:51Z","receivedAt":"2013-02-10T20:16:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Sun, Feb 10, 2013 at 06:17:32PM +0700, Duy Nguyen wrote:\n>> The following patch eliminates untracked search code. As we can see,\n>> open+getdents also disappears with this patch:\n>> \n>> 0.462909 40950 lstat   0.462909 40950 lstat\n>> 0.003417 129 brk       0.003417 129 brk\n>> 0.000762 53 read       0.000762 53 read\n>> 0.000720 36 open       0.000720 36 open\n>> 0.000544 12 munmap     0.000454 33 close\n>\n> .. and the patch is missing:\n>\n> -- 8< --\n> diff --git a/dir.c b/dir.c\n> index 57394e4..1963c6f 100644\n> --- a/dir.c\n> +++ b/dir.c\n> @@ -1439,8 +1439,10 @@ int read_directory(struct dir_struct *dir, const char *path, int len, const char\n>  \t\treturn dir->nr;\n>  \n>  \tsimplify = create_simplify(pathspec);\n> +#if 0\n>  \tif (!len || treat_leading_path(dir, path, len, simplify))\n>  \t\tread_directory_recursive(dir, path, len, 0, simplify);\n> +#endif\n\nThe other \"lstat()\" experiment was a very interesting one, but this\nis not yet an interesting experiment to see where in the \"ignore\"\ncodepath we are spending times.\n\nWe know that we can tell wt_status_collect_untracked() not to bother\nwith the untracked or ignored files with !s->show_untracked_files\nalready, but I think the more interesting question is if we can show\nthe untracked files with less overhead.\n\nIf we want to show untrackedd files, it is a given that we need to\nread directories to see what paths there are on the filesystem. Is\nthe opendir/readdir cost dominating in the process? Are we spending\na lot of time sifting the result of opendir/readdir via the ignore\nmechanism? Is reading the \"ignore\" files costing us much to prime\nthe ignore mechanism?\n\nIf readdir cost is dominant, then that makes \"cache gitignore\" a\nnonsense proposition, I think.  If you really want to \"cache\"\nsomething, you need to have somebody (i.e. a daemon) who constantly\nkeeps an eye on the filesystem changes and can respond with the up\nto date result directly to fill_directory().  I somehow doubt that\nit is a direction we would want to go in, though.\n"},{"id":"209148","messageId":"53A99181-32FC-402F-8ADD-49076131B891@gmail.com","threadId":"32862","inReplyTo":"201302101226.12646.mfick@codeaurora.org","subject":"Re: inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-02-10T20:18:36Z","receivedAt":"2013-02-10T20:18:36Z","isPatch":false,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"\n\nOn Feb 10, 2013, at 1:26 PM, Martin Fick <mfick@codeaurora.org> wrote:\n\n> On Sunday, February 10, 2013 12:03:00 pm Robert Zeh wrote:\n>> On Sat, Feb 9, 2013 at 1:35 PM, Junio C Hamano\n> <gitster@pobox.com> wrote:\n>>> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>>>> This is much better than Junio's suggestion to study\n>>>> possible implementations on all platforms and\n>>>> designing a generic daemon/ communication channel. \n>>>> That's no weekend project.\n>>> \n>>> It appears that you misunderstood what I wrote.  That\n>>> was not \"here is a design; I want it in my system.  Go\n>>> implemment it\".\n>>> \n>>> It was \"If somebody wants to discuss it but does not\n>>> know where to begin, doing a small experiment like\n>>> this and reporting how well it worked here may be one\n>>> way to do so.\", nothing more.\n>> \n>> What if instead of communicating over a socket, the\n>> daemon dumped a file containing all of the lstat\n>> information after git wrote a file? By definition the\n>> daemon should know about file writes.\n> \n> But git doesn't, how will it know when the file is written?\n> Will it use inotify, or poll (kind of defeats the point)?\n> \n> -Martin\n\nI was thinking it would loop on calls to stat for the file with a timeout; this is no different than what we would want to do over a socket in that we would need timeouts for network reads.  But we would only be calling stat on one file, instead of the entire repo. \n\nI think we can set things up so the file read is atomic, which means we can ignore the case of a daemon crashing midway through a conversation. \n\nRobert\n"},{"id":"209193","messageId":"CACsJy8DnvAjQPL4aP_LRC7aqx6OC4M5dMtj-OUot76qET2z08Q@mail.gmail.com","threadId":"32862","inReplyTo":"7vhaljudos.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-11T02:56:27Z","receivedAt":"2013-02-11T02:56:27Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Feb 11, 2013 at 3:16 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> The other \"lstat()\" experiment was a very interesting one, but this\n> is not yet an interesting experiment to see where in the \"ignore\"\n> codepath we are spending times.\n>\n> We know that we can tell wt_status_collect_untracked() not to bother\n> with the untracked or ignored files with !s->show_untracked_files\n> already, but I think the more interesting question is if we can show\n> the untracked files with less overhead.\n>\n> If we want to show untrackedd files, it is a given that we need to\n> read directories to see what paths there are on the filesystem. Is\n> the opendir/readdir cost dominating in the process? Are we spending\n> a lot of time sifting the result of opendir/readdir via the ignore\n> mechanism? Is reading the \"ignore\" files costing us much to prime\n> the ignore mechanism?\n>\n> If readdir cost is dominant, then that makes \"cache gitignore\" a\n> nonsense proposition, I think.  If you really want to \"cache\"\n> something, you need to have somebody (i.e. a daemon) who constantly\n> keeps an eye on the filesystem changes and can respond with the up\n> to date result directly to fill_directory().  I somehow doubt that\n> it is a direction we would want to go in, though.\n\nYeah, it did not cut out syscall cost, I also cut a lot of user-space\nprocessing (plus .gitignore content access). From the timings I posted\nearlier,\n\n>         unmodified  dir.c\n> real    0m0.550s    0m0.287s\n> user    0m0.305s    0m0.201s\n> sys     0m0.240s    0m0.084s\n\nsys time is reduced from 0.24s to 0.08s, so readdir+opendir definitely\nhas something to do with it (and perhaps reading .gitignore). But it\nalso reduces user time from 0.305 to 0.201s. I don't think avoiding\nreaddir+openddir will bring us this gain. It's probably the cost of\nmatching .gitignore. I'll try to replace opendir+readdir with a\nno-syscall version. At this point \"untracked caching\" sounds more\nfeasible (and less complex) than \".gitignore cachine\".\n-- \nDuy\n"},{"id":"209194","messageId":"CACsJy8Aw1GpKXGwjMdzjXBxAMrC-q6HDSyi2u6EoXCYDV8fJ4Q@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0kLieAfPihX3j=BzD+ndo-g-2210Za2xN=HbHcRVwgMtA@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-11T03:03:52Z","receivedAt":"2013-02-11T03:03:52Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 10, 2013 at 11:45 PM, Ramkumar Ramachandra\n<artagnon@gmail.com> wrote:\n> So you're skipping the rest of refresh_cache_ent(), which contains our\n> lstat() and returning immediately.  Instead of marking paths with the\n> \"assume unchanged\" bit, as core.ignoreStat does, you're directly\n> attacking the function that refreshes the index and bypassing the\n> lstat() call.  How are they different?  read-cache.c:1030 checks\n> ce->flags & CE_VALID (which is set in read-cache.c:88 if\n> assume_unchanged) and bypasses the lstat() call anyway.  So why didn't\n> you just set core.ignoreStat for your test?\n\nIt just did not occur to me that core.ignoreStat does the same.\n\n> Ah, read_directory(), from the .gitignore/ exclude angle.  Yes,\n> read_directory() seems to be the main culprit there, from my reading\n> of Documentation/technical/api-directory-listing.txt.\n>\n> So, what did you do?  You short-circuited the function into never\n> executing read_directory_recursive(), so the opendir() and readdir()\n> are gone.  I'm confused about what this means: will new directories\n> fail to appear as \"untracked\" now?\n\nNo, read_directory returns the list of untracked/ignored files.\nReturning empty lists means no untracked nor ignored files.\n-- \nDuy\n"},{"id":"209195","messageId":"CACsJy8BvN0xX_=fx78hVLw=2Wyk=RUHYs_x9r5RJ0TvBAoA83g@mail.gmail.com","threadId":"32862","inReplyTo":"CAKXa9=qQwJqxZLxhAS35QeF1+dwH+ukod0NfFggVCuUZHz-USg@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-11T03:21:33Z","receivedAt":"2013-02-11T03:21:33Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Feb 11, 2013 at 2:03 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n> On Sat, Feb 9, 2013 at 1:35 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>>\n>>> This is much better than Junio's suggestion to study possible\n>>> implementations on all platforms and designing a generic daemon/\n>>> communication channel.  That's no weekend project.\n>>\n>> It appears that you misunderstood what I wrote.  That was not \"here\n>> is a design; I want it in my system.  Go implemment it\".\n>>\n>> It was \"If somebody wants to discuss it but does not know where to\n>> begin, doing a small experiment like this and reporting how well it\n>> worked here may be one way to do so.\", nothing more.\n>\n> What if instead of communicating over a socket, the daemon\n> dumped a file containing all of the lstat information after git\n> wrote a file? By definition the daemon should know about file writes.\n>\n> There would be no network communication, which I think would make\n> things more secure. It would simplify the rendezvous by insisting on\n> well known locations in $GIT_DIR.\n\nWe need some sort of interactive communication to the daemon anyway,\nto validate that the information is uptodate. Assume that a user makes\nsome changes to his worktree before starting the daemon, git needs to\nknow that what the daemon provides does not represent a complete\nfile-change picture and it better refreshes the index the old way\nonce, then trust the daemon.\n\nI think we could solve that by storing a \"session id\", provided by the\ndaemon, in .git/index. If the session id is not present (or does not\nmatch what the current daemon gives), refresh the old way. After\nrefreshing, it may ask the daemon for new session id and store it.\nNext time if the session id is still valid, trust the daemon's data.\nThis session id should be different every time the daemon restarts for\nthis to work.\n-- \nDuy\n"},{"id":"209197","messageId":"CACsJy8AWyJ=dW5f44huWyPPe4X62xyi+R9CNM5Tg6u6TYf+thQ@mail.gmail.com","threadId":"32862","inReplyTo":"CABPQNSZ282Lre=sy-+ZQdJA9JnGqQguq2bQDOwvjb0fP+1-w8Q@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-11T03:53:12Z","receivedAt":"2013-02-11T03:53:12Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 10, 2013 at 11:58 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n> Karsten Blees has done something similar-ish on Windows, and he posted\n> the results here:\n>\n> https://groups.google.com/forum/#!topic/msysgit/fL_jykUmUNE/discussion\n>\n> I also seem to remember he doing a ReadDirectoryChangesW version, but\n> I don't remember what happened with that.\n\nThanks. I came across that but did not remember. For one thing, we\nknow the inotify alternative for Windows: ReadDirectoryChangesW.\n\nBut the meat of the patch is not about that function. In fact it's\ndropped in fscache-v3 [1]. It seems that doing\nFindFirstFile/FindNextFile for an entire directory, cache the results\nand use it to simulate lstat() is faster on Windows. Sounds similar to\npreload-index. And because directory listing is cached anyway,\nopendir/readdir is replaced to read from cache instead of opening the\ndirectory again.\n\nSo it is orthogonal with using ReadDirectoryChangesW/inotify to\nfurther reduce the system calls.\n\nI copy \"git status\"'s (impressive) numbers from fscache-v0 for those\nwho are interested in:\n\npreload | -u  | normal | cached | gain\n--------+-----+--------+--------+------\nfalse   | all | 25.144 | 3.055  |  8.2\nfalse   | no  | 22.822 | 1.748  | 12.8\ntrue    | all |  9.234 | 2.179  |  4.2\ntrue    | no  |  6.833 | 0.955  |  7.2\n\n[1] https://github.com/kblees/git/commit/35f319609aa046d2350db32d3afa1fa44920e880\n-- \nDuy\n"},{"id":"209244","messageId":"CACsJy8B3rOROJ2HoQ0iLjfvxfcEGjJgP7MH0yhq+8Gjnm-EGAg@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8DnvAjQPL4aP_LRC7aqx6OC4M5dMtj-OUot76qET2z08Q@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-11T11:12:38Z","receivedAt":"2013-02-11T11:12:38Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Feb 11, 2013 at 9:56 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n> Yeah, it did not cut out syscall cost, I also cut a lot of user-space\n> processing (plus .gitignore content access). From the timings I posted\n> earlier,\n>\n>>         unmodified  dir.c\n>> real    0m0.550s    0m0.287s\n>> user    0m0.305s    0m0.201s\n>> sys     0m0.240s    0m0.084s\n>\n> sys time is reduced from 0.24s to 0.08s, so readdir+opendir definitely\n> has something to do with it (and perhaps reading .gitignore). But it\n> also reduces user time from 0.305 to 0.201s. I don't think avoiding\n> readdir+openddir will bring us this gain. It's probably the cost of\n> matching .gitignore. I'll try to replace opendir+readdir with a\n> no-syscall version. At this point \"untracked caching\" sounds more\n> feasible (and less complex) than \".gitignore cachine\".\n\nAnd this is read_directory's timing breakdown (again, \"git status\" on\ngentoo-x86,git, built with -O2 on x86-64 if I did not mention before)\n\nopendir   = 0.030s\nreaddir   = 0.083s\nclosedir  = 0.020s\n{open,read,close}dir = 0.132s\ntreat_path           = 0.094s (172534 times)\ndir_add_name         = 0.050s (101917 times)\nread_directory       = 0.292s\n# On branch master\nnothing to commit, working directory clean\n\nreal    0m0.511s\nuser    0m0.347s\nsys     0m0.157s\n\nInstrumentation is done with gettimeofday. Without gettimeofday calls\ninside read_directory_recursive, read_directory takes 0.267s (iow,\ngettimeofday cost is about 0.30s). {open,read,close}dir + treat_path +\ndir_add_name + gettimeofday add up quite close to 0.292s (strbuf_*\ntakes just about 0.005s)\n\nEliminating xxxdir syscalls may save us 0.132s (or less, we need to\npay to get the information elsewhere).\n\nBecause my worktree is clean, dir_add_name spends all 0.05s in\ncache_name_exists(). If we somehow know the input path is not a\ntracked entry, we could avoid cache_name_exists() and save 0.05s.\n\nIf we do the \"untracked cache\", the number of treat_path calls should\nbe much lower. In this particular case of gentoo-x86, I'd expect no\nmore than a dozen of untracked files, which cuts down treat_path and\ndir_add_name's time to near zero. On a normal repository like git.git,\nuntracked files are about 1075 files with 2552 tracked files, we\nshould be able to save 2/3 to 1/2 of treat_path calls.\n-- \nDuy\n"},{"id":"209250","messageId":"CAKXa9=pCSWtXq+5x_LcZ9gsSpa1yT0QD5DsBguTqosoH0cj-nw@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8BvN0xX_=fx78hVLw=2Wyk=RUHYs_x9r5RJ0TvBAoA83g@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-02-11T14:13:53Z","receivedAt":"2013-02-11T14:13:53Z","isPatch":false,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"On Sun, Feb 10, 2013 at 9:21 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Mon, Feb 11, 2013 at 2:03 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n>> On Sat, Feb 9, 2013 at 1:35 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>>> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>>>\n>>>> This is much better than Junio's suggestion to study possible\n>>>> implementations on all platforms and designing a generic daemon/\n>>>> communication channel.  That's no weekend project.\n>>>\n>>> It appears that you misunderstood what I wrote.  That was not \"here\n>>> is a design; I want it in my system.  Go implemment it\".\n>>>\n>>> It was \"If somebody wants to discuss it but does not know where to\n>>> begin, doing a small experiment like this and reporting how well it\n>>> worked here may be one way to do so.\", nothing more.\n>>\n>> What if instead of communicating over a socket, the daemon\n>> dumped a file containing all of the lstat information after git\n>> wrote a file? By definition the daemon should know about file writes.\n>>\n>> There would be no network communication, which I think would make\n>> things more secure. It would simplify the rendezvous by insisting on\n>> well known locations in $GIT_DIR.\n>\n> We need some sort of interactive communication to the daemon anyway,\n> to validate that the information is uptodate. Assume that a user makes\n> some changes to his worktree before starting the daemon, git needs to\n> know that what the daemon provides does not represent a complete\n> file-change picture and it better refreshes the index the old way\n> once, then trust the daemon.\n>\n> I think we could solve that by storing a \"session id\", provided by the\n> daemon, in .git/index. If the session id is not present (or does not\n> match what the current daemon gives), refresh the old way. After\n> refreshing, it may ask the daemon for new session id and store it.\n> Next time if the session id is still valid, trust the daemon's data.\n> This session id should be different every time the daemon restarts for\n> this to work.\n\nI think we could do this without interactive communication,\nif we did the following:\n   1) The Daemon waits to see $GIT_DIR/lstat_request, and atomically\n       writes out $GIT_DIR/lstat_cache.  By atomically I mean that it writes\n       things out to a temporary file, and then does a rename.\n\n   2) The client erases $GIT_DIR/lstat_cache, and writes\n      $GIT_DIR/lstat_request\n\nI think this is better than socket based communication because there\nare fewer places to check\nfor failures.\n\nRobert\n"},{"id":"209409","messageId":"511AAA92.4030508@gmail.com","threadId":"32862","inReplyTo":"CACsJy8AWyJ=dW5f44huWyPPe4X62xyi+R9CNM5Tg6u6TYf+thQ@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-12T20:48:18Z","receivedAt":"2013-02-12T20:48:18Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 11.02.2013 04:53, schrieb Duy Nguyen:\n> On Sun, Feb 10, 2013 at 11:58 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n>> Karsten Blees has done something similar-ish on Windows, and he posted\n>> the results here:\n>>\n>> https://groups.google.com/forum/#!topic/msysgit/fL_jykUmUNE/discussion\n>>\n\nThe new hashtable implementation in fscache [1] supports O(1) removal and has no mingw dependencies - might come in handy for anyone trying to implement an inotify daemon.\n\n[1] https://github.com/kblees/git/commit/f7eb85c2\n\n>> I also seem to remember he doing a ReadDirectoryChangesW version, but\n>> I don't remember what happened with that.\n> \n> Thanks. I came across that but did not remember. For one thing, we\n> know the inotify alternative for Windows: ReadDirectoryChangesW.\n> \n\nI dropped ReadDirectoryChangesW because maintaining a 'live' file system cache became more and more complicated. For example, according to MSDN docs, ReadDirectoryChangesW *may* report short DOS 8.3 names (i.e. \"PROGRA~1\" instead of \"Program Files\"), so a correct and fast cache implementation would have to be indexed by long *and* short names...\n\nAnother problem was that the 'live' cache had quite negative performance impact on mutating git commands (checkout, reset...). An inotify daemon running as a background process (not in-process as fscache) will probably affect everyone that modifies the working copy, e.g. running 'make' or the test-suite. This should be considered in the design.\n\n> I copy \"git status\"'s (impressive) numbers from fscache-v0 for those\n> who are interested in:\n> \n> preload | -u  | normal | cached | gain\n> --------+-----+--------+--------+------\n> false   | all | 25.144 | 3.055  |  8.2\n> false   | no  | 22.822 | 1.748  | 12.8\n> true    | all |  9.234 | 2.179  |  4.2\n> true    | no  |  6.833 | 0.955  |  7.2\n> \n\nNote that I wasn't able to reproduce such bad 'normal' values in later tests, I guess disk fragmentation and/or virus scanner must have tricked me on that day...gain factors of 2.5 - 5 are more appropriate.\n\n\nHowever, the difference between git status -uall and -uno was always about 1.3 s in all fscache versions, even though opendir/readdir/closedir was served entirely from the cache. I added a bit of performance tracing to find the cause, and I think most of the time spent in wt_status_collect_untracked can be eliminated:\n\n1.) 0.939 s is spent in dir.c/excluded (i.e. checking .gitignore). This check is done for *every* file in the working copy, including files in the index. Checking the index first could eliminate most of that, i.e.:\n\n(Note: patches are for discussion only, I'm aware that they may have unintended side effects...)\n\n@@ -1097,6 +1097,8 @@ static enum path_treatment treat_path(struct dir_struct *dir,\n                return path_ignored;\n        strbuf_setlen(path, baselen);\n        strbuf_addstr(path, de->d_name);\n+       if (cache_name_exists(path->buf, path->len, ignore_case))\n+               return path_ignored;\n        if (simplify_away(path->buf, path->len, simplify))\n                return path_ignored;\n---\n\n\n2.) 0.135 s is spent in name-hash.c/hash_index_entry_directories, reindexing the same directories over and over again. In the end, the hashtable contains 939k directory entries, even though the WebKit test repo only has 7k directories. Checking if a directory entry already exists could reduce that, i.e.:\n\n@@ -53,14 +55,23 @@ static void hash_index_entry_directories(struct index_state *istate, struct cach\n \tunsigned int hash;\n \tvoid **pos;\n \tdouble t = ticks();\n+\tstruct cache_entry *ce2;\n+\tint len = ce_namelen(ce);\n \n-\tconst char *ptr = ce->name;\n-\twhile (*ptr) {\n-\t\twhile (*ptr && *ptr != '/')\n-\t\t\t++ptr;\n-\t\tif (*ptr == '/') {\n-\t\t\t++ptr;\n-\t\t\thash = hash_name(ce->name, ptr - ce->name);\n+\twhile (len > 0) {\n+\t\twhile (len > 0 && ce->name[len - 1] != '/')\n+\t\t\tlen--;\n+\t\tif (len > 0) {\n+\t\t\thash = hash_name(ce->name, len);\n+\t\t\tce2 = lookup_hash(hash, &istate->name_hash);\n+\t\t\twhile (ce2) {\n+\t\t\t\tif (same_name(ce2, ce->name, len, ignore_case)) {\n+\t\t\t\t\tadd_since(t, &hash_dirs);\n+\t\t\t\t\treturn;\n+\t\t\t\t}\n+\t\t\t\tce2 = ce2->dir_next;\n+\t\t\t}\n+\t\t\tlen--;\n \t\t\tpos = insert_hash(hash, ce, &istate->name_hash);\n \t\t\tif (pos) {\n \t\t\t\tce->dir_next = *pos;\n---\n\n\nTests were done with the WebKit repo (~200k files, ~7k directories, 15 .gitignore files, ~100 entries in root .gitignore). Instrumented code can be found here: https://github.com/kblees/git/tree/kb/git-status-performance-tracing\n\nHere's the performance traces of 'git status -s -uall'\n\nBefore patches:\n\ntrace: at builtin/commit.c:1221, time: 0.523429 s: cmd_status/read_cache_preload\ntrace: at builtin/commit.c:1223, time: 0.00403477 s: cmd_status/refresh_index\ntrace: at builtin/commit.c:1231, time: 0.00318494 s: cmd_status/hold_locked_index\ntrace: at wt-status.c:539, time: 0.00527396 s: wt_status_collect_changes_worktree\ntrace: at wt-status.c:544, time: 0.00545771 s: wt_status_collect_changes\ntrace: at wt-status.c:546, time: 1.286 s: wt_status_collect_untracked\ntrace: at builtin/commit.c:1233, time: 1.29852 s: cmd_status/wt_status_collect\ntrace: at dir.c:1540, time: 0.00170986 s: read_directory_recursive/strbuf_add\ntrace: at dir.c:1541, time: 0.00623972 s: read_directory_recursive/opendir\ntrace: at dir.c:1542, time: 0.00517881 s: read_directory_recursive/readdir\ntrace: at dir.c:1543, time: 0.992936 s: read_directory_recursive/treat_path\ntrace: at dir.c:1544, time: 0.277942 s: read_directory_recursive/dir_add_name\ntrace: at dir.c:1545, time: 0.0014594 s: read_directory_recursive/close\ntrace: at dir.c:1546, time: 0.939349 s: treat_one_path/excluded\ntrace: at dir.c:1547, time: 0.0050811 s: treat_one_path/dir_add_ignored\ntrace: at dir.c:1548, time: 0.00515875 s: treat_one_path/get_dtype\ntrace: at dir.c:1549, time: 0.00329322 s: treat_one_path/treat_directory\ntrace: at dir.c:1550, time: 0.222969 s: excluded/prep_exclude\ntrace: at dir.c:1551, time: 0.00443398 s: excluded/excluded_from_list[EXC_CMDL]\ntrace: at dir.c:1552, time: 0.699602 s: excluded/excluded_from_list[EXC_DIRS]\ntrace: at dir.c:1553, time: 0.00475736 s: excluded/excluded_from_list[EXC_FILE]\ntrace: at read-cache.c:460, time: 0.00967987 s: index_name_pos\ntrace: at name-hash.c:213, time: 0.190481 s: lazy_init_name_hash\ntrace: at name-hash.c:216, time: 0.135248 s: hash_index_entry_directories (938865 entries)\ntrace: at name-hash.c:217, time: 0.0806647 s: index_name_exists\ntrace: at compat/mingw.c:2137, time: 1.97424 s: command: c:\\git\\msysgit\\git\\git-status.exe -s -uall\n\n\nAfter patches:\n\ntrace: at builtin/commit.c:1221, time: 0.517511 s: cmd_status/read_cache_preload\ntrace: at builtin/commit.c:1223, time: 0.00405227 s: cmd_status/refresh_index\ntrace: at builtin/commit.c:1231, time: 0.00322796 s: cmd_status/hold_locked_index\ntrace: at wt-status.c:539, time: 0.00530057 s: wt_status_collect_changes_worktree\ntrace: at wt-status.c:544, time: 0.00546062 s: wt_status_collect_changes\ntrace: at wt-status.c:546, time: 0.322799 s: wt_status_collect_untracked\ntrace: at builtin/commit.c:1233, time: 0.33536 s: cmd_status/wt_status_collect\ntrace: at dir.c:1542, time: 0.00120529 s: read_directory_recursive/strbuf_add\ntrace: at dir.c:1543, time: 0.00476647 s: read_directory_recursive/opendir\ntrace: at dir.c:1544, time: 0.00502022 s: read_directory_recursive/readdir\ntrace: at dir.c:1545, time: 0.310515 s: read_directory_recursive/treat_path\ntrace: at dir.c:1546, time: 0 s: read_directory_recursive/dir_add_name\ntrace: at dir.c:1547, time: 0.000831234 s: read_directory_recursive/close\ntrace: at dir.c:1548, time: 0.0668582 s: treat_one_path/excluded\ntrace: at dir.c:1549, time: 0.000173174 s: treat_one_path/dir_add_ignored\ntrace: at dir.c:1550, time: 0.000174267 s: treat_one_path/get_dtype\ntrace: at dir.c:1551, time: 0.00315468 s: treat_one_path/treat_directory\ntrace: at dir.c:1552, time: 0.039733 s: excluded/prep_exclude\ntrace: at dir.c:1553, time: 0.000185205 s: excluded/excluded_from_list[EXC_CMDL]\ntrace: at dir.c:1554, time: 0.0264496 s: excluded/excluded_from_list[EXC_DIRS]\ntrace: at dir.c:1555, time: 0.000170622 s: excluded/excluded_from_list[EXC_FILE]\ntrace: at read-cache.c:460, time: 0.00260636 s: index_name_pos\ntrace: at name-hash.c:224, time: 0.126637 s: lazy_init_name_hash\ntrace: at name-hash.c:227, time: 0.0500866 s: hash_index_entry_directories (7152 entries)\ntrace: at name-hash.c:228, time: 0.0790143 s: index_name_exists\ntrace: at compat/mingw.c:2137, time: 1.00595 s: command: c:\\git\\msysgit\\git\\git-status.exe -s -uall\n"},{"id":"209457","messageId":"20130213100646.GA24993@lanh","threadId":"32862","inReplyTo":"511AAA92.4030508@gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-13T10:06:46Z","receivedAt":"2013-02-13T10:06:46Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Feb 12, 2013 at 09:48:18PM +0100, Karsten Blees wrote:\n\n> However, the difference between git status -uall and -uno was always\n> about 1.3 s in all fscache versions, even though\n> opendir/readdir/closedir was served entirely from the cache. I added\n> a bit of performance tracing to find the cause, and I think most of\n> the time spent in wt_status_collect_untracked can be eliminated:\n> \n> 1.) 0.939 s is spent in dir.c/excluded (i.e. checking\n> .gitignore). This check is done for *every* file in the working\n> copy, including files in the index. Checking the index first could\n> eliminate most of that, i.e.:\n> \n> (Note: patches are for discussion only, I'm aware that they may have\n> unintended side effects...)\n>\n> @@ -1097,6 +1097,8 @@ static enum path_treatment treat_path(struct dir_struct *dir,\n>                 return path_ignored;\n>         strbuf_setlen(path, baselen);\n>         strbuf_addstr(path, de->d_name);\n> +       if (cache_name_exists(path->buf, path->len, ignore_case))\n> +               return path_ignored;\n>         if (simplify_away(path->buf, path->len, simplify))\n>                 return path_ignored;\n\nThe below patch passes the test suite for me and still does the same\nthing. On my Linux box, running \"git status\" on gentoo-x86.git with\nthis patch saves 0.05s (0.548s without the patch, 0.505s with the\npatch, best number of 20 runs).\n\nAnd I just realized gentoo-x86.git does not have any .gitignore. On\nwebkit.git, it cuts \"git status\" time from 1.121s down to\n0.762s. Unless I'm mistaken, \"git add\" should have the same benefit on\nnormal case too. Good finding!\n\n-- 8< --\ndiff --git a/dir.c b/dir.c\nindex 57394e4..4b4cf60 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -1244,7 +1244,19 @@ static enum path_treatment treat_one_path(struct dir_struct *dir,\n \t\t\t\t\t  const struct path_simplify *simplify,\n \t\t\t\t\t  int dtype, struct dirent *de)\n {\n-\tint exclude = is_excluded(dir, path->buf, &dtype);\n+\tint exclude;\n+\n+\tif (dtype == DT_UNKNOWN)\n+\t\tdtype = get_dtype(de, path->buf, path->len);\n+\n+\tif (!(dir->flags & DIR_SHOW_IGNORED) &&\n+\t    !(dir->flags & DIR_COLLECT_IGNORED) &&\n+\t    dtype == DT_REG &&\n+\t    cache_name_exists(path->buf, path->len, ignore_case))\n+\t\treturn path_ignored;\n+\n+\texclude = is_excluded(dir, path->buf, &dtype);\n+\n \tif (exclude && (dir->flags & DIR_COLLECT_IGNORED)\n \t    && exclude_matches_pathspec(path->buf, path->len, simplify))\n \t\tdir_add_ignored(dir, path->buf, path->len);\n@@ -1256,9 +1268,6 @@ static enum path_treatment treat_one_path(struct dir_struct *dir,\n \tif (exclude && !(dir->flags & DIR_SHOW_IGNORED))\n \t\treturn path_ignored;\n \n-\tif (dtype == DT_UNKNOWN)\n-\t\tdtype = get_dtype(de, path->buf, path->len);\n-\n \tswitch (dtype) {\n \tdefault:\n \t\treturn path_ignored;\n-- 8< --\n"},{"id":"209464","messageId":"CACsJy8C=2xKcsby048WWCFNhgKObGwrzeCOJPVVqgj88AfSHQw@mail.gmail.com","threadId":"32862","inReplyTo":"511AAA92.4030508@gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-13T12:15:47Z","receivedAt":"2013-02-13T12:15:47Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Feb 13, 2013 at 3:48 AM, Karsten Blees <karsten.blees@gmail.com> wrote:\n> 2.) 0.135 s is spent in name-hash.c/hash_index_entry_directories, reindexing the same directories over and over again. In the end, the hashtable contains 939k directory entries, even though the WebKit test repo only has 7k directories. Checking if a directory entry already exists could reduce that, i.e.:\n\nThis function is only used when core.ignorecase = true. I probably\nwon't be able to test this, so I'll leave this to other people who\ncare about ignorecase.\n\nThis function used to have lookup_hash, but it was removed by Jeff in\n2548183 (fix phantom untracked files when core.ignorecase is set -\n2011-10-06). There's a looong commit message which I'm too lazy to\nread. Anybody who works on this should though.\n\n\n> @@ -53,14 +55,23 @@ static void hash_index_entry_directories(struct index_state *istate, struct cach\n>         unsigned int hash;\n>         void **pos;\n>         double t = ticks();\n> +       struct cache_entry *ce2;\n> +       int len = ce_namelen(ce);\n>\n> -       const char *ptr = ce->name;\n> -       while (*ptr) {\n> -               while (*ptr && *ptr != '/')\n> -                       ++ptr;\n> -               if (*ptr == '/') {\n> -                       ++ptr;\n> -                       hash = hash_name(ce->name, ptr - ce->name);\n> +       while (len > 0) {\n> +               while (len > 0 && ce->name[len - 1] != '/')\n> +                       len--;\n> +               if (len > 0) {\n> +                       hash = hash_name(ce->name, len);\n> +                       ce2 = lookup_hash(hash, &istate->name_hash);\n> +                       while (ce2) {\n> +                               if (same_name(ce2, ce->name, len, ignore_case)) {\n> +                                       add_since(t, &hash_dirs);\n> +                                       return;\n> +                               }\n> +                               ce2 = ce2->dir_next;\n> +                       }\n> +                       len--;\n>                         pos = insert_hash(hash, ce, &istate->name_hash);\n>                         if (pos) {\n>                                 ce->dir_next = *pos;\n-- \nDuy\n"},{"id":"209485","messageId":"20130213181851.GA5603@sigill.intra.peff.net","threadId":"32862","inReplyTo":"CACsJy8C=2xKcsby048WWCFNhgKObGwrzeCOJPVVqgj88AfSHQw@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-02-13T18:18:51Z","receivedAt":"2013-02-13T18:18:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 13, 2013 at 07:15:47PM +0700, Nguyen Thai Ngoc Duy wrote:\n\n> On Wed, Feb 13, 2013 at 3:48 AM, Karsten Blees <karsten.blees@gmail.com> wrote:\n> > 2.) 0.135 s is spent in name-hash.c/hash_index_entry_directories, reindexing the same directories over and over again. In the end, the hashtable contains 939k directory entries, even though the WebKit test repo only has 7k directories. Checking if a directory entry already exists could reduce that, i.e.:\n> \n> This function is only used when core.ignorecase = true. I probably\n> won't be able to test this, so I'll leave this to other people who\n> care about ignorecase.\n> \n> This function used to have lookup_hash, but it was removed by Jeff in\n> 2548183 (fix phantom untracked files when core.ignorecase is set -\n> 2011-10-06). There's a looong commit message which I'm too lazy to\n> read. Anybody who works on this should though.\n\nYeah, the problem that commit tried to solve is that linking to a single\ncache entry through the hash is not enough, because we may remove cache\nitems. Imagine you have \"dir/one\" and \"dir/two\", and you add them to the\nin-memory index in that order. The original code hashed \"dir/\" and\ninserted a link to the \"dir/one\" cache entry. When it came time to put\nin the \"dir/two\" entry, we noticed that there was already a \"dir/\" entry\nand did nothing. Then later, if we remove \"dir/one\", we do so by marking\nit with CE_UNHASHED. So a later query for \"dir/\" will see \"nope, nothing\nhere that wasn't CE_UNHASHED\", which is wrong. We never recorded that\n\"dir/two\" existed under the hash for \"dir/\", so we can't know about it.\n\nMy patch just stores the cache_entry for both under the \"dir/\" hash.\nAs Karsten noticed, that can lead to a large number of hash entries,\nbecause adding \"some/deep/hierarchy/with/files\" will add 4 directory\nentries for just that single file. Moreover, looking at it again, I\ndon't think my patch produces the right behavior: we have a single\ndir_next pointer, even though the same ce_entry may appear under many\ndirectory hashes. So the cache_entries that has to \"dir/foo/\" and those\nthat hash to \"dir/bar/\" may get confused, because they will also both be\nfound under \"dir/\", and both try to create a linked list from the\ndir_next pointer.\n\nLooking at Karsten's patch, it seems like it will not add a cache entry\nif there is one of the same name. But I'm not sure if that is right, as\nthe old one might be CE_UNHASHED (or it might get removed later). You\nactually want to be able to find each cache_entry that has a file under\nthe directory at the hash of that directory, so you can make sure it is\nstill valid.\n\nAnd of course that still leaves the existing correctness problem I\nmentioned above.\n\nI think the best way forward is to actually create a separate hash table\nfor the directory lookups. I note that we only care about these entries\nin directory_exists_in_index_icase, which is really about whether\nsomething is there, versus what exactly is there. So could we maybe get\nby with a separate hash table that stores a count of entries at each\ndirectory, and increment/decrement the count when we add/remove entries?\n\nThe biggest problem I see with that is that we do indeed care a little\nbit what is at the directory: we check the mode to see if it is a gitdir\nor not. But I think we can maybe sneak around that: gitdirs have actual\nentries in the index, whereas the directories do not. So we would find\nthem via index_name_exists; anything that is not there, but _is_ in the\nspecial directory hash would therefore be a directory.\n\nI realize it got pretty esoteric there in the middle. I'll see if I can\nwork up a patch that expresses what I'm thinking.\n\n-Peff\n"},{"id":"209487","messageId":"20130213194741.GA20712@sigill.intra.peff.net","threadId":"32862","inReplyTo":"20130213181851.GA5603@sigill.intra.peff.net","subject":"Re: inotify to minimize stat() calls","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-02-13T19:47:41Z","receivedAt":"2013-02-13T19:47:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 13, 2013 at 01:18:51PM -0500, Jeff King wrote:\n\n> I think the best way forward is to actually create a separate hash table\n> for the directory lookups. I note that we only care about these entries\n> in directory_exists_in_index_icase, which is really about whether\n> something is there, versus what exactly is there. So could we maybe get\n> by with a separate hash table that stores a count of entries at each\n> directory, and increment/decrement the count when we add/remove entries?\n> \n> The biggest problem I see with that is that we do indeed care a little\n> bit what is at the directory: we check the mode to see if it is a gitdir\n> or not. But I think we can maybe sneak around that: gitdirs have actual\n> entries in the index, whereas the directories do not. So we would find\n> them via index_name_exists; anything that is not there, but _is_ in the\n> special directory hash would therefore be a directory.\n> \n> I realize it got pretty esoteric there in the middle. I'll see if I can\n> work up a patch that expresses what I'm thinking.\n\nSo here's a patch. It's mostly meant to illustrate what I'm thinking,\nand I have no clue if it introduces regressions. It does pass the test\nsuite, but we have virtually no ignorecase tests.  It seems to behave\nsanely when I set core.ignorecase on my Linux box, but I have no idea\nwhat it will do on a real case-insensitive system (nor even, to be\nhonest, what kinds of scenarios should be tested for the dir-hashing\nstuff).\n\n---\ndiff --git a/cache.h b/cache.h\nindex e493563..6630a35 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -131,7 +131,6 @@ struct cache_entry {\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n \tstruct cache_entry *next;\n-\tstruct cache_entry *dir_next;\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n@@ -267,26 +266,14 @@ extern void add_name_hash(struct index_state *istate, struct cache_entry *ce);\n \tunsigned name_hash_initialized : 1,\n \t\t initialized : 1;\n \tstruct hash_table name_hash;\n+\tstruct hash_table dir_hash;\n };\n \n extern struct index_state the_index;\n \n /* Name hashing */\n extern void add_name_hash(struct index_state *istate, struct cache_entry *ce);\n-/*\n- * We don't actually *remove* it, we can just mark it invalid so that\n- * we won't find it in lookups.\n- *\n- * Not only would we have to search the lists (simple enough), but\n- * we'd also have to rehash other hash buckets in case this makes the\n- * hash bucket empty (common). So it's much better to just mark\n- * it.\n- */\n-static inline void remove_name_hash(struct cache_entry *ce)\n-{\n-\tce->ce_flags |= CE_UNHASHED;\n-}\n-\n+extern void remove_name_hash(struct index_state *istate, struct cache_entry *ce);\n \n #ifndef NO_THE_INDEX_COMPATIBILITY_MACROS\n #define active_cache (the_index.cache)\n@@ -443,6 +430,7 @@ extern struct cache_entry *index_name_exists(struct index_state *istate, const c\n extern int unmerged_index(const struct index_state *);\n extern int verify_path(const char *path);\n extern struct cache_entry *index_name_exists(struct index_state *istate, const char *name, int namelen, int igncase);\n+extern int index_icase_dir_exists(struct index_state *istate, const char *name, int namelen);\n extern int index_name_pos(const struct index_state *, const char *name, int namelen);\n #define ADD_CACHE_OK_TO_ADD 1\t\t/* Ok to add */\n #define ADD_CACHE_OK_TO_REPLACE 2\t/* Ok to replace file/directory */\ndiff --git a/dir.c b/dir.c\nindex 57394e4..f73ac34 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -927,29 +927,27 @@ static enum exist_status directory_exists_in_index_icase(const char *dirname, in\n  */\n static enum exist_status directory_exists_in_index_icase(const char *dirname, int len)\n {\n-\tstruct cache_entry *ce = index_name_exists(&the_index, dirname, len + 1, ignore_case);\n-\tunsigned char endchar;\n-\n-\tif (!ce)\n-\t\treturn index_nonexistent;\n-\tendchar = ce->name[len];\n+\tstruct cache_entry *ce = index_name_exists(&the_index, dirname, len, ignore_case);\n \n \t/*\n-\t * The cache_entry structure returned will contain this dirname\n-\t * and possibly additional path components.\n+\t * We found something in the index, which means it is either an actual\n+\t * file, or a gitdir.\n \t */\n-\tif (endchar == '/')\n-\t\treturn index_directory;\n+\tif (ce) {\n+\t    if (S_ISGITLINK(ce->ce_mode))\n+\t\t    return index_gitdir;\n+\t    /* We call a file \"index_nonexistent\" here, because the caller is\n+\t     * asking about a directory.  */\n+\t    return index_nonexistent;\n+\t}\n \n \t/*\n-\t * If there are no additional path components, then this cache_entry\n-\t * represents a submodule.  Submodules, despite being directories,\n-\t * are stored in the cache without a closing slash.\n+\t * Otherwise, it might be a leading path of something that is in the\n+\t * index. We can look it up in the special dir hash.\n \t */\n-\tif (!endchar && S_ISGITLINK(ce->ce_mode))\n-\t\treturn index_gitdir;\n+\tif (index_icase_dir_exists(&the_index, dirname, len))\n+\t\treturn index_directory;\n \n-\t/* This should never be hit, but it exists just in case. */\n \treturn index_nonexistent;\n }\n \ndiff --git a/name-hash.c b/name-hash.c\nindex d8d25c2..de8239f 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -32,37 +32,88 @@ static void hash_index_entry_directories(struct index_state *istate, struct cach\n \treturn hash;\n }\n \n-static void hash_index_entry_directories(struct index_state *istate, struct cache_entry *ce)\n+struct dir_hash_entry {\n+\tstruct dir_hash_entry *next;\n+\tint nr;\n+\tunsigned int namelen;\n+\tchar name[FLEX_ARRAY];\n+};\n+\n+static struct dir_hash_entry *find_dir_hash(struct hash_table *t,\n+\t\t\t\t\t    const char *name,\n+\t\t\t\t\t    unsigned int namelen)\n+{\n+\tunsigned int hash = hash_name(name, namelen);\n+\tstruct dir_hash_entry *ent;\n+\n+\tfor (ent = lookup_hash(hash, t); ent; ent = ent->next) {\n+\t\tif (ent->namelen == namelen &&\n+\t\t    !strncasecmp(ent->name, name, namelen))\n+\t\t\treturn ent;\n+\t}\n+\treturn NULL;\n+}\n+\n+static struct dir_hash_entry *find_or_create_dir_hash(struct hash_table *t,\n+\t\t\t\t\t\t      const char *name,\n+\t\t\t\t\t\t      unsigned int namelen)\n+{\n+\tstruct dir_hash_entry *ent;\n+\n+\tent = find_dir_hash(t, name, namelen);\n+\tif (!ent) {\n+\t\tvoid **pos;\n+\n+\t\tent = xcalloc(sizeof(*ent) + namelen + 1, 1);\n+\t\tmemcpy(ent->name, name, namelen);\n+\t\tent->namelen = namelen;\n+\n+\t\tpos = insert_hash(hash_name(name, namelen), ent, t);\n+\t\tif (pos) {\n+\t\t\tent->next = *pos;\n+\t\t\t*pos = ent;\n+\t\t}\n+\t}\n+\n+\treturn ent;\n+}\n+\n+static void hash_index_entry_directories(struct index_state *istate,\n+\t\t\t\t\t struct cache_entry *ce,\n+\t\t\t\t\t int add)\n {\n \t/*\n-\t * Throw each directory component in the hash for quick lookup\n+\t * Throw each directory component into a hash for quick lookup\n \t * during a git status. Directory components are stored with their\n \t * closing slash.  Despite submodules being a directory, they never\n \t * reach this point, because they are stored without a closing slash\n-\t * in the cache.\n-\t *\n-\t * Note that the cache_entry stored with the directory does not\n-\t * represent the directory itself.  It is a pointer to an existing\n-\t * filename, and its only purpose is to represent existence of the\n-\t * directory in the cache.  It is very possible multiple directory\n-\t * hash entries may point to the same cache_entry.\n+\t * in the cache. This means we don't need to know anything about\n+\t * what is stored at a particular directory, just that it is a leading\n+\t * directory component of something else. Which means we can get away\n+\t * with storing a count instead of a complete\n \t */\n-\tunsigned int hash;\n-\tvoid **pos;\n-\n \tconst char *ptr = ce->name;\n \twhile (*ptr) {\n \t\twhile (*ptr && *ptr != '/')\n \t\t\t++ptr;\n \t\tif (*ptr == '/') {\n-\t\t\t++ptr;\n-\t\t\thash = hash_name(ce->name, ptr - ce->name);\n-\t\t\tpos = insert_hash(hash, ce, &istate->name_hash);\n-\t\t\tif (pos) {\n-\t\t\t\tce->dir_next = *pos;\n-\t\t\t\t*pos = ce;\n+\t\t\tstruct dir_hash_entry *ent;\n+\n+\t\t\tif (add) {\n+\t\t\t\tent = find_or_create_dir_hash(&istate->dir_hash,\n+\t\t\t\t\t\t\t      ce->name,\n+\t\t\t\t\t\t\t      ptr - ce->name);\n+\t\t\t\tent->nr++;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\tent = find_dir_hash(&istate->dir_hash,\n+\t\t\t\t\t\t    ce->name,\n+\t\t\t\t\t\t    ptr - ce->name);\n+\t\t\t\tif (ent)\n+\t\t\t\t\tent->nr--;\n \t\t\t}\n \t\t}\n+\t\tptr++;\n \t}\n }\n \n@@ -74,7 +125,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \tif (ce->ce_flags & CE_HASHED)\n \t\treturn;\n \tce->ce_flags |= CE_HASHED;\n-\tce->next = ce->dir_next = NULL;\n+\tce->next = NULL;\n \thash = hash_name(ce->name, ce_namelen(ce));\n \tpos = insert_hash(hash, ce, &istate->name_hash);\n \tif (pos) {\n@@ -83,7 +134,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \t}\n \n \tif (ignore_case)\n-\t\thash_index_entry_directories(istate, ce);\n+\t\thash_index_entry_directories(istate, ce, 1);\n }\n \n static void lazy_init_name_hash(struct index_state *istate)\n@@ -104,6 +155,22 @@ void add_name_hash(struct index_state *istate, struct cache_entry *ce)\n \t\thash_index_entry(istate, ce);\n }\n \n+/*\n+ * We don't actually *remove* it, we can just mark it invalid so that\n+ * we won't find it in lookups.\n+ *\n+ * Not only would we have to search the lists (simple enough), but\n+ * we'd also have to rehash other hash buckets in case this makes the\n+ * hash bucket empty (common). So it's much better to just mark\n+ * it.\n+ */\n+void remove_name_hash(struct index_state *istate, struct cache_entry *ce)\n+{\n+\tce->ce_flags |= CE_UNHASHED;\n+\tif (istate->dir_hash.nr)\n+\t\thash_index_entry_directories(istate, ce, 0);\n+}\n+\n static int slow_same_name(const char *name1, int len1, const char *name2, int len2)\n {\n \tif (len1 != len2)\n@@ -137,18 +204,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n \tif (!icase)\n \t\treturn 0;\n \n-\t/*\n-\t * If the entry we're comparing is a filename (no trailing slash), then compare\n-\t * the lengths exactly.\n-\t */\n-\tif (name[namelen - 1] != '/')\n-\t\treturn slow_same_name(name, namelen, ce->name, len);\n-\n-\t/*\n-\t * For a directory, we point to an arbitrary cache_entry filename.  Just\n-\t * make sure the directory portion matches.\n-\t */\n-\treturn slow_same_name(name, namelen, ce->name, namelen < len ? namelen : len);\n+\treturn slow_same_name(name, namelen, ce->name, len);\n }\n \n struct cache_entry *index_name_exists(struct index_state *istate, const char *name, int namelen, int icase)\n@@ -164,10 +220,7 @@ struct cache_entry *index_name_exists(struct index_state *istate, const char *na\n \t\t\tif (same_name(ce, name, namelen, icase))\n \t\t\t\treturn ce;\n \t\t}\n-\t\tif (icase && name[namelen - 1] == '/')\n-\t\t\tce = ce->dir_next;\n-\t\telse\n-\t\t\tce = ce->next;\n+\t\tce = ce->next;\n \t}\n \n \t/*\n@@ -188,3 +241,11 @@ struct cache_entry *index_name_exists(struct index_state *istate, const char *na\n \t}\n \treturn NULL;\n }\n+\n+int index_icase_dir_exists(struct index_state *istate, const char *name, int namelen)\n+{\n+\tstruct dir_hash_entry *ent;\n+\n+\tent = find_dir_hash(&istate->dir_hash, name, namelen);\n+\treturn ent && ent->nr;\n+}\ndiff --git a/read-cache.c b/read-cache.c\nindex 827ae55..116c25c 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -46,7 +46,7 @@ static void replace_index_entry(struct index_state *istate, int nr, struct cache\n {\n \tstruct cache_entry *old = istate->cache[nr];\n \n-\tremove_name_hash(old);\n+\tremove_name_hash(istate, old);\n \tset_index_entry(istate, nr, ce);\n \tistate->cache_changed = 1;\n }\n@@ -460,7 +460,7 @@ int remove_index_entry_at(struct index_state *istate, int pos)\n \tstruct cache_entry *ce = istate->cache[pos];\n \n \trecord_resolve_undo(istate, ce);\n-\tremove_name_hash(ce);\n+\tremove_name_hash(istate, ce);\n \tistate->cache_changed = 1;\n \tistate->cache_nr--;\n \tif (pos >= istate->cache_nr)\n@@ -483,7 +483,7 @@ void remove_marked_cache_entries(struct index_state *istate)\n \n \tfor (i = j = 0; i < istate->cache_nr; i++) {\n \t\tif (ce_array[i]->ce_flags & CE_REMOVE)\n-\t\t\tremove_name_hash(ce_array[i]);\n+\t\t\tremove_name_hash(istate, ce_array[i]);\n \t\telse\n \t\t\tce_array[j++] = ce_array[i];\n \t}\n"},{"id":"209494","messageId":"511BF6D7.3030404@gmail.com","threadId":"32862","inReplyTo":"20130213181851.GA5603@sigill.intra.peff.net","subject":"Re: inotify to minimize stat() calls","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-13T20:25:59Z","receivedAt":"2013-02-13T20:25:59Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 13.02.2013 19:18, schrieb Jeff King:\n> Moreover, looking at it again, I\n> don't think my patch produces the right behavior: we have a single\n> dir_next pointer, even though the same ce_entry may appear under many\n> directory hashes. So the cache_entries that has to \"dir/foo/\" and those\n> that hash to \"dir/bar/\" may get confused, because they will also both be\n> found under \"dir/\", and both try to create a linked list from the\n> dir_next pointer.\n> \n\nIndeed. In the worst case, this causes an endless loop if ce->dir_next == ce\n---8<---\nmkdir V\nmkdir V/XQANY\nmkdir WURZAUP\ntouch V/XQANY/test\ngit init\ngit config core.ignorecase true\ngit add .\ngit status\n---8<---\nNote: \"V/\", \"V/XQANY/\" and \"WURZAUP/\" all have the same hash_name. Although I found those strange values by brute force, hash collisions in 32 bit values are not that uncommon in real life :-)\n\n> Looking at Karsten's patch, it seems like it will not add a cache entry\n> if there is one of the same name. But I'm not sure if that is right, as\n> the old one might be CE_UNHASHED (or it might get removed later). You\n> actually want to be able to find each cache_entry that has a file under\n> the directory at the hash of that directory, so you can make sure it is\n> still valid.\n> \n\nYes, the patch was just to show potential performance savings, I didn't consider CE_UNHASHED at all.\n\n> I think the best way forward is to actually create a separate hash table\n> for the directory lookups. I note that we only care about these entries\n> in directory_exists_in_index_icase, which is really about whether\n> something is there, versus what exactly is there. So could we maybe get\n> by with a separate hash table that stores a count of entries at each\n> directory, and increment/decrement the count when we add/remove entries?\n> \n\nAlternatively, we could simply create normal cache_entries for the directories that are linked via ce->next, but have a trailing '/' in their name?\n\nReference counting sounds good to me, at least better than allocating directory entries per cache entry * parent dirs.\n"},{"id":"209500","messageId":"20130213225529.GA25353@sigill.intra.peff.net","threadId":"32862","inReplyTo":"511BF6D7.3030404@gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-02-13T22:55:29Z","receivedAt":"2013-02-13T22:55:29Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 13, 2013 at 09:25:59PM +0100, Karsten Blees wrote:\n\n> Am 13.02.2013 19:18, schrieb Jeff King:\n> > Moreover, looking at it again, I\n> > don't think my patch produces the right behavior: we have a single\n> > dir_next pointer, even though the same ce_entry may appear under many\n> > directory hashes. So the cache_entries that has to \"dir/foo/\" and those\n> > that hash to \"dir/bar/\" may get confused, because they will also both be\n> > found under \"dir/\", and both try to create a linked list from the\n> > dir_next pointer.\n> > \n> \n> Indeed. In the worst case, this causes an endless loop if ce->dir_next == ce\n> ---8<---\n> mkdir V\n> mkdir V/XQANY\n> mkdir WURZAUP\n> touch V/XQANY/test\n> git init\n> git config core.ignorecase true\n> git add .\n> git status\n> ---8<---\n\nGreat, thanks for the test case. I can trivially replicate the endless\nloop. The patch I sent earlier fixes that. So it's at least a step in\nthe (possible) right direction. I'm slightly concerned that there is\nsome other case that is expecting the directories in the main hash, but\nI think I got them all.\n\n> Note: \"V/\", \"V/XQANY/\" and \"WURZAUP/\" all have the same hash_name.\n> Although I found those strange values by brute force, hash collisions\n> in 32 bit values are not that uncommon in real life :-)\n\nCute. :)\n\n> Alternatively, we could simply create normal cache_entries for the\n> directories that are linked via ce->next, but have a trailing '/' in\n> their name?\n>\n> Reference counting sounds good to me, at least better than allocating\n> directory entries per cache entry * parent dirs.\n\nI think that is more or less what my patch does, but it splits the\nref-counted fake cache_entries out into a separate hash of \"struct\ndir_hash_entry\" rather than storing it in the regular hash. Which IMHO\nis a bit cleaner for two reasons:\n\n  1. You do not have to pay the memory price of storing fake\n     cache_entries (the name+refcount struct for each directory is much\n     smaller than a real cache_entry).\n\n  2. It makes the code a bit simpler, as you do not have to do any\n     \"check for trailing /\" magic on the result of index_name_exists to\n     determine if it is a \"real\" name or just a fake dir entry.\n\n-Peff\n"},{"id":"209504","messageId":"511C3454.6080204@gmail.com","threadId":"32862","inReplyTo":"20130213225529.GA25353@sigill.intra.peff.net","subject":"Re: inotify to minimize stat() calls","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-14T00:48:20Z","receivedAt":"2013-02-14T00:48:20Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 13.02.2013 23:55, schrieb Jeff King:\n> On Wed, Feb 13, 2013 at 09:25:59PM +0100, Karsten Blees wrote:\n> \n>> Alternatively, we could simply create normal cache_entries for the\n>> directories that are linked via ce->next, but have a trailing '/' in\n>> their name?\n>>\n>> Reference counting sounds good to me, at least better than allocating\n>> directory entries per cache entry * parent dirs.\n> \n> I think that is more or less what my patch does, but it splits the\n> ref-counted fake cache_entries out into a separate hash of \"struct\n> dir_hash_entry\" rather than storing it in the regular hash. Which IMHO\n> is a bit cleaner for two reasons:\n> \n>   1. You do not have to pay the memory price of storing fake\n>      cache_entries (the name+refcount struct for each directory is much\n>      smaller than a real cache_entry).\n> \n\nYes, but considering the small number of directories compared to files, I think this is a relatively small price to pay.\n\n>   2. It makes the code a bit simpler, as you do not have to do any\n>      \"check for trailing /\" magic on the result of index_name_exists to\n>      determine if it is a \"real\" name or just a fake dir entry.\n> \n\nTrue for dir.c. On the other hand, you need a lot of new find / find_or_create logic in name-hash.c.\n\nJust to illustrate what I mean, here's a quick sketch (there's still a segfault somewhere, but I don't have time to debug right now...).\n\nNote that hash_index_entry_directories works from right to left - if the immediate parent directory is there, there's no need to check the parent's parent.\n\ncache_entry.dir points to the parent directory so that we don't need to lookup all path components for reference counting when adding / removing entries.\n\nAs directory entries are 'real' cache_entries, we can reuse the existing index_name_exists and hash_index_entry code.\n\nI feel slightly guilty for abusing ce_size as reference counter...well :-)\n\n---\n cache.h     |  4 +++-\n name-hash.c | 80 ++++++++++++++++++++++++++++---------------------------------\n 2 files changed, 39 insertions(+), 45 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex 665b512..2bc1372 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -131,7 +131,7 @@ struct cache_entry {\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n \tstruct cache_entry *next;\n-\tstruct cache_entry *dir_next;\n+\tstruct cache_entry *dir;\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n@@ -285,6 +285,8 @@ extern void add_name_hash(struct index_state *istate, struct cache_entry *ce);\n static inline void remove_name_hash(struct cache_entry *ce)\n {\n \tce->ce_flags |= CE_UNHASHED;\n+\tif (ce->dir && !(--ce->dir->ce_size))\n+\t\tremove_name_hash(ce->dir);\n }\n \n \ndiff --git a/name-hash.c b/name-hash.c\nindex d8d25c2..01e8320 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -32,6 +32,9 @@ static unsigned int hash_name(const char *name, int namelen)\n \treturn hash;\n }\n \n+static struct cache_entry *lookup_index_entry(struct index_state *istate, const char *name, int namelen, int icase);\n+static void hash_index_entry(struct index_state *istate, struct cache_entry *ce);\n+\n static void hash_index_entry_directories(struct index_state *istate, struct cache_entry *ce)\n {\n \t/*\n@@ -40,30 +43,25 @@ static void hash_index_entry_directories(struct index_state *istate, struct cach\n \t * closing slash.  Despite submodules being a directory, they never\n \t * reach this point, because they are stored without a closing slash\n \t * in the cache.\n-\t *\n-\t * Note that the cache_entry stored with the directory does not\n-\t * represent the directory itself.  It is a pointer to an existing\n-\t * filename, and its only purpose is to represent existence of the\n-\t * directory in the cache.  It is very possible multiple directory\n-\t * hash entries may point to the same cache_entry.\n \t */\n-\tunsigned int hash;\n-\tvoid **pos;\n+\tint len = ce_namelen(ce);\n+\tif (len && ce->name[len - 1] == '/')\n+\t\tlen--;\n+\twhile (len && ce->name[len - 1] != '/')\n+\t\tlen--;\n+\tif (!len)\n+\t\treturn;\n \n-\tconst char *ptr = ce->name;\n-\twhile (*ptr) {\n-\t\twhile (*ptr && *ptr != '/')\n-\t\t\t++ptr;\n-\t\tif (*ptr == '/') {\n-\t\t\t++ptr;\n-\t\t\thash = hash_name(ce->name, ptr - ce->name);\n-\t\t\tpos = insert_hash(hash, ce, &istate->name_hash);\n-\t\t\tif (pos) {\n-\t\t\t\tce->dir_next = *pos;\n-\t\t\t\t*pos = ce;\n-\t\t\t}\n-\t\t}\n+\tce->dir = lookup_index_entry(istate, ce->name, len, ignore_case);\n+\tif (!ce->dir) {\n+\t\tce->dir = xcalloc(1, cache_entry_size(len));\n+\t\tmemcpy(ce->dir->name, ce->name, len);\n+\t\tce->dir->ce_namelen = len;\n+\t\tce->dir->name[len] = 0;\n+\t\thash_index_entry(istate, ce->dir);\n \t}\n+\tce->dir->ce_flags &= ~CE_UNHASHED;\n+\tce->dir->ce_size++;\n }\n \n static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n@@ -74,7 +72,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \tif (ce->ce_flags & CE_HASHED)\n \t\treturn;\n \tce->ce_flags |= CE_HASHED;\n-\tce->next = ce->dir_next = NULL;\n+\tce->next = ce->dir = NULL;\n \thash = hash_name(ce->name, ce_namelen(ce));\n \tpos = insert_hash(hash, ce, &istate->name_hash);\n \tif (pos) {\n@@ -137,38 +135,32 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n \tif (!icase)\n \t\treturn 0;\n \n-\t/*\n-\t * If the entry we're comparing is a filename (no trailing slash), then compare\n-\t * the lengths exactly.\n-\t */\n-\tif (name[namelen - 1] != '/')\n-\t\treturn slow_same_name(name, namelen, ce->name, len);\n-\n-\t/*\n-\t * For a directory, we point to an arbitrary cache_entry filename.  Just\n-\t * make sure the directory portion matches.\n-\t */\n-\treturn slow_same_name(name, namelen, ce->name, namelen < len ? namelen : len);\n+\treturn slow_same_name(name, namelen, ce->name, len);\n }\n \n-struct cache_entry *index_name_exists(struct index_state *istate, const char *name, int namelen, int icase)\n+static struct cache_entry *lookup_index_entry(struct index_state *istate, const char *name, int namelen, int icase)\n {\n \tunsigned int hash = hash_name(name, namelen);\n-\tstruct cache_entry *ce;\n-\n-\tlazy_init_name_hash(istate);\n-\tce = lookup_hash(hash, &istate->name_hash);\n+\tstruct cache_entry *ce = lookup_hash(hash, &istate->name_hash);\n \n \twhile (ce) {\n \t\tif (!(ce->ce_flags & CE_UNHASHED)) {\n \t\t\tif (same_name(ce, name, namelen, icase))\n \t\t\t\treturn ce;\n \t\t}\n-\t\tif (icase && name[namelen - 1] == '/')\n-\t\t\tce = ce->dir_next;\n-\t\telse\n-\t\t\tce = ce->next;\n+\t\tce = ce->next;\n \t}\n+\treturn NULL;\n+}\n+\n+struct cache_entry *index_name_exists(struct index_state *istate, const char *name, int namelen, int icase)\n+{\n+\tstruct cache_entry *ce;\n+\n+\tlazy_init_name_hash(istate);\n+\tce = lookup_index_entry(istate, name, namelen, icase);\n+\tif (ce)\n+\t\treturn ce;\n \n \t/*\n \t * Might be a submodule.  Despite submodules being directories,\n@@ -182,7 +174,7 @@ struct cache_entry *index_name_exists(struct index_state *istate, const char *na\n \t * true.\n \t */\n \tif (icase && name[namelen - 1] == '/') {\n-\t\tce = index_name_exists(istate, name, namelen - 1, icase);\n+\t\tce = lookup_index_entry(istate, name, namelen - 1, icase);\n \t\tif (ce && S_ISGITLINK(ce->ce_mode))\n \t\t\treturn ce;\n \t}\n"},{"id":"209516","messageId":"20130214143558.GA671@google.com","threadId":"32862","inReplyTo":"CANgJU+WYSD8RHb19EP0M89=Y_XskfjDtFWf51qjg4ur+rDb3ug@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Magnus Bäck","fromEmail":"baeck@google.com","sentAt":"2013-02-14T14:36:01Z","receivedAt":"2013-02-14T14:36:01Z","isPatch":false,"sender":{"key":"baeck@google.com","avatar":null},"body":"On Sunday, February 10, 2013 at 08:26 EST,\n     demerphq <demerphq@gmail.com> wrote:\n\n> Is windows stat really so slow?\n\nWell, the problem is that there is no such thing as \"Windows stat\" :-)\n\n> I encountered this perception in windows Perl in the past, and I know\n> that on windows Perl stat *appears* slow compared to *nix, because in\n> order to satisfy the full *nix stat interface, specifically the nlink\n> field, it must open and close the file*. As of 5.10 this can be\n> disabled by setting a magic var ${^WIN32_SLOPPY_STAT} to a true value,\n> which makes a significant improvement to the performance of the Perl\n> level stat implementation.  I would not be surprised if the cygwin\n> implementation of stat() has the same issue as Perl did, and that stat\n> appears much slower than it actually need be if you don't care about\n> the nlink field.\n\nI suggested a few years ago that FindFirstFile() be used to implement\nstat() since it's way faster than opening and closing the file, but\nFindFirstFile() apparently produces unreliable mtime results when DST\nshifts are involved.\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/114041\n(The reference link in Johannes Sixt's first email is broken, but I'm\nsure the information can be dug up.)\n\nBased on a quick look it seems GetFileAttributesEx() is still used for\nmingw and cygwin Git.\n\n-- \nMagnus Bäck\nbaeck@google.com\n"},{"id":"209517","messageId":"CACBZZX6BVuQWtrLuTVXZo+77sT4yZQ3pvN=_fMma24-zd0NNqA@mail.gmail.com","threadId":"32862","inReplyTo":"CALkWK0=EP0Lv1F_BArub7SpL9rgFhmPtpMOCgwFqfJmVE=oa=A@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2013-02-14T15:16:24Z","receivedAt":"2013-02-14T15:16:24Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Fri, Feb 8, 2013 at 10:10 PM, Ramkumar Ramachandra\n<artagnon@gmail.com> wrote:\n> For large repositories, many simple git commands like `git status`\n> take a while to respond.  I understand that this is because of large\n> number of stat() calls to figure out which files were changed.  I\n> overheard that Mercurial wants to solve this problem using itnotify,\n> but the idea bothers me because it's not portable.  Will Git ever\n> consider using inotify on Linux?  What is the downside?\n\nThere's one relatively easy sub-task of this that I haven't seen\nmentioned: Improving the speed of interactive rebase on large (as in\nlots of checked out files) repositories.\n\nThat's the single biggest thing that bothers me when I use Git with\nlarge repos, not the speed of \"git status\". When you \"git rebase -i\nHEAD~100\" re-arrange some patches and save the TODO list it takes say\n0.5-1s for each patch to be applied, but at least 10x less than that\non a small repository. E.g. try this on linux-2.6.git v.s. some small\nproject with a few dozen files.\n\nI looked into this a long while ago and remembered that rebase was\ndoing something like a git status for every commit that it made to\ncheck the dirtyness.\n\nThis could be vastly improved by having an unsafe option to git-rebase\nwhere it just assumes that the starting state + whatever it wrote out\nis the current state, i.e. it would break if someone stuck up on your\ncheckout during an interactive rebase and changed a file, but the\ncommon case of the user having exclusive access to the repo and\nwaiting for the rebase would be much faster.\n"},{"id":"209527","messageId":"7vpq023lhd.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"CACBZZX6BVuQWtrLuTVXZo+77sT4yZQ3pvN=_fMma24-zd0NNqA@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-14T16:31:58Z","receivedAt":"2013-02-14T16:31:58Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> I looked into this a long while ago and remembered that rebase was\n> doing something like a git status for every commit that it made to\n> check the dirtyness.\n>\n> This could be vastly improved by having an unsafe option to git-rebase\n> where it just assumes that the starting state + whatever it wrote out\n> is the current state, i.e. it would break if someone stuck up on your\n> checkout during an interactive rebase and changed a file,...\n\nYou could make it a lot safer than \"just assumes\", and the result\nmay become generally usable, I think.  For example, you can set a\n\"magic\" bit somewhere in $GIT_DIR/rebase-i while you are in \"I am\ndoing pick/pick/pick and the user will not interfere me\" mode, and\nclear that bit upon \"rebase --continue\".  And you cheat only while\nthat \"magic\" bit is set.\n"},{"id":"209788","messageId":"CALkWK0=_AoWwAd8FN+GGvogT+p7PmTsm+KHNk0F09ymi2Snywg@mail.gmail.com","threadId":"32862","inReplyTo":"CACBZZX6BVuQWtrLuTVXZo+77sT4yZQ3pvN=_fMma24-zd0NNqA@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-19T09:40:21Z","receivedAt":"2013-02-19T09:40:21Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Ævar Arnfjörð Bjarmason wrote:\n> On Fri, Feb 8, 2013 at 10:10 PM, Ramkumar Ramachandra\n> <artagnon@gmail.com> wrote:\n>> For large repositories, many simple git commands like `git status`\n>> take a while to respond.  I understand that this is because of large\n>> number of stat() calls to figure out which files were changed.  I\n>> overheard that Mercurial wants to solve this problem using itnotify,\n>> but the idea bothers me because it's not portable.  Will Git ever\n>> consider using inotify on Linux?  What is the downside?\n>\n> There's one relatively easy sub-task of this that I haven't seen\n> mentioned: Improving the speed of interactive rebase on large (as in\n> lots of checked out files) repositories.\n>\n> That's the single biggest thing that bothers me when I use Git with\n> large repos, not the speed of \"git status\". When you \"git rebase -i\n> HEAD~100\" re-arrange some patches and save the TODO list it takes say\n> 0.5-1s for each patch to be applied, but at least 10x less than that\n> on a small repository. E.g. try this on linux-2.6.git v.s. some small\n> project with a few dozen files.\n>\n> I looked into this a long while ago and remembered that rebase was\n> doing something like a git status for every commit that it made to\n> check the dirtyness.\n\nWhat is it really doing?  I think the main culprit is\nrequire_clean_work_tree() from git-sh-setup.sh, and that is only run\nin the `--continue` and `exec` codepaths.\n"},{"id":"209791","messageId":"CALkWK0=XFBfZjO3oCJ8Jxya1ud79MQcQFm6pmZpOU8c3MxVtqQ@mail.gmail.com","threadId":"32862","inReplyTo":"511AAA92.4030508@gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-19T09:49:07Z","receivedAt":"2013-02-19T09:49:07Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Karsten Blees wrote:\n> Am 11.02.2013 04:53, schrieb Duy Nguyen:\n>> On Sun, Feb 10, 2013 at 11:58 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n>>> Karsten Blees has done something similar-ish on Windows, and he posted\n>>> the results here:\n>>>\n>>> https://groups.google.com/forum/#!topic/msysgit/fL_jykUmUNE/discussion\n>>>\n>\n> The new hashtable implementation in fscache [1] supports O(1) removal and has no mingw dependencies - might come in handy for anyone trying to implement an inotify daemon.\n>\n> [1] https://github.com/kblees/git/commit/f7eb85c2\n\nThanks!  I'm cherry-picking this.  Why didn't you propose replacing\nhash.{c,h} with this in git.git though?\n\n>>> I also seem to remember he doing a ReadDirectoryChangesW version, but\n>>> I don't remember what happened with that.\n>>\n>> Thanks. I came across that but did not remember. For one thing, we\n>> know the inotify alternative for Windows: ReadDirectoryChangesW.\n>>\n>\n> I dropped ReadDirectoryChangesW because maintaining a 'live' file system cache became more and more complicated. For example, according to MSDN docs, ReadDirectoryChangesW *may* report short DOS 8.3 names (i.e. \"PROGRA~1\" instead of \"Program Files\"), so a correct and fast cache implementation would have to be indexed by long *and* short names...\n>\n> Another problem was that the 'live' cache had quite negative performance impact on mutating git commands (checkout, reset...). An inotify daemon running as a background process (not in-process as fscache) will probably affect everyone that modifies the working copy, e.g. running 'make' or the test-suite. This should be considered in the design.\n\nYes, an external daemon would report creation of *.o files, from the\ncompile, for instance.  We need a way for it to be filtered at the\ndaemon itself, so git isn't burdened with the filtering.\n"},{"id":"209792","messageId":"CALkWK0mgzzGTqjMxaEm1t+f69X=U7R203BiVugnnXE_MN2zFZw@mail.gmail.com","threadId":"32862","inReplyTo":"CAKXa9=pCSWtXq+5x_LcZ9gsSpa1yT0QD5DsBguTqosoH0cj-nw@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-02-19T09:57:56Z","receivedAt":"2013-02-19T09:57:56Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Robert Zeh wrote:\n> On Sun, Feb 10, 2013 at 9:21 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n>> On Mon, Feb 11, 2013 at 2:03 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n>>> On Sat, Feb 9, 2013 at 1:35 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>>>> Ramkumar Ramachandra <artagnon@gmail.com> writes:\n>>>>\n>>>>> This is much better than Junio's suggestion to study possible\n>>>>> implementations on all platforms and designing a generic daemon/\n>>>>> communication channel.  That's no weekend project.\n>>>>\n>>>> It appears that you misunderstood what I wrote.  That was not \"here\n>>>> is a design; I want it in my system.  Go implemment it\".\n>>>>\n>>>> It was \"If somebody wants to discuss it but does not know where to\n>>>> begin, doing a small experiment like this and reporting how well it\n>>>> worked here may be one way to do so.\", nothing more.\n>>>\n>>> What if instead of communicating over a socket, the daemon\n>>> dumped a file containing all of the lstat information after git\n>>> wrote a file? By definition the daemon should know about file writes.\n>>>\n>>> There would be no network communication, which I think would make\n>>> things more secure. It would simplify the rendezvous by insisting on\n>>> well known locations in $GIT_DIR.\n>>\n>> We need some sort of interactive communication to the daemon anyway,\n>> to validate that the information is uptodate. Assume that a user makes\n>> some changes to his worktree before starting the daemon, git needs to\n>> know that what the daemon provides does not represent a complete\n>> file-change picture and it better refreshes the index the old way\n>> once, then trust the daemon.\n>>\n>> I think we could solve that by storing a \"session id\", provided by the\n>> daemon, in .git/index. If the session id is not present (or does not\n>> match what the current daemon gives), refresh the old way. After\n>> refreshing, it may ask the daemon for new session id and store it.\n>> Next time if the session id is still valid, trust the daemon's data.\n>> This session id should be different every time the daemon restarts for\n>> this to work.\n>\n> I think we could do this without interactive communication,\n> if we did the following:\n>    1) The Daemon waits to see $GIT_DIR/lstat_request, and atomically\n>        writes out $GIT_DIR/lstat_cache.  By atomically I mean that it writes\n>        things out to a temporary file, and then does a rename.\n>\n>    2) The client erases $GIT_DIR/lstat_cache, and writes\n>       $GIT_DIR/lstat_request\n>\n> I think this is better than socket based communication because there\n> are fewer places to check\n> for failures.\n\nMy main problem with file-based solutions is this: how will the daemon\naccumulate inotify change events over time, and report it in a batch\nto a git application that is spawned?  Will it append to the\n.git/inotify_changes file everytime there's a change?  Wouldn't you\nprefer to accumulate the events in-memory and report it over a socket\nupon explicit request, to minimize IO?\n"},{"id":"209817","messageId":"CAM9Z-nnWE9LeHaefKdju_p=_he7aJcOunGf1Ms6K=vEXgxS25w@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8DeM5--WVXg3b65RxLBS7Jho-7KmcGwWk7B5uAx77yOEw@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Drew Northup","fromEmail":"n1xim.email@gmail.com","sentAt":"2013-02-19T13:16:58Z","receivedAt":"2013-02-19T13:16:58Z","isPatch":false,"sender":{"key":"n1xim.email@gmail.com","avatar":null},"body":"On Sun, Feb 10, 2013 at 12:24 AM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Sun, Feb 10, 2013 at 12:10 AM, Ramkumar Ramachandra\n> <artagnon@gmail.com> wrote:\n>> Finn notes in the commit message that it offers no speedup, because\n>> .gitignore files in every directory still have to be read.  I think\n>> this is silly: we really should be caching .gitignore, and touching it\n>> only when lstat() reports that the file has changed.\n>> ...\n>> Really, the elephant in the room right now seems to be .gitignore.\n>> Until that is fixed, there is really no use of writing this inotify\n>> daemon, no?  Can someone enlighten me on how exactly .gitignore files\n>> are processed?\n>\n> .gitignore is a different issue. I think it's mainly used with\n> read_directory/fill_directory to collect ignored files (or not-ignored\n> files). And it's not always used (well, status and add does, but diff\n> should not). I think wee need to measure how much mass lstat\n> elimination gains us (especially on big repos) and how much\n> .gitignore/.gitattributes caching does. I don't think .gitignore has\n> such a big impact though. strace on git.git tells me \"git status\"\n> issues about 2500 lstat calls, and just 740 open+getdents calls (on\n> total 3800 syscalls). I will think if we can do something about\n> .gitignore/.gitattributes.\n> --\n> Duy\n\nDuy,\nDid your testing turn up anything about the amount of time spent\nparsing the .gitignore/.gitattributes files? Not the syscall count,\nbut the actual time spent running the parser (which I presume is\nlargely CPU-bound). The other notable bit of information to know would\nbe how much time is spent applying what has been parsed out of those\nfiles to the content of the tree. Both will give a clear signal of the\nprominence of those segments of code versus others elsewhere in the\n\"git stat\" flow path. That information will tell us more clearly what,\nif anything, it is worth keeping a cache of and what form that cache\nshould take.\n\n-- \n-Drew Northup\n--------------------------------------------------------------\n\"As opposed to vegetable or mineral error?\"\n-John Pescatore, SANS NewsBites Vol. 12 Num. 59\n"},{"id":"209821","messageId":"CACsJy8DmPV+61kP6TPCVir-b_xMOWJxoWXunWmvK--7_3cW-iA@mail.gmail.com","threadId":"32862","inReplyTo":"CAM9Z-nnWE9LeHaefKdju_p=_he7aJcOunGf1Ms6K=vEXgxS25w@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-02-19T13:47:02Z","receivedAt":"2013-02-19T13:47:02Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Feb 19, 2013 at 8:16 PM, Drew Northup <n1xim.email@gmail.com> wrote:\n> Did your testing turn up anything about the amount of time spent\n> parsing the .gitignore/.gitattributes files? Not the syscall count,\n> but the actual time spent running the parser (which I presume is\n> largely CPU-bound). The other notable bit of information to know would\n> be how much time is spent applying what has been parsed out of those\n> files to the content of the tree. Both will give a clear signal of the\n> prominence of those segments of code versus others elsewhere in the\n> \"git stat\" flow path. That information will tell us more clearly what,\n> if anything, it is worth keeping a cache of and what form that cache\n> should take.\n\nNot specifically parsing, but we do waste CPU on\n.gitignore/.gitattributes stuff. See\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/216347/focus=216381\n\nOther measurements (which led to the above patch):\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/215820/focus=215900\nhttp://thread.gmane.org/gmane.comp.version-control.git/215820/focus=216029\nhttp://thread.gmane.org/gmane.comp.version-control.git/215820/focus=216195\n\nSo far we could reduce lstat, {open,read,close}dir syscalls with the\nhelp of inotify, which saves time. I'm not sure if we should cache the\nlist of untracked-but-not-ignored files. It cuts down cpu time on\n.gitignore but invalidation could be complicated.\n-- \nDuy\n"},{"id":"209824","messageId":"51238B59.1050106@dcon.de","threadId":"32862","inReplyTo":"CALkWK0=XFBfZjO3oCJ8Jxya1ud79MQcQFm6pmZpOU8c3MxVtqQ@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-19T14:25:29Z","receivedAt":"2013-02-19T14:25:29Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 19.02.2013 10:49, schrieb Ramkumar Ramachandra:\n> Karsten Blees wrote:\n>> Am 11.02.2013 04:53, schrieb Duy Nguyen:\n>>> On Sun, Feb 10, 2013 at 11:58 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n>>>> Karsten Blees has done something similar-ish on Windows, and he posted\n>>>> the results here:\n>>>>\n>>>> https://groups.google.com/forum/#!topic/msysgit/fL_jykUmUNE/discussion\n>>>>\n>>\n>> The new hashtable implementation in fscache [1] supports O(1) removal and has no mingw dependencies - might come in handy for anyone trying to implement an inotify daemon.\n>>\n>> [1] https://github.com/kblees/git/commit/f7eb85c2\n> \n> Thanks!  I'm cherry-picking this.  Why didn't you propose replacing\n> hash.{c,h} with this in git.git though?\n> \n\nI was planning to, but didn't find the time yet to adapt existing hash.[ch] uses to the new version, and there's not much use adding four more files of dead code. If someone else could jump in here that would be great.\n\nNote that there's another t0007 now, so t/t0007-hashmap.sh needs to be renamed.\n\n>>>> I also seem to remember he doing a ReadDirectoryChangesW version, but\n>>>> I don't remember what happened with that.\n>>>\n>>> Thanks. I came across that but did not remember. For one thing, we\n>>> know the inotify alternative for Windows: ReadDirectoryChangesW.\n>>>\n>>\n>> I dropped ReadDirectoryChangesW because maintaining a 'live' file system cache became more and more complicated. For example, according to MSDN docs, ReadDirectoryChangesW *may* report short DOS 8.3 names (i.e. \"PROGRA~1\" instead of \"Program Files\"), so a correct and fast cache implementation would have to be indexed by long *and* short names...\n>>\n>> Another problem was that the 'live' cache had quite negative performance impact on mutating git commands (checkout, reset...). An inotify daemon running as a background process (not in-process as fscache) will probably affect everyone that modifies the working copy, e.g. running 'make' or the test-suite. This should be considered in the design.\n> \n> Yes, an external daemon would report creation of *.o files, from the\n> compile, for instance.  We need a way for it to be filtered at the\n> daemon itself, so git isn't burdened with the filtering.\n> \n\n...and this filtering should affect foreground processes as little as possible. For example, gaining 1 s per git-status is counter-productive if compile time increases by 10 s because the daemon re-reads .gitignore files for every new *.o.\n"},{"id":"210395","messageId":"512E1C0F.3050506@gmail.com","threadId":"32862","inReplyTo":"511C3454.6080204@gmail.com","subject":"[PATCH] name-hash.c: fix endless loop with core.ignorecase=true","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-27T14:45:35Z","receivedAt":"2013-02-27T14:45:35Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"With core.ignorecase=true, name-hash.c builds a case insensitive index of\nall tracked directories. Currently, the existing cache entry structures are\nadded multiple times to the same hashtable (with different name lengths and\nhash codes). However, there's only one dir_next pointer, which gets\ncompletely messed up in case of hash collisions. In the worst case, this\ncauses an endless loop if ce == ce->dir_next:\n\n---8<---\n# \"V/\", \"V/XQANY/\" and \"WURZAUP/\" all have the same hash_name\nmkdir V\nmkdir V/XQANY\nmkdir WURZAUP\ntouch V/XQANY/test\ngit init\ngit config core.ignorecase true\ngit add .\ngit status\n---8<---\n\nUse a separate hashtable and separate structures for the directory index\nso that each directory entry has its own next pointer. Use reference\ncounting to track which directory entry contains files.\n\nThere are only slight changes to the name-hash.c API:\n- new free_name_hash() used by read_cache.c::discard_index()\n- remove_name_hash() takes an additional index_state parameter\n- index_name_exists() for a directory (trailing '/') may return a cache\n  entry that has been removed (CE_UNHASHED). This is not a problem as the\n  return value is only used to check if the directory exists (dir.c) or to\n  normalize casing of directory names (read-cache.c).\n\nGetting rid of cache_entry.dir_next reduces memory consumption, especially\nwith core.ignorecase=false (which doesn't use that member at all).\n\nWith core.ignorecase=true, building the directory index is slightly faster\nas we add / check the parent directory first (instead of going through all\ndirectory levels for each file in the index). E.g. with WebKit (~200k\nfiles, ~7k dirs), time taken in lazy_init_name_hash is reduced from 176ms\nto 130ms.\n\nSigned-off-by: Karsten Blees <blees@dcon.de>\n---\nAlso available here:\nhttps://github.com/kblees/git/tree/kb/name-hash-fix-endless-loop\ngit pull git://github.com/kblees/git.git kb/name-hash-fix-endless-loop\n\nThis combines the pros of the patches suggested by Jeff and me:\n- reduced memory usage due to smaller dir_entry and cache_entry structs\n- faster indexing due to right-to-left directory lookup\n- marginal API changes, i.e. less impact on the rest of git\n\nTest suite runs clean on msysgit and Linux.\n\nHave fun,\nKarsten\n\n cache.h      |  17 ++-----\n name-hash.c  | 164 +++++++++++++++++++++++++++++++++++++++++++----------------\n read-cache.c |   9 ++--\n 3 files changed, 126 insertions(+), 64 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex e493563..898e346 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -131,7 +131,6 @@ struct cache_entry {\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n \tstruct cache_entry *next;\n-\tstruct cache_entry *dir_next;\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n@@ -267,25 +266,15 @@ struct index_state {\n \tunsigned name_hash_initialized : 1,\n \t\t initialized : 1;\n \tstruct hash_table name_hash;\n+\tstruct hash_table dir_hash;\n };\n \n extern struct index_state the_index;\n \n /* Name hashing */\n extern void add_name_hash(struct index_state *istate, struct cache_entry *ce);\n-/*\n- * We don't actually *remove* it, we can just mark it invalid so that\n- * we won't find it in lookups.\n- *\n- * Not only would we have to search the lists (simple enough), but\n- * we'd also have to rehash other hash buckets in case this makes the\n- * hash bucket empty (common). So it's much better to just mark\n- * it.\n- */\n-static inline void remove_name_hash(struct cache_entry *ce)\n-{\n-\tce->ce_flags |= CE_UNHASHED;\n-}\n+extern void remove_name_hash(struct index_state *istate, struct cache_entry *ce);\n+extern void free_name_hash(struct index_state *istate);\n \n \n #ifndef NO_THE_INDEX_COMPATIBILITY_MACROS\ndiff --git a/name-hash.c b/name-hash.c\nindex 942c459..6b130e1 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -32,38 +32,75 @@ static unsigned int hash_name(const char *name, int namelen)\n \treturn hash;\n }\n \n-static void hash_index_entry_directories(struct index_state *istate, struct cache_entry *ce)\n+struct dir_entry {\n+\tstruct dir_entry *next;\n+\tstruct dir_entry *parent;\n+\tstruct cache_entry *ce;\n+\tint nr;\n+\tunsigned int namelen;\n+};\n+\n+static struct dir_entry *find_dir_entry(struct index_state *istate,\n+\t\tconst char *name, unsigned int namelen)\n+{\n+\tunsigned int hash = hash_name(name, namelen);\n+\tstruct dir_entry *dir;\n+\n+\tfor (dir = lookup_hash(hash, &istate->dir_hash); dir; dir = dir->next)\n+\t\tif (dir->namelen == namelen &&\n+\t\t    !strncasecmp(dir->ce->name, name, namelen))\n+\t\t\treturn dir;\n+\treturn NULL;\n+}\n+\n+static struct dir_entry *hash_dir_entry(struct index_state *istate,\n+\t\tstruct cache_entry *ce, int namelen, int add)\n {\n \t/*\n \t * Throw each directory component in the hash for quick lookup\n \t * during a git status. Directory components are stored with their\n-\t * closing slash.  Despite submodules being a directory, they never\n-\t * reach this point, because they are stored without a closing slash\n-\t * in the cache.\n-\t *\n-\t * Note that the cache_entry stored with the directory does not\n-\t * represent the directory itself.  It is a pointer to an existing\n-\t * filename, and its only purpose is to represent existence of the\n-\t * directory in the cache.  It is very possible multiple directory\n-\t * hash entries may point to the same cache_entry.\n+\t * closing slash.\n \t */\n-\tunsigned int hash;\n-\tvoid **pos;\n+\tstruct dir_entry *dir, *p;\n+\n+\t/* get length of parent directory */\n+\twhile (namelen > 0 && !is_dir_sep(ce->name[namelen - 1]))\n+\t\tnamelen--;\n+\tif (namelen <= 0)\n+\t\treturn NULL;\n+\n+\t/* lookup existing entry for that directory */\n+\tdir = find_dir_entry(istate, ce->name, namelen);\n+\tif (add && !dir) {\n+\t\t/* not found, create it and add to hash table */\n+\t\tvoid **pdir;\n+\t\tunsigned int hash = hash_name(ce->name, namelen);\n \n-\tconst char *ptr = ce->name;\n-\twhile (*ptr) {\n-\t\twhile (*ptr && *ptr != '/')\n-\t\t\t++ptr;\n-\t\tif (*ptr == '/') {\n-\t\t\t++ptr;\n-\t\t\thash = hash_name(ce->name, ptr - ce->name);\n-\t\t\tpos = insert_hash(hash, ce, &istate->name_hash);\n-\t\t\tif (pos) {\n-\t\t\t\tce->dir_next = *pos;\n-\t\t\t\t*pos = ce;\n-\t\t\t}\n+\t\tdir = xcalloc(1, sizeof(struct dir_entry));\n+\t\tdir->namelen = namelen;\n+\t\tdir->ce = ce;\n+\n+\t\tpdir = insert_hash(hash, dir, &istate->dir_hash);\n+\t\tif (pdir) {\n+\t\t\tdir->next = *pdir;\n+\t\t\t*pdir = dir;\n \t\t}\n+\n+\t\t/* recursively add missing parent directories */\n+\t\tdir->parent = hash_dir_entry(istate, ce, namelen - 1, add);\n \t}\n+\n+\t/* add or release reference to this entry (and parents if 0) */\n+\tp = dir;\n+\tif (add) {\n+\t\twhile (p && !(p->nr++))\n+\t\t\tp = p->parent;\n+\t} else {\n+\t\twhile (p && p->nr && !(--p->nr))\n+\t\t\tp = p->parent;\n+\t}\n+\n+\treturn dir;\n }\n \n static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n@@ -74,7 +111,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \tif (ce->ce_flags & CE_HASHED)\n \t\treturn;\n \tce->ce_flags |= CE_HASHED;\n-\tce->next = ce->dir_next = NULL;\n+\tce->next = NULL;\n \thash = hash_name(ce->name, ce_namelen(ce));\n \tpos = insert_hash(hash, ce, &istate->name_hash);\n \tif (pos) {\n@@ -82,8 +119,8 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \t\t*pos = ce;\n \t}\n \n-\tif (ignore_case)\n-\t\thash_index_entry_directories(istate, ce);\n+\tif (ignore_case && !(ce->ce_flags & CE_UNHASHED))\n+\t\thash_dir_entry(istate, ce, ce_namelen(ce), 1);\n }\n \n static void lazy_init_name_hash(struct index_state *istate)\n@@ -99,11 +136,33 @@ static void lazy_init_name_hash(struct index_state *istate)\n \n void add_name_hash(struct index_state *istate, struct cache_entry *ce)\n {\n+\t/* if already hashed, add reference to directory entries */\n+\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_STATE_MASK)\n+\t\thash_dir_entry(istate, ce, ce_namelen(ce), 1);\n+\n \tce->ce_flags &= ~CE_UNHASHED;\n \tif (istate->name_hash_initialized)\n \t\thash_index_entry(istate, ce);\n }\n \n+/*\n+ * We don't actually *remove* it, we can just mark it invalid so that\n+ * we won't find it in lookups.\n+ *\n+ * Not only would we have to search the lists (simple enough), but\n+ * we'd also have to rehash other hash buckets in case this makes the\n+ * hash bucket empty (common). So it's much better to just mark\n+ * it.\n+ */\n+void remove_name_hash(struct index_state *istate, struct cache_entry *ce)\n+{\n+\t/* if already hashed, release reference to directory entries */\n+\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_HASHED)\n+\t\thash_dir_entry(istate, ce, ce_namelen(ce), 0);\n+\n+\tce->ce_flags |= CE_UNHASHED;\n+}\n+\n static int slow_same_name(const char *name1, int len1, const char *name2, int len2)\n {\n \tif (len1 != len2)\n@@ -137,18 +196,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n \tif (!icase)\n \t\treturn 0;\n \n-\t/*\n-\t * If the entry we're comparing is a filename (no trailing slash), then compare\n-\t * the lengths exactly.\n-\t */\n-\tif (name[namelen - 1] != '/')\n-\t\treturn slow_same_name(name, namelen, ce->name, len);\n-\n-\t/*\n-\t * For a directory, we point to an arbitrary cache_entry filename.  Just\n-\t * make sure the directory portion matches.\n-\t */\n-\treturn slow_same_name(name, namelen, ce->name, namelen < len ? namelen : len);\n+\treturn slow_same_name(name, namelen, ce->name, len);\n }\n \n struct cache_entry *index_name_exists(struct index_state *istate, const char *name, int namelen, int icase)\n@@ -164,16 +212,14 @@ struct cache_entry *index_name_exists(struct index_state *istate, const char *na\n \t\t\tif (same_name(ce, name, namelen, icase))\n \t\t\t\treturn ce;\n \t\t}\n-\t\tif (icase && name[namelen - 1] == '/')\n-\t\t\tce = ce->dir_next;\n-\t\telse\n-\t\t\tce = ce->next;\n+\t\tce = ce->next;\n \t}\n \n \t/*\n-\t * Might be a submodule.  Despite submodules being directories,\n+\t * When looking for a directory (trailing '/'), it might be a\n+\t * submodule or a directory. Despite submodules being directories,\n \t * they are stored in the name hash without a closing slash.\n-\t * When ignore_case is 1, directories are stored in the name hash\n+\t * When ignore_case is 1, directories are stored in a separate hash\n \t * with their closing slash.\n \t *\n \t * The side effect of this storage technique is we have need to\n@@ -182,9 +228,37 @@ struct cache_entry *index_name_exists(struct index_state *istate, const char *na\n \t * true.\n \t */\n \tif (icase && name[namelen - 1] == '/') {\n+\t\tstruct dir_entry *dir = find_dir_entry(istate, name, namelen);\n+\t\tif (dir && dir->nr)\n+\t\t\treturn dir->ce;\n+\n \t\tce = index_name_exists(istate, name, namelen - 1, icase);\n \t\tif (ce && S_ISGITLINK(ce->ce_mode))\n \t\t\treturn ce;\n \t}\n \treturn NULL;\n }\n+\n+static int free_dir_entry(void *entry, void *unused)\n+{\n+\tstruct dir_entry *dir = entry;\n+\twhile (dir) {\n+\t\tstruct dir_entry *next = dir->next;\n+\t\tfree(dir);\n+\t\tdir = next;\n+\t}\n+\treturn 0;\n+}\n+\n+void free_name_hash(struct index_state *istate)\n+{\n+\tif (!istate->name_hash_initialized)\n+\t\treturn;\n+\tistate->name_hash_initialized = 0;\n+\tif (ignore_case)\n+\t\t/* free directory entries */\n+\t\tfor_each_hash(&istate->dir_hash, free_dir_entry, NULL);\n+\n+\tfree_hash(&istate->name_hash);\n+\tfree_hash(&istate->dir_hash);\n+}\ndiff --git a/read-cache.c b/read-cache.c\nindex 827ae55..47eb9d8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -46,7 +46,7 @@ static void replace_index_entry(struct index_state *istate, int nr, struct cache\n {\n \tstruct cache_entry *old = istate->cache[nr];\n \n-\tremove_name_hash(old);\n+\tremove_name_hash(istate, old);\n \tset_index_entry(istate, nr, ce);\n \tistate->cache_changed = 1;\n }\n@@ -460,7 +460,7 @@ int remove_index_entry_at(struct index_state *istate, int pos)\n \tstruct cache_entry *ce = istate->cache[pos];\n \n \trecord_resolve_undo(istate, ce);\n-\tremove_name_hash(ce);\n+\tremove_name_hash(istate, ce);\n \tistate->cache_changed = 1;\n \tistate->cache_nr--;\n \tif (pos >= istate->cache_nr)\n@@ -483,7 +483,7 @@ void remove_marked_cache_entries(struct index_state *istate)\n \n \tfor (i = j = 0; i < istate->cache_nr; i++) {\n \t\tif (ce_array[i]->ce_flags & CE_REMOVE)\n-\t\t\tremove_name_hash(ce_array[i]);\n+\t\t\tremove_name_hash(istate, ce_array[i]);\n \t\telse\n \t\t\tce_array[j++] = ce_array[i];\n \t}\n@@ -1515,8 +1515,7 @@ int discard_index(struct index_state *istate)\n \tistate->cache_changed = 0;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n-\tistate->name_hash_initialized = 0;\n-\tfree_hash(&istate->name_hash);\n+\tfree_name_hash(istate);\n \tcache_tree_free(&(istate->cache_tree));\n \tistate->initialized = 0;\n \n-- \n1.8.1.2.7986.g6e98809.dirty\n"},{"id":"210408","messageId":"7v621dk8aa.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"512E1C0F.3050506@gmail.com","subject":"Re: [PATCH] name-hash.c: fix endless loop with core.ignorecase=true","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-27T16:53:33Z","receivedAt":"2013-02-27T16:53:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Karsten Blees <karsten.blees@gmail.com> writes:\n\n> With core.ignorecase=true, name-hash.c builds a case insensitive index of\n> all tracked directories. Currently, the existing cache entry structures are\n> added multiple times to the same hashtable (with different name lengths and\n> hash codes). However, there's only one dir_next pointer, which gets\n> completely messed up in case of hash collisions. In the worst case, this\n> causes an endless loop if ce == ce->dir_next:\n>\n> ---8<---\n> # \"V/\", \"V/XQANY/\" and \"WURZAUP/\" all have the same hash_name\n> mkdir V\n> mkdir V/XQANY\n> mkdir WURZAUP\n> touch V/XQANY/test\n> git init\n> git config core.ignorecase true\n> git add .\n> git status\n> ---8<---\n\nInstead of using the scissors mark to confuse \"am -c\", indenting\nthis block would have been easier to later readers.\n\nAlso it is somewhat a shame that we do not use the above sample\ncollisions in a new test case.\n\n> +static struct dir_entry *hash_dir_entry(struct index_state *istate,\n> +\t\tstruct cache_entry *ce, int namelen, int add)\n>  {\n>  \t/*\n>  \t * Throw each directory component in the hash for quick lookup\n>  \t * during a git status. Directory components are stored with their\n> -\t * closing slash.  Despite submodules being a directory, they never\n> -\t * reach this point, because they are stored without a closing slash\n> -\t * in the cache.\n\nIs the description of submodule no longer relevant?\n\n> -\t * Note that the cache_entry stored with the directory does not\n> -\t * represent the directory itself.  It is a pointer to an existing\n> -\t * filename, and its only purpose is to represent existence of the\n> -\t * directory in the cache.  It is very possible multiple directory\n> -\t * hash entries may point to the same cache_entry.\n\nIs this paragraph no longer relevant?  It seems to me that it still\nholds true, given the way how dir->ce points at the given ce.\n\n> +\t * closing slash.\n>  \t */\n> +\tstruct dir_entry *dir, *p;\n> +\n> +\t/* get length of parent directory */\n> +\twhile (namelen > 0 && !is_dir_sep(ce->name[namelen - 1]))\n> +\t\tnamelen--;\n> +\tif (namelen <= 0)\n> +\t\treturn NULL;\n> +\n> +\t/* lookup existing entry for that directory */\n> +\tdir = find_dir_entry(istate, ce->name, namelen);\n> +\tif (add && !dir) {\n> ...\n>  \t}\n> +\n> +\t/* add or release reference to this entry (and parents if 0) */\n> +\tp = dir;\n> +\tif (add) {\n> +\t\twhile (p && !(p->nr++))\n> +\t\t\tp = p->parent;\n> +\t} else {\n> +\t\twhile (p && p->nr && !(--p->nr))\n> +\t\t\tp = p->parent;\n> +\t}\n\nCan we free the entry when its refcnt goes down to zero?  If yes, is\nit worth doing so?\n\n> +\n> +\treturn dir;\n>  }\n>  \n>  static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n> @@ -74,7 +111,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n>  \tif (ce->ce_flags & CE_HASHED)\n>  \t\treturn;\n>  \tce->ce_flags |= CE_HASHED;\n> -\tce->next = ce->dir_next = NULL;\n> +\tce->next = NULL;\n>  \thash = hash_name(ce->name, ce_namelen(ce));\n>  \tpos = insert_hash(hash, ce, &istate->name_hash);\n>  \tif (pos) {\n> @@ -82,8 +119,8 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n>  \t\t*pos = ce;\n>  \t}\n>  \n> -\tif (ignore_case)\n> -\t\thash_index_entry_directories(istate, ce);\n> +\tif (ignore_case && !(ce->ce_flags & CE_UNHASHED))\n> +\t\thash_dir_entry(istate, ce, ce_namelen(ce), 1);\n>  }\n>  \n>  static void lazy_init_name_hash(struct index_state *istate)\n> @@ -99,11 +136,33 @@ static void lazy_init_name_hash(struct index_state *istate)\n>  \n>  void add_name_hash(struct index_state *istate, struct cache_entry *ce)\n>  {\n> +\t/* if already hashed, add reference to directory entries */\n> +\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_STATE_MASK)\n> +\t\thash_dir_entry(istate, ce, ce_namelen(ce), 1);\n\nInstead of a single function with \"are we adding or removing?\"\nparameter, it would be a lot easier to read the callers if they are\nwrapped in two helpers, add_dir_entry() and del_dir_entry() or\nsomething, especially when the add=[0|1] parameter is constant for\neach and every callsite (i.e. the codeflow determines it, not the\ndata).\n\n>  \tce->ce_flags &= ~CE_UNHASHED;\n>  \tif (istate->name_hash_initialized)\n>  \t\thash_index_entry(istate, ce);\n>  }\n>  \n> +/*\n> + * We don't actually *remove* it, we can just mark it invalid so that\n> + * we won't find it in lookups.\n> + *\n> + * Not only would we have to search the lists (simple enough), but\n> + * we'd also have to rehash other hash buckets in case this makes the\n> + * hash bucket empty (common). So it's much better to just mark\n> + * it.\n> + */\n> +void remove_name_hash(struct index_state *istate, struct cache_entry *ce)\n> +{\n> +\t/* if already hashed, release reference to directory entries */\n> +\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_HASHED)\n> +\t\thash_dir_entry(istate, ce, ce_namelen(ce), 0);\n\nAnd here as well.\n\n> +\n> +\tce->ce_flags |= CE_UNHASHED;\n> +}\n> +\n>  static int slow_same_name(const char *name1, int len1, const char *name2, int len2)\n>  {\n>  \tif (len1 != len2)\n> @@ -137,18 +196,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n>  \tif (!icase)\n>  \t\treturn 0;\n>  \n> -\t/*\n> -\t * If the entry we're comparing is a filename (no trailing slash), then compare\n> -\t * the lengths exactly.\n> -\t */\n> -\tif (name[namelen - 1] != '/')\n> -\t\treturn slow_same_name(name, namelen, ce->name, len);\n> -\n> -\t/*\n> -\t * For a directory, we point to an arbitrary cache_entry filename.  Just\n> -\t * make sure the directory portion matches.\n> -\t */\n> -\treturn slow_same_name(name, namelen, ce->name, namelen < len ? namelen : len);\n> +\treturn slow_same_name(name, namelen, ce->name, len);\n\nHmph, what is this change about?  Nobody calls same_name() with a\ndirectory name anymore or something?\n\nThanks.\n"},{"id":"210412","messageId":"512E8014.3090107@gmail.com","threadId":"32862","inReplyTo":"7v621dk8aa.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] name-hash.c: fix endless loop with core.ignorecase=true","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-27T21:52:20Z","receivedAt":"2013-02-27T21:52:20Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 27.02.2013 17:53, schrieb Junio C Hamano:\n> Karsten Blees <karsten.blees@gmail.com> writes:\n> \n>> With core.ignorecase=true, name-hash.c builds a case insensitive index of\n>> all tracked directories. Currently, the existing cache entry structures are\n>> added multiple times to the same hashtable (with different name lengths and\n>> hash codes). However, there's only one dir_next pointer, which gets\n>> completely messed up in case of hash collisions. In the worst case, this\n>> causes an endless loop if ce == ce->dir_next:\n>>\n>> ---8<---\n>> # \"V/\", \"V/XQANY/\" and \"WURZAUP/\" all have the same hash_name\n>> mkdir V\n>> mkdir V/XQANY\n>> mkdir WURZAUP\n>> touch V/XQANY/test\n>> git init\n>> git config core.ignorecase true\n>> git add .\n>> git status\n>> ---8<---\n> \n> Instead of using the scissors mark to confuse \"am -c\", indenting\n> this block would have been easier to later readers.\n> \n> Also it is somewhat a shame that we do not use the above sample\n> collisions in a new test case.\n> \n\nDuly noted.\n\nIs there a way to run 'git status' with timeout? A test that doesn't complete (instead of failing) isn't nice...\n\n>> +static struct dir_entry *hash_dir_entry(struct index_state *istate,\n>> +\t\tstruct cache_entry *ce, int namelen, int add)\n>>  {\n>>  \t/*\n>>  \t * Throw each directory component in the hash for quick lookup\n>>  \t * during a git status. Directory components are stored with their\n>> -\t * closing slash.  Despite submodules being a directory, they never\n>> -\t * reach this point, because they are stored without a closing slash\n>> -\t * in the cache.\n> \n> Is the description of submodule no longer relevant?\n> \n>> -\t * Note that the cache_entry stored with the directory does not\n>> -\t * represent the directory itself.  It is a pointer to an existing\n>> -\t * filename, and its only purpose is to represent existence of the\n>> -\t * directory in the cache.  It is very possible multiple directory\n>> -\t * hash entries may point to the same cache_entry.\n> \n> Is this paragraph no longer relevant?  It seems to me that it still\n> holds true, given the way how dir->ce points at the given ce.\n> \n\nI interpreted this as an explanation why it was safe to add the same cache_entry to the same name_hash multiple times...now that we have separate dir_entries and index_state.dir_hash, that's no longer a problem. But rereading that paragraph again, it is still mostly true (except for the 'existance' part, which is solved by reference counting).\n\n>> +\t * closing slash.\n>>  \t */\n>> +\tstruct dir_entry *dir, *p;\n>> +\n>> +\t/* get length of parent directory */\n>> +\twhile (namelen > 0 && !is_dir_sep(ce->name[namelen - 1]))\n>> +\t\tnamelen--;\n>> +\tif (namelen <= 0)\n>> +\t\treturn NULL;\n>> +\n>> +\t/* lookup existing entry for that directory */\n>> +\tdir = find_dir_entry(istate, ce->name, namelen);\n>> +\tif (add && !dir) {\n>> ...\n>>  \t}\n>> +\n>> +\t/* add or release reference to this entry (and parents if 0) */\n>> +\tp = dir;\n>> +\tif (add) {\n>> +\t\twhile (p && !(p->nr++))\n>> +\t\t\tp = p->parent;\n>> +\t} else {\n>> +\t\twhile (p && p->nr && !(--p->nr))\n>> +\t\t\tp = p->parent;\n>> +\t}\n> \n> Can we free the entry when its refcnt goes down to zero?  If yes, is\n> it worth doing so?\n> \n\nThere's no remove_hash in hash.[ch], and dir_entry.next may point to another dir_entry with the same hash code, so we must not free the memory (same problem as CE_UNHASHED).\n\n>> +\n>> +\treturn dir;\n>>  }\n>>  \n>>  static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n>> @@ -74,7 +111,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n>>  \tif (ce->ce_flags & CE_HASHED)\n>>  \t\treturn;\n>>  \tce->ce_flags |= CE_HASHED;\n>> -\tce->next = ce->dir_next = NULL;\n>> +\tce->next = NULL;\n>>  \thash = hash_name(ce->name, ce_namelen(ce));\n>>  \tpos = insert_hash(hash, ce, &istate->name_hash);\n>>  \tif (pos) {\n>> @@ -82,8 +119,8 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n>>  \t\t*pos = ce;\n>>  \t}\n>>  \n>> -\tif (ignore_case)\n>> -\t\thash_index_entry_directories(istate, ce);\n>> +\tif (ignore_case && !(ce->ce_flags & CE_UNHASHED))\n>> +\t\thash_dir_entry(istate, ce, ce_namelen(ce), 1);\n>>  }\n>>  \n>>  static void lazy_init_name_hash(struct index_state *istate)\n>> @@ -99,11 +136,33 @@ static void lazy_init_name_hash(struct index_state *istate)\n>>  \n>>  void add_name_hash(struct index_state *istate, struct cache_entry *ce)\n>>  {\n>> +\t/* if already hashed, add reference to directory entries */\n>> +\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_STATE_MASK)\n>> +\t\thash_dir_entry(istate, ce, ce_namelen(ce), 1);\n> \n> Instead of a single function with \"are we adding or removing?\"\n> parameter, it would be a lot easier to read the callers if they are\n> wrapped in two helpers, add_dir_entry() and del_dir_entry() or\n> something, especially when the add=[0|1] parameter is constant for\n> each and every callsite (i.e. the codeflow determines it, not the\n> data).\n> \n\nOK\n\n>>  \tce->ce_flags &= ~CE_UNHASHED;\n>>  \tif (istate->name_hash_initialized)\n>>  \t\thash_index_entry(istate, ce);\n>>  }\n>>  \n>> +/*\n>> + * We don't actually *remove* it, we can just mark it invalid so that\n>> + * we won't find it in lookups.\n>> + *\n>> + * Not only would we have to search the lists (simple enough), but\n>> + * we'd also have to rehash other hash buckets in case this makes the\n>> + * hash bucket empty (common). So it's much better to just mark\n>> + * it.\n>> + */\n>> +void remove_name_hash(struct index_state *istate, struct cache_entry *ce)\n>> +{\n>> +\t/* if already hashed, release reference to directory entries */\n>> +\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_HASHED)\n>> +\t\thash_dir_entry(istate, ce, ce_namelen(ce), 0);\n> \n> And here as well.\n> \n>> +\n>> +\tce->ce_flags |= CE_UNHASHED;\n>> +}\n>> +\n>>  static int slow_same_name(const char *name1, int len1, const char *name2, int len2)\n>>  {\n>>  \tif (len1 != len2)\n>> @@ -137,18 +196,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n>>  \tif (!icase)\n>>  \t\treturn 0;\n>>  \n>> -\t/*\n>> -\t * If the entry we're comparing is a filename (no trailing slash), then compare\n>> -\t * the lengths exactly.\n>> -\t */\n>> -\tif (name[namelen - 1] != '/')\n>> -\t\treturn slow_same_name(name, namelen, ce->name, len);\n>> -\n>> -\t/*\n>> -\t * For a directory, we point to an arbitrary cache_entry filename.  Just\n>> -\t * make sure the directory portion matches.\n>> -\t */\n>> -\treturn slow_same_name(name, namelen, ce->name, namelen < len ? namelen : len);\n>> +\treturn slow_same_name(name, namelen, ce->name, len);\n> \n> Hmph, what is this change about?  Nobody calls same_name() with a\n> directory name anymore or something?\n> \n\ndir_entries (with trailing /) are in index_state.dir_hash, so we wouldn't expect to find anything in index_state.name_hash, especially not a cache_entry. find_dir_entry simply uses strncasecmp, as we only do directory indexing with core.ignorecase=true.\n\n> Thanks.\n> \n"},{"id":"210416","messageId":"512E9D7C.2030803@gmail.com","threadId":"32862","inReplyTo":"512E8014.3090107@gmail.com","subject":"[PATCH v2] name-hash.c: fix endless loop with core.ignorecase=true","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2013-02-27T23:57:48Z","receivedAt":"2013-02-27T23:57:48Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"With core.ignorecase=true, name-hash.c builds a case insensitive index of\nall tracked directories. Currently, the existing cache entry structures are\nadded multiple times to the same hashtable (with different name lengths and\nhash codes). However, there's only one dir_next pointer, which gets\ncompletely messed up in case of hash collisions. In the worst case, this\ncauses an endless loop if ce == ce->dir_next (see t7062).\n\nUse a separate hashtable and separate structures for the directory index\nso that each directory entry has its own next pointer. Use reference\ncounting to track which directory entry contains files.\n\nThere are only slight changes to the name-hash.c API:\n- new free_name_hash() used by read_cache.c::discard_index()\n- remove_name_hash() takes an additional index_state parameter\n- index_name_exists() for a directory (trailing '/') may return a cache\n  entry that has been removed (CE_UNHASHED). This is not a problem as the\n  return value is only used to check if the directory exists (dir.c) or to\n  normalize casing of directory names (read-cache.c).\n\nGetting rid of cache_entry.dir_next reduces memory consumption, especially\nwith core.ignorecase=false (which doesn't use that member at all).\n\nWith core.ignorecase=true, building the directory index is slightly faster\nas we add / check the parent directory first (instead of going through all\ndirectory levels for each file in the index). E.g. with WebKit (~200k\nfiles, ~7k dirs), time spent in lazy_init_name_hash is reduced from 176ms\nto 130ms.\n\nSigned-off-by: Karsten Blees <blees@dcon.de>\n---\nAlso available here:\nhttps://github.com/kblees/git/tree/kb/name-hash-fix-endless-loop-v2\ngit pull git://github.com/kblees/git.git kb/name-hash-fix-endless-loop-v2\n\n cache.h                        |  17 +---\n name-hash.c                    | 182 +++++++++++++++++++++++++++++++----------\n read-cache.c                   |   9 +-\n t/t7062-wtstatus-ignorecase.sh |  20 +++++\n 4 files changed, 166 insertions(+), 62 deletions(-)\n create mode 100755 t/t7062-wtstatus-ignorecase.sh\n\ndiff --git a/cache.h b/cache.h\nindex e493563..898e346 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -131,7 +131,6 @@ struct cache_entry {\n \tunsigned int ce_namelen;\n \tunsigned char sha1[20];\n \tstruct cache_entry *next;\n-\tstruct cache_entry *dir_next;\n \tchar name[FLEX_ARRAY]; /* more */\n };\n \n@@ -267,25 +266,15 @@ struct index_state {\n \tunsigned name_hash_initialized : 1,\n \t\t initialized : 1;\n \tstruct hash_table name_hash;\n+\tstruct hash_table dir_hash;\n };\n \n extern struct index_state the_index;\n \n /* Name hashing */\n extern void add_name_hash(struct index_state *istate, struct cache_entry *ce);\n-/*\n- * We don't actually *remove* it, we can just mark it invalid so that\n- * we won't find it in lookups.\n- *\n- * Not only would we have to search the lists (simple enough), but\n- * we'd also have to rehash other hash buckets in case this makes the\n- * hash bucket empty (common). So it's much better to just mark\n- * it.\n- */\n-static inline void remove_name_hash(struct cache_entry *ce)\n-{\n-\tce->ce_flags |= CE_UNHASHED;\n-}\n+extern void remove_name_hash(struct index_state *istate, struct cache_entry *ce);\n+extern void free_name_hash(struct index_state *istate);\n \n \n #ifndef NO_THE_INDEX_COMPATIBILITY_MACROS\ndiff --git a/name-hash.c b/name-hash.c\nindex 942c459..6d7e198 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -32,38 +32,96 @@ static unsigned int hash_name(const char *name, int namelen)\n \treturn hash;\n }\n \n-static void hash_index_entry_directories(struct index_state *istate, struct cache_entry *ce)\n+struct dir_entry {\n+\tstruct dir_entry *next;\n+\tstruct dir_entry *parent;\n+\tstruct cache_entry *ce;\n+\tint nr;\n+\tunsigned int namelen;\n+};\n+\n+static struct dir_entry *find_dir_entry(struct index_state *istate,\n+\t\tconst char *name, unsigned int namelen)\n+{\n+\tunsigned int hash = hash_name(name, namelen);\n+\tstruct dir_entry *dir;\n+\n+\tfor (dir = lookup_hash(hash, &istate->dir_hash); dir; dir = dir->next)\n+\t\tif (dir->namelen == namelen &&\n+\t\t    !strncasecmp(dir->ce->name, name, namelen))\n+\t\t\treturn dir;\n+\treturn NULL;\n+}\n+\n+static struct dir_entry *hash_dir_entry(struct index_state *istate,\n+\t\tstruct cache_entry *ce, int namelen)\n {\n \t/*\n \t * Throw each directory component in the hash for quick lookup\n \t * during a git status. Directory components are stored with their\n \t * closing slash.  Despite submodules being a directory, they never\n \t * reach this point, because they are stored without a closing slash\n-\t * in the cache.\n+\t * in index_state.name_hash (as ordinary cache_entries).\n \t *\n-\t * Note that the cache_entry stored with the directory does not\n-\t * represent the directory itself.  It is a pointer to an existing\n-\t * filename, and its only purpose is to represent existence of the\n-\t * directory in the cache.  It is very possible multiple directory\n-\t * hash entries may point to the same cache_entry.\n+\t * Note that the cache_entry stored with the dir_entry merely\n+\t * supplies the name of the directory (up to dir_entry.namelen). We\n+\t * track the number of 'active' files in a directory in dir_entry.nr,\n+\t * so we can tell if the directory is still relevant, e.g. for git\n+\t * status. However, if cache_entries are removed, we cannot pinpoint\n+\t * an exact cache_entry that's still active. It is very possible that\n+\t * multiple dir_entries point to the same cache_entry.\n \t */\n-\tunsigned int hash;\n-\tvoid **pos;\n+\tstruct dir_entry *dir;\n+\n+\t/* get length of parent directory */\n+\twhile (namelen > 0 && !is_dir_sep(ce->name[namelen - 1]))\n+\t\tnamelen--;\n+\tif (namelen <= 0)\n+\t\treturn NULL;\n+\n+\t/* lookup existing entry for that directory */\n+\tdir = find_dir_entry(istate, ce->name, namelen);\n+\tif (!dir) {\n+\t\t/* not found, create it and add to hash table */\n+\t\tvoid **pdir;\n+\t\tunsigned int hash = hash_name(ce->name, namelen);\n \n-\tconst char *ptr = ce->name;\n-\twhile (*ptr) {\n-\t\twhile (*ptr && *ptr != '/')\n-\t\t\t++ptr;\n-\t\tif (*ptr == '/') {\n-\t\t\t++ptr;\n-\t\t\thash = hash_name(ce->name, ptr - ce->name);\n-\t\t\tpos = insert_hash(hash, ce, &istate->name_hash);\n-\t\t\tif (pos) {\n-\t\t\t\tce->dir_next = *pos;\n-\t\t\t\t*pos = ce;\n-\t\t\t}\n+\t\tdir = xcalloc(1, sizeof(struct dir_entry));\n+\t\tdir->namelen = namelen;\n+\t\tdir->ce = ce;\n+\n+\t\tpdir = insert_hash(hash, dir, &istate->dir_hash);\n+\t\tif (pdir) {\n+\t\t\tdir->next = *pdir;\n+\t\t\t*pdir = dir;\n \t\t}\n+\n+\t\t/* recursively add missing parent directories */\n+\t\tdir->parent = hash_dir_entry(istate, ce, namelen - 1);\n \t}\n+\treturn dir;\n+}\n+\n+static void add_dir_entry(struct index_state *istate, struct cache_entry *ce)\n+{\n+\t/* Add reference to the directory entry (and parents if 0). */\n+\tstruct dir_entry *dir = hash_dir_entry(istate, ce, ce_namelen(ce));\n+\twhile (dir && !(dir->nr++))\n+\t\tdir = dir->parent;\n+}\n+\n+static void remove_dir_entry(struct index_state *istate, struct cache_entry *ce)\n+{\n+\t/*\n+\t * Release reference to the directory entry (and parents if 0).\n+\t *\n+\t * Note: we do not remove / free the entry because there's no\n+\t * hash.[ch]::remove_hash and dir->next may point to other entries\n+\t * that are still valid, so we must not free the memory.\n+\t */\n+\tstruct dir_entry *dir = hash_dir_entry(istate, ce, ce_namelen(ce));\n+\twhile (dir && dir->nr && !(--dir->nr))\n+\t\tdir = dir->parent;\n }\n \n static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n@@ -74,7 +132,7 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \tif (ce->ce_flags & CE_HASHED)\n \t\treturn;\n \tce->ce_flags |= CE_HASHED;\n-\tce->next = ce->dir_next = NULL;\n+\tce->next = NULL;\n \thash = hash_name(ce->name, ce_namelen(ce));\n \tpos = insert_hash(hash, ce, &istate->name_hash);\n \tif (pos) {\n@@ -82,8 +140,8 @@ static void hash_index_entry(struct index_state *istate, struct cache_entry *ce)\n \t\t*pos = ce;\n \t}\n \n-\tif (ignore_case)\n-\t\thash_index_entry_directories(istate, ce);\n+\tif (ignore_case && !(ce->ce_flags & CE_UNHASHED))\n+\t\tadd_dir_entry(istate, ce);\n }\n \n static void lazy_init_name_hash(struct index_state *istate)\n@@ -99,11 +157,33 @@ static void lazy_init_name_hash(struct index_state *istate)\n \n void add_name_hash(struct index_state *istate, struct cache_entry *ce)\n {\n+\t/* if already hashed, add reference to directory entries */\n+\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_STATE_MASK)\n+\t\tadd_dir_entry(istate, ce);\n+\n \tce->ce_flags &= ~CE_UNHASHED;\n \tif (istate->name_hash_initialized)\n \t\thash_index_entry(istate, ce);\n }\n \n+/*\n+ * We don't actually *remove* it, we can just mark it invalid so that\n+ * we won't find it in lookups.\n+ *\n+ * Not only would we have to search the lists (simple enough), but\n+ * we'd also have to rehash other hash buckets in case this makes the\n+ * hash bucket empty (common). So it's much better to just mark\n+ * it.\n+ */\n+void remove_name_hash(struct index_state *istate, struct cache_entry *ce)\n+{\n+\t/* if already hashed, release reference to directory entries */\n+\tif (ignore_case && (ce->ce_flags & CE_STATE_MASK) == CE_HASHED)\n+\t\tremove_dir_entry(istate, ce);\n+\n+\tce->ce_flags |= CE_UNHASHED;\n+}\n+\n static int slow_same_name(const char *name1, int len1, const char *name2, int len2)\n {\n \tif (len1 != len2)\n@@ -137,18 +217,7 @@ static int same_name(const struct cache_entry *ce, const char *name, int namelen\n \tif (!icase)\n \t\treturn 0;\n \n-\t/*\n-\t * If the entry we're comparing is a filename (no trailing slash), then compare\n-\t * the lengths exactly.\n-\t */\n-\tif (name[namelen - 1] != '/')\n-\t\treturn slow_same_name(name, namelen, ce->name, len);\n-\n-\t/*\n-\t * For a directory, we point to an arbitrary cache_entry filename.  Just\n-\t * make sure the directory portion matches.\n-\t */\n-\treturn slow_same_name(name, namelen, ce->name, namelen < len ? namelen : len);\n+\treturn slow_same_name(name, namelen, ce->name, len);\n }\n \n struct cache_entry *index_name_exists(struct index_state *istate, const char *name, int namelen, int icase)\n@@ -164,27 +233,54 @@ struct cache_entry *index_name_exists(struct index_state *istate, const char *na\n \t\t\tif (same_name(ce, name, namelen, icase))\n \t\t\t\treturn ce;\n \t\t}\n-\t\tif (icase && name[namelen - 1] == '/')\n-\t\t\tce = ce->dir_next;\n-\t\telse\n-\t\t\tce = ce->next;\n+\t\tce = ce->next;\n \t}\n \n \t/*\n-\t * Might be a submodule.  Despite submodules being directories,\n+\t * When looking for a directory (trailing '/'), it might be a\n+\t * submodule or a directory. Despite submodules being directories,\n \t * they are stored in the name hash without a closing slash.\n-\t * When ignore_case is 1, directories are stored in the name hash\n-\t * with their closing slash.\n+\t * When ignore_case is 1, directories are stored in a separate hash\n+\t * table *with* their closing slash.\n \t *\n \t * The side effect of this storage technique is we have need to\n+\t * lookup the directory in a separate hash table, and if not found\n \t * remove the slash from name and perform the lookup again without\n \t * the slash.  If a match is made, S_ISGITLINK(ce->mode) will be\n \t * true.\n \t */\n \tif (icase && name[namelen - 1] == '/') {\n+\t\tstruct dir_entry *dir = find_dir_entry(istate, name, namelen);\n+\t\tif (dir && dir->nr)\n+\t\t\treturn dir->ce;\n+\n \t\tce = index_name_exists(istate, name, namelen - 1, icase);\n \t\tif (ce && S_ISGITLINK(ce->ce_mode))\n \t\t\treturn ce;\n \t}\n \treturn NULL;\n }\n+\n+static int free_dir_entry(void *entry, void *unused)\n+{\n+\tstruct dir_entry *dir = entry;\n+\twhile (dir) {\n+\t\tstruct dir_entry *next = dir->next;\n+\t\tfree(dir);\n+\t\tdir = next;\n+\t}\n+\treturn 0;\n+}\n+\n+void free_name_hash(struct index_state *istate)\n+{\n+\tif (!istate->name_hash_initialized)\n+\t\treturn;\n+\tistate->name_hash_initialized = 0;\n+\tif (ignore_case)\n+\t\t/* free directory entries */\n+\t\tfor_each_hash(&istate->dir_hash, free_dir_entry, NULL);\n+\n+\tfree_hash(&istate->name_hash);\n+\tfree_hash(&istate->dir_hash);\n+}\ndiff --git a/read-cache.c b/read-cache.c\nindex 827ae55..47eb9d8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -46,7 +46,7 @@ static void replace_index_entry(struct index_state *istate, int nr, struct cache\n {\n \tstruct cache_entry *old = istate->cache[nr];\n \n-\tremove_name_hash(old);\n+\tremove_name_hash(istate, old);\n \tset_index_entry(istate, nr, ce);\n \tistate->cache_changed = 1;\n }\n@@ -460,7 +460,7 @@ int remove_index_entry_at(struct index_state *istate, int pos)\n \tstruct cache_entry *ce = istate->cache[pos];\n \n \trecord_resolve_undo(istate, ce);\n-\tremove_name_hash(ce);\n+\tremove_name_hash(istate, ce);\n \tistate->cache_changed = 1;\n \tistate->cache_nr--;\n \tif (pos >= istate->cache_nr)\n@@ -483,7 +483,7 @@ void remove_marked_cache_entries(struct index_state *istate)\n \n \tfor (i = j = 0; i < istate->cache_nr; i++) {\n \t\tif (ce_array[i]->ce_flags & CE_REMOVE)\n-\t\t\tremove_name_hash(ce_array[i]);\n+\t\t\tremove_name_hash(istate, ce_array[i]);\n \t\telse\n \t\t\tce_array[j++] = ce_array[i];\n \t}\n@@ -1515,8 +1515,7 @@ int discard_index(struct index_state *istate)\n \tistate->cache_changed = 0;\n \tistate->timestamp.sec = 0;\n \tistate->timestamp.nsec = 0;\n-\tistate->name_hash_initialized = 0;\n-\tfree_hash(&istate->name_hash);\n+\tfree_name_hash(istate);\n \tcache_tree_free(&(istate->cache_tree));\n \tistate->initialized = 0;\n \ndiff --git a/t/t7062-wtstatus-ignorecase.sh b/t/t7062-wtstatus-ignorecase.sh\nnew file mode 100755\nindex 0000000..73709db\n--- /dev/null\n+++ b/t/t7062-wtstatus-ignorecase.sh\n@@ -0,0 +1,20 @@\n+#!/bin/sh\n+\n+test_description='git-status with core.ignorecase=true'\n+\n+. ./test-lib.sh\n+\n+test_expect_success 'status with hash collisions' '\n+\t# note: \"V/\", \"V/XQANY/\" and \"WURZAUP/\" produce the same hash code\n+\t# in name-hash.c::hash_name\n+\tmkdir V &&\n+\tmkdir V/XQANY &&\n+\tmkdir WURZAUP &&\n+\ttouch V/XQANY/test &&\n+\tgit config core.ignorecase true &&\n+\tgit add . &&\n+\t# test is successful if git status completes (no endless loop)\n+\tgit status\n+'\n+\n+test_done\n-- \n1.8.1.2.7987.g4a34b82\n"},{"id":"210418","messageId":"7vsj4hi8pf.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"512E9D7C.2030803@gmail.com","subject":"Re: [PATCH v2] name-hash.c: fix endless loop with core.ignorecase=true","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-28T00:27:24Z","receivedAt":"2013-02-28T00:27:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Karsten Blees <karsten.blees@gmail.com> writes:\n\n> With core.ignorecase=true, name-hash.c builds a case insensitive index of\n> all tracked directories. Currently, the existing cache entry structures are\n> added multiple times to the same hashtable (with different name lengths and\n> hash codes). However, there's only one dir_next pointer, which gets\n> completely messed up in case of hash collisions. In the worst case, this\n> causes an endless loop if ce == ce->dir_next (see t7062).\n>\n> Use a separate hashtable and separate structures for the directory index\n> so that each directory entry has its own next pointer. Use reference\n> counting to track which directory entry contains files.\n>\n> There are only slight changes to the name-hash.c API:\n> - new free_name_hash() used by read_cache.c::discard_index()\n> - remove_name_hash() takes an additional index_state parameter\n> - index_name_exists() for a directory (trailing '/') may return a cache\n>   entry that has been removed (CE_UNHASHED). This is not a problem as the\n>   return value is only used to check if the directory exists (dir.c) or to\n>   normalize casing of directory names (read-cache.c).\n>\n> Getting rid of cache_entry.dir_next reduces memory consumption, especially\n> with core.ignorecase=false (which doesn't use that member at all).\n>\n> With core.ignorecase=true, building the directory index is slightly faster\n> as we add / check the parent directory first (instead of going through all\n> directory levels for each file in the index). E.g. with WebKit (~200k\n> files, ~7k dirs), time spent in lazy_init_name_hash is reduced from 176ms\n> to 130ms.\n>\n> Signed-off-by: Karsten Blees <blees@dcon.de>\n> ---\n\nOne thing that still puzzles me is what guarantee we have on the\nliftime of these ce's that are borrowed by these dir_hash entries.\nThere are a few places where we call free(ce) around \"aliased\"\nentries (only happens with ignore_case set).  I do not think it is a\nnew issue (we used to borrow a ce to represent a directory in the\nname_hash by using the leading prefix of its name anyway, and this\npatch only changes which hash table is used to hold it), and I do\nnot think it will be an issue for case sensitive systems, so I would\nstop being worried about it for now, though ;-)\n\nThanks, will replace and queue.\n\n\n> diff --git a/t/t7062-wtstatus-ignorecase.sh b/t/t7062-wtstatus-ignorecase.sh\n> new file mode 100755\n> index 0000000..73709db\n> --- /dev/null\n> +++ b/t/t7062-wtstatus-ignorecase.sh\n> @@ -0,0 +1,20 @@\n> +#!/bin/sh\n> +\n> +test_description='git-status with core.ignorecase=true'\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'status with hash collisions' '\n> +\t# note: \"V/\", \"V/XQANY/\" and \"WURZAUP/\" produce the same hash code\n> +\t# in name-hash.c::hash_name\n> +\tmkdir V &&\n> +\tmkdir V/XQANY &&\n> +\tmkdir WURZAUP &&\n> +\ttouch V/XQANY/test &&\n> +\tgit config core.ignorecase true &&\n> +\tgit add . &&\n> +\t# test is successful if git status completes (no endless loop)\n> +\tgit status\n> +'\n> +\n> +test_done\n"},{"id":"210793","messageId":"513911B3.7010903@web.de","threadId":"32862","inReplyTo":"CACsJy8DnvAjQPL4aP_LRC7aqx6OC4M5dMtj-OUot76qET2z08Q@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-07T22:16:19Z","receivedAt":"2013-03-07T22:16:19Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 11.02.13 03:56, Duy Nguyen wrote:\n> On Mon, Feb 11, 2013 at 3:16 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>> The other \"lstat()\" experiment was a very interesting one, but this\n>> is not yet an interesting experiment to see where in the \"ignore\"\n>> codepath we are spending times.\n>>\n>> We know that we can tell wt_status_collect_untracked() not to bother\n>> with the untracked or ignored files with !s->show_untracked_files\n>> already, but I think the more interesting question is if we can show\n>> the untracked files with less overhead.\n>>\n>> If we want to show untrackedd files, it is a given that we need to\n>> read directories to see what paths there are on the filesystem. Is\n>> the opendir/readdir cost dominating in the process? Are we spending\n>> a lot of time sifting the result of opendir/readdir via the ignore\n>> mechanism? Is reading the \"ignore\" files costing us much to prime\n>> the ignore mechanism?\n>>\n>> If readdir cost is dominant, then that makes \"cache gitignore\" a\n>> nonsense proposition, I think.  If you really want to \"cache\"\n>> something, you need to have somebody (i.e. a daemon) who constantly\n>> keeps an eye on the filesystem changes and can respond with the up\n>> to date result directly to fill_directory().  I somehow doubt that\n>> it is a direction we would want to go in, though.\n> \n> Yeah, it did not cut out syscall cost, I also cut a lot of user-space\n> processing (plus .gitignore content access). From the timings I posted\n> earlier,\n> \n>>         unmodified  dir.c\n>> real    0m0.550s    0m0.287s\n>> user    0m0.305s    0m0.201s\n>> sys     0m0.240s    0m0.084s\n> \n> sys time is reduced from 0.24s to 0.08s, so readdir+opendir definitely\n> has something to do with it (and perhaps reading .gitignore). But it\n> also reduces user time from 0.305 to 0.201s. I don't think avoiding\n> readdir+openddir will bring us this gain. It's probably the cost of\n> matching .gitignore. I'll try to replace opendir+readdir with a\n> no-syscall version. At this point \"untracked caching\" sounds more\n> feasible (and less complex) than \".gitignore cachine\".\n> \nThanks for Duy for the measurements, and patches.\nI took the freedom to convert the patched dir.c into a \n\"runtime configurable\" git status option.\nI'm not sure if the following copy-and-paste work applies,\n(it is based on Git 1.8.1.3), but the time spend for \n\"git status --changed-only\" is basically half the time of\n\"git status\", similar to what Duy has measured.\nI did a test both on a Linux box and Mac OS.\n\nAnd the speedup is so impressive, that I am tempted to submit a patch simlar\nto the following, what do we think about it?\n/Torsten\n\n\n\n\n-- >8 --\n\n[PATCH] git status: add option changed-only\ngit status may be run faster if\n- we only check if files are changed which are already known to git.\n- we don't check if there are untracked files.\n\n\"git status --changed-only\" (or the short form \"git status -c\")\n\nwill only check for changed files which are already known to git,\nand which are in the index.\n\nThe call to read_directory_recursive() is skipped and untracked files\nin the working tree are not reported.\n\nInspired-by: Duy Nguyen <pclouds@gmail.com>\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n builtin/commit.c | 2 ++\n dir.c            | 5 +++--\n dir.h            | 3 ++-\n wt-status.c      | 3 +++\n wt-status.h      | 1 +\n 5 files changed, 11 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex d6dd3df..6a5ba11 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -1158,6 +1158,8 @@ int cmd_status(int argc, const char **argv, const char *prefix)\n \tunsigned char sha1[20];\n \tstatic struct option builtin_status_options[] = {\n \t\tOPT__VERBOSE(&verbose, N_(\"be verbose\")),\n+\t\tOPT_BOOLEAN('c', \"changed-only\", &s.check_changed_only,\n+\t\t\t    N_(\"Ignore untracked files. Check if files known to git are modified\")),\n \t\tOPT_SET_INT('s', \"short\", &status_format,\n \t\t\t    N_(\"show status concisely\"), STATUS_FORMAT_SHORT),\n \t\tOPT_BOOLEAN('b', \"branch\", &s.show_branch,\ndiff --git a/dir.c b/dir.c\nindex a473ca2..555b652 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -1274,8 +1274,9 @@ int read_directory(struct dir_struct *dir, const char *path, int len, const char\n \t\treturn dir->nr;\n \n \tsimplify = create_simplify(pathspec);\n-\tif (!len || treat_leading_path(dir, path, len, simplify))\n-\t\tread_directory_recursive(dir, path, len, 0, simplify);\n+\tif ((!(dir->flags & DIR_CHECK_CHANGED_ONLY)) &&\n+\t\t\t(!len || treat_leading_path(dir, path, len, simplify))) o\n+\t\t\tread_directory_recursive(dir, path, len, 0, simplify);\n \tfree_simplify(simplify);\n \tqsort(dir->entries, dir->nr, sizeof(struct dir_entry *), cmp_name);\n \tqsort(dir->ignored, dir->ignored_nr, sizeof(struct dir_entry *), cmp_name);\ndiff --git a/dir.h b/dir.h\nindex f5c89e3..1a915a7 100644\n--- a/dir.h\n+++ b/dir.h\n@@ -41,7 +41,8 @@ struct dir_struct {\n \t\tDIR_SHOW_OTHER_DIRECTORIES = 1<<1,\n \t\tDIR_HIDE_EMPTY_DIRECTORIES = 1<<2,\n \t\tDIR_NO_GITLINKS = 1<<3,\n-\t\tDIR_COLLECT_IGNORED = 1<<4\n+\t\tDIR_COLLECT_IGNORED = 1<<4,\n+\t\tDIR_CHECK_CHANGED_ONLY = 1<<5\n \t} flags;\n \tstruct dir_entry **entries;\n \tstruct dir_entry **ignored;\ndiff --git a/wt-status.c b/wt-status.c\nindex d7cfe8f..b315785 100644\n--- a/wt-status.c\n+++ b/wt-status.c\n@@ -503,6 +503,9 @@ static void wt_status_collect_untracked(struct wt_status *s)\n \tif (s->show_untracked_files != SHOW_ALL_UNTRACKED_FILES)\n \t\tdir.flags |=\n \t\t\tDIR_SHOW_OTHER_DIRECTORIES | DIR_HIDE_EMPTY_DIRECTORIES;\n+\tif (s->check_changed_only)\n+\t\tdir.flags |= DIR_CHECK_CHANGED_ONLY;\n+\n \tsetup_standard_excludes(&dir);\n \n \tfill_directory(&dir, s->pathspec);\ndiff --git a/wt-status.h b/wt-status.h\nindex 236b41f..7eb0115 100644\n--- a/wt-status.h\n+++ b/wt-status.h\n@@ -47,6 +47,7 @@ struct wt_status {\n \tconst char **pathspec;\n \tint verbose;\n \tint amend;\n+\tint check_changed_only;\n \tenum commit_whence whence;\n \tint nowarn;\n \tint use_color;\n-- \n1.8.2.rc2\n"},{"id":"210804","messageId":"7vr4jqkb9g.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"513911B3.7010903@web.de","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-08T00:04:11Z","receivedAt":"2013-03-08T00:04:11Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> diff --git a/builtin/commit.c b/builtin/commit.c\n> index d6dd3df..6a5ba11 100644\n> --- a/builtin/commit.c\n> +++ b/builtin/commit.c\n> @@ -1158,6 +1158,8 @@ int cmd_status(int argc, const char **argv, const char *prefix)\n>  \tunsigned char sha1[20];\n>  \tstatic struct option builtin_status_options[] = {\n>  \t\tOPT__VERBOSE(&verbose, N_(\"be verbose\")),\n> +\t\tOPT_BOOLEAN('c', \"changed-only\", &s.check_changed_only,\n> +\t\t\t    N_(\"Ignore untracked files. Check if files known to git are modified\")),\n\nDoesn't this make one wonder why a separate bit and implementation\nis necessary to say \"I am not interested in untracked files\" when\n\"-uno\" option is already there?\n"},{"id":"210812","messageId":"51398CD5.1070603@web.de","threadId":"32862","inReplyTo":"7vr4jqkb9g.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-08T07:01:41Z","receivedAt":"2013-03-08T07:01:41Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 08.03.13 01:04, Junio C Hamano wrote:\n> Torsten Bögershausen <tboegi@web.de> writes:\n> \n>> diff --git a/builtin/commit.c b/builtin/commit.c\n>> index d6dd3df..6a5ba11 100644\n>> --- a/builtin/commit.c\n>> +++ b/builtin/commit.c\n>> @@ -1158,6 +1158,8 @@ int cmd_status(int argc, const char **argv, const char *prefix)\n>>  \tunsigned char sha1[20];\n>>  \tstatic struct option builtin_status_options[] = {\n>>  \t\tOPT__VERBOSE(&verbose, N_(\"be verbose\")),\n>> +\t\tOPT_BOOLEAN('c', \"changed-only\", &s.check_changed_only,\n>> +\t\t\t    N_(\"Ignore untracked files. Check if files known to git are modified\")),\n> \n> Doesn't this make one wonder why a separate bit and implementation\n> is necessary to say \"I am not interested in untracked files\" when\n> \"-uno\" option is already there?\nThanks Junio,\nthis is good news.\nI need to admit that I wasn't aware about \"git status -uno\".\n\nThinking about it, how many git users are aware of the speed penalty\nwhen running git status to find out which (tracked) files they had changed?\n\nOr to put it the other way, when a developer wants a quick overview\nabout the files she changed, then git status -uno may be a good and fast friend.\n\nDoes it make sence to stress put that someway in the documentation?\n\ndiff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\nindex 9f1ef9a..360d813 100644\n--- a/Documentation/git-status.txt\n+++ b/Documentation/git-status.txt\n@@ -51,13 +51,18 @@ default is 'normal', i.e. show untracked files and directori\n +\n The possible options are:\n +\n-       - 'no'     - Show no untracked files\n+       - 'no'     - Show no untracked files (this is fastest)\n        - 'normal' - Shows untracked files and directories\n        - 'all'    - Also shows individual files in untracked directories.\n +\n The default can be changed using the status.showUntrackedFiles\n configuration variable documented in linkgit:git-config[1].\n \n++\n+Note: Searching for untracked files or directories may take some time.\n+A fast way to get a status of files tracked by git is to use\n+'git status -uno'\n+\n\n\n\n\n\n\n\n\n\n\n\n\n> \n> \n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n"},{"id":"210816","messageId":"7v7glijoiy.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"51398CD5.1070603@web.de","subject":"Re: inotify to minimize stat() calls","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-08T08:15:17Z","receivedAt":"2013-03-08T08:15:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n>> Doesn't this make one wonder why a separate bit and implementation\n>> is necessary to say \"I am not interested in untracked files\" when\n>> \"-uno\" option is already there?\n> ...\n> I need to admit that I wasn't aware about \"git status -uno\".\n\nNot so fast.  I did not ask you \"Why do you need a new one to solve\nthe same problem -uno already solves?\"\n\n> Thinking about it, how many git users are aware of the speed penalty\n> when running git status to find out which (tracked) files they had changed?\n>\n> Or to put it the other way, when a developer wants a quick overview\n> about the files she changed, then git status -uno may be a good and fast friend.\n>\n> Does it make sence to stress put that someway in the documentation?\n>\n> diff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\n> index 9f1ef9a..360d813 100644\n> --- a/Documentation/git-status.txt\n> +++ b/Documentation/git-status.txt\n> @@ -51,13 +51,18 @@ default is 'normal', i.e. show untracked files and directori\n>  +\n>  The possible options are:\n>  +\n> -       - 'no'     - Show no untracked files\n> +       - 'no'     - Show no untracked files (this is fastest)\n\nThere is a trade-off around the use of -uno between safety and\nperformance.  The default is not to use -uno so that you will not\nforget to add a file you newly created (i.e safety).  You would pay\nfor the safety with the cost to find such untracked files (i.e.\nperformance).\n\nI suspect that the documentation was written with the assumption\nthat at least for the people who are reading this part of the\ndocumentation, the trade-off is obvious.  In order to find more\ninformation, you naturally need to spend more cycles.\n\nIf the trade-off is not so obvious, however, I do not object at all\nto describing it. But if we are to do so, I do object to mentioning\nonly one side of the trade-off.  People who choose \"fastest\" needs\nto be made very aware that they are disabling \"safety\".\n\nThat brings us back to the \"Why a separate implementation when -uno\nis there?\" question.\n\nYour patch adds new code; it does not just enables the same logic as\nwhat the existing -uno does with a new flag.  Does the new compute\ndifferent things?  Does it find more stuff by spending extra cycles?\nDoes it find less stuff by being extra faster?\n\nThese questions are important.\n\nIf the new option strikes the trade-off between safety and\nperformance at a point different from the point where the existing\n-uno option does, it _might_ still be worth adding as a separate\noption.  I didn't get that impression when I saw the patch, but I\nadmit that I did not follow the code carefully myself.\n\nThat is the reason why I was wondering why a separate bit and\nimplementation had to be added by the patch.\n"},{"id":"210819","messageId":"5139AE30.6010200@web.de","threadId":"32862","inReplyTo":"7v7glijoiy.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-08T09:24:00Z","receivedAt":"2013-03-08T09:24:00Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 08.03.13 09:15, Junio C Hamano wrote:\n> Torsten Bögershausen <tboegi@web.de> writes:\n> \n>>> Doesn't this make one wonder why a separate bit and implementation\n>>> is necessary to say \"I am not interested in untracked files\" when\n>>> \"-uno\" option is already there?\n>> ...\n>> I need to admit that I wasn't aware about \"git status -uno\".\n> \n> Not so fast.  I did not ask you \"Why do you need a new one to solve\n> the same problem -uno already solves?\"\n> \n>> Thinking about it, how many git users are aware of the speed penalty\n>> when running git status to find out which (tracked) files they had changed?\n>>\n>> Or to put it the other way, when a developer wants a quick overview\n>> about the files she changed, then git status -uno may be a good and fast friend.\n>>\n>> Does it make sence to stress put that someway in the documentation?\n>>\n>> diff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\n>> index 9f1ef9a..360d813 100644\n>> --- a/Documentation/git-status.txt\n>> +++ b/Documentation/git-status.txt\n>> @@ -51,13 +51,18 @@ default is 'normal', i.e. show untracked files and directori\n>>  +\n>>  The possible options are:\n>>  +\n>> -       - 'no'     - Show no untracked files\n>> +       - 'no'     - Show no untracked files (this is fastest)\n> \n> There is a trade-off around the use of -uno between safety and\n> performance.  The default is not to use -uno so that you will not\n> forget to add a file you newly created (i.e safety).  You would pay\n> for the safety with the cost to find such untracked files (i.e.\n> performance).\n> \n> I suspect that the documentation was written with the assumption\n> that at least for the people who are reading this part of the\n> documentation, the trade-off is obvious.  In order to find more\n> information, you naturally need to spend more cycles.\n> \n> If the trade-off is not so obvious, however, I do not object at all\n> to describing it. But if we are to do so, I do object to mentioning\n> only one side of the trade-off.  People who choose \"fastest\" needs\n> to be made very aware that they are disabling \"safety\".\n> \n> That brings us back to the \"Why a separate implementation when -uno\n> is there?\" question.\n[...]\nThe short version:\nThe -uno option does exactly what the -c option intended to do ;-)\n(The code path to disable the \"expensive\" call to read_directory_recursive()\nin dir.c is slightly different).\nMaking benchmarks (again, sorry for the noise) shows that -uno and -c are equally fast,\nmaking 5 git status on a linux tree, take the best of 5:\n\ngit status\nreal    0m0.697s\n\ngit status -uno\nreal    0m0.291s\n\n(with the patch) git status -c\nreal    0m0.289s\n\n\nThese are not really scientific numbers, but all in all we have motivation enough to drop\nthe \"git status -c\" patch completely.\n\nMy feeling is still that the suggested documentation \"this is fastest\" is not a good choice either.\nLet me try to come up with a better suggestion.\n/Torsten\n \n"},{"id":"210829","messageId":"CACsJy8DZm153Tu_3GTOnxF8bFrYPh7_DP6Rn6rr3n6tfuVuv2Q@mail.gmail.com","threadId":"32862","inReplyTo":"7v7glijoiy.fsf@alter.siamese.dyndns.org","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-03-08T10:53:56Z","receivedAt":"2013-03-08T10:53:56Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Mar 8, 2013 at 3:15 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>>  The possible options are:\n>>  +\n>> -       - 'no'     - Show no untracked files\n>> +       - 'no'     - Show no untracked files (this is fastest)\n>\n> There is a trade-off around the use of -uno between safety and\n> performance.  The default is not to use -uno so that you will not\n> forget to add a file you newly created (i.e safety).  You would pay\n> for the safety with the cost to find such untracked files (i.e.\n> performance).\n>\n> I suspect that the documentation was written with the assumption\n> that at least for the people who are reading this part of the\n> documentation, the trade-off is obvious.  In order to find more\n> information, you naturally need to spend more cycles.\n>\n> If the trade-off is not so obvious, however, I do not object at all\n> to describing it. But if we are to do so, I do object to mentioning\n> only one side of the trade-off.  People who choose \"fastest\" needs\n> to be made very aware that they are disabling \"safety\".\n\nOn the topic of trading off, I was thinking about new -uauto as\ndefault that is like -uall if it takes less than a certan amount of\ntime (e.g. 0.5 seconds), if it exceeds that limit, the operation is\naborted (i.e. it turns to -uno). The safety net is still there, \"git\nstatus\" advices to use -u to show full information.\n\nOr a less intrusive approach: measure the time and advice the user to\n(read doc and) use -uno.\n\nBut it's probably worth waiting for the first cut of inotify support\nfrom Ram. It's better with inotify anyway.\n-- \nDuy\n"},{"id":"210961","messageId":"CALkWK0n_9wy575U5os287J+dnc_LqrJ0JwP3ur6kA+4S_=yMag@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8DZm153Tu_3GTOnxF8bFrYPh7_DP6Rn6rr3n6tfuVuv2Q@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-03-10T08:23:27Z","receivedAt":"2013-03-10T08:23:27Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Duy Nguyen wrote:\n> On Fri, Mar 8, 2013 at 3:15 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>>>  The possible options are:\n>>>  +\n>>> -       - 'no'     - Show no untracked files\n>>> +       - 'no'     - Show no untracked files (this is fastest)\n>>\n>> There is a trade-off around the use of -uno between safety and\n>> performance.  The default is not to use -uno so that you will not\n>> forget to add a file you newly created (i.e safety).  You would pay\n>> for the safety with the cost to find such untracked files (i.e.\n>> performance).\n>>\n>> I suspect that the documentation was written with the assumption\n>> that at least for the people who are reading this part of the\n>> documentation, the trade-off is obvious.  In order to find more\n>> information, you naturally need to spend more cycles.\n>>\n>> If the trade-off is not so obvious, however, I do not object at all\n>> to describing it. But if we are to do so, I do object to mentioning\n>> only one side of the trade-off.  People who choose \"fastest\" needs\n>> to be made very aware that they are disabling \"safety\".\n>\n> On the topic of trading off, I was thinking about new -uauto as\n> default that is like -uall if it takes less than a certan amount of\n> time (e.g. 0.5 seconds), if it exceeds that limit, the operation is\n> aborted (i.e. it turns to -uno). The safety net is still there, \"git\n> status\" advices to use -u to show full information.\n\nUgh, this is too opaque; the user has no idea whether untracked files\nare being counted or not.\n\n> Or a less intrusive approach: measure the time and advice the user to\n> (read doc and) use -uno.\n\nI just learnt about -uno myself, from this thread.  At best, it's a\nstopgap until we get inotify support.\n\n> But it's probably worth waiting for the first cut of inotify support\n> from Ram. It's better with inotify anyway.\n\nThis is quite urgent in my opinion.  One of git's primary tasks is to\nquickly tell me what changed in the repository, and inotify is the\nperfect way to do this.\nI'll try to get the first cut out quickly, so we can immediately\ncorrect any fundamental design flaws.\n"},{"id":"211213","messageId":"1363179556-4144-1-git-send-email-pclouds@gmail.com","threadId":"32862","inReplyTo":"CACsJy8DZm153Tu_3GTOnxF8bFrYPh7_DP6Rn6rr3n6tfuVuv2Q@mail.gmail.com","subject":"[PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2013-03-13T12:59:16Z","receivedAt":"2013-03-13T12:59:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This patch attempts to advertise -uno to the users who tolerate slow\n\"git status\" on large repositories (or slow machines/disks). The 2\nseconds limit is quite arbitrary but is probably long enough to start\nusing -uno.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/config.txt |  4 ++++\n advice.c                 |  2 ++\n advice.h                 |  1 +\n t/t7060-wtstatus.sh      |  2 ++\n t/t7508-status.sh        |  4 ++++\n t/t7512-status-help.sh   |  1 +\n wt-status.c              | 20 +++++++++++++++++++-\n wt-status.h              |  1 +\n 8 files changed, 34 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex bbba728..e91d06f 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -178,6 +178,10 @@ advice.*::\n \t\tthe template shown when writing commit messages in\n \t\tlinkgit:git-commit[1], and in the help message shown\n \t\tby linkgit:git-checkout[1] when switching branch.\n+\tstatusUno::\n+\t\tIf collecting untracked files in linkgit:git-status[1]\n+\t\ttakes more than 2 seconds, hint the user that the option\n+\t\t`-uno` could be used to stop collecting untracked files.\n \tcommitBeforeMerge::\n \t\tAdvice shown when linkgit:git-merge[1] refuses to\n \t\tmerge to avoid overwriting local changes.\ndiff --git a/advice.c b/advice.c\nindex 780f58d..72b5c66 100644\n--- a/advice.c\n+++ b/advice.c\n@@ -8,6 +8,7 @@ int advice_push_already_exists = 1;\n int advice_push_fetch_first = 1;\n int advice_push_needs_force = 1;\n int advice_status_hints = 1;\n+int advice_status_uno = 1;\n int advice_commit_before_merge = 1;\n int advice_resolve_conflict = 1;\n int advice_implicit_identity = 1;\n@@ -25,6 +26,7 @@ static struct {\n \t{ \"pushfetchfirst\", &advice_push_fetch_first },\n \t{ \"pushneedsforce\", &advice_push_needs_force },\n \t{ \"statushints\", &advice_status_hints },\n+\t{ \"statusuno\", &advice_status_uno },\n \t{ \"commitbeforemerge\", &advice_commit_before_merge },\n \t{ \"resolveconflict\", &advice_resolve_conflict },\n \t{ \"implicitidentity\", &advice_implicit_identity },\ndiff --git a/advice.h b/advice.h\nindex fad36df..d7e03be 100644\n--- a/advice.h\n+++ b/advice.h\n@@ -11,6 +11,7 @@ extern int advice_push_already_exists;\n extern int advice_push_fetch_first;\n extern int advice_push_needs_force;\n extern int advice_status_hints;\n+extern int advice_status_uno;\n extern int advice_commit_before_merge;\n extern int advice_resolve_conflict;\n extern int advice_implicit_identity;\ndiff --git a/t/t7060-wtstatus.sh b/t/t7060-wtstatus.sh\nindex f4f38a5..dd340d5 100755\n--- a/t/t7060-wtstatus.sh\n+++ b/t/t7060-wtstatus.sh\n@@ -5,6 +5,7 @@ test_description='basic work tree status reporting'\n . ./test-lib.sh\n \n test_expect_success setup '\n+\tgit config advice.statusuno false &&\n \ttest_commit A &&\n \ttest_commit B oneside added &&\n \tgit checkout A^0 &&\n@@ -46,6 +47,7 @@ test_expect_success 'M/D conflict does not segfault' '\n \t(\n \t\tcd mdconflict &&\n \t\tgit init &&\n+\t\tgit config advice.statusuno false\n \t\ttest_commit initial foo \"\" &&\n \t\ttest_commit modify foo foo &&\n \t\tgit checkout -b side HEAD^ &&\ndiff --git a/t/t7508-status.sh b/t/t7508-status.sh\nindex a79c032..9d6e4db 100755\n--- a/t/t7508-status.sh\n+++ b/t/t7508-status.sh\n@@ -8,11 +8,13 @@ test_description='git status'\n . ./test-lib.sh\n \n test_expect_success 'status -h in broken repository' '\n+\tgit config advice.statusuno false &&\n \tmkdir broken &&\n \ttest_when_finished \"rm -fr broken\" &&\n \t(\n \t\tcd broken &&\n \t\tgit init &&\n+\t\tgit config advice.statusuno false &&\n \t\techo \"[status] showuntrackedfiles = CORRUPT\" >>.git/config &&\n \t\ttest_expect_code 129 git status -h >usage 2>&1\n \t) &&\n@@ -25,6 +27,7 @@ test_expect_success 'commit -h in broken repository' '\n \t(\n \t\tcd broken &&\n \t\tgit init &&\n+\t\tgit config advice.statusuno false &&\n \t\techo \"[status] showuntrackedfiles = CORRUPT\" >>.git/config &&\n \t\ttest_expect_code 129 git commit -h >usage 2>&1\n \t) &&\n@@ -780,6 +783,7 @@ test_expect_success 'status refreshes the index' '\n test_expect_success 'setup status submodule summary' '\n \ttest_create_repo sm && (\n \t\tcd sm &&\n+\t\tgit config advice.statusuno false &&\n \t\t>foo &&\n \t\tgit add foo &&\n \t\tgit commit -m \"Add foo\"\ndiff --git a/t/t7512-status-help.sh b/t/t7512-status-help.sh\nindex d2da89a..033a1b3 100755\n--- a/t/t7512-status-help.sh\n+++ b/t/t7512-status-help.sh\n@@ -14,6 +14,7 @@ test_description='git status advice'\n set_fake_editor\n \n test_expect_success 'prepare for conflicts' '\n+\tgit config advice.statusuno false &&\n \ttest_commit init main.txt init &&\n \tgit branch conflicts &&\n \ttest_commit on_master main.txt on_master &&\ndiff --git a/wt-status.c b/wt-status.c\nindex ef405d0..6fde08b 100644\n--- a/wt-status.c\n+++ b/wt-status.c\n@@ -540,7 +540,16 @@ void wt_status_collect(struct wt_status *s)\n \t\twt_status_collect_changes_initial(s);\n \telse\n \t\twt_status_collect_changes_index(s);\n-\twt_status_collect_untracked(s);\n+\tif (s->show_untracked_files && advice_status_uno) {\n+\t\tstruct timeval tv1, tv2;\n+\t\tgettimeofday(&tv1, NULL);\n+\t\twt_status_collect_untracked(s);\n+\t\tgettimeofday(&tv2, NULL);\n+\t\ts->untracked_in_ms =\n+\t\t\t(uint64_t)tv2.tv_sec * 1000 + tv2.tv_usec / 1000 -\n+\t\t\t((uint64_t)tv1.tv_sec * 1000 + tv1.tv_usec / 1000);\n+\t} else\n+\t\twt_status_collect_untracked(s);\n }\n \n static void wt_status_print_unmerged(struct wt_status *s)\n@@ -1097,6 +1106,15 @@ void wt_status_print(struct wt_status *s)\n \t\twt_status_print_other(s, &s->untracked, _(\"Untracked files\"), \"add\");\n \t\tif (s->show_ignored_files)\n \t\t\twt_status_print_other(s, &s->ignored, _(\"Ignored files\"), \"add -f\");\n+\t\tif (advice_status_uno && s->untracked_in_ms > 2000) {\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n+\t\t\t\t\t _(\"It took %.2f seconds to collect untracked files.\"),\n+\t\t\t\t\t (float)s->untracked_in_ms / 1000);\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n+\t\t\t\t\t _(\"If it happens often, you may want to use option -uno\"));\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n+\t\t\t\t\t _(\"to speed up by stopping displaying untracked files\"));\n+\t\t}\n \t} else if (s->commitable)\n \t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, _(\"Untracked files not listed%s\"),\n \t\t\tadvice_status_hints\ndiff --git a/wt-status.h b/wt-status.h\nindex 81e1dcf..74208c0 100644\n--- a/wt-status.h\n+++ b/wt-status.h\n@@ -69,6 +69,7 @@ struct wt_status {\n \tstruct string_list change;\n \tstruct string_list untracked;\n \tstruct string_list ignored;\n+\tuint32_t untracked_in_ms;\n };\n \n struct wt_status_state {\n-- \n1.8.1.2.536.gf441e6d\n"},{"id":"211216","messageId":"5140995E.3070907@web.de","threadId":"32862","inReplyTo":"1363179556-4144-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-13T15:21:02Z","receivedAt":"2013-03-13T15:21:02Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 13.03.13 13:59, Nguyễn Thái Ngọc Duy wrote:\n> This patch attempts to advertise -uno to the users who tolerate slow\n> \"git status\" on large repositories (or slow machines/disks). The 2\n> seconds limit is quite arbitrary but is probably long enough to start\n> using -uno.\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  Documentation/config.txt |  4 ++++\n>  advice.c                 |  2 ++\n>  advice.h                 |  1 +\n>  t/t7060-wtstatus.sh      |  2 ++\n>  t/t7508-status.sh        |  4 ++++\n>  t/t7512-status-help.sh   |  1 +\n>  wt-status.c              | 20 +++++++++++++++++++-\n>  wt-status.h              |  1 +\n>  8 files changed, 34 insertions(+), 1 deletion(-)\n>\n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index bbba728..e91d06f 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -178,6 +178,10 @@ advice.*::\n>  \t\tthe template shown when writing commit messages in\n>  \t\tlinkgit:git-commit[1], and in the help message shown\n>  \t\tby linkgit:git-checkout[1] when switching branch.\n> +\tstatusUno::\n> +\t\tIf collecting untracked files in linkgit:git-status[1]\n> +\t\ttakes more than 2 seconds, hint the user that the option\n> +\t\t`-uno` could be used to stop collecting untracked files.\nThanks, I like the idea\ncould we make a \"de-Luxe\" version where\n\nstatusUno is an integer, counting in milliseconds?\n\n/Torsten\n"},{"id":"211223","messageId":"7vehfj46mu.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"1363179556-4144-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-13T16:16:41Z","receivedAt":"2013-03-13T16:16:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index bbba728..e91d06f 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -178,6 +178,10 @@ advice.*::\n>  \t\tthe template shown when writing commit messages in\n>  \t\tlinkgit:git-commit[1], and in the help message shown\n>  \t\tby linkgit:git-checkout[1] when switching branch.\n> +\tstatusUno::\n> +\t\tIf collecting untracked files in linkgit:git-status[1]\n> +\t\ttakes more than 2 seconds, hint the user that the option\n> +\t\t`-uno` could be used to stop collecting untracked files.\n\nIt looks to me that the way this paragraph conveys information is\nvastly different from all the others in the section.  The section\nbegins with \"by setting their corresponding variables to false\nvarious advice messages can be squelched; here are the list of\nvariables and which advice message each of them controls\", so the\ndescription should be in \"variable:: which advice message\" form.\n\nThe noise this introduces to the test suite is a bit irritating and\nmakes us think twice if this really a good change.\n\n> diff --git a/wt-status.c b/wt-status.c\n> index ef405d0..6fde08b 100644\n> --- a/wt-status.c\n> +++ b/wt-status.c\n> @@ -540,7 +540,16 @@ void wt_status_collect(struct wt_status *s)\n>  \t\twt_status_collect_changes_initial(s);\n>  \telse\n>  \t\twt_status_collect_changes_index(s);\n> -\twt_status_collect_untracked(s);\n> +\tif (s->show_untracked_files && advice_status_uno) {\n> +\t\tstruct timeval tv1, tv2;\n> +\t\tgettimeofday(&tv1, NULL);\n> +\t\twt_status_collect_untracked(s);\n> +\t\tgettimeofday(&tv2, NULL);\n> +\t\ts->untracked_in_ms =\n> +\t\t\t(uint64_t)tv2.tv_sec * 1000 + tv2.tv_usec / 1000 -\n> +\t\t\t((uint64_t)tv1.tv_sec * 1000 + tv1.tv_usec / 1000);\n> +\t} else\n> +\t\twt_status_collect_untracked(s);\n>  }\n\nThis is not wrong per-se but it took me two reads to spot that this\nis not \"if advise is active, do the timer but do not collect;\notherwise do just collect as before\".  I wonder if we can structure\nthe code a bit better to make the timing bit less loud.\n\n>  static void wt_status_print_unmerged(struct wt_status *s)\n> @@ -1097,6 +1106,15 @@ void wt_status_print(struct wt_status *s)\n>  \t\twt_status_print_other(s, &s->untracked, _(\"Untracked files\"), \"add\");\n>  \t\tif (s->show_ignored_files)\n>  \t\t\twt_status_print_other(s, &s->ignored, _(\"Ignored files\"), \"add -f\");\n> +\t\tif (advice_status_uno && s->untracked_in_ms > 2000) {\n> +\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n> +\t\t\t\t\t _(\"It took %.2f seconds to collect untracked files.\"),\n> +\t\t\t\t\t (float)s->untracked_in_ms / 1000);\n> +\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n> +\t\t\t\t\t _(\"If it happens often, you may want to use option -uno\"));\n> +\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n> +\t\t\t\t\t _(\"to speed up by stopping displaying untracked files\"));\n> +\t\t}\n\n\"to speed up by stopping displaying untracked files\" does not look\nlike giving a balanced suggestion.  It is increasing the risk of\nforgetting about newly created files the user may want to add, but\nthe risk is not properly warned.\n\nI tend to agree that the new advice would help users if phrased in a\nright way.  Do we want them in COLOR_NORMAL, or do we want to make\nthem stand out a bit more (do we have COLOR_BLINK ;-)?\n"},{"id":"211284","messageId":"CACsJy8BixM-9bPB3G_WO+W3cTHBFxLQ=YCU2NDEzHmCYW73ZPQ@mail.gmail.com","threadId":"32862","inReplyTo":"7vehfj46mu.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-03-14T10:22:35Z","receivedAt":"2013-03-14T10:22:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 13, 2013 at 10:21 PM, Torsten Bögershausen <tboegi@web.de> wrote:\n>> +     statusUno::\n>> +             If collecting untracked files in linkgit:git-status[1]\n>> +             takes more than 2 seconds, hint the user that the option\n>> +             `-uno` could be used to stop collecting untracked files.\n> Thanks, I like the idea\n> could we make a \"de-Luxe\" version where\n>\n> statusUno is an integer, counting in milliseconds?\n\nNo problem.\n\nOn Wed, Mar 13, 2013 at 11:16 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> The noise this introduces to the test suite is a bit irritating and\n> makes us think twice if this really a good change.\n\nI originally thought of two options, this or add an env flag in git\nbinary that turns this off in the test suite. The latter did not sound\ngood. But I forgot that we set a fake $HOME in the test suite, we\ncould disable this in $HOME/.gitconfig, less clutter in individual\ntests.\n\n>>  static void wt_status_print_unmerged(struct wt_status *s)\n>> +             if (advice_status_uno && s->untracked_in_ms > 2000) {\n>> +                     status_printf_ln(s, GIT_COLOR_NORMAL,\n>> +                                      _(\"It took %.2f seconds to collect untracked files.\"),\n>> +                                      (float)s->untracked_in_ms / 1000);\n>> +                     status_printf_ln(s, GIT_COLOR_NORMAL,\n>> +                                      _(\"If it happens often, you may want to use option -uno\"));\n>> +                     status_printf_ln(s, GIT_COLOR_NORMAL,\n>> +                                      _(\"to speed up by stopping displaying untracked files\"));\n>> +             }\n>\n> \"to speed up by stopping displaying untracked files\" does not look\n> like giving a balanced suggestion.  It is increasing the risk of\n> forgetting about newly created files the user may want to add, but\n> the risk is not properly warned.\n\nHow about \"It took X ms to collect untracked files.\\nCheck out the\noption -u for a potential speedup\"? I deliberately hide \"no\" so that\nthe user cannot blindly type and run it without reading document\nfirst. We can give full explanation and warning there in the document.\n\n> I tend to agree that the new advice would help users if phrased in a\n> right way.  Do we want them in COLOR_NORMAL, or do we want to make\n> them stand out a bit more (do we have COLOR_BLINK ;-)?\n\nThere will be false positives (cold cache for example). So yeah\nsomething more standing out is good but it should catch too much\nattention. We're currently using red and green in status output. Maybe\nthis one can take blue.\n\nPS. What about advertising index v4? I sent a patch some time ago to\nput an advice in git-clone. I think it's a good place, but we could\nplace it somewhere else..\n-- \nDuy\n"},{"id":"211300","messageId":"7vmwu6yqbd.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"CACsJy8BixM-9bPB3G_WO+W3cTHBFxLQ=YCU2NDEzHmCYW73ZPQ@mail.gmail.com","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-14T15:05:42Z","receivedAt":"2013-03-14T15:05:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Wed, Mar 13, 2013 at 10:21 PM, Torsten Bögershausen <tboegi@web.de> wrote:\n>>> +     statusUno::\n>>> +             If collecting untracked files in linkgit:git-status[1]\n>>> +             takes more than 2 seconds, hint the user that the option\n>>> +             `-uno` could be used to stop collecting untracked files.\n>> Thanks, I like the idea\n>> could we make a \"de-Luxe\" version where\n>>\n>> statusUno is an integer, counting in milliseconds?\n>\n> No problem.\n\nA huge problem, as it breaks consistency and more importantly, the\nsuggestion misses the entire point of what \"advice.*\" variables are.\n\n\"advise.*\" variables are bools that indicate \"Have I learned this\nsomewhat tricky feature and/or characteristics of Git yet or do I\nstill need a reminder?\"  There is no room for \"I still need a\nreminder if it takes more than N seconds\".  You either already have\ngot it, or you haven't.\n\n>> \"to speed up by stopping displaying untracked files\" does not look\n>> like giving a balanced suggestion.  It is increasing the risk of\n>> forgetting about newly created files the user may want to add, but\n>> the risk is not properly warned.\n>\n> How about \"It took X ms to collect untracked files.\\nCheck out the\n> option -u for a potential speedup\"? I deliberately hide \"no\" so that\n> the user cannot blindly type and run it without reading document\n> first. We can give full explanation and warning there in the document.\n\nBut it makes the advise much less useful to introduce more levels of\nindirections, no?\n"},{"id":"211395","messageId":"CACsJy8BruzR=EGnwA5nc_aCJ5pO4FHyQKxd-9_36U48Ci_FFew@mail.gmail.com","threadId":"32862","inReplyTo":"7vmwu6yqbd.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-03-15T12:30:34Z","receivedAt":"2013-03-15T12:30:34Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Mar 14, 2013 at 10:05 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>>> \"to speed up by stopping displaying untracked files\" does not look\n>>> like giving a balanced suggestion.  It is increasing the risk of\n>>> forgetting about newly created files the user may want to add, but\n>>> the risk is not properly warned.\n>>\n>> How about \"It took X ms to collect untracked files.\\nCheck out the\n>> option -u for a potential speedup\"? I deliberately hide \"no\" so that\n>> the user cannot blindly type and run it without reading document\n>> first. We can give full explanation and warning there in the document.\n>\n> But it makes the advise much less useful to introduce more levels of\n> indirections, no?\n\nTo me the message's value is the pointer to -uno that not many people\nknow about. And I don't want it to be too verbose as there'll be false\npositives (cold cache, busy disks, low memory..), 2-3 lines should be\nmax. So indirections are not a concern. You want to speed up, you need\nto pay some time. Anyway how do you put it to suggest -uno in\ngit-status with all the implications?\n-- \nDuy\n"},{"id":"211401","messageId":"514343BA.3030405@web.de","threadId":"32862","inReplyTo":"CACsJy8BruzR=EGnwA5nc_aCJ5pO4FHyQKxd-9_36U48Ci_FFew@mail.gmail.com","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-15T15:52:26Z","receivedAt":"2013-03-15T15:52:26Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 03/15/2013 01:30 PM, Duy Nguyen wrote:\n> On Thu, Mar 14, 2013 at 10:05 PM, Junio C Hamano<gitster@pobox.com>  wrote:\n>>>> \"to speed up by stopping displaying untracked files\" does not look\n>>>> like giving a balanced suggestion.  It is increasing the risk of\n>>>> forgetting about newly created files the user may want to add, but\n>>>> the risk is not properly warned.\n>>> How about \"It took X ms to collect untracked files.\\nCheck out the\n>>> option -u for a potential speedup\"? I deliberately hide \"no\" so that\n>>> the user cannot blindly type and run it without reading document\n>>> first. We can give full explanation and warning there in the document.\n>> But it makes the advise much less useful to introduce more levels of\n>> indirections, no?\n> To me the message's value is the pointer to -uno that not many people\n> know about. And I don't want it to be too verbose as there'll be false\n> positives (cold cache, busy disks, low memory..), 2-3 lines should be\n> max. So indirections are not a concern. You want to speed up, you need\n> to pay some time. Anyway how do you put it to suggest -uno in\n> git-status with all the implications?\nI was thinking about the documentation, the best patch so far may look\nlike this:\nWhat we think?\n/Torsten\n\n\n-- >8 --\n\n[PATCH] git status: Document that git status -uno is faster\n\nIn some repostories users expere that \"git status\" command takes long time.\nThe command spends some time searching the file system for untracked files.\nDocument that searching for untracked file may take some time, and docuemnt\nthe option -uno better.\n\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n  Documentation/git-status.txt | 7 +++++++\n  1 file changed, 7 insertions(+)\n\ndiff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\nindex 0412c40..fd36bbd 100644\n--- a/Documentation/git-status.txt\n+++ b/Documentation/git-status.txt\n@@ -58,6 +58,13 @@ The possible options are:\n  The default can be changed using the status.showUntrackedFiles\n  configuration variable documented in linkgit:git-config[1].\n\n++\n+Note: Searching the file system for untracked files may take some time.\n+git status -uno is faster than git status -uall.\n+There is a trade-off around the use of -uno between safety and performance.\n+The default is not to use -uno so that you will not forget to add a \nfile you newly created (i.e safety).\n+You would pay for the safety with the cost to find such untracked files \n(i.e. performance).\n+\n  --ignore-submodules[=<when>]::\n      Ignore changes to submodules when looking for changes. <when> can be\n      either \"none\", \"untracked\", \"dirty\" or \"all\", which is the default.\n-- \n1.8.2.rc3.16.gce432ca\n"},{"id":"211402","messageId":"CALkWK0kaAVE=sH+iMgm8cVid5wLCpvkN7GJMfKh3O60OBaFOiw@mail.gmail.com","threadId":"32862","inReplyTo":"514343BA.3030405@web.de","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-03-15T15:57:50Z","receivedAt":"2013-03-15T15:57:50Z","isPatch":true,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Torsten Bögershausen wrote:\n> [PATCH] git status: Document that git status -uno is faster\n\nYes.  I like this patch.\n\n> In some repostories users expere that \"git status\" command takes long time.\n> The command spends some time searching the file system for untracked files.\n> Document that searching for untracked file may take some time, and docuemnt\n> the option -uno better.\n\nPlease correct the typos in the commit message.\n\n> Signed-off-by: Torsten Bögershausen <tboegi@web.de>\n> ---\n>  Documentation/git-status.txt | 7 +++++++\n>  1 file changed, 7 insertions(+)\n>\n> diff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\n> index 0412c40..fd36bbd 100644\n> --- a/Documentation/git-status.txt\n> +++ b/Documentation/git-status.txt\n> @@ -58,6 +58,13 @@ The possible options are:\n>  The default can be changed using the status.showUntrackedFiles\n>  configuration variable documented in linkgit:git-config[1].\n>\n> ++\n> +Note: Searching the file system for untracked files may take some time.\n> +git status -uno is faster than git status -uall.\n> +There is a trade-off around the use of -uno between safety and performance.\n> +The default is not to use -uno so that you will not forget to add a file\n> you newly created (i.e safety).\n> +You would pay for the safety with the cost to find such untracked files\n> (i.e. performance).\n\nGood writeup.  What -uno does is already documented, so you've\nexplained the trade-off.\nWhy didn't you just wrap the paragraph to 80 columns though?\n"},{"id":"211407","messageId":"7vvc8svc2r.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"514343BA.3030405@web.de","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-15T16:53:48Z","receivedAt":"2013-03-15T16:53:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> [PATCH] git status: Document that git status -uno is faster\n>\n> In some repostories users expere that \"git status\" command takes long time.\n\nexpere???  Certainly you did not mean \"expect\".  \"observe\",\n\"experience\", or \"see\", perhaps?\n\n> The command spends some time searching the file system for untracked files.\n> Document that searching for untracked file may take some time, and docuemnt\n> the option -uno better.\n\nGood intentions.\n\n> Signed-off-by: Torsten Bögershausen <tboegi@web.de>\n> ---\n>  Documentation/git-status.txt | 7 +++++++\n>  1 file changed, 7 insertions(+)\n>\n> diff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\n> index 0412c40..fd36bbd 100644\n> --- a/Documentation/git-status.txt\n> +++ b/Documentation/git-status.txt\n> @@ -58,6 +58,13 @@ The possible options are:\n>  The default can be changed using the status.showUntrackedFiles\n>  configuration variable documented in linkgit:git-config[1].\n>\n> ++\n> +Note: Searching the file system for untracked files may take some time.\n> +git status -uno is faster than git status -uall.\n> +There is a trade-off around the use of -uno between safety and performance.\n> +The default is not to use -uno so that you will not forget to add a\n> file you newly created (i.e safety).\n> +You would pay for the safety with the cost to find such untracked\n> files (i.e. performance).\n> +\n\nThe second sentence looks out of flow, and the last sentence, while\ntechnically not incorrect, is unclear what it is trying to convey in\nthe larger picture.\n\nPerhaps it is just me.\n\nIn any case, I think it is a good idea to explain the reason why the\nuser might want to use a non-default setting, and the criteria the\nuser may want to base the choice on (which is the gist of your\naddition), and it is a good idea to do so _before_ saying \"The\ndefault can be changed using ...\".\n\nHow about this?\n\n Documentation/git-status.txt | 14 ++++++++++----\n 1 file changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\nindex 0412c40..9046df9 100644\n--- a/Documentation/git-status.txt\n+++ b/Documentation/git-status.txt\n@@ -46,15 +46,21 @@ OPTIONS\n \tShow untracked files.\n +\n The mode parameter is optional (defaults to 'all'), and is used to\n-specify the handling of untracked files; when -u is not used, the\n-default is 'normal', i.e. show untracked files and directories.\n+specify the handling of untracked files.\n +\n The possible options are:\n +\n-\t- 'no'     - Show no untracked files\n-\t- 'normal' - Shows untracked files and directories\n+\t- 'no'     - Show no untracked files.\n+\t- 'normal' - Shows untracked files and directories.\n \t- 'all'    - Also shows individual files in untracked directories.\n +\n+When `-u` option is not used, untracked files and directories are\n+shown (i.e. the same as specifying `normal`), to help you avoid\n+forgetting to add newly created files.  Because it takes extra work\n+to find untracked files in the filesystem, this mode may take some\n+time in a large working tree.  You can use `no` to have `git status`\n+return more quickly without showing untracked files.\n++\n The default can be changed using the status.showUntrackedFiles\n configuration variable documented in linkgit:git-config[1].\n \n"},{"id":"211410","messageId":"51435D49.6040005@web.de","threadId":"32862","inReplyTo":"7vvc8svc2r.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-15T17:41:29Z","receivedAt":"2013-03-15T17:41:29Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 03/15/2013 05:53 PM, Junio C Hamano wrote:\n> Torsten Bögershausen<tboegi@web.de>  writes:\n>\n>> [PATCH] git status: Document that git status -uno is faster\n>>\n>> In some repostories users expere that \"git status\" command takes long time.\n> expere???  Certainly you did not mean \"expect\".  \"observe\",\n> \"experience\", or \"see\", perhaps?\n>\n>> The command spends some time searching the file system for untracked files.\n>> Document that searching for untracked file may take some time, and docuemnt\n>> the option -uno better.\n> Good intentions.\n>\n>> Signed-off-by: Torsten Bögershausen<tboegi@web.de>\n>> ---\n>>   Documentation/git-status.txt | 7 +++++++\n>>   1 file changed, 7 insertions(+)\n>>\n>> diff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\n>> index 0412c40..fd36bbd 100644\n>> --- a/Documentation/git-status.txt\n>> +++ b/Documentation/git-status.txt\n>> @@ -58,6 +58,13 @@ The possible options are:\n>>   The default can be changed using the status.showUntrackedFiles\n>>   configuration variable documented in linkgit:git-config[1].\n>>\n>> ++\n>> +Note: Searching the file system for untracked files may take some time.\n>> +git status -uno is faster than git status -uall.\n>> +There is a trade-off around the use of -uno between safety and performance.\n>> +The default is not to use -uno so that you will not forget to add a\n>> file you newly created (i.e safety).\n>> +You would pay for the safety with the cost to find such untracked\n>> files (i.e. performance).\n>> +\n> The second sentence looks out of flow, and the last sentence, while\n> technically not incorrect, is unclear what it is trying to convey in\n> the larger picture.\n>\n> Perhaps it is just me.\n>\n> In any case, I think it is a good idea to explain the reason why the\n> user might want to use a non-default setting, and the criteria the\n> user may want to base the choice on (which is the gist of your\n> addition), and it is a good idea to do so _before_ saying \"The\n> default can be changed using ...\".\n>\n> How about this?\n>\n>   Documentation/git-status.txt | 14 ++++++++++----\n>   1 file changed, 10 insertions(+), 4 deletions(-)\n>\n> diff --git a/Documentation/git-status.txt b/Documentation/git-status.txt\n> index 0412c40..9046df9 100644\n> --- a/Documentation/git-status.txt\n> +++ b/Documentation/git-status.txt\n> @@ -46,15 +46,21 @@ OPTIONS\n>   \tShow untracked files.\n>   +\n>   The mode parameter is optional (defaults to 'all'), and is used to\n> -specify the handling of untracked files; when -u is not used, the\n> -default is 'normal', i.e. show untracked files and directories.\n> +specify the handling of untracked files.\n>   +\n>   The possible options are:\n>   +\n> -\t- 'no'     - Show no untracked files\n> -\t- 'normal' - Shows untracked files and directories\n> +\t- 'no'     - Show no untracked files.\n> +\t- 'normal' - Shows untracked files and directories.\n>   \t- 'all'    - Also shows individual files in untracked directories.\n>   +\n> +When `-u` option is not used, untracked files and directories are\n> +shown (i.e. the same as specifying `normal`), to help you avoid\n> +forgetting to add newly created files.  Because it takes extra work\n> +to find untracked files in the filesystem, this mode may take some\n> +time in a large working tree.  You can use `no` to have `git status`\n(Small nit: extra space before the \"You\" in the line above)\n\nThanks, I like that much better than mine\n(and expere is probably a word not yet invented)\n/Torsten\n"},{"id":"211414","messageId":"7v4ngcv35l.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"51435D49.6040005@web.de","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-15T20:06:30Z","receivedAt":"2013-03-15T20:06:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> Thanks, I like that much better than mine\n> (and expere is probably a word not yet invented)\n\nOK, then how about redoing Duy's patch like this on top?\n\nI've moved the timing collection from the caller to callee, and I\nthink the result is more readable.  The message looked easier to see\nwith a leading blank line, so I added one.\n\n-- >8 --\nFrom: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nDate: Wed, 13 Mar 2013 19:59:16 +0700\nSubject: [PATCH] status: advise to consider use of -u when read_directory takes too long\n\nIntroduce advice.statusUoption to suggest considering use of -u to\nstrike different trade-off when it took more than 2 seconds to\nenumerate untracked/ignored files.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n Documentation/config.txt |  4 ++++\n advice.c                 |  2 ++\n advice.h                 |  1 +\n t/t7060-wtstatus.sh      |  1 +\n t/t7508-status.sh        |  1 +\n t/t7512-status-help.sh   |  1 +\n wt-status.c              | 21 +++++++++++++++++++++\n wt-status.h              |  1 +\n 8 files changed, 32 insertions(+)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex d1de857..a16eda5 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -163,6 +163,10 @@ advice.*::\n \t\tstate in the output of linkgit:git-status[1] and in\n \t\tthe template shown when writing commit messages in\n \t\tlinkgit:git-commit[1].\n+\tstatusUoption::\n+\t\tAdvise to consider using the `-u` option to linkgit:git-status[1]\n+\t\twhen the command takes more than 2 seconds to enumerate untracked\n+\t\tfiles.\n \tcommitBeforeMerge::\n \t\tAdvice shown when linkgit:git-merge[1] refuses to\n \t\tmerge to avoid overwriting local changes.\ndiff --git a/advice.c b/advice.c\nindex edfbd4a..015011f 100644\n--- a/advice.c\n+++ b/advice.c\n@@ -5,6 +5,7 @@ int advice_push_non_ff_current = 1;\n int advice_push_non_ff_default = 1;\n int advice_push_non_ff_matching = 1;\n int advice_status_hints = 1;\n+int advice_status_u_option = 1;\n int advice_commit_before_merge = 1;\n int advice_resolve_conflict = 1;\n int advice_implicit_identity = 1;\n@@ -19,6 +20,7 @@ static struct {\n \t{ \"pushnonffdefault\", &advice_push_non_ff_default },\n \t{ \"pushnonffmatching\", &advice_push_non_ff_matching },\n \t{ \"statushints\", &advice_status_hints },\n+\t{ \"statusuoption\", &advice_status_u_option },\n \t{ \"commitbeforemerge\", &advice_commit_before_merge },\n \t{ \"resolveconflict\", &advice_resolve_conflict },\n \t{ \"implicitidentity\", &advice_implicit_identity },\ndiff --git a/advice.h b/advice.h\nindex f3cdbbf..e3e665d 100644\n--- a/advice.h\n+++ b/advice.h\n@@ -8,6 +8,7 @@ extern int advice_push_non_ff_current;\n extern int advice_push_non_ff_default;\n extern int advice_push_non_ff_matching;\n extern int advice_status_hints;\n+extern int advice_status_u_option;\n extern int advice_commit_before_merge;\n extern int advice_resolve_conflict;\n extern int advice_implicit_identity;\ndiff --git a/t/t7060-wtstatus.sh b/t/t7060-wtstatus.sh\nindex f4f38a5..52ef06b 100755\n--- a/t/t7060-wtstatus.sh\n+++ b/t/t7060-wtstatus.sh\n@@ -5,6 +5,7 @@ test_description='basic work tree status reporting'\n . ./test-lib.sh\n \n test_expect_success setup '\n+\tgit config --global advice.statusuoption false &&\n \ttest_commit A &&\n \ttest_commit B oneside added &&\n \tgit checkout A^0 &&\ndiff --git a/t/t7508-status.sh b/t/t7508-status.sh\nindex e313ef1..15e063a 100755\n--- a/t/t7508-status.sh\n+++ b/t/t7508-status.sh\n@@ -8,6 +8,7 @@ test_description='git status'\n . ./test-lib.sh\n \n test_expect_success 'status -h in broken repository' '\n+\tgit config --global advice.statusuoption false &&\n \tmkdir broken &&\n \ttest_when_finished \"rm -fr broken\" &&\n \t(\ndiff --git a/t/t7512-status-help.sh b/t/t7512-status-help.sh\nindex b3f6eb9..2d53e03 100755\n--- a/t/t7512-status-help.sh\n+++ b/t/t7512-status-help.sh\n@@ -14,6 +14,7 @@ test_description='git status advices'\n set_fake_editor\n \n test_expect_success 'prepare for conflicts' '\n+\tgit config --global advice.statusuoption false &&\n \ttest_commit init main.txt init &&\n \tgit branch conflicts &&\n \ttest_commit on_master main.txt on_master &&\ndiff --git a/wt-status.c b/wt-status.c\nindex 2a9658b..6e75468 100644\n--- a/wt-status.c\n+++ b/wt-status.c\n@@ -496,9 +496,14 @@ static void wt_status_collect_untracked(struct wt_status *s)\n {\n \tint i;\n \tstruct dir_struct dir;\n+\tstruct timeval t_begin;\n \n \tif (!s->show_untracked_files)\n \t\treturn;\n+\n+\tif (advice_status_u_option)\n+\t\tgettimeofday(&t_begin, NULL);\n+\n \tmemset(&dir, 0, sizeof(dir));\n \tif (s->show_untracked_files != SHOW_ALL_UNTRACKED_FILES)\n \t\tdir.flags |=\n@@ -528,6 +533,14 @@ static void wt_status_collect_untracked(struct wt_status *s)\n \t}\n \n \tfree(dir.entries);\n+\n+\tif (advice_status_u_option) {\n+\t\tstruct timeval t_end;\n+\t\tgettimeofday(&t_end, NULL);\n+\t\ts->untracked_in_ms =\n+\t\t\t(uint64_t)t_end.tv_sec * 1000 + t_end.tv_usec / 1000 -\n+\t\t\t((uint64_t)t_begin.tv_sec * 1000 + t_begin.tv_usec / 1000);\n+\t}\n }\n \n void wt_status_collect(struct wt_status *s)\n@@ -1011,6 +1024,14 @@ void wt_status_print(struct wt_status *s)\n \t\twt_status_print_other(s, &s->untracked, _(\"Untracked files\"), \"add\");\n \t\tif (s->show_ignored_files)\n \t\t\twt_status_print_other(s, &s->ignored, _(\"Ignored files\"), \"add -f\");\n+\t\tif (advice_status_u_option && 2000 < s->untracked_in_ms) {\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, \"\");\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n+\t\t\t\t _(\"It took %.2f seconds to enumerate untracked files.\"),\n+\t\t\t\t s->untracked_in_ms / 1000.0);\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n+\t\t\t\t _(\"Consider the -u option for a possible speed-up?\"));\n+\t\t}\n \t} else if (s->commitable)\n \t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, _(\"Untracked files not listed%s\"),\n \t\t\tadvice_status_hints\ndiff --git a/wt-status.h b/wt-status.h\nindex 236b41f..09420d0 100644\n--- a/wt-status.h\n+++ b/wt-status.h\n@@ -69,6 +69,7 @@ struct wt_status {\n \tstruct string_list change;\n \tstruct string_list untracked;\n \tstruct string_list ignored;\n+\tuint32_t untracked_in_ms;\n };\n \n struct wt_status_state {\n-- \n1.8.2-279-g744670c\n"},{"id":"211416","messageId":"51438F33.3080607@web.de","threadId":"32862","inReplyTo":"7v4ngcv35l.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-15T21:14:27Z","receivedAt":"2013-03-15T21:14:27Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 15.03.13 21:06, Junio C Hamano wrote:\n> Torsten Bögershausen <tboegi@web.de> writes:\n> \n>> > Thanks, I like that much better than mine\n>> > (and expere is probably a word not yet invented)\n> OK, then how about redoing Duy's patch like this on top?\n> \n> I've moved the timing collection from the caller to callee, and I\n> think the result is more readable.  The message looked easier to see\n> with a leading blank line, so I added one.\n> \n> -- >8 --\n> From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> Date: Wed, 13 Mar 2013 19:59:16 +0700\n> Subject: [PATCH] status: advise to consider use of -u when read_directory takes too long\n> \n> Introduce advice.statusUoption to suggest considering use of -u to\n> strike different trade-off when it took more than 2 seconds to\n> enumerate untracked/ignored files.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  Documentation/config.txt |  4 ++++\n>  advice.c                 |  2 ++\n>  advice.h                 |  1 +\n>  t/t7060-wtstatus.sh      |  1 +\n>  t/t7508-status.sh        |  1 +\n>  t/t7512-status-help.sh   |  1 +\n>  wt-status.c              | 21 +++++++++++++++++++++\n>  wt-status.h              |  1 +\n>  8 files changed, 32 insertions(+)\n> \n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index d1de857..a16eda5 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -163,6 +163,10 @@ advice.*::\n>  \t\tstate in the output of linkgit:git-status[1] and in\n>  \t\tthe template shown when writing commit messages in\n>  \t\tlinkgit:git-commit[1].\n> +\tstatusUoption::\n> +\t\tAdvise to consider using the `-u` option to linkgit:git-status[1]\n> +\t\twhen the command takes more than 2 seconds to enumerate untracked\n> +\t\tfiles.\n>  \tcommitBeforeMerge::\n>  \t\tAdvice shown when linkgit:git-merge[1] refuses to\n>  \t\tmerge to avoid overwriting local changes.\n> diff --git a/advice.c b/advice.c\n> index edfbd4a..015011f 100644\n> --- a/advice.c\n> +++ b/advice.c\n> @@ -5,6 +5,7 @@ int advice_push_non_ff_current = 1;\n>  int advice_push_non_ff_default = 1;\n>  int advice_push_non_ff_matching = 1;\n>  int advice_status_hints = 1;\n> +int advice_status_u_option = 1;\n>  int advice_commit_before_merge = 1;\n>  int advice_resolve_conflict = 1;\n>  int advice_implicit_identity = 1;\n> @@ -19,6 +20,7 @@ static struct {\n>  \t{ \"pushnonffdefault\", &advice_push_non_ff_default },\n>  \t{ \"pushnonffmatching\", &advice_push_non_ff_matching },\n>  \t{ \"statushints\", &advice_status_hints },\n> +\t{ \"statusuoption\", &advice_status_u_option },\n>  \t{ \"commitbeforemerge\", &advice_commit_before_merge },\n>  \t{ \"resolveconflict\", &advice_resolve_conflict },\n>  \t{ \"implicitidentity\", &advice_implicit_identity },\n> diff --git a/advice.h b/advice.h\n> index f3cdbbf..e3e665d 100644\n> --- a/advice.h\n> +++ b/advice.h\n> @@ -8,6 +8,7 @@ extern int advice_push_non_ff_current;\n>  extern int advice_push_non_ff_default;\n>  extern int advice_push_non_ff_matching;\n>  extern int advice_status_hints;\n> +extern int advice_status_u_option;\n>  extern int advice_commit_before_merge;\n>  extern int advice_resolve_conflict;\n>  extern int advice_implicit_identity;\n> diff --git a/t/t7060-wtstatus.sh b/t/t7060-wtstatus.sh\n> index f4f38a5..52ef06b 100755\n> --- a/t/t7060-wtstatus.sh\n> +++ b/t/t7060-wtstatus.sh\n> @@ -5,6 +5,7 @@ test_description='basic work tree status reporting'\n>  . ./test-lib.sh\n>  \n>  test_expect_success setup '\n> +\tgit config --global advice.statusuoption false &&\n>  \ttest_commit A &&\n>  \ttest_commit B oneside added &&\n>  \tgit checkout A^0 &&\n> diff --git a/t/t7508-status.sh b/t/t7508-status.sh\n> index e313ef1..15e063a 100755\n> --- a/t/t7508-status.sh\n> +++ b/t/t7508-status.sh\n> @@ -8,6 +8,7 @@ test_description='git status'\n>  . ./test-lib.sh\n>  \n>  test_expect_success 'status -h in broken repository' '\n> +\tgit config --global advice.statusuoption false &&\n>  \tmkdir broken &&\n>  \ttest_when_finished \"rm -fr broken\" &&\n>  \t(\n> diff --git a/t/t7512-status-help.sh b/t/t7512-status-help.sh\n> index b3f6eb9..2d53e03 100755\n> --- a/t/t7512-status-help.sh\n> +++ b/t/t7512-status-help.sh\n> @@ -14,6 +14,7 @@ test_description='git status advices'\n>  set_fake_editor\n>  \n>  test_expect_success 'prepare for conflicts' '\n> +\tgit config --global advice.statusuoption false &&\n>  \ttest_commit init main.txt init &&\n>  \tgit branch conflicts &&\n>  \ttest_commit on_master main.txt on_master &&\n> diff --git a/wt-status.c b/wt-status.c\n> index 2a9658b..6e75468 100644\n> --- a/wt-status.c\n> +++ b/wt-status.c\n> @@ -496,9 +496,14 @@ static void wt_status_collect_untracked(struct wt_status *s)\n>  {\n>  \tint i;\n>  \tstruct dir_struct dir;\n> +\tstruct timeval t_begin;\n>  \n>  \tif (!s->show_untracked_files)\n>  \t\treturn;\n> +\n> +\tif (advice_status_u_option)\n> +\t\tgettimeofday(&t_begin, NULL);\n> +\n>  \tmemset(&dir, 0, sizeof(dir));\n>  \tif (s->show_untracked_files != SHOW_ALL_UNTRACKED_FILES)\n>  \t\tdir.flags |=\n> @@ -528,6 +533,14 @@ static void wt_status_collect_untracked(struct wt_status *s)\n>  \t}\n>  \n>  \tfree(dir.entries);\n> +\n> +\tif (advice_status_u_option) {\n> +\t\tstruct timeval t_end;\n> +\t\tgettimeofday(&t_end, NULL);\n> +\t\ts->untracked_in_ms =\n> +\t\t\t(uint64_t)t_end.tv_sec * 1000 + t_end.tv_usec / 1000 -\n> +\t\t\t((uint64_t)t_begin.tv_sec * 1000 + t_begin.tv_usec / 1000);\n> +\t}\n>  }\n>  \n>  void wt_status_collect(struct wt_status *s)\n> @@ -1011,6 +1024,14 @@ void wt_status_print(struct wt_status *s)\n>  \t\twt_status_print_other(s, &s->untracked, _(\"Untracked files\"), \"add\");\n>  \t\tif (s->show_ignored_files)\n>  \t\t\twt_status_print_other(s, &s->ignored, _(\"Ignored files\"), \"add -f\");\n> +\t\tif (advice_status_u_option && 2000 < s->untracked_in_ms) {\n> +\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, \"\");\n> +\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n> +\t\t\t\t _(\"It took %.2f seconds to enumerate untracked files.\"),\n> +\t\t\t\t s->untracked_in_ms / 1000.0);\n> +\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n> +\t\t\t\t _(\"Consider the -u option for a possible speed-up?\"));\n> +\t\t}\n>  \t} else if (s->commitable)\n>  \t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, _(\"Untracked files not listed%s\"),\n>  \t\t\tadvice_status_hints\n> diff --git a/wt-status.h b/wt-status.h\n> index 236b41f..09420d0 100644\n> --- a/wt-status.h\n> +++ b/wt-status.h\n> @@ -69,6 +69,7 @@ struct wt_status {\n>  \tstruct string_list change;\n>  \tstruct string_list untracked;\n>  \tstruct string_list ignored;\n> +\tuint32_t untracked_in_ms;\n>  };\n>  \n>  struct wt_status_state {\n> -- 1.8.2-279-g744670c -- To unsubscribe from this list: send the line \"unsubscribe git\" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html\n> \nThanks, that looks good to me:\n\n# It took 2.58 seconds to enumerate untracked files.\n# Consider the -u option for a possible speed-up?\n\nBut:\nIf I follow the advice as is given and use \"git status -u\", the result is the same.\n\n\nIf I think loud, would it be better to say:\n\n# It took 2.58 seconds to search for untracked files.\n# Consider the -uno option for a possible speed-up?\n\nor\n\n# It took 2.58 seconds to search for untracked files.\n# Consider the -u option for a possible speed-up?\n# Please see git help status\n\n/Torsten\n"},{"id":"211418","messageId":"7vzjy4tjd5.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"51438F33.3080607@web.de","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-15T21:59:18Z","receivedAt":"2013-03-15T21:59:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> Thanks, that looks good to me:\n>\n> # It took 2.58 seconds to enumerate untracked files.\n> # Consider the -u option for a possible speed-up?\n>\n> But:\n> If I follow the advice as is given and use \"git status -u\", the result is the same.\n\nYeah, that was taken from\n\n    http://thread.gmane.org/gmane.comp.version-control.git/215820/focus=218125\n\nto which I said something about \"more levels of indirections\".  This\nepisode shows that even a user who was very well aware of the issue\ndid not follow a single level of indirection.\n\n> If I think loud, would it be better to say:\n>\n> # It took 2.58 seconds to search for untracked files.\n> # Consider the -uno option for a possible speed-up?\n>\n> or\n>\n> # It took 2.58 seconds to search for untracked files.\n> # Consider the -u option for a possible speed-up?\n> # Please see git help status\n\nThe former actively hurts the users, but the latter would be good,\ngiven that your documentation updates clarifies the trade off.\n\nOr we can be more explicit and say\n\n# It took 2.58 seconds to search for untracked files.  'status -uno'\n# may speed it up, but you have to be careful not to forget to add\n# new files yourself (see 'git help status').\n\nor something.\n"},{"id":"211426","messageId":"CACsJy8CoHJ-rGf8LLQZMgVN4rcZOA2eFRy=7oBDtpDbFksVSeg@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8BixM-9bPB3G_WO+W3cTHBFxLQ=YCU2NDEzHmCYW73ZPQ@mail.gmail.com","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-03-16T01:51:22Z","receivedAt":"2013-03-16T01:51:22Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Mar 14, 2013 at 5:22 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Wed, Mar 13, 2013 at 11:16 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>> The noise this introduces to the test suite is a bit irritating and\n>> makes us think twice if this really a good change.\n>\n> I originally thought of two options, this or add an env flag in git\n> binary that turns this off in the test suite. The latter did not sound\n> good. But I forgot that we set a fake $HOME in the test suite, we\n> could disable this in $HOME/.gitconfig, less clutter in individual\n> tests.\n\nfwiw, adding to $HOME/.gitconfig by default in test-libs.sh does not\nwork. Else where we check \"git config --list\" and the new global\nconfig key will fail them.\n-- \nDuy\n"},{"id":"211441","messageId":"51441D8E.7090303@web.de","threadId":"32862","inReplyTo":"7vzjy4tjd5.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-03-16T07:21:50Z","receivedAt":"2013-03-16T07:21:50Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 15.03.13 22:59, Junio C Hamano wrote:\n> Torsten Bögershausen <tboegi@web.de> writes:\n> \n>> Thanks, that looks good to me:\n>>\n>> # It took 2.58 seconds to enumerate untracked files.\n>> # Consider the -u option for a possible speed-up?\n>>\n>> But:\n>> If I follow the advice as is given and use \"git status -u\", the result is the same.\n> \n> Yeah, that was taken from\n> \n>     http://thread.gmane.org/gmane.comp.version-control.git/215820/focus=218125\n> \n> to which I said something about \"more levels of indirections\".  This\n> episode shows that even a user who was very well aware of the issue\n> did not follow a single level of indirection.\n> \n>> If I think loud, would it be better to say:\n>>\n>> # It took 2.58 seconds to search for untracked files.\n>> # Consider the -uno option for a possible speed-up?\n>>\n>> or\n>>\n>> # It took 2.58 seconds to search for untracked files.\n>> # Consider the -u option for a possible speed-up?\n>> # Please see git help status\n> \n> The former actively hurts the users, but the latter would be good,\n> given that your documentation updates clarifies the trade off.\n> \n> Or we can be more explicit and say\n> \n> # It took 2.58 seconds to search for untracked files.  'status -uno'\n> # may speed it up, but you have to be careful not to forget to add\n> # new files yourself (see 'git help status').\n> \nThanks, that looks good for me\n"},{"id":"211481","messageId":"7vobeityxq.fsf@alter.siamese.dyndns.org","threadId":"32862","inReplyTo":"51441D8E.7090303@web.de","subject":"Re: [PATCH] status: hint the user about -uno if read_directory takes too long","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-03-17T04:47:29Z","receivedAt":"2013-03-17T04:47:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n>> Or we can be more explicit and say\n>> \n>> # It took 2.58 seconds to search for untracked files.  'status -uno'\n>> # may speed it up, but you have to be careful not to forget to add\n>> # new files yourself (see 'git help status').\n>> \n> Thanks, that looks good for me\n\nOK, then I'll squash this in to the version queued to 'pu', but we\nshould start thinking about merging these multi-line messages into\none multi-line strings that is split by the output layer to help the\nlocalization folks, using something like strbuf_commented_addf() and\nstrbuf_add_commented_lines().\n\ndiff --git a/wt-status.c b/wt-status.c\nindex 6e75468..53c2222 100644\n--- a/wt-status.c\n+++ b/wt-status.c\n@@ -1027,10 +1027,14 @@ void wt_status_print(struct wt_status *s)\n \t\tif (advice_status_u_option && 2000 < s->untracked_in_ms) {\n \t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, \"\");\n \t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n-\t\t\t\t _(\"It took %.2f seconds to enumerate untracked files.\"),\n+\t\t\t\t _(\"It took %.2f seconds to enumerate untracked files.\"\n+\t\t\t\t   \"  'status -uno'\"),\n \t\t\t\t s->untracked_in_ms / 1000.0);\n \t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n-\t\t\t\t _(\"Consider the -u option for a possible speed-up?\"));\n+\t\t\t\t _(\"may speed it up, but you have to be careful not\"\n+\t\t\t\t   \" to forget to add\"));\n+\t\t\tstatus_printf_ln(s, GIT_COLOR_NORMAL,\n+\t\t\t\t _(\"new files yourself (see 'git help status').\"));\n \t\t}\n \t} else if (s->commitable)\n \t\tstatus_printf_ln(s, GIT_COLOR_NORMAL, _(\"Untracked files not listed%s\"),\n"},{"id":"215339","messageId":"51781455.9090600@gmail.com","threadId":"32862","inReplyTo":"CAKXa9=qQwJqxZLxhAS35QeF1+dwH+ukod0NfFggVCuUZHz-USg@mail.gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-04-24T17:20:21Z","receivedAt":"2013-04-24T17:20:21Z","isPatch":true,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"Here is a patch that creates a daemon that tracks file\nstate with inotify, writes it out to a file upon request,\nand changes most of the calls to stat to use said cache.\n\nIt has bugs, but I figured it would be smarter to see\nif the approach was acceptable at all before spending the\ntime to root the bugs out.\n\nI've implemented the communication with a file, and not a socket, \nbecause I think implementing a socket is going to create\nsecurity issues on multiuser systems.  For example, would a\nsocket allow stat information to cross user boundaries?\n\n\nMost stat calls are redirected to a cache that is maintained by a\ndaemon that maintains file system state via inotify.\n\nSigned-off-by: Robert Zeh <robert.allan.zeh@gmail.com>\n---\n  abspath.c            |   9 ++-\n  bisect.c             |   3 +-\n  check-racy.c         |   2 +-\n  combine-diff.c       |   3 +-\n  command-list.txt     |   1 +\n  config.c             |   3 +-\n  copy.c               |   3 +-\n  diff-lib.c           |   3 +-\n  diff-no-index.c      |   3 +-\n  diff.c               |   9 ++-\n  diffcore-order.c     |   3 +-\n  dir.c                |   4 +-\n  filechange-cache.c   | 203 \n+++++++++++++++++++++++++++++++++++++++++++++++++++\n  filechange-cache.h   |  20 +++++\n  filechange-daemon.c  | 164 +++++++++++++++++++++++++++++++++++++++++\n  filechange-printer.c |  13 ++++\n  git.c                |  27 +++++++\n  ll-merge.c           |   3 +-\n  merge-recursive.c    |   5 +-\n  name-hash.c          |   3 +-\n  name-hash.h          |   1 +\n  notes-merge.c        |   3 +-\n  path.c               |   5 +-\n  read-cache.c         |  11 +--\n  rerere.c             |   7 +-\n  setup.c              |   5 +-\n  test-chmtime.c       |   2 +-\n  test-wildmatch.c     |   2 +-\n  unpack-trees.c       |   6 +-\n  29 files changed, 486 insertions(+), 40 deletions(-)\n  create mode 100644 filechange-cache.c\n  create mode 100644 filechange-cache.h\n  create mode 100644 filechange-daemon.c\n  create mode 100644 filechange-printer.c\n  create mode 100644 name-hash.h\n\ndiff --git a/abspath.c b/abspath.c\nindex 40cdc46..798c005 100644\n--- a/abspath.c\n+++ b/abspath.c\n@@ -1,3 +1,4 @@\n+#include \"filechange-cache.h\"\n  #include \"cache.h\"\n\n  /*\n@@ -8,7 +9,7 @@\n  int is_directory(const char *path)\n  {\n  \tstruct stat st;\n-\treturn (!stat(path, &st) && S_ISDIR(st.st_mode));\n+\treturn (!cached_stat(path, &st) && S_ISDIR(st.st_mode));\n  }\n\n  /* We allow \"recursive\" symbolic links. Only within reason, though. */\n@@ -117,7 +118,7 @@ static const char *real_path_internal(const char \n*path, int die_on_error)\n  \t\t\tlast_elem = NULL;\n  \t\t}\n\n-\t\tif (!lstat(buf, &st) && S_ISLNK(st.st_mode)) {\n+\t\tif (!cached_lstat(buf, &st) && S_ISLNK(st.st_mode)) {\n  \t\t\tssize_t len = readlink(buf, next_buf, PATH_MAX);\n  \t\t\tif (len < 0) {\n  \t\t\t\tif (die_on_error)\n@@ -167,9 +168,9 @@ static const char *get_pwd_cwd(void)\n  \t\treturn NULL;\n  \tpwd = getenv(\"PWD\");\n  \tif (pwd && strcmp(pwd, cwd)) {\n-\t\tstat(cwd, &cwd_stat);\n+\t\tcached_stat(cwd, &cwd_stat);\n  \t\tif ((cwd_stat.st_dev || cwd_stat.st_ino) &&\n-\t\t    !stat(pwd, &pwd_stat) &&\n+\t\t    !cached_stat(pwd, &pwd_stat) &&\n  \t\t    pwd_stat.st_dev == cwd_stat.st_dev &&\n  \t\t    pwd_stat.st_ino == cwd_stat.st_ino) {\n  \t\t\tstrlcpy(cwd, pwd, PATH_MAX);\ndiff --git a/bisect.c b/bisect.c\nindex bd1b7b5..d4b1af7 100644\n--- a/bisect.c\n+++ b/bisect.c\n@@ -1,6 +1,7 @@\n  #include \"cache.h\"\n  #include \"commit.h\"\n  #include \"diff.h\"\n+#include \"filechange-cache.h\"\n  #include \"revision.h\"\n  #include \"refs.h\"\n  #include \"list-objects.h\"\n@@ -649,7 +650,7 @@ static int is_expected_rev(const unsigned char *sha1)\n  \tFILE *fp;\n  \tint res = 0;\n\n-\tif (stat(filename, &st) || !S_ISREG(st.st_mode))\n+\tif (cached_stat(filename, &st) || !S_ISREG(st.st_mode))\n  \t\treturn 0;\n\n  \tfp = fopen(filename, \"r\");\ndiff --git a/check-racy.c b/check-racy.c\nindex 00d92a1..c54be01 100644\n--- a/check-racy.c\n+++ b/check-racy.c\n@@ -11,7 +11,7 @@ int main(int ac, char **av)\n  \t\tstruct cache_entry *ce = active_cache[i];\n  \t\tstruct stat st;\n\n-\t\tif (lstat(ce->name, &st)) {\n+\t\tif (cached_lstat(ce->name, &st)) {\n  \t\t\terror(\"lstat(%s): %s\", ce->name, strerror(errno));\n  \t\t\tcontinue;\n  \t\t}\ndiff --git a/combine-diff.c b/combine-diff.c\nindex 35d41cd..b6a09a5 100644\n--- a/combine-diff.c\n+++ b/combine-diff.c\n@@ -3,6 +3,7 @@\n  #include \"blob.h\"\n  #include \"diff.h\"\n  #include \"diffcore.h\"\n+#include \"filechange-cache.h\"\n  #include \"quote.h\"\n  #include \"xdiff-interface.h\"\n  #include \"log-tree.h\"\n@@ -806,7 +807,7 @@ static void show_patch_diff(struct combine_diff_path \n*elem, int num_parent,\n  \t\tstruct stat st;\n  \t\tint fd = -1;\n\n-\t\tif (lstat(elem->path, &st) < 0)\n+\t\tif (cached_lstat(elem->path, &st) < 0)\n  \t\t\tgoto deleted_file;\n\n  \t\tif (S_ISLNK(st.st_mode)) {\ndiff --git a/command-list.txt b/command-list.txt\nindex bf83303..9dec5e1 100644\n--- a/command-list.txt\n+++ b/command-list.txt\n@@ -29,6 +29,7 @@ git-count-objects \nancillaryinterrogators\n  git-credential                          purehelpers\n  git-credential-cache                    purehelpers\n  git-credential-store                    purehelpers\n+git-filechange-daemon\t\t\tpurehelpers\n  git-cvsexportcommit                     foreignscminterface\n  git-cvsimport                           foreignscminterface\n  git-cvsserver                           foreignscminterface\ndiff --git a/config.c b/config.c\nindex aefd80b..99749fe 100644\n--- a/config.c\n+++ b/config.c\n@@ -7,6 +7,7 @@\n   */\n  #include \"cache.h\"\n  #include \"exec_cmd.h\"\n+#include \"filechange-cache.h\"\n  #include \"strbuf.h\"\n  #include \"quote.h\"\n\n@@ -1436,7 +1437,7 @@ int git_config_set_multivar_in_file(const char \n*config_filename,\n  \t\t\tgoto out_free;\n  \t\t}\n\n-\t\tfstat(in_fd, &st);\n+\t\tcached_fstat(in_fd, &st);\n  \t\tcontents_sz = xsize_t(st.st_size);\n  \t\tcontents = xmmap(NULL, contents_sz, PROT_READ,\n  \t\t\tMAP_PRIVATE, in_fd, 0);\ndiff --git a/copy.c b/copy.c\nindex a7f58fd..972fabe 100644\n--- a/copy.c\n+++ b/copy.c\n@@ -1,3 +1,4 @@\n+#include \"filechange-cache.h\"\n  #include \"cache.h\"\n\n  int copy_fd(int ifd, int ofd)\n@@ -39,7 +40,7 @@ static int copy_times(const char *dst, const char *src)\n  {\n  \tstruct stat st;\n  \tstruct utimbuf times;\n-\tif (stat(src, &st) < 0)\n+\tif (cached_stat(src, &st) < 0)\n  \t\treturn -1;\n  \ttimes.actime = st.st_atime;\n  \ttimes.modtime = st.st_mtime;\ndiff --git a/diff-lib.c b/diff-lib.c\nindex f35de0f..8d5a005 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -2,6 +2,7 @@\n   * Copyright (C) 2005 Junio C Hamano\n   */\n  #include \"cache.h\"\n+#include \"filechange-cache.h\"\n  #include \"quote.h\"\n  #include \"commit.h\"\n  #include \"diff.h\"\n@@ -27,7 +28,7 @@\n   */\n  static int check_removed(const struct cache_entry *ce, struct stat *st)\n  {\n-\tif (lstat(ce->name, st) < 0) {\n+\tif (cached_lstat(ce->name, st) < 0) {\n  \t\tif (errno != ENOENT && errno != ENOTDIR)\n  \t\t\treturn -1;\n  \t\treturn 1;\ndiff --git a/diff-no-index.c b/diff-no-index.c\nindex 74da659..d3fb354 100644\n--- a/diff-no-index.c\n+++ b/diff-no-index.c\n@@ -7,6 +7,7 @@\n  #include \"cache.h\"\n  #include \"color.h\"\n  #include \"commit.h\"\n+#include \"filechange-cache.h\"\n  #include \"blob.h\"\n  #include \"tag.h\"\n  #include \"diff.h\"\n@@ -51,7 +52,7 @@ static int get_mode(const char *path, int *mode)\n  #endif\n  \telse if (path == file_from_standard_input)\n  \t\t*mode = create_ce_mode(0666);\n-\telse if (lstat(path, &st))\n+\telse if (cached_lstat(path, &st))\n  \t\treturn error(\"Could not access '%s'\", path);\n  \telse\n  \t\t*mode = st.st_mode;\ndiff --git a/diff.c b/diff.c\nindex 156fec4..a5be122 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -5,6 +5,7 @@\n  #include \"quote.h\"\n  #include \"diff.h\"\n  #include \"diffcore.h\"\n+#include \"filechange-cache.h\"\n  #include \"delta.h\"\n  #include \"xdiff-interface.h\"\n  #include \"color.h\"\n@@ -2629,7 +2630,7 @@ static int reuse_worktree_file(const char *name, \nconst unsigned char *sha1, int\n  \t * If ce matches the file in the work tree, we can reuse it.\n  \t */\n  \tif (ce_uptodate(ce) ||\n-\t    (!lstat(name, &st) && !ce_match_stat(ce, &st, 0)))\n+\t    (!cached_lstat(name, &st) && !ce_match_stat(ce, &st, 0)))\n  \t\treturn 1;\n\n  \treturn 0;\n@@ -2684,7 +2685,7 @@ int diff_populate_filespec(struct diff_filespec \n*s, int size_only)\n  \t\tstruct stat st;\n  \t\tint fd;\n\n-\t\tif (lstat(s->path, &st) < 0) {\n+\t\tif (cached_lstat(s->path, &st) < 0) {\n  \t\t\tif (errno == ENOENT) {\n  \t\t\terr_empty:\n  \t\t\t\terr = -1;\n@@ -2826,7 +2827,7 @@ static struct diff_tempfile \n*prepare_temp_file(const char *name,\n  \tif (!one->sha1_valid ||\n  \t    reuse_worktree_file(name, one->sha1, 1)) {\n  \t\tstruct stat st;\n-\t\tif (lstat(name, &st) < 0) {\n+\t\tif (cached_lstat(name, &st) < 0) {\n  \t\t\tif (errno == ENOENT)\n  \t\t\t\tgoto not_a_valid_file;\n  \t\t\tdie_errno(\"stat(%s)\", name);\n@@ -3043,7 +3044,7 @@ static void diff_fill_sha1_info(struct \ndiff_filespec *one)\n  \t\t\t\thashcpy(one->sha1, null_sha1);\n  \t\t\t\treturn;\n  \t\t\t}\n-\t\t\tif (lstat(one->path, &st) < 0)\n+\t\t\tif (cached_lstat(one->path, &st) < 0)\n  \t\t\t\tdie_errno(\"stat '%s'\", one->path);\n  \t\t\tif (index_path(one->sha1, one->path, &st, 0))\n  \t\t\t\tdie(\"cannot hash %s\", one->path);\ndiff --git a/diffcore-order.c b/diffcore-order.c\nindex 23e9385..636be01 100644\n--- a/diffcore-order.c\n+++ b/diffcore-order.c\n@@ -4,6 +4,7 @@\n  #include \"cache.h\"\n  #include \"diff.h\"\n  #include \"diffcore.h\"\n+#include \"filechange-cache.h\"\n\n  static char **order;\n  static int order_cnt;\n@@ -22,7 +23,7 @@ static void prepare_order(const char *orderfile)\n  \tfd = open(orderfile, O_RDONLY);\n  \tif (fd < 0)\n  \t\treturn;\n-\tif (fstat(fd, &st)) {\n+\tif (cached_fstat(fd, &st)) {\n  \t\tclose(fd);\n  \t\treturn;\n  \t}\ndiff --git a/dir.c b/dir.c\nindex 57394e4..a67a592 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -476,7 +476,7 @@ int add_excludes_from_file_to_list(const char *fname,\n  \tchar *buf, *entry;\n\n  \tfd = open(fname, O_RDONLY);\n-\tif (fd < 0 || fstat(fd, &st) < 0) {\n+\tif (fd < 0 || cached_fstat(fd, &st) < 0) {\n  \t\tif (errno != ENOENT)\n  \t\t\twarn_on_inaccessible(fname);\n  \t\tif (0 <= fd)\n@@ -1551,7 +1551,7 @@ static int remove_dir_recurse(struct strbuf *path, \nint flag, int *kept_up)\n\n  \t\tstrbuf_setlen(path, len);\n  \t\tstrbuf_addstr(path, e->d_name);\n-\t\tif (lstat(path->buf, &st))\n+\t\tif (cached_lstat(path->buf, &st))\n  \t\t\t; /* fall thru */\n  \t\telse if (S_ISDIR(st.st_mode)) {\n  \t\t\tif (!remove_dir_recurse(path, flag, &kept_down))\ndiff --git a/filechange-cache.c b/filechange-cache.c\nnew file mode 100644\nindex 0000000..80c698f\n--- /dev/null\n+++ b/filechange-cache.c\n@@ -0,0 +1,203 @@\n+#include <unistd.h>\n+#include <stdio.h>\n+#include \"builtin.h\"\n+#include \"hash.h\"\n+#include \"name-hash.h\"\n+#include \"strbuf.h\"\n+#include \"filechange-cache.h\"\n+\n+\n+static struct hash_table stat_cache;\n+static const int CACHE_ENTRY_FILE_SIZE =\n+\tsizeof(struct stat) + /* sizeof(stat_cache_entry.st) */\n+\tsizeof(int); /* sizeof(stat_cache_entry.stat_return) */\n+\n+static void insert_stat_cache_entry(const char *path,\n+\t\t\t\t    struct stat_cache_entry *new_entry);\n+\n+void setup_stat_cache()\n+{\n+\tinit_hash(&stat_cache);\n+}\n+\n+static int write_stat_cache_entry(void *void_stat_cache_entry, void \n*void_fp)\n+{\n+\tFILE *fp = (FILE*)(void_fp);\n+\tconst struct stat_cache_entry *entry =\n+\t\t(struct stat_cache_entry*)(void_stat_cache_entry);\n+\n+\tfor (; entry; entry = entry->next) {\n+\t\tif (fprintf(fp, \"%s\\n\", entry->path) < 0)\n+\t\t\tdie_errno(\"Unable to write to %s\",\n+\t\t\t\t  git_path(\"WT_STATUS_TMP\"));\n+\t\tif (fwrite(&entry->stat_return,\n+\t\t\t   CACHE_ENTRY_FILE_SIZE, 1, fp) < 0)\n+\t\t\tdie_errno(\"Unable to write to %s\",\n+\t\t\t\t  git_path(\"WT_STATUS_TMP\"));\n+\t}\n+\treturn 0;\n+}\n+\n+void write_stat_cache()\n+{\n+\tconst char *status_tmp = git_path(\"WT_STATUS_TMP\");\n+\tconst char *status_output = git_path(\"WT_STATUS\");\n+\tFILE *fp = fopen(status_tmp, \"w\");\n+\tif (!fp)\n+\t\tdie_errno(\"Unable to create %s\", status_tmp);\n+\tif (fprintf(fp, \"version_format=1\\n\") < 0)\n+\t\tdie_errno(\"Unable to write to %s\", status_tmp);\n+\tfor_each_hash(&stat_cache, write_stat_cache_entry, fp);\n+\tif (fclose(fp) < 0)\n+\t\tdie_errno(\"Unable to close %s\", status_tmp);\n+\tif (rename(status_tmp, status_output) < 0)\n+\t\tdie_errno(\"Unable to rename %s to %s\", status_tmp,\n+\t\t\t  status_output);\n+}\n+\n+static void read_stat_cache_file()\n+{\n+\tstruct strbuf line = STRBUF_INIT;\n+\tconst char *status_output = git_path(\"WT_STATUS\");\n+\tint read_version = 0;\n+\n+\tFILE *fp = fopen(status_output, \"r\");\n+\tif (!fp)\n+\t\tdie_errno(\"Unable to read %s\", status_output);\n+\t\n+\tif (strbuf_getline(&line, fp, '\\n') != EOF) {\n+\t\tsscanf(line.buf, \"version_format=%d\\n\", &read_version);\n+\t\tif (read_version != 1) {\n+\t\t\tdie(\"Expected version 1 of stat_cache file\");\n+\t\t}\n+\t}\n+\n+\twhile (strbuf_getline(&line, fp, '\\n') != EOF) {\n+\t\tstruct stat_cache_entry *entry =\n+\t\t\t(struct stat_cache_entry*)(xcalloc(1, sizeof(*entry)));\n+\t\tentry->path = xstrdup(line.buf);\n+\t\tif (fread(&entry->stat_return,\n+\t\t\t  CACHE_ENTRY_FILE_SIZE, 1, fp) != 1) {\n+\t\t\tdie_errno(\"Unable to read stat_cache file\");\n+\t\t}\n+\t\tinsert_stat_cache_entry(entry->path, entry);\n+\t}\n+\n+\tstrbuf_release(&line);\n+}\n+\n+static int request_stat_cache_file()\n+{\n+\tint count = 0;\n+\tint stat_return_code = 0;\n+\tconst char *request_path = git_path(\"REQUEST_WT_STATUS\");\n+\tconst char *status_output = git_path(\"WT_STATUS\");\n+\tstruct stat stat_buf;\n+\n+\tchar buffer[1] = { 0 };\n+\tFILE *fp = NULL;\n+\n+\tif (0 != stat(request_path, &stat_buf))\n+\t\treturn 0;\n+\t\n+\tif (unlink(status_output) != 0 && (errno != ENOENT))\n+\t\tdie_errno(\"Unable to remove %s\", status_output);\n+\n+\tfp = fopen(request_path, \"w\");\n+\tif (!fp) {\n+\t\tdie_errno(\"Unable to open %s\", request_path);\n+\t}\n+\n+\tif (fwrite(&buffer, 0, 0, fp) != 0)\n+\t\tdie_errno(\"Unable to write to %s\", request_path);\n+\t\n+\tfor (count = 0;\n+\t     (count < 10) &&\n+\t\t     ((stat_return_code = stat(status_output, &stat_buf)) != 0) &&\n+\t\t     (errno == ENOENT);\n+\t     count++) {\n+\t\tusleep(1000);\n+\t}\n+\n+\treturn stat_return_code == 0;\n+}\n+\n+void read_stat_cache()\n+{\n+\tstatic int read_cache = 1;\n+\tif (read_cache && request_stat_cache_file()) {\n+\t\tread_stat_cache_file();\n+\t\tread_cache = 0;\n+\t}\n+\tread_cache = 0;\n+}\n+\n+struct stat_cache_entry *get_stat_cache_entry(const char *path)\n+{\n+\tconst unsigned int hash = hash_name(path, strlen(path));\n+\tstruct stat_cache_entry *entry = NULL;\n+\tfor(entry = lookup_hash(hash, &stat_cache); entry;\n+\t    entry = entry->next) {\n+\t\tif (!strcmp(path, entry->path)) return entry;\n+\t}\n+\treturn NULL;\n+}\n+\n+static void insert_stat_cache_entry(const char *path,\n+\t\t\t\t    struct stat_cache_entry *new_entry)\n+{\n+\tassert(get_stat_cache_entry(path) == NULL);\n+\n+\tvoid **insert_result =\n+\t\tinsert_hash(hash_name(path, strlen(path)), (void*)new_entry,\n+\t\t\t    &stat_cache);\n+\tif (!insert_result) return;\n+\tstruct stat_cache_entry *existing_entry =\n+\t\t(struct stat_cache_entry*)(*insert_result);\n+\twhile(existing_entry->next) {\n+\t\texisting_entry = existing_entry->next;\n+\t}\n+\tassert(!existing_entry->next);\n+\texisting_entry->next = new_entry;\n+}\n+\n+void update_stat_cache(const char *path)\n+{\n+\tstruct stat_cache_entry *entry = get_stat_cache_entry(path);\n+\tif (!entry) {\n+\t\tentry = (struct stat_cache_entry*)(xcalloc(1, sizeof(*entry)));\n+\t\tentry->path = xstrdup(path);\n+\t\tinsert_stat_cache_entry(path, entry);\n+\t}\n+\t\n+\tentry->stat_return = lstat(path, &entry->st);\n+}\n+\n+int cached_stat(const char *path, struct stat *buf)\n+{\n+\treturn stat(path, buf);\n+}\n+\n+int cached_fstat(int fd, struct stat *buf)\n+{\n+\treturn fstat(fd, buf);\n+}\n+\n+int cached_lstat(const char *path, struct stat *buf)\n+{\n+\tint stat_return_value = 0;\n+\tstruct stat_cache_entry *entry = 0;\n+\n+\tread_stat_cache();\n+\n+\tentry = get_stat_cache_entry(path);\n+\n+\tstat_return_value = lstat(path, buf);\n+\t\n+\tif (entry && (stat_return_value != entry->stat_return) &&\n+\t    (memcpy(&entry->st, buf, sizeof(*buf)))) {\n+\t\tabort();\n+\t}\n+\t\n+\treturn stat_return_value;\n+}\ndiff --git a/filechange-cache.h b/filechange-cache.h\nnew file mode 100644\nindex 0000000..75a9f79\n--- /dev/null\n+++ b/filechange-cache.h\n@@ -0,0 +1,20 @@\n+#include <sys/types.h>\n+#include <sys/stat.h>\n+#include <unistd.h>\n+\n+struct stat_cache_entry {\n+\tconst char *path;\n+\tstruct stat_cache_entry *next;\n+\tint stat_return;\n+\tstruct stat st;\n+};\n+\n+extern void write_stat_cache();\n+extern void read_stat_cache();\n+extern void setup_stat_cache();\n+extern struct stat_cache_entry *get_stat_cache_entry(const char *path);\n+extern void update_stat_cache(const char *path);\n+\n+extern int cached_stat(const char *path, struct stat *buf);\n+extern int cached_fstat(int fd, struct stat *buf);\n+extern int cached_lstat(const char *path, struct stat *buf);\ndiff --git a/filechange-daemon.c b/filechange-daemon.c\nnew file mode 100644\nindex 0000000..df6f0d3\n--- /dev/null\n+++ b/filechange-daemon.c\n@@ -0,0 +1,164 @@\n+#include <stdio.h>\n+#include <libgen.h>\n+#include <x86_64-linux-gnu/sys/inotify.h>\n+\n+#include \"filechange-cache.h\"\n+#include \"builtin.h\"\n+#include \"dir.h\"\n+#include \"hash.h\"\n+\n+static int request_watch_descriptor = -1;\n+static int root_directory_watch_descriptor = -1;\n+\n+static void setup_environment()\n+{\n+\tsetup_stat_cache();\n+}\n+\n+static int setup_inotify()\n+{\n+\tint inotify_fd = inotify_init();\n+\tif (inotify_fd < 0) {\n+\t\tdie_errno(\"Unable to create inotify watch\");\n+\t}\n+\treturn inotify_fd;\n+}\n+\n+static void restart()\n+{\n+\n+}\n+\n+\n+static void watch_control(int inotify_fd)\n+{\n+\tstruct stat stat_buf;\n+\tconst char *request_path = git_path(\"REQUEST_WT_STATUS\");\n+\n+\tif ((stat(request_path, &stat_buf) == -1) && (errno == ENOENT)) {\n+\t\tFILE *out = fopen(request_path, \"w\");\n+\t\tif (out == NULL)\n+\t\t\tdie_errno(\"Unable to create %s\", request_path);\n+\t}\n+\n+\trequest_watch_descriptor = inotify_add_watch(inotify_fd,\n+\t\t\t\t\t\t     request_path, IN_MODIFY);\n+\t\n+\tif (request_watch_descriptor < 0)\n+\t\tdie_errno(\"Unable to watch %s\", get_git_dir());\n+}\n+\n+static void watch_file(int inotify_fd, const char *path)\n+{\n+\tint watch_descriptor = 0;\n+\tchar *path_copy = xstrdup(path);\n+\tchar *dir = dirname(path_copy);\n+\tconst int interest_set =\n+\t\tIN_MODIFY  | IN_DELETE | IN_CREATE  |\n+\t\tIN_DELETE_SELF | IN_MOVE_SELF |\n+\t\tIN_MOVED_TO;\n+\n+\twatch_descriptor = inotify_add_watch(inotify_fd, dir, interest_set);\n+\tif (watch_descriptor < 0)\n+\t\tdie_errno(\"Unable to create inotify watch for %s\", dir);\n+\n+\twatch_descriptor = inotify_add_watch(inotify_fd, path, interest_set);\n+\tif (watch_descriptor < 0)\n+\t\tdie_errno(\"Unable to create inotify watch for %s\", dir);\n+\tupdate_stat_cache(path);\n+\n+\tfree(path_copy);\n+}\n+\n+static void watch_directory(int inotify_fd)\n+{\n+\tchar buf[PATH_MAX];\n+\n+\tif (!getcwd(buf, sizeof(buf)))\n+\t\tdie_errno(\"Unable to get current directory\");\n+\n+\tint i = 0;\n+\tstruct dir_struct dir;\n+\tconst char *pathspec[1] = { buf, NULL };\n+\n+\tmemset(&dir, 0, sizeof(dir));\n+\tsetup_standard_excludes(&dir);\n+\n+\tfill_directory(&dir, pathspec);\n+\tfor(i = 0; i < dir.nr; i++) {\n+\t\tstruct dir_entry *ent = dir.entries[i];\n+\t\twatch_file(inotify_fd, ent->name);\n+\t\tfree(ent);\n+\t}\n+\n+\tfree(dir.entries);\n+\tfree(dir.ignored);\n+}\n+\n+static void watch_root_directory(int inotify_fd)\n+{\n+\tchar buf[PATH_MAX];\n+\n+\tif (!getcwd(buf, sizeof(buf)))\n+\t\tdie_errno(\"Unable to get current directory\");\n+\n+\troot_directory_watch_descriptor =\n+\t\tinotify_add_watch(inotify_fd, buf, IN_DELETE);\n+\tif (root_directory_watch_descriptor < 0)\n+\t\tdie_errno(\"Unable to watch %s directory\", buf);\n+}\n+\n+#define INOTIFY_EVENT_SIZE  (sizeof (struct inotify_event)  + PATH_MAX + 1)\n+\n+static void remove_request_file(void)\n+{\n+\tconst char *request_path = git_path(\"REQUEST_WT_STATUS\");\n+\tif (unlink(request_path)) {\n+\t\tdie_errno(\"Unable to remove %s on exit\",\n+\t\t\t  request_path);\n+\t}\n+}\n+\n+static void loop(int inotify_fd)\n+{\n+\tchar buffer[INOTIFY_EVENT_SIZE * 10];\n+\tint length = 0;\n+\t\n+\twhile (1) {\n+\t\tint i = 0;\n+\t\tlength = read(inotify_fd, buffer, sizeof(buffer));\n+\t\tfor(i = 0; i < length; ) {\n+\t\t\tstruct inotify_event *event =\n+\t\t\t\t(struct inotify_event*)(buffer+i);\n+\t\t\t/* printf(\"event: %d %x %d %s\\n\", event->wd, event->mask,\n+\t\t\t   event->len, event->name); */\n+\t\t\tif (request_watch_descriptor == event->wd) {\n+\t\t\t\twrite_stat_cache();\n+\t\t\t} else if (root_directory_watch_descriptor\n+\t\t\t\t   == event->wd) {\n+\t\t\t\tprintf(\"root directory died!\\n\");\n+\t\t\t\texit(0);\n+\t\t\t} else if (event->mask & IN_Q_OVERFLOW) {\n+\t\t\t\trestart();\n+\t\t\t} else if (event->mask & IN_MODIFY) {\n+\t\t\t\tif (event->len)\n+\t\t\t\t\tupdate_stat_cache(event->name);\n+\t\t\t}\n+\t\t\t\n+\t\t\ti += sizeof(struct inotify_event) + event->len;\n+\t\t}\n+\t}\n+}\n+\n+int main(int argc, const char **argv)\n+{\n+\tconst int inotify_fd = setup_inotify();\n+\n+\tatexit(remove_request_file);\n+\tsetup_environment();\n+\twatch_control(inotify_fd);\n+\twatch_root_directory(inotify_fd);\n+\twatch_directory(inotify_fd);\n+\tloop(inotify_fd);\n+\treturn 0;\n+}\ndiff --git a/filechange-printer.c b/filechange-printer.c\nnew file mode 100644\nindex 0000000..fe43d80\n--- /dev/null\n+++ b/filechange-printer.c\n@@ -0,0 +1,13 @@\n+#include <stdio.h>\n+#include \"filechange-cache.h\"\n+\n+int main()\n+{\n+\tstruct stat_cache_entry *entry = NULL;\n+\tconst char *missing = \"t/t7201-co.sh\";\n+\tread_stat_cache();\n+\t\n+\tentry = get_stat_cache_entry(missing);\n+\tprintf(\"%p\\n\", entry);\n+\treturn 0;\n+}\ndiff --git a/git.c b/git.c\nindex b10c18b..ea92a65 100644\n--- a/git.c\n+++ b/git.c\n@@ -504,6 +504,31 @@ static int run_argv(int *argcp, const char ***argv)\n  }\n\n\n+static void fork_filechange_daemon()\n+{\n+\tstruct stat stat_buf;\n+\tFILE *log = fopen(\"/tmp/foo.txt\", \"a\");\n+\tfprintf(log, \"cwd = %s\\n\", get_current_dir_name());\n+\n+\tif (stat(git_path(\"REQUEST_WT_STATUS\"), &stat_buf) == -1) {\n+\t\tpid_t child = 0;\n+\n+\t\tchild = fork();\n+\t\tfprintf(log, \"starting %d\\n\", (int)child);\n+\t\tif (!child) {\n+\t\t\tfclose(log);\n+\t\t\texecl(\"/home/razeh/src/git/git-filechange-daemon\",\n+\t\t\t      \"/home/razeh/src/git/git-filechange-daemon\",\n+\t\t\t      get_current_dir_name(),\n+\t\t\t      (char*) NULL);\n+\t\t\tdie_errno(\"Unable to launch file change daemon\");\n+\t\t}\n+\t} else {\n+\t\tfprintf(log, \"already running\\n\");\n+\t}\n+\n+}\n+\n  int main(int argc, const char **argv)\n  {\n  \tconst char *cmd;\n@@ -558,6 +583,8 @@ int main(int argc, const char **argv)\n  \t */\n  \tsetup_path();\n\n+\tfork_filechange_daemon();\n+\n  \twhile (1) {\n  \t\tstatic int done_help = 0;\n  \t\tstatic int was_alias = 0;\ndiff --git a/ll-merge.c b/ll-merge.c\nindex fb61ea6..7ced2bb 100644\n--- a/ll-merge.c\n+++ b/ll-merge.c\n@@ -6,6 +6,7 @@\n\n  #include \"cache.h\"\n  #include \"attr.h\"\n+#include \"filechange-cache.h\"\n  #include \"xdiff-interface.h\"\n  #include \"run-command.h\"\n  #include \"ll-merge.h\"\n@@ -195,7 +196,7 @@ static int ll_ext_merge(const struct ll_merge_driver \n*fn,\n  \tfd = open(temp[1], O_RDONLY);\n  \tif (fd < 0)\n  \t\tgoto bad;\n-\tif (fstat(fd, &st))\n+\tif (cached_fstat(fd, &st))\n  \t\tgoto close_bad;\n  \tresult->size = st.st_size;\n  \tresult->ptr = xmalloc(result->size + 1);\ndiff --git a/merge-recursive.c b/merge-recursive.c\nindex ea9dbd3..7d371d6 100644\n--- a/merge-recursive.c\n+++ b/merge-recursive.c\n@@ -12,6 +12,7 @@\n  #include \"tree-walk.h\"\n  #include \"diff.h\"\n  #include \"diffcore.h\"\n+#include \"filechange-cache.h\"\n  #include \"tag.h\"\n  #include \"unpack-trees.h\"\n  #include \"string-list.h\"\n@@ -606,7 +607,7 @@ static char *unique_path(struct merge_options *o, \nconst char *path, const char *\n  \t\t\t*p = '_';\n  \twhile (string_list_has_string(&o->current_file_set, newpath) ||\n  \t       string_list_has_string(&o->current_directory_set, newpath) ||\n-\t       lstat(newpath, &st) == 0)\n+\t       cached_lstat(newpath, &st) == 0)\n  \t\tsprintf(p, \"_%d\", suffix++);\n\n  \tstring_list_insert(&o->current_file_set, newpath);\n@@ -634,7 +635,7 @@ static int dir_in_way(const char *path, int \ncheck_working_copy)\n  \t}\n\n  \tfree(dirpath);\n-\treturn check_working_copy && !lstat(path, &st) && S_ISDIR(st.st_mode);\n+\treturn check_working_copy && !cached_lstat(path, &st) && \nS_ISDIR(st.st_mode);\n  }\n\n  static int was_tracked(const char *path)\ndiff --git a/name-hash.c b/name-hash.c\nindex d8d25c2..d88185f 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -7,6 +7,7 @@\n   */\n  #define NO_THE_INDEX_COMPATIBILITY_MACROS\n  #include \"cache.h\"\n+#include \"name-hash.h\"\n\n  /*\n   * This removes bit 5 if bit 6 is set.\n@@ -20,7 +21,7 @@ static inline unsigned char icase_hash(unsigned char c)\n  \treturn c & ~((c & 0x40) >> 1);\n  }\n\n-static unsigned int hash_name(const char *name, int namelen)\n+unsigned int hash_name(const char *name, int namelen)\n  {\n  \tunsigned int hash = 0x123;\n\ndiff --git a/name-hash.h b/name-hash.h\nnew file mode 100644\nindex 0000000..3355d94\n--- /dev/null\n+++ b/name-hash.h\n@@ -0,0 +1 @@\n+extern unsigned int hash_name(const char *name, int namelen);\ndiff --git a/notes-merge.c b/notes-merge.c\nindex 0f67bd3..f792f83 100644\n--- a/notes-merge.c\n+++ b/notes-merge.c\n@@ -3,6 +3,7 @@\n  #include \"refs.h\"\n  #include \"diff.h\"\n  #include \"diffcore.h\"\n+#include \"filechange-cache.h\"\n  #include \"xdiff-interface.h\"\n  #include \"ll-merge.h\"\n  #include \"dir.h\"\n@@ -731,7 +732,7 @@ int notes_merge_commit(struct notes_merge_options *o,\n\n  \t\tstrbuf_addstr(&path, e->d_name);\n  \t\t/* write file as blob, and add to partial_tree */\n-\t\tif (stat(path.buf, &st))\n+\t\tif (cached_stat(path.buf, &st))\n  \t\t\tdie_errno(\"Failed to stat '%s'\", path.buf);\n  \t\tif (index_path(blob_sha1, path.buf, &st, HASH_WRITE_OBJECT))\n  \t\t\tdie(\"Failed to write blob object from '%s'\", path.buf);\ndiff --git a/path.c b/path.c\nindex d3d3f8b..6844d2d 100644\n--- a/path.c\n+++ b/path.c\n@@ -11,6 +11,7 @@\n   * which is what it's designed for.\n   */\n  #include \"cache.h\"\n+#include \"filechange-cache.h\"\n  #include \"strbuf.h\"\n  #include \"string-list.h\"\n\n@@ -360,7 +361,7 @@ const char *enter_repo(const char *path, int strict)\n  \t\tfor (i = 0; suffix[i]; i++) {\n  \t\t\tstruct stat st;\n  \t\t\tstrcpy(used_path + len, suffix[i]);\n-\t\t\tif (!stat(used_path, &st) &&\n+\t\t\tif (!cached_stat(used_path, &st) &&\n  \t\t\t    (S_ISREG(st.st_mode) ||\n  \t\t\t    (S_ISDIR(st.st_mode) && is_git_directory(used_path)))) {\n  \t\t\t\tstrcat(validated_path, suffix[i]);\n@@ -400,7 +401,7 @@ int set_shared_perm(const char *path, int mode)\n  \t\treturn 0;\n  \t}\n  \tif (!mode) {\n-\t\tif (lstat(path, &st) < 0)\n+\t\tif (cached_lstat(path, &st) < 0)\n  \t\t\treturn -1;\n  \t\tmode = st.st_mode;\n  \t\torig_mode = mode;\ndiff --git a/read-cache.c b/read-cache.c\nindex 827ae55..508ddc1 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -8,6 +8,7 @@\n  #include \"cache-tree.h\"\n  #include \"refs.h\"\n  #include \"dir.h\"\n+#include \"filechange-cache.h\"\n  #include \"tree.h\"\n  #include \"commit.h\"\n  #include \"blob.h\"\n@@ -672,7 +673,7 @@ int add_to_index(struct index_state *istate, const \nchar *path, struct stat *st,\n  int add_file_to_index(struct index_state *istate, const char *path, \nint flags)\n  {\n  \tstruct stat st;\n-\tif (lstat(path, &st))\n+\tif (cached_lstat(path, &st))\n  \t\tdie_errno(\"unable to stat '%s'\", path);\n  \treturn add_to_index(istate, path, &st, flags);\n  }\n@@ -1032,7 +1033,7 @@ static struct cache_entry \n*refresh_cache_ent(struct index_state *istate,\n  \t\treturn ce;\n  \t}\n\n-\tif (lstat(ce->name, &st) < 0) {\n+\tif (cached_lstat(ce->name, &st) < 0) {\n  \t\tif (err)\n  \t\t\t*err = errno;\n  \t\treturn NULL;\n@@ -1430,7 +1431,7 @@ int read_index_from(struct index_state *istate, \nconst char *path)\n  \t\tdie_errno(\"index file open failed\");\n  \t}\n\n-\tif (fstat(fd, &st))\n+\tif (cached_fstat(fd, &st))\n  \t\tdie_errno(\"cannot stat the open index\");\n\n  \tmmap_size = xsize_t(st.st_size);\n@@ -1618,7 +1619,7 @@ static void ce_smudge_racily_clean_entry(struct \ncache_entry *ce)\n  \t */\n  \tstruct stat st;\n\n-\tif (lstat(ce->name, &st) < 0)\n+\tif (cached_lstat(ce->name, &st) < 0)\n  \t\treturn;\n  \tif (ce_match_stat_basic(ce, &st))\n  \t\treturn;\n@@ -1830,7 +1831,7 @@ int write_index(struct index_state *istate, int newfd)\n  \t\t\treturn -1;\n  \t}\n\n-\tif (ce_flush(&c, newfd) || fstat(newfd, &st))\n+\tif (ce_flush(&c, newfd) || cached_fstat(newfd, &st))\n  \t\treturn -1;\n  \tistate->timestamp.sec = (unsigned int)st.st_mtime;\n  \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\ndiff --git a/rerere.c b/rerere.c\nindex a6a5cd5..5115d0e 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -1,4 +1,5 @@\n  #include \"cache.h\"\n+#include \"filechange-cache.h\"\n  #include \"string-list.h\"\n  #include \"rerere.h\"\n  #include \"xdiff-interface.h\"\n@@ -28,7 +29,7 @@ const char *rerere_path(const char *hex, const char *file)\n  static int has_rerere_resolution(const char *hex)\n  {\n  \tstruct stat st;\n-\treturn !stat(rerere_path(hex, \"postimage\"), &st);\n+\treturn !cached_stat(rerere_path(hex, \"postimage\"), &st);\n  }\n\n  static void read_rr(struct string_list *rr)\n@@ -681,13 +682,13 @@ int rerere_forget(const char **pathspec)\n  static time_t rerere_created_at(const char *name)\n  {\n  \tstruct stat st;\n-\treturn stat(rerere_path(name, \"preimage\"), &st) ? (time_t) 0 : \nst.st_mtime;\n+\treturn cached_stat(rerere_path(name, \"preimage\"), &st) ? (time_t) 0 : \nst.st_mtime;\n  }\n\n  static time_t rerere_last_used_at(const char *name)\n  {\n  \tstruct stat st;\n-\treturn stat(rerere_path(name, \"postimage\"), &st) ? (time_t) 0 : \nst.st_mtime;\n+\treturn cached_stat(rerere_path(name, \"postimage\"), &st) ? (time_t) 0 : \nst.st_mtime;\n  }\n\n  static void unlink_rr_item(const char *name)\ndiff --git a/setup.c b/setup.c\nindex 2e1521b..690987a 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1,5 +1,6 @@\n  #include \"cache.h\"\n  #include \"dir.h\"\n+#include \"filechange-cache.h\"\n  #include \"string-list.h\"\n\n  static int inside_git_dir = -1;\n@@ -74,7 +75,7 @@ int check_filename(const char *prefix, const char *arg)\n  \t\tname = prefix_filename(prefix, strlen(prefix), arg);\n  \telse\n  \t\tname = arg;\n-\tif (!lstat(name, &st))\n+\tif (!cached_lstat(name, &st))\n  \t\treturn 1; /* file exists */\n  \tif (errno == ENOENT || errno == ENOTDIR)\n  \t\treturn 0; /* file does not exist */\n@@ -638,7 +639,7 @@ static const char *setup_nongit(const char *cwd, int \n*nongit_ok)\n  static dev_t get_device_or_die(const char *path, const char *prefix, \nint prefix_len)\n  {\n  \tstruct stat buf;\n-\tif (stat(path, &buf)) {\n+\tif (cached_stat(path, &buf)) {\n  \t\tdie_errno(\"failed to stat '%*s%s%s'\",\n  \t\t\t\tprefix_len,\n  \t\t\t\tprefix ? prefix : \"\",\ndiff --git a/test-chmtime.c b/test-chmtime.c\nindex 92713d1..bb5f22a 100644\n--- a/test-chmtime.c\n+++ b/test-chmtime.c\n@@ -81,7 +81,7 @@ int main(int argc, const char *argv[])\n  \t\tstruct stat sb;\n  \t\tstruct utimbuf utb;\n\n-\t\tif (stat(argv[i], &sb) < 0) {\n+\t\tif (cached_stat(argv[i], &sb) < 0) {\n  \t\t\tfprintf(stderr, \"Failed to stat %s: %s\\n\",\n  \t\t\t        argv[i], strerror(errno));\n  \t\t\treturn -1;\ndiff --git a/test-wildmatch.c b/test-wildmatch.c\nindex a3e2643..838ff69 100644\n--- a/test-wildmatch.c\n+++ b/test-wildmatch.c\n@@ -19,7 +19,7 @@ static int perf(int ac, char **av)\n  \tif (lang && strcmp(lang, \"C\"))\n  \t\tdie(\"Please test it on C locale.\");\n\n-\tif ((fd = open(file, O_RDONLY)) == -1 || fstat(fd, &st))\n+\tif ((fd = open(file, O_RDONLY)) == -1 || cached_fstat(fd, &st))\n  \t\tdie_errno(\"file open\");\n\n  \tbuffer = xmalloc(st.st_size + 2);\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 09e53df..fc20be4 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1430,13 +1430,13 @@ static int verify_absent_1(struct cache_entry *ce,\n  \t\tchar path[PATH_MAX + 1];\n  \t\tmemcpy(path, ce->name, len);\n  \t\tpath[len] = 0;\n-\t\tif (lstat(path, &st))\n+\t\tif (cached_lstat(path, &st))\n  \t\t\treturn error(\"cannot stat '%s': %s\", path,\n  \t\t\t\t\tstrerror(errno));\n\n  \t\treturn check_ok_to_remove(path, len, DT_UNKNOWN, NULL, &st,\n  \t\t\t\terror_type, o);\n-\t} else if (lstat(ce->name, &st)) {\n+\t} else if (cached_lstat(ce->name, &st)) {\n  \t\tif (errno != ENOENT)\n  \t\t\treturn error(\"cannot stat '%s': %s\", ce->name,\n  \t\t\t\t     strerror(errno));\n@@ -1838,7 +1838,7 @@ int oneway_merge(struct cache_entry **src, struct \nunpack_trees_options *o)\n  \t\tint update = 0;\n  \t\tif (o->reset && o->update && !ce_uptodate(old) && \n!ce_skip_worktree(old)) {\n  \t\t\tstruct stat st;\n-\t\t\tif (lstat(old->name, &st) ||\n+\t\t\tif (cached_lstat(old->name, &st) ||\n  \t\t\t    ie_match_stat(o->src_index, old, &st, \nCE_MATCH_IGNORE_VALID|CE_MATCH_IGNORE_SKIP_WORKTREE))\n  \t\t\t\tupdate |= CE_UPDATE;\n  \t\t}\n-- \n1.8.2.rc0.29.g3a0aba8.dirty\n"},{"id":"215390","messageId":"CACsJy8AuQFGCwOBTXU48T65+7DTmCw31RZc0Z-2YBpkKYcoAoA@mail.gmail.com","threadId":"32862","inReplyTo":"51781455.9090600@gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-04-24T21:32:38Z","receivedAt":"2013-04-24T21:32:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Apr 25, 2013 at 3:20 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n> Here is a patch that creates a daemon that tracks file\n> state with inotify, writes it out to a file upon request,\n> and changes most of the calls to stat to use said cache.\n>\n> It has bugs, but I figured it would be smarter to see\n> if the approach was acceptable at all before spending the\n> time to root the bugs out.\n\nAny preliminary performance numbers? How does it do compared to\nno-inotify version? When only a few files are changed? When half the\nrepo is changed?\n\n> I've implemented the communication with a file, and not a socket, because I\n> think implementing a socket is going to create\n> security issues on multiuser systems.  For example, would a\n> socket allow stat information to cross user boundaries?\n\nI think UNIX socket on Linux at least respects file permissions. But\nunix(7) follows with \"This behavior differs from many BSD-derived\nsystems which ignore permissions for Unix sockets\". Sighh\n\n>  abspath.c            |   9 ++-\n>  bisect.c             |   3 +-\n>  check-racy.c         |   2 +-\n>  combine-diff.c       |   3 +-\n>  command-list.txt     |   1 +\n>  config.c             |   3 +-\n>  copy.c               |   3 +-\n>  diff-lib.c           |   3 +-\n>  diff-no-index.c      |   3 +-\n>  diff.c               |   9 ++-\n>  diffcore-order.c     |   3 +-\n>  dir.c                |   4 +-\n>  filechange-cache.c   | 203\n> +++++++++++++++++++++++++++++++++++++++++++++++++++\n>  filechange-cache.h   |  20 +++++\n>  filechange-daemon.c  | 164 +++++++++++++++++++++++++++++++++++++++++\n>  filechange-printer.c |  13 ++++\n>  git.c                |  27 +++++++\n>  ll-merge.c           |   3 +-\n>  merge-recursive.c    |   5 +-\n>  name-hash.c          |   3 +-\n>  name-hash.h          |   1 +\n>  notes-merge.c        |   3 +-\n>  path.c               |   5 +-\n>  read-cache.c         |  11 +--\n>  rerere.c             |   7 +-\n>  setup.c              |   5 +-\n>  test-chmtime.c       |   2 +-\n>  test-wildmatch.c     |   2 +-\n>  unpack-trees.c       |   6 +-\n>  29 files changed, 486 insertions(+), 40 deletions(-)\n>  create mode 100644 filechange-cache.c\n>  create mode 100644 filechange-cache.h\n>  create mode 100644 filechange-daemon.c\n>  create mode 100644 filechange-printer.c\n>  create mode 100644 name-hash.h\n\nCan you just replace lstat/stat with cached_lstat/stat inside\ngit-compat-util.h and not touch all files at once? I think you may\nneed to deal with paths outside working directory. But because you're\nusing lookup table, that should be no problem.\n--\nDuy\n"},{"id":"215414","messageId":"87sj2f6n1u.fsf@linux-k42r.v.cablecom.net","threadId":"32862","inReplyTo":"51781455.9090600@gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-04-25T08:18:21Z","receivedAt":"2013-04-25T08:18:21Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Robert Zeh <robert.allan.zeh@gmail.com> writes:\n\n> Here is a patch that creates a daemon that tracks file\n> state with inotify, writes it out to a file upon request,\n> and changes most of the calls to stat to use said cache.\n>\n> It has bugs, but I figured it would be smarter to see\n> if the approach was acceptable at all before spending the\n> time to root the bugs out.\n\nThanks for tackling this; it's probably about time we got a inotify\nsupport :-(\n\n> I've implemented the communication with a file, and not a socket,\n> because I think implementing a socket is going to create\n> security issues on multiuser systems.  For example, would a\n> socket allow stat information to cross user boundaries?\n\nThis ties in with an issue discussed in an earlier thread:\n\n  http://thread.gmane.org/gmane.comp.version-control.git/217817/focus=218307\n\nThe conclusion there was that the default limits are set such that it is\nnot feasible to run one daemon per repository (that would quickly hit\nthe limits when e.g. iterating all repos in a typical android tree using\nrepo).\n\nSo whatever you use for communication needs to work as a global daemon.\n\nI'd just trust the SSH folks to know about security; on my system\nssh-agent creates\n\n  /tmp/ssh-RANDOMSTRING/agent.PID\n\nwhere the directory has mode 0700, and the file is a unit socket with\nmode 0600.  That should make doubly sure that no other user can open the\nsocket.\n\n>  filechange-cache.c   | 203\n> +++++++++++++++++++++++++++++++++++++++++++++++++++\n\nIs your MUA wrapping the patch?\n\n> +static void watch_directory(int inotify_fd)\n> +{\n> +\tchar buf[PATH_MAX];\n> +\n> +\tif (!getcwd(buf, sizeof(buf)))\n> +\t\tdie_errno(\"Unable to get current directory\");\n> +\n> +\tint i = 0;\n> +\tstruct dir_struct dir;\n> +\tconst char *pathspec[1] = { buf, NULL };\n> +\n> +\tmemset(&dir, 0, sizeof(dir));\n> +\tsetup_standard_excludes(&dir);\n> +\n> +\tfill_directory(&dir, pathspec);\n> +\tfor(i = 0; i < dir.nr; i++) {\n> +\t\tstruct dir_entry *ent = dir.entries[i];\n> +\t\twatch_file(inotify_fd, ent->name);\n> +\t\tfree(ent);\n> +\t}\n\nI don't get this bit.  The lstat() are run over all files listed in the\nindex.  So shouldn't your daemon watch exactly those (or rather, all\ndirnames of such files)?\n\nThe actual directory contents are only needed to find untracked files,\nand there would be a lot of complication surrounding that, so I suggest\nsaving that for later (and for now measuring the speedup with 'git\nstatus -uno'!).\n\nFor example, you'd have to actually watch and re-read all .gitignore\nfiles, and the .git/info/exclude, and the core.excludesfile, to see if\nyour notion of an ignored file became stale.\n\nAlso, you seem to call watch_directory() only on the current(?) dir, but\nyou need to recursively set up watches for all directories in the\nrepository.\n\n> +\twhile (1) {\n> +\t\tint i = 0;\n> +\t\tlength = read(inotify_fd, buffer, sizeof(buffer));\n> +\t\tfor(i = 0; i < length; ) {\n> +\t\t\tstruct inotify_event *event =\n> +\t\t\t\t(struct inotify_event*)(buffer+i);\n> +\t\t\t/* printf(\"event: %d %x %d %s\\n\", event->wd, event->mask,\n> +\t\t\t   event->len, event->name); */\n> +\t\t\tif (request_watch_descriptor == event->wd) {\n> +\t\t\t\twrite_stat_cache();\n> +\t\t\t} else if (root_directory_watch_descriptor\n> +\t\t\t\t   == event->wd) {\n> +\t\t\t\tprintf(\"root directory died!\\n\");\n> +\t\t\t\texit(0);\n> +\t\t\t} else if (event->mask & IN_Q_OVERFLOW) {\n> +\t\t\t\trestart();\n\nGood.\n\n> +\t\t\t} else if (event->mask & IN_MODIFY) {\n> +\t\t\t\tif (event->len)\n> +\t\t\t\t\tupdate_stat_cache(event->name);\n> +\t\t\t}\n\nSo whenever a file changes, you stat() it.  That's good for simplicity\nnow, but I suspect it will provide some optimization opportunities\nlater.\n\n\nOn some design aspects, I'd want:\n\n* a toggle to run the test suite with the daemons, or without\n\n* if you go with a user-wide daemon, a way to ensure that the test-suite\n  daemon is not the same as my \"real\" daemon, and make sure it is killed\n  after the test runs finish\n\n* a test that triggers IN_Q_OVERFLOW, e.g. by sending SIGSTOP and doing\n  a large repository operation\n\n* a test that renames directories\n\nThe last one is just based on my personal experience with messing with\ninotify; renaming directories is the \"hard\" case for that API.  We may\nalready cover this in the test suite, or we may not; but it must be\ntested.\n\nOther than that last point, focus your tests not on small tests but on\nthe test suite.  It would seem rather unlikely to me that you could\nmanage to pass the entire test suite with this daemon active but broken.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"215482","messageId":"CAKXa9=rvDQ7DXwCiTp9PTc55gNTW2UDZ4auaYG5tbboomrDAGQ@mail.gmail.com","threadId":"32862","inReplyTo":"87sj2f6n1u.fsf@linux-k42r.v.cablecom.net","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-04-25T19:37:58Z","receivedAt":"2013-04-25T19:37:58Z","isPatch":true,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"On Thu, Apr 25, 2013 at 3:18 AM, Thomas Rast <trast@inf.ethz.ch> wrote:\n>\n> Robert Zeh <robert.allan.zeh@gmail.com> writes:\n>\n> > Here is a patch that creates a daemon that tracks file\n> > state with inotify, writes it out to a file upon request,\n> > and changes most of the calls to stat to use said cache.\n> >\n> > It has bugs, but I figured it would be smarter to see\n> > if the approach was acceptable at all before spending the\n> > time to root the bugs out.\n>\n> Thanks for tackling this; it's probably about time we got a inotify\n> support :-(\n\n> > I've implemented the communication with a file, and not a socket,\n> > because I think implementing a socket is going to create\n> > security issues on multiuser systems.  For example, would a\n> > socket allow stat information to cross user boundaries?\n>\n> This ties in with an issue discussed in an earlier thread:\n>\n>   http://thread.gmane.org/gmane.comp.version-control.git/217817/focus=218307\n>\n> The conclusion there was that the default limits are set such that it is\n> not feasible to run one daemon per repository (that would quickly hit\n> the limits when e.g. iterating all repos in a typical android tree using\n> repo).\n>\n> So whatever you use for communication needs to work as a global daemon.\n>\n> I'd just trust the SSH folks to know about security; on my system\n> ssh-agent creates\n>\n>   /tmp/ssh-RANDOMSTRING/agent.PID\n>\n> where the directory has mode 0700, and the file is a unit socket with\n> mode 0600.  That should make doubly sure that no other user can open the\n> socket.\n>\n> >  filechange-cache.c   | 203\n> > +++++++++++++++++++++++++++++++++++++++++++++++++++\n>\n> Is your MUA wrapping the patch?\n\nAlmost certainly.  I'll double check before I send off the next patch.\n\n> > +static void watch_directory(int inotify_fd)\n> > +{\n> > +     char buf[PATH_MAX];\n> > +\n> > +     if (!getcwd(buf, sizeof(buf)))\n> > +             die_errno(\"Unable to get current directory\");\n> > +\n> > +     int i = 0;\n> > +     struct dir_struct dir;\n> > +     const char *pathspec[1] = { buf, NULL };\n> > +\n> > +     memset(&dir, 0, sizeof(dir));\n> > +     setup_standard_excludes(&dir);\n> > +\n> > +     fill_directory(&dir, pathspec);\n> > +     for(i = 0; i < dir.nr; i++) {\n> > +             struct dir_entry *ent = dir.entries[i];\n> > +             watch_file(inotify_fd, ent->name);\n> > +             free(ent);\n> > +     }\n>\n> I don't get this bit.  The lstat() are run over all files listed in the\n> index.  So shouldn't your daemon watch exactly those (or rather, all\n> dirnames of such files)?\nI believe that fill_directory is handling watching only files in the index.\nI had some problems a while back when I was only watching the\ndirectory with some of the inotify structures coming back empty, which\nis why I started watching each individual file.\n\n> The actual directory contents are only needed to find untracked files,\n> and there would be a lot of complication surrounding that, so I suggest\n> saving that for later (and for now measuring the speedup with 'git\n> status -uno'!).\nThe speed up test is a good idea.\n\n> For example, you'd have to actually watch and re-read all .gitignore\n> files, and the .git/info/exclude, and the core.excludesfile, to see if\n> your notion of an ignored file became stale.\nThe thought in the back of my head was to simple have the daemon\nrestart if one of those files changed, under the assumption that a\nrestart wasn't that expensive, and that it would be complicated to check.\n\n\n> Also, you seem to call watch_directory() only on the current(?) dir, but\n> you need to recursively set up watches for all directories in the\n> repository.\n\nI'm calling fill_directory to get the list of files to watch; it appears to\nbe handling the recursion for me.  It also appears to be handling filtering\nout all of the untracked files, etc.\n\n> > +     while (1) {\n> > +             int i = 0;\n> > +             length = read(inotify_fd, buffer, sizeof(buffer));\n> > +             for(i = 0; i < length; ) {\n> > +                     struct inotify_event *event =\n> > +                             (struct inotify_event*)(buffer+i);\n> > +                     /* printf(\"event: %d %x %d %s\\n\", event->wd, event->mask,\n> > +                        event->len, event->name); */\n> > +                     if (request_watch_descriptor == event->wd) {\n> > +                             write_stat_cache();\n> > +                     } else if (root_directory_watch_descriptor\n> > +                                == event->wd) {\n> > +                             printf(\"root directory died!\\n\");\n> > +                             exit(0);\n> > +                     } else if (event->mask & IN_Q_OVERFLOW) {\n> > +                             restart();\n>\n> Good.\n>\n> > +                     } else if (event->mask & IN_MODIFY) {\n> > +                             if (event->len)\n> > +                                     update_stat_cache(event->name);\n> > +                     }\n>\n> So whenever a file changes, you stat() it.  That's good for simplicity\n> now, but I suspect it will provide some optimization opportunities\n> later.\nI figured it would be a good idea to get things working, and then worry\nabout optimization later :-)\n\n>\n> On some design aspects, I'd want:\n>\n> * a toggle to run the test suite with the daemons, or without\nYeap.\n> * if you go with a user-wide daemon, a way to ensure that the test-suite\n>   daemon is not the same as my \"real\" daemon, and make sure it is killed\n>   after the test runs finish\nI'm assuming a command line argument that points a daemon at a port would be the\nway to handle that.\n\n> * a test that triggers IN_Q_OVERFLOW, e.g. by sending SIGSTOP and doing\n>   a large repository operation\nYeap.  I think you'd want some way to verify (through a log file?)\nthat the overflow\nhappened.\n\n> * a test that renames directories\nYeap.\n\n> The last one is just based on my personal experience with messing with\n> inotify; renaming directories is the \"hard\" case for that API.  We may\n> already cover this in the test suite, or we may not; but it must be\n> tested.\n>\n> Other than that last point, focus your tests not on small tests but on\n> the test suite.  It would seem rather unlikely to me that you could\n> manage to pass the entire test suite with this daemon active but broken.\nI've had some experiences where the test suite passes with the daemon active,\nbut not populating the cache.\n> --\n> Thomas Rast\n> trast@{inf,student}.ethz.ch\n"},{"id":"215486","messageId":"CAKXa9=pt2mxwFtepoOLZ-Atw3Ey5_OHh6rzk43kVTs8=vcVuRw@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8AuQFGCwOBTXU48T65+7DTmCw31RZc0Z-2YBpkKYcoAoA@mail.gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-04-25T19:44:28Z","receivedAt":"2013-04-25T19:44:28Z","isPatch":true,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"On Wed, Apr 24, 2013 at 4:32 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Thu, Apr 25, 2013 at 3:20 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n>> Here is a patch that creates a daemon that tracks file\n>> state with inotify, writes it out to a file upon request,\n>> and changes most of the calls to stat to use said cache.\n>>\n>> It has bugs, but I figured it would be smarter to see\n>> if the approach was acceptable at all before spending the\n>> time to root the bugs out.\n>\n> Any preliminary performance numbers? How does it do compared to\n> no-inotify version? When only a few files are changed? When half the\n> repo is changed?\n\nNo numbers yet; I'm still working on correctness.  What I posted does\nnot pass all of the tests.\n\nI like your ideas for performance tests.  My testing setup is\na VirtualBox instance on MacOS, so I'm not convinced that my numbers\nwill be meaningful.  The one thing I can report is that running the daemon\ndoesn't affect compilation performance.\n\nThe real win for this type of cache is Windows, where the file system\nis slow.\n\n>> I've implemented the communication with a file, and not a socket, because I\n>> think implementing a socket is going to create\n>> security issues on multiuser systems.  For example, would a\n>> socket allow stat information to cross user boundaries?\n>\n> I think UNIX socket on Linux at least respects file permissions. But\n> unix(7) follows with \"This behavior differs from many BSD-derived\n> systems which ignore permissions for Unix sockets\". Sighh\n>\n>>  abspath.c            |   9 ++-\n>>  bisect.c             |   3 +-\n>>  check-racy.c         |   2 +-\n>>  combine-diff.c       |   3 +-\n>>  command-list.txt     |   1 +\n>>  config.c             |   3 +-\n>>  copy.c               |   3 +-\n>>  diff-lib.c           |   3 +-\n>>  diff-no-index.c      |   3 +-\n>>  diff.c               |   9 ++-\n>>  diffcore-order.c     |   3 +-\n>>  dir.c                |   4 +-\n>>  filechange-cache.c   | 203\n>> +++++++++++++++++++++++++++++++++++++++++++++++++++\n>>  filechange-cache.h   |  20 +++++\n>>  filechange-daemon.c  | 164 +++++++++++++++++++++++++++++++++++++++++\n>>  filechange-printer.c |  13 ++++\n>>  git.c                |  27 +++++++\n>>  ll-merge.c           |   3 +-\n>>  merge-recursive.c    |   5 +-\n>>  name-hash.c          |   3 +-\n>>  name-hash.h          |   1 +\n>>  notes-merge.c        |   3 +-\n>>  path.c               |   5 +-\n>>  read-cache.c         |  11 +--\n>>  rerere.c             |   7 +-\n>>  setup.c              |   5 +-\n>>  test-chmtime.c       |   2 +-\n>>  test-wildmatch.c     |   2 +-\n>>  unpack-trees.c       |   6 +-\n>>  29 files changed, 486 insertions(+), 40 deletions(-)\n>>  create mode 100644 filechange-cache.c\n>>  create mode 100644 filechange-cache.h\n>>  create mode 100644 filechange-daemon.c\n>>  create mode 100644 filechange-printer.c\n>>  create mode 100644 name-hash.h\n>\n> Can you just replace lstat/stat with cached_lstat/stat inside\n> git-compat-util.h and not touch all files at once? I think you may\n> need to deal with paths outside working directory. But because you're\n> using lookup table, that should be no problem.\n\nThat's a good idea; but there are a few places where you want to call\nthe uncached stat because calling the cache leads to recursion or\nyou bump into things that haven't been setup yet.  Any ideas how to\nhandle that?\n"},{"id":"215487","messageId":"87bo92nzzi.fsf@hexa.v.cablecom.net","threadId":"32862","inReplyTo":"CAKXa9=rvDQ7DXwCiTp9PTc55gNTW2UDZ4auaYG5tbboomrDAGQ@mail.gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-04-25T19:59:13Z","receivedAt":"2013-04-25T19:59:13Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Robert Zeh <robert.allan.zeh@gmail.com> writes:\n\n> On Thu, Apr 25, 2013 at 3:18 AM, Thomas Rast <trast@inf.ethz.ch> wrote:\n>>\n>> I don't get this bit.  The lstat() are run over all files listed in the\n>> index.  So shouldn't your daemon watch exactly those (or rather, all\n>> dirnames of such files)?\n> I believe that fill_directory is handling watching only files in the index.\n> I had some problems a while back when I was only watching the\n> directory with some of the inotify structures coming back empty, which\n> is why I started watching each individual file.\n\nThis probably doesn't scale well enough.  For example on my system the\nmaximum number of watches I can set[1] is 64k.  linux.git contains 38k\nfiles and the total number of files in all repos of an android clone I\nhave lying around is almost 300k.\n\nCan you clarify what went wrong if you only watch directories?  After\nall the events should be the same, except that you need to reassemble\nthe actual filename from the 'name' field in inotify_event and the\ndirectory name associated with the watch descriptor.\n\nI'll keep the rest of your mail for another reply ;-)\n\n[1]  /proc/sys/fs/inotify/max_user_watches\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"215511","messageId":"CACsJy8ChXRMR93r2R5NoTL7Ly1HqWCXq=t=Kj4ma5+MyYvESpg@mail.gmail.com","threadId":"32862","inReplyTo":"CAKXa9=pt2mxwFtepoOLZ-Atw3Ey5_OHh6rzk43kVTs8=vcVuRw@mail.gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-04-25T21:20:26Z","receivedAt":"2013-04-25T21:20:26Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Apr 26, 2013 at 2:44 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n>> Can you just replace lstat/stat with cached_lstat/stat inside\n>> git-compat-util.h and not touch all files at once? I think you may\n>> need to deal with paths outside working directory. But because you're\n>> using lookup table, that should be no problem.\n>\n> That's a good idea; but there are a few places where you want to call\n> the uncached stat because calling the cache leads to recursion or\n> you bump into things that haven't been setup yet.  Any ideas how to\n> handle that?\n\nOn second thought, no my idea was stupid. We only need to optimize\nlstat for certain cases and naming cached_lstat is much clearer. I\nsuspect read-cache.c and maybe dir.c and unpack-trees.c are the only\nplaces that need cached_lstat. Other places should not issue many\nlstats and we don't need to touch them.\n--\nDuy\n"},{"id":"215593","messageId":"CAKXa9=rjQJMwgcAs9ic3XSqFh40NYrQt217QS_VUw7ifKihqdA@mail.gmail.com","threadId":"32862","inReplyTo":"CACsJy8ChXRMR93r2R5NoTL7Ly1HqWCXq=t=Kj4ma5+MyYvESpg@mail.gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Robert Zeh","fromEmail":"robert.allan.zeh@gmail.com","sentAt":"2013-04-26T15:35:43Z","receivedAt":"2013-04-26T15:35:43Z","isPatch":true,"sender":{"key":"robert.allan.zeh@gmail.com","avatar":null},"body":"On Thu, Apr 25, 2013 at 4:20 PM, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Fri, Apr 26, 2013 at 2:44 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n>>> Can you just replace lstat/stat with cached_lstat/stat inside\n>>> git-compat-util.h and not touch all files at once? I think you may\n>>> need to deal with paths outside working directory. But because you're\n>>> using lookup table, that should be no problem.\n>>\n>> That's a good idea; but there are a few places where you want to call\n>> the uncached stat because calling the cache leads to recursion or\n>> you bump into things that haven't been setup yet.  Any ideas how to\n>> handle that?\n>\n> On second thought, no my idea was stupid. We only need to optimize\n> lstat for certain cases and naming cached_lstat is much clearer. I\n> suspect read-cache.c and maybe dir.c and unpack-trees.c are the only\n> places that need cached_lstat. Other places should not issue many\n> lstats and we don't need to touch them.\n\nok.  The only reason I did it for all of them was the it was a simple search\nand replace, and I didn't know how often lstat was called from various\nlocations.\n"},{"id":"215707","messageId":"87r4hwdqur.fsf@hexa.v.cablecom.net","threadId":"32862","inReplyTo":"87bo92nzzi.fsf@hexa.v.cablecom.net","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-04-27T13:51:08Z","receivedAt":"2013-04-27T13:51:08Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Thomas Rast <trast@inf.ethz.ch> writes:\n\n> Robert Zeh <robert.allan.zeh@gmail.com> writes:\n>\n>> On Thu, Apr 25, 2013 at 3:18 AM, Thomas Rast <trast@inf.ethz.ch> wrote:\n>>>\n>>> I don't get this bit.  The lstat() are run over all files listed in the\n>>> index.  So shouldn't your daemon watch exactly those (or rather, all\n>>> dirnames of such files)?\n>> I believe that fill_directory is handling watching only files in the index.\n>> I had some problems a while back when I was only watching the\n>> directory with some of the inotify structures coming back empty, which\n>> is why I started watching each individual file.\n>\n> This probably doesn't scale well enough.  For example on my system the\n> maximum number of watches I can set[1] is 64k.  linux.git contains 38k\n> files and the total number of files in all repos of an android clone I\n> have lying around is almost 300k.\n\n[I just sent something similar as a reply to a mail that I then noticed\nwas sent off-list, but I meant it to be public.]\n\nI just had a change of heart.  It's probably better for the early work\nif you make a very controllable, single-repository daemon.  Perhaps one\nthat only starts on demand (by running some git command) and runs until\nagain killed on demand.\n\nThat way it's much easier to test, and integrate as an option in the\ntest suite.  And for single-repo-minded people, like (I guess?) the\nkernel and webkit folks, this should already provide some benefit.\n\nThe per-user daemon complication can come later; we know even at this\npoint that it will have to be done *eventually*, but let's go one step\nat a time.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"215773","messageId":"CACsJy8BuMsdAAxPoY_R0tOKJ9toTnDwAwOx_=vmbOOpFLWmS5A@mail.gmail.com","threadId":"32862","inReplyTo":"51781455.9090600@gmail.com","subject":"Re: [PATCH] inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-04-27T23:56:59Z","receivedAt":"2013-04-27T23:56:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Apr 25, 2013 at 12:20 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n> +int cached_lstat(const char *path, struct stat *buf)\n> +{\n> +       int stat_return_value = 0;\n> +       struct stat_cache_entry *entry = 0;\n> +\n> +       read_stat_cache();\n> +\n> +       entry = get_stat_cache_entry(path);\n> +\n> +       stat_return_value = lstat(path, buf);\n> +\n> +       if (entry && (stat_return_value != entry->stat_return) &&\n> +           (memcpy(&entry->st, buf, sizeof(*buf)))) {\n> +               abort();\n> +       }\n> +\n> +       return stat_return_value;\n> +}\n\nI must be missing something. If you always do lstat() in\ncached_lstat(), what's the point of the cache? If you worry about\nintegrity (in the abort case), it'll be easier if you just record and\nsend paths from the daemon to git. Then you do lstat at one place\n(git). This function may become more complex if still want to watch a\nworktree way bigger that inotify limit. But I guess for now we could\njust exit the daemon early in that case and fall back to normal lstat.\n--\nDuy\n"},{"id":"215943","messageId":"CACsJy8Bzd+K39MtiF3nWRbFfBo+kjAUmm-qLgpzeZZtSSxnqWg@mail.gmail.com","threadId":"32862","inReplyTo":"CAKXa9=r2A7UeBV2s2H3wVGdPkS1zZ9huNJhtvTC-p0S5Ed12xA@mail.gmail.com","subject":"Re: inotify to minimize stat() calls","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-04-30T00:27:12Z","receivedAt":"2013-04-30T00:27:12Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Apr 30, 2013 at 1:05 AM, Robert Zeh <robert.allan.zeh@gmail.com> wrote:\n> The call to lstat is only there for testing and should not be in there for\n> the final version. Is there an easy way to only enable it for tests?\n\nThe usual trick is invent a new GIT_ environment variable. Then check\nit and do something different. Then you can set the env in tests only.\n--\nDuy\n"}]}