{"thread":{"id":"20702","subject":"Commit performance, or lack thereof","startedAt":"2009-08-22T23:42:32Z","lastAt":"2009-08-23T02:20:22Z","messageCount":2,"participants":["James Cloos","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"121533","messageId":"m3r5v39zzz.fsf@lugabout.jhcloos.org","threadId":"20702","inReplyTo":null,"subject":"Commit performance, or lack thereof","fromName":"James Cloos","fromEmail":"cloos@jhcloos.com","sentAt":"2009-08-22T23:42:32Z","receivedAt":"2009-08-22T23:42:32Z","isPatch":false,"sender":{"key":"cloos@jhcloos.com","avatar":"https://gravatar.com/avatar/ec9a05787d29afe41e243e4b60bd0e2f69d757688e8f0bfe5e78bc185a3e317f?d=mp&s=160"},"body":"Starting in the kernel tree, if one edits and adds a single file and\nthen commits it w/o specifying the file name as an argument to commit,\ngit uses 10 or so Megs of VM and one sees performace akin to this:\n\n,----« gtime -v git commit -m 'make oldconfig' »\n| [master 53d6af1] make oldconfig\n|  1 files changed, 5 insertions(+), 2 deletions(-)\n| \tCommand being timed: \"git commit -m make oldconfig\"\n| \tUser time (seconds): 0.26\n| \tSystem time (seconds): 1.06\n| \tPercent of CPU this job got: 2%\n| \tElapsed (wall clock) time (h:mm:ss or m:ss): 0:47.04\n| \tAverage shared text size (kbytes): 0\n| \tAverage unshared data size (kbytes): 0\n| \tAverage stack size (kbytes): 0\n| \tAverage total size (kbytes): 0\n| \tMaximum resident set size (kbytes): 0\n| \tAverage resident set size (kbytes): 0\n| \tMajor (requiring I/O) page faults: 86\n| \tMinor (reclaiming a frame) page faults: 4703\n| \tVoluntary context switches: 4805\n| \tInvoluntary context switches: 274\n| \tSwaps: 0\n| \tFile system inputs: 48384\n| \tFile system outputs: 5680\n| \tSocket messages sent: 0\n| \tSocket messages received: 0\n| \tSignals delivered: 0\n| \tPage size (bytes): 4096\n| \tExit status: 0\n`----\n\nOTOH, if one does specify the filename as an argument to commit, git\nuses almost 300 Megs of VM and the numbers look more like:\n\n,----« gtime -v git commit -m 'make oldconfig' .config »\n| [master 4db1e8b] make oldconfig\n|  1 files changed, 1 insertions(+), 1 deletions(-)\n| \tCommand being timed: \"git commit -m make oldconfig .config\"\n| \tUser time (seconds): 1.82\n| \tSystem time (seconds): 1.80\n| \tPercent of CPU this job got: 3%\n| \tElapsed (wall clock) time (h:mm:ss or m:ss): 1:45.72\n| \tAverage shared text size (kbytes): 0\n| \tAverage unshared data size (kbytes): 0\n| \tAverage stack size (kbytes): 0\n| \tAverage total size (kbytes): 0\n| \tMaximum resident set size (kbytes): 0\n| \tAverage resident set size (kbytes): 0\n| \tMajor (requiring I/O) page faults: 1609\n| \tMinor (reclaiming a frame) page faults: 21363\n| \tVoluntary context switches: 10707\n| \tInvoluntary context switches: 620\n| \tSwaps: 0\n| \tFile system inputs: 361192\n| \tFile system outputs: 11296\n| \tSocket messages sent: 0\n| \tSocket messages received: 0\n| \tSignals delivered: 0\n| \tPage size (bytes): 4096\n| \tExit status: 0\n`----\n\nGit should be able to do the latter operation as efficiently as it can\ndo the former operation.\n\n-JimC\n-- \nJames Cloos <cloos@jhcloos.com>         OpenPGP: 1024D/ED7DAEA6\n"},{"id":"121548","messageId":"7vvdkf1dax.fsf@alter.siamese.dyndns.org","threadId":"20702","inReplyTo":"m3r5v39zzz.fsf@lugabout.jhcloos.org","subject":"Re: Commit performance, or lack thereof","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-23T02:20:22Z","receivedAt":"2009-08-23T02:20:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"James Cloos <cloos@jhcloos.com> writes:\n\n> Starting in the kernel tree, if one edits and adds a single file and\n> then commits it w/o specifying the file name as an argument to commit,\n> ...\n> OTOH, if one does specify the filename as an argument to commit,...\n\nWhen you do:\n\n    git add something && git commit\n\nthe commit step is just \"write the index out as a tree, create a commit\nobject to wrap that tree object\".  It never has to look at your work tree\nand is a very cheap operation.\n\nOn the other hand, if you do:\n\n    some other random things && git commit pathspec\n\nthe commit step involves a lot more operations.  It has to do:\n\n    - create a temporary index from HEAD (i.e. ignore the modification to\n      the real index you did so far with your earlier \"git add\");\n\n    - rehash all the paths in the work tree that matches the pathspec and\n      add them to that temporary index; and finally\n\n    - write the temporary index out as a tree, create a commit object to\n      wrap that tree object.\n\nThe second step has to go to your work tree, and if you have a slow disk,\nor your buffer cache is cold, you naturally have to pay.\n\nSo it is not surprising at all if you observe that the latter takes more\ntime and resource than the former, although there could be something wrong\nin your set-up that you need to spend 300M for it.  It is fundamentally a\nmore expensive operation.\n\n\n    \n"}]}