{"thread":{"id":"33984","subject":"Poor performance of git describe in big repos","startedAt":"2013-05-30T10:38:32Z","lastAt":"2013-06-03T17:48:54Z","messageCount":33,"participants":["Alex Bennée","Ramkumar Ramachandra","John Keeping","Duy Nguyen","Thomas Rast","Antoine Pelisse","Jeff King","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"218964","messageId":"CAJ-05NPQLVFhtb9KMLNLc5MqguBYM1=gKEVrrtT3kSMiZKma_g@mail.gmail.com","threadId":"33984","inReplyTo":null,"subject":"Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T10:38:32Z","receivedAt":"2013-05-30T10:38:32Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"Hi,\n\nI'm a fairly heavy user of the magit Emacs extension for interacting\nwith my git repos. However I've noticed there are some cases where lag\nis very high. By analysing strace output of emacs calling git I found\ntwo commands that where particularly problematic when interrogating\nthe repo:\n\n11:00 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\ndescribe --long --tags\najb-build-test-5224-10-gfa296e6\n\nreal    0m5.016s\nuser    0m4.364s\nsys     0m0.444s\n\n11:34 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\ndescribe --contains HEAD\nfatal: cannot describe 'fa296e61f549a1252a65a13b2f734d7afbc7e88e'\n\nreal    0m4.805s\nuser    0m4.388s\nsys     0m0.400s\n\nRunning with first command with the --debug flag on gives:\n\n11:34 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\ndescribe --long --tags --debug\nsearching to describe HEAD\n lightweight       10 ajb-build-test-5224\n lightweight       41 ajb-build-test-5222\n annotated        146 vnms-2-1-36-32\n annotated        155 vnms-2-1-36-31\n annotated        174 vnms-2-1-36-30\n annotated        183 vnms-2-1-36-29\n lightweight      188 vnms-2-1-36-28\n annotated        193 vnms-2-1-36-27\n annotated        206 vnms-2-1-36-26\n annotated        215 vectastar-4-2-83-5\ntraversed 223 commits\nmore than 10 tags found; listed 10 most recent\ngave up search at 2b69df72d47be8440e3ce4cee91b9b7ceaf8b77c\najb-build-test-5224-10-gfa296e6\n\nreal    0m4.817s\nuser    0m4.320s\nsys     0m0.464s\n\nWhich has only traversed 223 before coming to a decision. This seems\nlike a very low number of commits given the time it's spent doing\nthis.\n\nOne factor might be the size of my repo (.git is around 2.4G). Could\nthis just be due to computational cost of searching through large\npacks to walk the commit chain? Is there any way to make this easier\nfor git to do?\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"218966","messageId":"CALkWK0ndKMZRuWgdg6djqPUGxbDAqZPcv2q0qPrv_2b=1NEM5g@mail.gmail.com","threadId":"33984","inReplyTo":"CAJ-05NPQLVFhtb9KMLNLc5MqguBYM1=gKEVrrtT3kSMiZKma_g@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-05-30T11:33:27Z","receivedAt":"2013-05-30T11:33:27Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Alex Bennée wrote:\n>>time /usr/bin/git --no-pager\n> traversed 223 commits\n>\n> real    0m4.817s\n> user    0m4.320s\n> sys     0m0.464s\n\nI'm quite clueless about why it is taking this long: I think it's IO\nbecause there's nothing to compute?  I really can't trace anything\nunless you can reproduce it on a public repository.  On linux.git with\nmy rotating hard disk:\n\n$ time git describe --debug --long --tags HEAD~10000\nsearching to describe HEAD~10000\n annotated       5445 v2.6.33\n annotated       5660 v2.6.33-rc8\n annotated       5884 v2.6.33-rc7\n annotated       6140 v2.6.33-rc6\n annotated       6467 v2.6.33-rc5\n annotated       6999 v2.6.33-rc4\n annotated       7430 v2.6.33-rc3\n annotated       7746 v2.6.33-rc2\n annotated       8212 v2.6.33-rc1\n annotated      13854 v2.6.32\ntraversed 18895 commits\nmore than 10 tags found; listed 10 most recent\ngave up search at 648f4e3e50c4793d9dbf9a09afa193631f76fa26\nv2.6.33-5445-ge7c84ee\n\nreal    0m0.509s\nuser    0m0.470s\nsys     0m0.037s\n\n18k+ commits traversed in half a second here, so I really don't know\nwhat is going on.\n"},{"id":"218976","messageId":"20130530114808.GD17475@serenity.lan","threadId":"33984","inReplyTo":"CAJ-05NPQLVFhtb9KMLNLc5MqguBYM1=gKEVrrtT3kSMiZKma_g@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2013-05-30T11:48:08Z","receivedAt":"2013-05-30T11:48:08Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Thu, May 30, 2013 at 11:38:32AM +0100, Alex Bennée wrote:\n> One factor might be the size of my repo (.git is around 2.4G). Could\n> this just be due to computational cost of searching through large\n> packs to walk the commit chain? Is there any way to make this easier\n> for git to do?\n\nWhat does \"git count-objects -v\" say for your repository?\n\nYou may find that performance improves if you repack with \"git gc\n--aggressive\".\n"},{"id":"218986","messageId":"CAJ-05NM9EhikDBP0izqWrnLbZW6RcHq_cH-20YTE08SZw5fjqA@mail.gmail.com","threadId":"33984","inReplyTo":"20130530114808.GD17475@serenity.lan","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T12:29:02Z","receivedAt":"2013-05-30T12:29:02Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"The repo is a fairly hairy one as it includes two historically\nun-related but content related repos which I'm the process of\ncherry-picking stuff across.\n\n11:58 ajb@sloy/x86_64 [work.git] >git count-objects -v\ncount: 493\nsize: 4572\nin-pack: 399307\npacks: 1\nsize-pack: 1930755\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0\n\nThis was after a repack which did have slight negative effect on\nperformance. The pack file is:\n\n13:27 ajb@sloy/x86_64 [work.git] >ls -lh ./.git/objects/pack/*\n-r--r--r-- 1 ajb cvs  11M May 30 11:56\n./.git/objects/pack/pack-a9ba133a6f25ffa74c3c407e09ab030f8745b201.idx\n-r--r--r-- 1 ajb cvs 1.9G May 30 11:56\n./.git/objects/pack/pack-a9ba133a6f25ffa74c3c407e09ab030f8745b201.pack\n\nI ran perf on it and the top items in the report where:\n\n 41.58%   git  libcrypto.so.1.0.0  [.] 0x6ae73\n 33.96%   git  libz.so.1.2.3.4     [.] 0xe0ec\n 10.39%   git  libz.so.1.2.3.4     [.] adler32\n  2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n\nSo I'm guessing it's spending a lot of non-cache efficient time\nun-packing and processing the deltas?\n\n--\nAlex.\n\nOn 30 May 2013 12:48, John Keeping <john@keeping.me.uk> wrote:\n> On Thu, May 30, 2013 at 11:38:32AM +0100, Alex Bennée wrote:\n>> One factor might be the size of my repo (.git is around 2.4G). Could\n>> this just be due to computational cost of searching through large\n>> packs to walk the commit chain? Is there any way to make this easier\n>> for git to do?\n>\n> What does \"git count-objects -v\" say for your repository?\n>\n> You may find that performance improves if you repack with \"git gc\n> --aggressive\".\n\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"218992","messageId":"CAJ-05NNAeLUfyk8+NU8PmjKqfTcZ1NT_NPAk3M1QROtzsQKJ8g@mail.gmail.com","threadId":"33984","inReplyTo":"CALkWK0ndKMZRuWgdg6djqPUGxbDAqZPcv2q0qPrv_2b=1NEM5g@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T13:09:42Z","receivedAt":"2013-05-30T13:09:42Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"It looks like it's a file caching effect combined with my repo being\nmore pathalogical in size and contents. Note run 1 (cold) vs run 2 on\nthe linux file tree:\n\n13:52 ajb@sloy/x86_64 [linux.git] >time git describe --debug --long\n--tags HEAD~10000\nsearching to describe HEAD~10000\n annotated         57 v2.6.34-rc2\n annotated       1688 v2.6.34-rc1\n annotated       7932 v2.6.33\n annotated       8157 v2.6.33-rc8\n annotated       8381 v2.6.33-rc7\n annotated       8637 v2.6.33-rc6\n annotated       8964 v2.6.33-rc5\n annotated       9493 v2.6.33-rc4\n annotated       9912 v2.6.33-rc3\n annotated      10202 v2.6.33-rc2\ntraversed 10547 commits\nmore than 10 tags found; listed 10 most recent\ngave up search at 55639353a0035052d9ea6cfe4dde0ac7fcbb2c9f\nv2.6.34-rc2-57-gef5da59\n\nreal    0m7.332s\nuser    0m0.308s\nsys     0m0.244s\n14:03 ajb@sloy/x86_64 [linux.git] >time git describe --debug --long\n--tags HEAD~10000\nsearching to describe HEAD~10000\n annotated         57 v2.6.34-rc2\n annotated       1688 v2.6.34-rc1\n annotated       7932 v2.6.33\n annotated       8157 v2.6.33-rc8\n annotated       8381 v2.6.33-rc7\n annotated       8637 v2.6.33-rc6\n annotated       8964 v2.6.33-rc5\n annotated       9493 v2.6.33-rc4\n annotated       9912 v2.6.33-rc3\n annotated      10202 v2.6.33-rc2\ntraversed 10547 commits\nmore than 10 tags found; listed 10 most recent\ngave up search at 55639353a0035052d9ea6cfe4dde0ac7fcbb2c9f\nv2.6.34-rc2-57-gef5da59\n\nreal    0m0.298s\nuser    0m0.244s\nsys     0m0.036s\n\nAlthough the perf profile looks subtly different.\n\nFirst through the linux tree:\n\n 22.35%   git  libz.so.1.2.3.4    [.] inflate\n 18.56%   git  libz.so.1.2.3.4    [.] inflate_fast\n 17.48%   git  libz.so.1.2.3.4    [.] inflate_table\n  7.84%   git  git                [.] hashcmp\n  3.93%   git  git                [.] get_sha1_hex\n  3.46%   git  libz.so.1.2.3.4    [.] adler32\n\nAnd through my \"special\" repo:\n\n 41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n 33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n 10.39%   git  libz.so.1.2.3.4     [.] adler32\n  2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n\n I'm not sure why libcrypto features so highly in the results\n\n\n --\n Alex.\n\nOn 30 May 2013 12:33, Ramkumar Ramachandra <artagnon@gmail.com> wrote:\n> Alex Bennée wrote:\n>>>time /usr/bin/git --no-pager\n>> traversed 223 commits\n>>\n>> real    0m4.817s\n>> user    0m4.320s\n>> sys     0m0.464s\n>\n> I'm quite clueless about why it is taking this long: I think it's IO\n> because there's nothing to compute?  I really can't trace anything\n> unless you can reproduce it on a public repository.  On linux.git with\n> my rotating hard disk:\n>\n> $ time git describe --debug --long --tags HEAD~10000\n> searching to describe HEAD~10000\n>  annotated       5445 v2.6.33\n>  annotated       5660 v2.6.33-rc8\n>  annotated       5884 v2.6.33-rc7\n>  annotated       6140 v2.6.33-rc6\n>  annotated       6467 v2.6.33-rc5\n>  annotated       6999 v2.6.33-rc4\n>  annotated       7430 v2.6.33-rc3\n>  annotated       7746 v2.6.33-rc2\n>  annotated       8212 v2.6.33-rc1\n>  annotated      13854 v2.6.32\n> traversed 18895 commits\n> more than 10 tags found; listed 10 most recent\n> gave up search at 648f4e3e50c4793d9dbf9a09afa193631f76fa26\n> v2.6.33-5445-ge7c84ee\n>\n> real    0m0.509s\n> user    0m0.470s\n> sys     0m0.037s\n>\n> 18k+ commits traversed in half a second here, so I really don't know\n> what is going on.\n\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"218993","messageId":"CAJ-05NOyQu7pfY7jwKTJ2ZS_h9pBtnyeMAxJVzjC7R0kVoLUBw@mail.gmail.com","threadId":"33984","inReplyTo":"20130530114808.GD17475@serenity.lan","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T13:16:48Z","receivedAt":"2013-05-30T13:16:48Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"> You may find that performance improves if you repack with \"git gc\n--aggressive\".\n\nIt seems that increases the time to get to where it wants to:\n\n14:12 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\ndescribe --long --tags --debug\nsearching to describe HEAD\n lightweight       10 ajb-build-test-5224\n lightweight       41 ajb-build-test-5222\n annotated        146 vnms-2-1-36-32\n annotated        155 vnms-2-1-36-31\n annotated        174 vnms-2-1-36-30\n annotated        183 vnms-2-1-36-29\n lightweight      188 vnms-2-1-36-28\n annotated        193 vnms-2-1-36-27\n annotated        206 vnms-2-1-36-26\n annotated        215 vectastar-4-2-83-5\ntraversed 223 commits\nmore than 10 tags found; listed 10 most recent\ngave up search at 2b69df72d47be8440e3ce4cee91b9b7ceaf8b77c\najb-build-test-5224-10-gfa296e6\n\nreal    0m14.658s\nuser    0m12.845s\nsys     0m1.776s\n\nOn 30 May 2013 12:48, John Keeping <john@keeping.me.uk> wrote:\n> On Thu, May 30, 2013 at 11:38:32AM +0100, Alex Bennée wrote:\n>> One factor might be the size of my repo (.git is around 2.4G). Could\n>> this just be due to computational cost of searching through large\n>> packs to walk the commit chain? Is there any way to make this easier\n>> for git to do?\n>\n> What does \"git count-objects -v\" say for your repository?\n>\n> You may find that performance improves if you repack with \"git gc\n> --aggressive\".\n\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"218994","messageId":"CACsJy8A1oEezNeFjXTrQ=+gJ6nxDheFYTU0xtiSRt0aOOvE=Vw@mail.gmail.com","threadId":"33984","inReplyTo":"CAJ-05NM9EhikDBP0izqWrnLbZW6RcHq_cH-20YTE08SZw5fjqA@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-30T13:20:01Z","receivedAt":"2013-05-30T13:20:01Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, May 30, 2013 at 7:29 PM, Alex Bennée <kernel-hacker@bennee.com> wrote:\n> I ran perf on it and the top items in the report where:\n>\n>  41.58%   git  libcrypto.so.1.0.0  [.] 0x6ae73\n>  33.96%   git  libz.so.1.2.3.4     [.] 0xe0ec\n>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n>\n> So I'm guessing it's spending a lot of non-cache efficient time\n> un-packing and processing the deltas?\n\nIf I'm not mistaken, commits are never deltified. They are usually\nsmall and packed close together for better I/O patterns. If you really\njust read hundreds of commits, it can't take that long. Maybe some\ncode paths accidentally open a tree, a blob or something..\n\nCan you try setting core.logpackaccess to a path on and rerun\ndescribe? Jugding from the code (I never actually tried it), it'll\ncreate a file at the given path with the accessed pack offsets. You\ncan check what offset corresponds to what object with verify-pack -v.\n--\nDuy\n"},{"id":"219000","messageId":"CACsJy8AuhbwkjGjeQRYe1XZFsAntNdpKYxBM9aeMwF3HpB16Ow@mail.gmail.com","threadId":"33984","inReplyTo":"CAJ-05NPacjAEC99Ntd9eMnTD9_PMMYFob-_tAx5CeSB79TkRSg@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-30T13:45:45Z","receivedAt":"2013-05-30T13:45:45Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, May 30, 2013 at 8:34 PM, Alex Bennée <kernel-hacker@bennee.com> wrote:\n> From the following run:\n>\n>\n> 14:31 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\n> describe --long --tags\n> ajb-build-test-5224-11-g9660048\n>\n> real    0m14.720s\n> user    0m12.985s\n> sys     0m1.700s\n> 14:31 ajb@sloy/x86_64 [work.git] >wc -l /tmp/log-pack.txt\n> 1610 /tmp/log-pack.txt\n>\n> The pack has been \"tuned\" with a gc --aggressive. Assuming the numbers\n> are offsets into the pack it looks fairly random access until the last\n> 100 or so.\n>\n> [snipped]\n\nThanks. Can you share \"verify-pack -v\" output of\npack-a9ba133a6f25ffa74c3c407e09ab030f8745b201.pack? I think you need\nto put it somewhere on Internet temporarily as it's likely to exceed\ngit@vger limits.\n--\nDuy\n"},{"id":"219010","messageId":"CAJ-05NNB7b5Bgc4B5OUOta+Q08+fgKf_NzuG44z8_teqWgnctQ@mail.gmail.com","threadId":"33984","inReplyTo":"CACsJy8AuhbwkjGjeQRYe1XZFsAntNdpKYxBM9aeMwF3HpB16Ow@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T14:02:50Z","receivedAt":"2013-05-30T14:02:50Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 30 May 2013 14:45, Duy Nguyen <pclouds@gmail.com> wrote:\n> On Thu, May 30, 2013 at 8:34 PM, Alex Bennée <kernel-hacker@bennee.com> wrote:\n> <snip>\n> Thanks. Can you share \"verify-pack -v\" output of\n> pack-a9ba133a6f25ffa74c3c407e09ab030f8745b201.pack? I think you need\n> to put it somewhere on Internet temporarily as it's likely to exceed\n> git@vger limits.\n> --\n> Duy\n\nhttp://www.bennee.com/~alex/stuff/git-pack-access.tar.bz2\n\n--\nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219011","messageId":"CALkWK0=bgM+fYcVEwjHHF8k2Q8wMmjdbM5bxXdPH6s9StDH_Ng@mail.gmail.com","threadId":"33984","inReplyTo":"CAJ-05NNAeLUfyk8+NU8PmjKqfTcZ1NT_NPAk3M1QROtzsQKJ8g@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-05-30T14:32:50Z","receivedAt":"2013-05-30T14:32:50Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Alex Bennée wrote:\n> And through my \"special\" repo:\n>\n>  41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n>  33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n>\n>  I'm not sure why libcrypto features so highly in the results\n\nWhile Duy churns on the delta chain, let me try to make a (rather\ncrude) observation:\n\nWhat does it mean for libcrypto to be so high in your perf report?\nsha1_block_data_order is ultimately by object.c:parse_object.  While\nit indicates that deltas are taking a long time to apply (or are\nsomehow not optimally organized for IO), I think it indicates either:\n\n1. Your history is very deep and there are an unusually high number of\ndeltas for each blob.  What are the total number of commits?\n\n2. You have have huge (binary) files checked into your repository.  Do\nyou?  If so, why isn't the code in streaming.c kicking in?\n"},{"id":"219017","messageId":"CAJ-05NMt6h=JFLLCP+LASKMcToENhF=BSsk1dPML0024hJTwTw@mail.gmail.com","threadId":"33984","inReplyTo":"CALkWK0=bgM+fYcVEwjHHF8k2Q8wMmjdbM5bxXdPH6s9StDH_Ng@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T15:01:58Z","receivedAt":"2013-05-30T15:01:58Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 30 May 2013 15:32, Ramkumar Ramachandra <artagnon@gmail.com> wrote:\n> Alex Bennée wrote:\n>> And through my \"special\" repo:\n>>\n>>  41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n>>  33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n>>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n>>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n>>\n>>  I'm not sure why libcrypto features so highly in the results\n>\n> While Duy churns on the delta chain, let me try to make a (rather\n> crude) observation:\n>\n> What does it mean for libcrypto to be so high in your perf report?\n> sha1_block_data_order is ultimately by object.c:parse_object.  While\n> it indicates that deltas are taking a long time to apply (or are\n> somehow not optimally organized for IO), I think it indicates either:\n>\n> 1. Your history is very deep and there are an unusually high number of\n> deltas for each blob.  What are the total number of commits?\n\nWell the history does en-compose about 10 years of product development\nand has a high number of files in the repo (including about 3 copies of\nthe kernel - sans upstream history).\n\n15:50 ajb@sloy/x86_64 [work.git] >time git log --pretty=oneline | wc -l\n24648\n\nreal    0m0.434s\nuser    0m0.388s\nsys     0m0.112s\n\nAlthough it doesn't take too long to walk the whole mainline history\n(obviously ignoring all the other branches).\n\n15:52 ajb@sloy/x86_64 [work.git] >git count-objects -v -H\ncount: 581\nsize: 5.09 MiB\nin-pack: 399307\npacks: 1\nsize-pack: 1.49 GiB\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0 bytes\n\nIt is a pick repo. The gc --aggressive nearly took out my machine keeping\naround 4gb resident for most of the half hour and using nearly 8gb of VM.\n\nOf course most of the history is not needed for day to day stuff. Maybe\nif I split the pack files up it wouldn't be quite such a strain to work\nthrough them?\n\n> 2. You have have huge (binary) files checked into your repository.  Do\n> you?  If so, why isn't the code in streaming.c kicking in?\n\nWe do have some binary blobs in the repository (mainly DSP and FPGA images)\nalthough not a huge number:\n\n15:58 ajb@sloy/x86_64 [work.git] >time git log --pretty=oneline -- xxx\nxxx/xxxxxx/*.out ./xxx/xxx/*.out ./xxx/xxxxxxx/*.out | wc -l\n234\n\nreal    0m0.590s\nuser    0m0.552s\nsys     0m0.040s\n\nHow can I tell if streaming is kicking in or now?\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219020","messageId":"CALkWK0=6xg8X7ZLXoaTk_ZgRnhgXWsftsTzr6JAWFBJvmVOgFw@mail.gmail.com","threadId":"33984","inReplyTo":"CAJ-05NMt6h=JFLLCP+LASKMcToENhF=BSsk1dPML0024hJTwTw@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2013-05-30T15:17:35Z","receivedAt":"2013-05-30T15:17:35Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Alex Bennée wrote:\n> 15:50 ajb@sloy/x86_64 [work.git] >time git log --pretty=oneline | wc -l\n> 24648\n>\n> real    0m0.434s\n> user    0m0.388s\n> sys     0m0.112s\n>\n> Although it doesn't take too long to walk the whole mainline history\n> (obviously ignoring all the other branches).\n\nDamn, non-starter.  linux.git has 361k+ commits in mainline history.\n\nNit: use git rev-list --count HEAD next time.\n\n> 15:52 ajb@sloy/x86_64 [work.git] >git count-objects -v -H\n> count: 581\n> size: 5.09 MiB\n> in-pack: 399307\n> packs: 1\n> size-pack: 1.49 GiB\n> prune-packable: 0\n> garbage: 0\n> size-garbage: 0 bytes\n\nlinux.git has 2.9m+ in-pack.  The pack-size is much lower at about\n800+ MiB, but I don't think 1.49 GiB is a problem in itself.  Looking\nforward to your big-files report to see why it's so big.\n\n> It is a pick repo. The gc --aggressive nearly took out my machine keeping\n> around 4gb resident for most of the half hour and using nearly 8gb of VM.\n>\n> Of course most of the history is not needed for day to day stuff. Maybe\n> if I split the pack files up it wouldn't be quite such a strain to work\n> through them?\n\nReally out of my depth here, sorry.  Let's see what Duy (or the\nothers) have to say.\n\n>> 2. You have have huge (binary) files checked into your repository.  Do\n>> you?  If so, why isn't the code in streaming.c kicking in?\n>\n> We do have some binary blobs in the repository (mainly DSP and FPGA images)\n> although not a huge number:\n>\n> 15:58 ajb@sloy/x86_64 [work.git] >time git log --pretty=oneline -- xxx\n> xxx/xxxxxx/*.out ./xxx/xxx/*.out ./xxx/xxxxxxx/*.out | wc -l\n> 234\n>\n> real    0m0.590s\n> user    0m0.552s\n> sys     0m0.040s\n\nlog is streaming, and is not a good measure: it doesn't even walk the\nentire commit graph.  How big are these files?\n\n> How can I tell if streaming is kicking in or now?\n\nI use callgrind (and kcachegrind to visualize).  Can you post\ncallgrind output?  It will be helpful in figuring out where exactly\ngit is spending time.\n"},{"id":"219023","messageId":"87ehcoeb3t.fsf@linux-k42r.v.cablecom.net","threadId":"33984","inReplyTo":"CAJ-05NNAeLUfyk8+NU8PmjKqfTcZ1NT_NPAk3M1QROtzsQKJ8g@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-05-30T15:33:26Z","receivedAt":"2013-05-30T15:33:26Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Alex Bennée <kernel-hacker@bennee.com> writes:\n\n>  41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n>  33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n\nDo you have any large blobs in the repo that are referenced directly by\na tag?\n\nBecause this just so happens to exactly reproduce your symptoms:\n\n  # in a random git.git\n  $ time git describe --debug\n  [...]\n  real    0m0.390s\n  user    0m0.037s\n  sys     0m0.011s\n  $ git tag big1 $(dd if=/dev/urandom bs=1M count=512 | git hash-object -w --stdin)\n  512+0 records in\n  512+0 records out\n  536870912 bytes (537 MB) copied, 45.5088 s, 11.8 MB/s\n  $ time git describe --debug\n  [...]\n  real    0m1.875s\n  user    0m1.738s\n  sys     0m0.129s\n  $ git tag big2 $(dd if=/dev/urandom bs=1M count=512 | git hash-object -w --stdin)\n  512+0 records in\n  512+0 records out\n  536870912 bytes (537 MB) copied, 44.972 s, 11.9 MB/s\n  $ time git describe --debugsuche zur Beschreibung von HEAD\n  [...]\n  real    0m3.620s\n  user    0m3.357s\n  sys     0m0.248s\n\n(I actually ran the git-describe invocations more than once to ensure\nthat they are again cache-hot.)\n\ngit-describe should probably be fixed to avoid loading blobs, though I'm\nnot sure off hand if we have any infrastructure to infer the type of a\nloose object without inflating it.  (This could probably be added by\ninflating only the first block.)  We do have this for packed objects, so\nat least for packed repos there's a speedup to be had.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"219026","messageId":"CAJ-05NOjVhb+3Cab7uQE8K3VE0Q2GhqR3FE=WzJZvSn8Djt6tw@mail.gmail.com","threadId":"33984","inReplyTo":"87ehcoeb3t.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-30T16:01:58Z","receivedAt":"2013-05-30T16:01:58Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n> Alex Bennée <kernel-hacker@bennee.com> writes:\n>\n>>  41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n>>  33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n>>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n>>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n>\n> Do you have any large blobs in the repo that are referenced directly by\n> a tag?\n\nMost probably. I've certainly done a bunch of releases (which are tagged) were\nthe last thing that was updated was an FPGA image.\n\n> Because this just so happens to exactly reproduce your symptoms:\n>\n>   # in a random git.git\n>   $ time git describe --debug\n>   [...]\n>   real    0m0.390s\n>   user    0m0.037s\n>   sys     0m0.011s\n>   $ git tag big1 $(dd if=/dev/urandom bs=1M count=512 | git hash-object -w --stdin)\n>   512+0 records in\n>   512+0 records out\n>   536870912 bytes (537 MB) copied, 45.5088 s, 11.8 MB/s\n>   $ time git describe --debug\n>   [...]\n>   real    0m1.875s\n>   user    0m1.738s\n>   sys     0m0.129s\n>   $ git tag big2 $(dd if=/dev/urandom bs=1M count=512 | git hash-object -w --stdin)\n>   512+0 records in\n>   512+0 records out\n>   536870912 bytes (537 MB) copied, 44.972 s, 11.9 MB/s\n>   $ time git describe --debugsuche zur Beschreibung von HEAD\n>   [...]\n>   real    0m3.620s\n>   user    0m3.357s\n>   sys     0m0.248s\n>\n> (I actually ran the git-describe invocations more than once to ensure\n> that they are again cache-hot.)\n\nThat looks pretty promising as a replication.\n\n> git-describe should probably be fixed to avoid loading blobs, though I'm\n> not sure off hand if we have any infrastructure to infer the type of a\n> loose object without inflating it.  (This could probably be added by\n> inflating only the first block.)  We do have this for packed objects, so\n> at least for packed repos there's a speedup to be had.\n\nWill it be loading the blob for every commit it traverses or just ones that hit\na tag? Why does it need to load the blob at all? Surely the commit\ntree state doesn't\nneed to be walked down?\n\n>\n> --\n> Thomas Rast\n> trast@{inf,student}.ethz.ch\n\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219027","messageId":"87ip20bfq4.fsf@linux-k42r.v.cablecom.net","threadId":"33984","inReplyTo":"CAJ-05NOjVhb+3Cab7uQE8K3VE0Q2GhqR3FE=WzJZvSn8Djt6tw@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-05-30T16:21:55Z","receivedAt":"2013-05-30T16:21:55Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Alex Bennée <kernel-hacker@bennee.com> writes:\n\n> On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n>> Alex Bennée <kernel-hacker@bennee.com> writes:\n>>\n>>>  41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n>>>  33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n>>>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n>>>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n>>\n>> Do you have any large blobs in the repo that are referenced directly by\n>> a tag?\n>\n> Most probably. I've certainly done a bunch of releases (which are tagged) were\n> the last thing that was updated was an FPGA image.\n[...]\n>> git-describe should probably be fixed to avoid loading blobs, though I'm\n>> not sure off hand if we have any infrastructure to infer the type of a\n>> loose object without inflating it.  (This could probably be added by\n>> inflating only the first block.)  We do have this for packed objects, so\n>> at least for packed repos there's a speedup to be had.\n>\n> Will it be loading the blob for every commit it traverses or just ones that hit\n> a tag? Why does it need to load the blob at all? Surely the commit\n> tree state doesn't\n> need to be walked down?\n\nNo, my theory is that you tagged *the blobs*.  Git supports this.\n\ngit-describe needs to look at the commit (if any) obtained by peeling\neach tag (i.e. dereferencing tags until it reaches a non-tag).  So to do\nthat, it resolves the tag's referent and loads it.  Usually this will be\na commit, in which case it is marked as reached by the tag.\n\nAs my example shows, it also resolves tags' referents if they refer to\nnon-commits, in particular, it will decompress large blobs that are\n(directly) referenced by a tag.\n\nNote that while annotated tags provide the type information themselves,\ne.g.\n\n  $ git cat-file tag junio-gpg-pub\n  object fe113d3f96636710600c6b02d5fd421fa7e87dd6\n  type blob\n  tag junio-gpg-pub\n  [...]\n\nunannotated tags are simply refs, so it is not enough to just look at\nthe tag objects' referent type.\n\nI had a brief look around sha1_file.c, in particular sha1_object_info,\nand it turns out we lack the \"deflate only early part\" logic as I\nsuspected.  So that'll have to be fixed first.  After that I *think* it\nshould automatically carry over into the tag readers.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"219031","messageId":"87bo7sbeoc.fsf@linux-k42r.v.cablecom.net","threadId":"33984","inReplyTo":"87ip20bfq4.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-05-30T16:44:35Z","receivedAt":"2013-05-30T16:44:35Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Thomas Rast <trast@inf.ethz.ch> writes:\n\n> I had a brief look around sha1_file.c, in particular sha1_object_info,\n> and it turns out we lack the \"deflate only early part\" logic as I\n> suspected.  So that'll have to be fixed first.  After that I *think* it\n> should automatically carry over into the tag readers.\n\nStrike that, I'm wrong.  sha1_object_info is fast even for these big\nloose objects.\n\nThe culprit, according to some callgrind investigation, is\nlookup_commit_reference_gently() [for the unannotated case] or\nderef_tag() [annotated case] calling parse_object().\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"219037","messageId":"CALWbr2xd4k_P6KUQOcRJotWj3=DfJNnhL4rsGSw0yE+53gdyWw@mail.gmail.com","threadId":"33984","inReplyTo":"87bo7sbeoc.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Antoine Pelisse","fromEmail":"apelisse@gmail.com","sentAt":"2013-05-30T19:01:21Z","receivedAt":"2013-05-30T19:01:21Z","isPatch":false,"sender":{"key":"apelisse@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1929644?v=4"},"body":"> The culprit, according to some callgrind investigation, is\n> lookup_commit_reference_gently() [for the unannotated case] or\n> deref_tag() [annotated case] calling parse_object().\n\nUsing the scenario you described earlier, I think it ends-up spending\nmost of its time in check_sha1_signature (both deref_tag and\nlookup_commit_reference_gently() go there) with 20% inflating, 80% in\nSHA1_Update(). Not much we can do about that, can we ?\n"},{"id":"219041","messageId":"20130530193046.GG17475@serenity.lan","threadId":"33984","inReplyTo":"87ip20bfq4.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2013-05-30T19:30:46Z","receivedAt":"2013-05-30T19:30:46Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Thu, May 30, 2013 at 06:21:55PM +0200, Thomas Rast wrote:\n> Alex Bennée <kernel-hacker@bennee.com> writes:\n> \n> > On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n> >>\n> >>>  41.58%   git  libcrypto.so.1.0.0  [.] sha1_block_data_order_ssse3\n> >>>  33.62%   git  libz.so.1.2.3.4     [.] inflate_fast\n> >>>  10.39%   git  libz.so.1.2.3.4     [.] adler32\n> >>>   2.03%   git  [kernel.kallsyms]   [k] clear_page_c\n> >>\n> >> Do you have any large blobs in the repo that are referenced directly by\n> >> a tag?\n> >\n> > Most probably. I've certainly done a bunch of releases (which are tagged) were\n> > the last thing that was updated was an FPGA image.\n> [...]\n> >> git-describe should probably be fixed to avoid loading blobs, though I'm\n> >> not sure off hand if we have any infrastructure to infer the type of a\n> >> loose object without inflating it.  (This could probably be added by\n> >> inflating only the first block.)  We do have this for packed objects, so\n> >> at least for packed repos there's a speedup to be had.\n> >\n> > Will it be loading the blob for every commit it traverses or just ones that hit\n> > a tag? Why does it need to load the blob at all? Surely the commit\n> > tree state doesn't\n> > need to be walked down?\n> \n> No, my theory is that you tagged *the blobs*.  Git supports this.\n\nYou can see if that is the case by doing something like this:\n\n    eval $(git for-each-ref --shell --format '\n        test $(git cat-file -t %(objectname)^{}) = commit ||\n        echo %(refname);')\n\nThat will print out the name of any ref that doesn't point at a commit.\n"},{"id":"219084","messageId":"CAJ-05NOEuxOVy7LFp_XRa_08G-Mj0x7q+RiR=u71-iyfOXpHow@mail.gmail.com","threadId":"33984","inReplyTo":"20130530193046.GG17475@serenity.lan","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-31T08:14:49Z","receivedAt":"2013-05-31T08:14:49Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 30 May 2013 20:30, John Keeping <john@keeping.me.uk> wrote:\n> On Thu, May 30, 2013 at 06:21:55PM +0200, Thomas Rast wrote:\n>> Alex Bennée <kernel-hacker@bennee.com> writes:\n>>\n>> > On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n>> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n> <snip>\n>> > Will it be loading the blob for every commit it traverses or just ones that hit\n>> > a tag? Why does it need to load the blob at all? Surely the commit\n>> > tree state doesn't\n>> > need to be walked down?\n>>\n>> No, my theory is that you tagged *the blobs*.  Git supports this.\n\nWait is this the difference between annotated and non-annotated tags?\nI thought a non-annotated just acted like references to a particular\ntree state?\n\n>\n> You can see if that is the case by doing something like this:\n>\n>     eval $(git for-each-ref --shell --format '\n>         test $(git cat-file -t %(objectname)^{}) = commit ||\n>         echo %(refname);')\n>\n> That will print out the name of any ref that doesn't point at a\n> commit.\n\nHmm that didn't seem to work. But looking at the output by hand I\ncertainly have a mix of tags that are commits vs tags:\n\n\n09:08 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n| grep \"commit\" | wc -l\n1345\n09:12 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n| grep -v \"commit\" | wc -l\n66\n\nUnfortunately I can't just delete those tags as they do refer to known\nreleases which we obviously care about. If I delete the tags on my\nlocal repo and test for a speed increase can I re-create them as\nannotated tag objects?\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219087","messageId":"87obbr5zg3.fsf@linux-k42r.v.cablecom.net","threadId":"33984","inReplyTo":"CAJ-05NOEuxOVy7LFp_XRa_08G-Mj0x7q+RiR=u71-iyfOXpHow@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-05-31T08:24:44Z","receivedAt":"2013-05-31T08:24:44Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Alex Bennée <kernel-hacker@bennee.com> writes:\n\n> On 30 May 2013 20:30, John Keeping <john@keeping.me.uk> wrote:\n>> On Thu, May 30, 2013 at 06:21:55PM +0200, Thomas Rast wrote:\n>>> Alex Bennée <kernel-hacker@bennee.com> writes:\n>>>\n>>> > On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n>>> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n>> <snip>\n>>> > Will it be loading the blob for every commit it traverses or just ones that hit\n>>> > a tag? Why does it need to load the blob at all? Surely the commit\n>>> > tree state doesn't\n>>> > need to be walked down?\n>>>\n>>> No, my theory is that you tagged *the blobs*.  Git supports this.\n>\n> Wait is this the difference between annotated and non-annotated tags?\n> I thought a non-annotated just acted like references to a particular\n> tree state?\n\nA tag is just a ref.  It can point at anything, in particular also a\nblob (= some file *contents*).\n\nAn annotated tag is just a tag pointing at a \"tag object\".  A tag object\ncontains tagger name/email/date, a reference to an object, and a tag\nmessage.\n\nThe slowness I found relates to having tags that point at blobs directly\n(unannotated).\n\n>> You can see if that is the case by doing something like this:\n>>\n>>     eval $(git for-each-ref --shell --format '\n>>         test $(git cat-file -t %(objectname)^{}) = commit ||\n>>         echo %(refname);')\n>>\n>> That will print out the name of any ref that doesn't point at a\n>> commit.\n>\n> Hmm that didn't seem to work. But looking at the output by hand I\n> certainly have a mix of tags that are commits vs tags:\n>\n>\n> 09:08 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n> | grep \"commit\" | wc -l\n> 1345\n> 09:12 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n> | grep -v \"commit\" | wc -l\n> 66\n>\n> Unfortunately I can't just delete those tags as they do refer to known\n> releases which we obviously care about. If I delete the tags on my\n> local repo and test for a speed increase can I re-create them as\n> annotated tag objects?\n\nI would be more interested in this:\n\n  git for-each-ref | grep ' blob'\n\nand\n\n  (git for-each-ref | grep ' blob' | cut -d\\  -f1 | xargs -n1 git cat-file blob) | wc -c\n\nThe first tells you if you have any refs pointing at blobs.  The second\ncomputes their total unpacked size.  My theory is that the second yields\nsome large number (hundreds of megabytes at least).\n\nIt would be nice if you checked, because if there turn out to be big\nblobs, we have all the pieces and just need to assemble the best\nsolution.  Otherwise, there's something else going on and the problem\nremains open.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"219090","messageId":"20130531083252.GA1072@serenity.lan","threadId":"33984","inReplyTo":"CAJ-05NOEuxOVy7LFp_XRa_08G-Mj0x7q+RiR=u71-iyfOXpHow@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2013-05-31T08:32:52Z","receivedAt":"2013-05-31T08:32:52Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Fri, May 31, 2013 at 09:14:49AM +0100, Alex Bennée wrote:\n> On 30 May 2013 20:30, John Keeping <john@keeping.me.uk> wrote:\n> > On Thu, May 30, 2013 at 06:21:55PM +0200, Thomas Rast wrote:\n> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n> >>\n> >> > On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n> >> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n> > <snip>\n> >> > Will it be loading the blob for every commit it traverses or just ones that hit\n> >> > a tag? Why does it need to load the blob at all? Surely the commit\n> >> > tree state doesn't\n> >> > need to be walked down?\n> >>\n> >> No, my theory is that you tagged *the blobs*.  Git supports this.\n> \n> Wait is this the difference between annotated and non-annotated tags?\n> I thought a non-annotated just acted like references to a particular\n> tree state?\n\nNo, this is something slightly different.  In Git there are four types\nof object: tag, commit, tree and blob.  When you have a heavyweight tag,\nthe tag reference points at a tag object (which in turn points at\nanother object).  With a lightweight tag, the tag reference typically\npoints at a commit object.\n\nHowever, there is no restriction that says that a tag object must point\nto a commit or that a lightweight tag must point at a commit - it is\nequally possible to point directly at a tree or a blob (although a lot\nless common).\n\nThomas is suggesting that you might have a tag that does not point at a\ncommit but instead points to a blob object.\n\n> > You can see if that is the case by doing something like this:\n> >\n> >     eval $(git for-each-ref --shell --format '\n> >         test $(git cat-file -t %(objectname)^{}) = commit ||\n> >         echo %(refname);')\n> >\n> > That will print out the name of any ref that doesn't point at a\n> > commit.\n> \n> Hmm that didn't seem to work.\n\nYou mean there was no output?  In that case it's likely that all your\nreferences do indeed point at commits.\n\n>                               But looking at the output by hand I\n> certainly have a mix of tags that are commits vs tags:\n> \n> \n> 09:08 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n> | grep \"commit\" | wc -l\n> 1345\n> 09:12 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n> | grep -v \"commit\" | wc -l\n> 66\n\nThis means that you have 1345 lightweight tags and 66 heavyweight tags,\nassuming that all of the lines that don't say \"commit\" do say \"tag\".\n\nBy the way, I don't remember if you said which version of Git you're\nusing.  If it's an older version then it's possible that something has\nchanged.\n"},{"id":"219091","messageId":"CAJ-05NOdg5TvjzEMrXaPgogU5z5W6kywZhD-82eTUmvE9Hp=Lw@mail.gmail.com","threadId":"33984","inReplyTo":"87obbr5zg3.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-31T08:40:01Z","receivedAt":"2013-05-31T08:40:01Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 31 May 2013 09:24, Thomas Rast <trast@inf.ethz.ch> wrote:\n> Alex Bennée <kernel-hacker@bennee.com> writes:\n>> On 30 May 2013 20:30, John Keeping <john@keeping.me.uk> wrote:\n>>> On Thu, May 30, 2013 at 06:21:55PM +0200, Thomas Rast wrote:\n>>>> Alex Bennée <kernel-hacker@bennee.com> writes:\n>>>> > On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n> <snip>\n>>>> No, my theory is that you tagged *the blobs*.  Git supports this.\n>>\n>> Wait is this the difference between annotated and non-annotated tags?\n>> I thought a non-annotated just acted like references to a particular\n>> tree state?\n>\n> A tag is just a ref.  It can point at anything, in particular also a\n> blob (= some file *contents*).\n>\n> An annotated tag is just a tag pointing at a \"tag object\".  A tag object\n> contains tagger name/email/date, a reference to an object, and a tag\n> message.\n>\n> The slowness I found relates to having tags that point at blobs directly\n> (unannotated).\n\nI think you are right. I was brave (well I assumed the tags would come\nback from the upstream repo) and ran:\n\ngit for-each-ref | grep \"refs/tags\" | grep \"commit\" | cut -d '/' -f 3\n| xargs git tag -d\n\nAnd boom:\n\n09:19 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\ndescribe --long --tags\najb-build-test-5225-2-gdc0b771\n\nreal    0m0.009s\nuser    0m0.008s\nsys     0m0.000s\n\nWhich is much better performance. So it does look like unannotated\ntags pointing at binary blobs is the failure case.\n\n<snip>\n>\n> I would be more interested in this:\n>\n>   git for-each-ref | grep ' blob'\n\nHmmm that gives nothing. All the refs are either tag or commit\n\n> and\n>\n>   (git for-each-ref | grep ' blob' | cut -d\\  -f1 | xargs -n1 git\n>cat-file blob) | wc -c\n\nHowever I have some big commits it seems:\n\n09:37 ajb@sloy/x86_64 [work.git] >(git for-each-ref | grep ' commit' |\ncut -d\\  -f1 | xargs -n1 git cat-file commit) | wc -c\n1147231984\n\n>\n> The first tells you if you have any refs pointing at blobs.  The second\n> computes their total unpacked size.  My theory is that the second yields\n> some large number (hundreds of megabytes at least).\n>\n> It would be nice if you checked, because if there turn out to be big\n> blobs, we have all the pieces and just need to assemble the best\n> solution.  Otherwise, there's something else going on and the problem\n> remains open.\n\nIf you want any other numbers I'm only too happy to help. Sorry I\ncan't share the repo though...\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219092","messageId":"87y5av4jvj.fsf@linux-k42r.v.cablecom.net","threadId":"33984","inReplyTo":"CAJ-05NOdg5TvjzEMrXaPgogU5z5W6kywZhD-82eTUmvE9Hp=Lw@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-05-31T08:46:24Z","receivedAt":"2013-05-31T08:46:24Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Alex Bennée <kernel-hacker@bennee.com> writes:\n\n> I think you are right. I was brave (well I assumed the tags would come\n> back from the upstream repo) and ran:\n>\n> git for-each-ref | grep \"refs/tags\" | grep \"commit\" | cut -d '/' -f 3\n> | xargs git tag -d\n\nSo that deleted all unannotated tags pointing at commits, and then it\nwas fast.  Curious.\n\n> However I have some big commits it seems:\n>\n> 09:37 ajb@sloy/x86_64 [work.git] >(git for-each-ref | grep ' commit' |\n> cut -d\\  -f1 | xargs -n1 git cat-file commit) | wc -c\n> 1147231984\n\nHow many unique entries are there in that list, i.e., what does\n\n  git for-each-ref | grep ' commit' | cut -d\\  -f1 | sort -u | wc -l\n\nsay?  Perhaps you can also find the biggest commit, e.g. like so:\n\n  git for-each-ref | grep ' commit' | cut -d\\  -f1 |\n  while read sha; do git cat-file commit $sha | wc -c; done |\n  sort -n\n\nHowever, if that turns out to be the culprit, it's not fixable\ncurrently[1].  Having commits with insanely long messages is just, well,\ninsane.\n\n\n[1]  unless we do a major rework of the loading infrastructure, so that\nwe can teach it to load only the beginning of a commit as long as we are\nonly interested in parents and such\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"219093","messageId":"CAJ-05NNgVPukJchskVv9oL7-9p+txC0g_SXfHne-mwF327Q_3Q@mail.gmail.com","threadId":"33984","inReplyTo":"20130531083252.GA1072@serenity.lan","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-31T08:49:57Z","receivedAt":"2013-05-31T08:49:57Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 31 May 2013 09:32, John Keeping <john@keeping.me.uk> wrote:\n> On Fri, May 31, 2013 at 09:14:49AM +0100, Alex Bennée wrote:\n>> On 30 May 2013 20:30, John Keeping <john@keeping.me.uk> wrote:\n>> > On Thu, May 30, 2013 at 06:21:55PM +0200, Thomas Rast wrote:\n>> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n>> >>\n>> >> > On 30 May 2013 16:33, Thomas Rast <trast@inf.ethz.ch> wrote:\n>> >> >> Alex Bennée <kernel-hacker@bennee.com> writes:\n>> > <snip>\n>> >> > Will it be loading the blob for every commit it traverses or just ones that hit\n>> >> > a tag? Why does it need to load the blob at all? Surely the commit\n>> >> > tree state doesn't\n>> >> > need to be walked down?\n>> >>\n>> >> No, my theory is that you tagged *the blobs*.  Git supports this.\n>>\n>> Wait is this the difference between annotated and non-annotated tags?\n>> I thought a non-annotated just acted like references to a particular\n>> tree state?\n>\n> No, this is something slightly different.  In Git there are four types\n> of object: tag, commit, tree and blob.  When you have a heavyweight tag,\n> the tag reference points at a tag object (which in turn points at\n> another object).  With a lightweight tag, the tag reference typically\n> points at a commit object.\n\nI think this is the case in my repo.\n\n> However, there is no restriction that says that a tag object must point\n> to a commit or that a lightweight tag must point at a commit - it is\n> equally possible to point directly at a tree or a blob (although a lot\n> less common).\n>\n> Thomas is suggesting that you might have a tag that does not point at a\n> commit but instead points to a blob object.\n\nIt's looking like I just have some very heavy commits. One data point\nI probably should have mentioned at the beginning is this was a\nconverted CVS repo and I'm wondering if some of the artifacts that\nintroduced has contributed to this.\n\n>> > You can see if that is the case by doing something like this:\n>> >\n>> >     eval $(git for-each-ref --shell --format '\n>> >         test $(git cat-file -t %(objectname)^{}) = commit ||\n>> >         echo %(refname);')\n>> >\n>> > That will print out the name of any ref that doesn't point at a\n>> > commit.\n>>\n>> Hmm that didn't seem to work.\n>\n> You mean there was no output?  In that case it's likely that all your\n> references do indeed point at commits.\n\nCorrect.\n\n>\n>>                               But looking at the output by hand I\n>> certainly have a mix of tags that are commits vs tags:\n>>\n>>\n>> 09:08 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n>> | grep \"commit\" | wc -l\n>> 1345\n>> 09:12 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep \"refs/tags\"\n>> | grep -v \"commit\" | wc -l\n>> 66\n>\n> This means that you have 1345 lightweight tags and 66 heavyweight tags,\n> assuming that all of the lines that don't say \"commit\" do say \"tag\".\n\nYep all commits and tags, nothing else\n\n> By the way, I don't remember if you said which version of Git you're\n> using.  If it's an older version then it's possible that something has\n> changed.\n\nI'm running the GIT stable PPA:\n\n09:38 ajb@sloy/x86_64 [work.git] >git --version\ngit version 1.8.3\n\nAlthough I have also tested with the latest git.git maint. I'm happy\nto try master if it's likely to have changed.\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219094","messageId":"20130531085959.GB1072@serenity.lan","threadId":"33984","inReplyTo":"CAJ-05NNgVPukJchskVv9oL7-9p+txC0g_SXfHne-mwF327Q_3Q@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2013-05-31T08:59:59Z","receivedAt":"2013-05-31T08:59:59Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Fri, May 31, 2013 at 09:49:57AM +0100, Alex Bennée wrote:\n> On 31 May 2013 09:32, John Keeping <john@keeping.me.uk> wrote:\n> > Thomas is suggesting that you might have a tag that does not point at a\n> > commit but instead points to a blob object.\n> \n> It's looking like I just have some very heavy commits. One data point\n> I probably should have mentioned at the beginning is this was a\n> converted CVS repo and I'm wondering if some of the artifacts that\n> introduced has contributed to this.\n\nYou can try another for-each-ref invocation to see if that's the case:\n\n    eval $(git for-each-ref --format 'printf \"%s %s\\n\" \\\n        $(git cat-file -s %(objectname)) %(refname);') | sort -n\n\nThat will print the size of each object followed by the ref that points\nto it, sorted by size.\n\n> I'm running the GIT stable PPA:\n> \n> 09:38 ajb@sloy/x86_64 [work.git] >git --version\n> git version 1.8.3\n> \n> Although I have also tested with the latest git.git maint. I'm happy\n> to try master if it's likely to have changed.\n\nmaster's still very close to 1.8.3 at the moment, so I don't think that\nwill make a difference.\n"},{"id":"219095","messageId":"CAJ-05NN8cARpPTnsCfHt3kY6gTnhZ=Vq55EzqxWBV_3ju-oczQ@mail.gmail.com","threadId":"33984","inReplyTo":"87y5av4jvj.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-05-31T09:57:08Z","receivedAt":"2013-05-31T09:57:08Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 31 May 2013 09:46, Thomas Rast <trast@inf.ethz.ch> wrote:\n> Alex Bennée <kernel-hacker@bennee.com> writes:\n>\n>> I think you are right. I was brave (well I assumed the tags would come\n>> back from the upstream repo) and ran:\n>>\n>> git for-each-ref | grep \"refs/tags\" | grep \"commit\" | cut -d '/' -f 3\n>> | xargs git tag -d\n>\n> So that deleted all unannotated tags pointing at commits, and then it\n> was fast.  Curious.\n>\n>> However I have some big commits it seems:\n>>\n>> 09:37 ajb@sloy/x86_64 [work.git] >(git for-each-ref | grep ' commit' |\n>> cut -d\\  -f1 | xargs -n1 git cat-file commit) | wc -c\n>> 1147231984\n>\n> How many unique entries are there in that list, i.e., what does\n>\n>   git for-each-ref | grep ' commit' | cut -d\\  -f1 | sort -u | wc -l\n\n09:49 ajb@sloy/x86_64 [work.git] >git for-each-ref | grep ' commit' |\ncut -d\\  -f1 | sort -u | wc -l\n1508\n\n> say?  Perhaps you can also find the biggest commit, e.g. like so:\n>\n>   git for-each-ref | grep ' commit' | cut -d\\  -f1 |\n>   while read sha; do git cat-file commit $sha | wc -c; done |\n>   sort -n\n\nYeah there is a range from a few hundred bytes to a large number of 3M\ncommits. I guess I need to identify which commits they are and remove\nthe tags or convert them to annotated reference tags.\n\n> However, if that turns out to be the culprit, it's not fixable\n> currently[1].  Having commits with insanely long messages is just, well,\n> insane.\n>\n>\n\n> [1]  unless we do a major rework of the loading infrastructure, so that\n> we can teach it to load only the beginning of a commit as long as we are\n> only interested in parents and such\n\nI'll do a bit of scripting to dig into the nature of these\nuber-commits and try and work out how they cam about. I suspect they\nare simply start of branch states in our broken and disparate history.\n\nI'll get back to you once I've dug a little deeper.\n\n>\n> --\n> Thomas Rast\n> trast@{inf,student}.ethz.ch\n\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219096","messageId":"87txlj30n4.fsf@linux-k42r.v.cablecom.net","threadId":"33984","inReplyTo":"87y5av4jvj.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Thomas Rast","fromEmail":"trast@inf.ethz.ch","sentAt":"2013-05-31T10:27:11Z","receivedAt":"2013-05-31T10:27:11Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Thomas Rast <trast@inf.ethz.ch> writes:\n\n> However, if that turns out to be the culprit, it's not fixable\n> currently[1].  Having commits with insanely long messages is just, well,\n> insane.\n>\n> [1]  unless we do a major rework of the loading infrastructure, so that\n> we can teach it to load only the beginning of a commit as long as we are\n> only interested in parents and such\n\nActually, Peff, doesn't your commit parent/tree pointer caching give us\nthis for free?\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"219104","messageId":"20130531161710.GB1365@sigill.intra.peff.net","threadId":"33984","inReplyTo":"87txlj30n4.fsf@linux-k42r.v.cablecom.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-05-31T16:17:10Z","receivedAt":"2013-05-31T16:17:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, May 31, 2013 at 12:27:11PM +0200, Thomas Rast wrote:\n\n> Thomas Rast <trast@inf.ethz.ch> writes:\n> \n> > However, if that turns out to be the culprit, it's not fixable\n> > currently[1].  Having commits with insanely long messages is just, well,\n> > insane.\n> >\n> > [1]  unless we do a major rework of the loading infrastructure, so that\n> > we can teach it to load only the beginning of a commit as long as we are\n> > only interested in parents and such\n> \n> Actually, Peff, doesn't your commit parent/tree pointer caching give us\n> this for free?\n\nIt does. You can test it from the \"jk/metapacks\" branch at\ngit://github.com/peff/git. After building, you'd need to do:\n\n  $ git gc\n  $ git metapack --all --commits\n\nin the target repository. You can check that it's working because \"git\nrev-list --all --count\" should be an order of magnitude faster. You may\nneed to add \"save_commit_buffer = 0\" in any commands you are checking,\nthough, as the optimization can only kick in if parse_commit does not\nwant to save the buffer as a side effect.\n\nI also looked into trying to just read the beginning part of a commit[1],\nbut it turned out not to be all that much of an improvement.\n\n-Peff\n\n[1] http://article.gmane.org/gmane.comp.version-control.git/212301\n"},{"id":"219213","messageId":"CAJ-05NNgcj_pPer2Tw4HvKkVib7N1ZFo7rZOrR9z8NMV1WHQsQ@mail.gmail.com","threadId":"33984","inReplyTo":"CAJ-05NN8cARpPTnsCfHt3kY6gTnhZ=Vq55EzqxWBV_3ju-oczQ@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-06-03T08:02:55Z","receivedAt":"2013-06-03T08:02:55Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 31 May 2013 10:57, Alex Bennée <kernel-hacker@bennee.com> wrote:\n> On 31 May 2013 09:46, Thomas Rast <trast@inf.ethz.ch> wrote:\n>>\n>> So that deleted all unannotated tags pointing at commits, and then it\n>> was fast.  Curious.\n>>\n>> However, if that turns out to be the culprit, it's not fixable\n>> currently[1].  Having commits with insanely long messages is just, well,\n>> insane.\n>>\n>>\n>> [1]  unless we do a major rework of the loading infrastructure, so that\n>> we can teach it to load only the beginning of a commit as long as we are\n>> only interested in parents and such\n>\n> I'll do a bit of scripting to dig into the nature of these\n> uber-commits and try and work out how they cam about. I suspect they\n> are simply start of branch states in our broken and disparate history.\n>\n> I'll get back to you once I've dug a little deeper.\n\nSo I wrote a little script [1] which I ran to remove all tags that did\nnot exist on any branches:\n\ngit-tag-cleaner.py -d no-branch\n\nAfter a lot of churning:\n\n17:26 ajb@sloy/x86_64 [work.git] >time /usr/bin/git --no-pager\ndescribe --long --tags\najb-build-test-5225-2-gdc0b771\n\nreal    0m0.799s\nuser    0m0.024s\nsys     0m0.052s\n\nSo at least I can fix up my repo. All the big ones look at least as\nthough they were weird cvs2svn creations that exist to represent the\ndetached state of a strange CVS tag from the converted repository.\nHowever it does raise one question.\n\nWhy is git attempting to parse a commit not on the DAG for the branch\nI'm attempting to describe?\n\nAnyway as I have a work around I'm going to do a slightly more\nconservative clean of the repo with my script and move on.\n\n[1] https://github.com/stsquad/git-tag-cleaner\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219214","messageId":"CAJ-05NO2reGkboet1c2kYy0Y7xzkb9K45mTdCLq_AU7dp1OTNw@mail.gmail.com","threadId":"33984","inReplyTo":"20130531161710.GB1365@sigill.intra.peff.net","subject":"Re: Poor performance of git describe in big repos","fromName":"Alex Bennée","fromEmail":"kernel-hacker@bennee.com","sentAt":"2013-06-03T08:39:21Z","receivedAt":"2013-06-03T08:39:21Z","isPatch":false,"sender":{"key":"kernel-hacker@bennee.com","avatar":null},"body":"On 31 May 2013 17:17, Jeff King <peff@peff.net> wrote:\n> On Fri, May 31, 2013 at 12:27:11PM +0200, Thomas Rast wrote:\n>\n>> Thomas Rast <trast@inf.ethz.ch> writes:\n>>\n>> > However, if that turns out to be the culprit, it's not fixable\n>> > currently[1].  Having commits with insanely long messages is just, well,\n>> > insane.\n>> >\n>> > [1]  unless we do a major rework of the loading infrastructure, so that\n>> > we can teach it to load only the beginning of a commit as long as we are\n>> > only interested in parents and such\n>>\n>> Actually, Peff, doesn't your commit parent/tree pointer caching give us\n>> this for free?\n>\n> It does. You can test it from the \"jk/metapacks\" branch at\n> git://github.com/peff/git. After building, you'd need to do:\n>\n>   $ git gc\n>   $ git metapack --all --commits\n>\n> in the target repository. You can check that it's working because \"git\n> rev-list --all --count\" should be an order of magnitude faster. You may\n> need to add \"save_commit_buffer = 0\" in any commands you are checking,\n> though, as the optimization can only kick in if parse_commit does not\n> want to save the buffer as a side effect.\n\nIs this a command line argument? The tools don't seem to think so.\n\nAnyway it seems to make a marginal difference to my case:\n\n09:08 ajb@sloy/x86_64 [work.git] >time git --no-pager describe --long --tags\najb-build-test-5225-2-gdc0b771\n\nreal    0m14.105s\nuser    0m12.409s\nsys     0m1.660s\n09:11 ajb@sloy/x86_64 [work.git] >git gc\nCounting objects: 399436, done.\nDelta compression using up to 4 threads.\nCompressing objects: 100% (110874/110874), done.\nWriting objects: 100% (399436/399436), done.\nTotal 399436 (delta 281538), reused 398357 (delta 280493)\nChecking connectivity: 399436, done.\n09:12 ajb@sloy/x86_64 [work.git] >git metapack --all --commits\n09:13 ajb@sloy/x86_64 [work.git] >time git --no-pager describe --long --tags\najb-build-test-5225-2-gdc0b771\n\nreal    0m12.781s\nuser    0m11.669s\nsys     0m1.080s\n09:32 ajb@sloy/x86_64 [work.git] >time git --no-pager describe --long --tags\najb-build-test-5225-2-gdc0b771\n\nreal    0m12.768s\nuser    0m11.817s\nsys     0m0.908s\n09:33 ajb@sloy/x86_64 [work.git] >time git --no-pager describe --long --tags\najb-build-test-5225-2-gdc0b771\n\nreal    0m12.642s\nuser    0m11.705s\nsys     0m0.904s\n\n\n>\n> I also looked into trying to just read the beginning part of a commit[1],\n> but it turned out not to be all that much of an improvement.\n>\n> -Peff\n>\n> [1] http://article.gmane.org/gmane.comp.version-control.git/212301\n\n\n\n-- \nAlex, homepage: http://www.bennee.com/~alex/\n"},{"id":"219227","messageId":"20130603144907.GA5938@sigill.intra.peff.net","threadId":"33984","inReplyTo":"CAJ-05NO2reGkboet1c2kYy0Y7xzkb9K45mTdCLq_AU7dp1OTNw@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-06-03T14:49:08Z","receivedAt":"2013-06-03T14:49:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 03, 2013 at 09:39:21AM +0100, Alex Bennée wrote:\n\n> > in the target repository. You can check that it's working because \"git\n> > rev-list --all --count\" should be an order of magnitude faster. You may\n> > need to add \"save_commit_buffer = 0\" in any commands you are checking,\n> > though, as the optimization can only kick in if parse_commit does not\n> > want to save the buffer as a side effect.\n> \n> Is this a command line argument? The tools don't seem to think so.\n\nIf you mean the \"save_commit_buffer = 0\", no; I mean you would have to\ninsert it somewhere in builtin/$CMD.c, and then recompile. However,\ngit-describe already has it, so it should work.\n\n> Anyway it seems to make a marginal difference to my case:\n\nI get much better results:\n\n  $ cd linux-2.6\n  $ time git --no-pager describe --long --tags HEAD~800\n  v3.5-6956-gaa0b3b2\n\n  real    0m0.261s\n  user    0m0.248s\n  sys     0m0.012s\n\n  $ git metapack --commits --all\n  $ time git --no-pager describe --long --tags HEAD~800\n  v3.5-6956-gaa0b3b2\n\n  real    0m0.057s\n  user    0m0.032s\n  sys     0m0.024s\n\nwhich implies that your time is being spent elsewhere. That topic\nwouldn't avoid inflating tag objects from disk. Do you have really big\ntag objects (or unannotated tags pointing to blobs)? What does:\n\n  git for-each-ref --format='%(object)' refs/tags |\n  git cat-file --batch-check |\n  sort -k 3nr |\n  head\n\nsay?\n\n-Peff\n"},{"id":"219242","messageId":"7vfvwzi29f.fsf@alter.siamese.dyndns.org","threadId":"33984","inReplyTo":"CAJ-05NNgcj_pPer2Tw4HvKkVib7N1ZFo7rZOrR9z8NMV1WHQsQ@mail.gmail.com","subject":"Re: Poor performance of git describe in big repos","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-03T16:32:12Z","receivedAt":"2013-06-03T16:32:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Alex Bennée <kernel-hacker@bennee.com> writes:\n\n> Why is git attempting to parse a commit not on the DAG for the branch\n> I'm attempting to describe?\n\nI think that is because you need to parse the objects at the tip of\nrefs to see if they are on the DAG in the first place.\n\nIf there weren't any annotated tag, conceivably you could do without\nparsing these objects.  You would:\n\n - First read the refs without parsing anything to learn the object\n   name of the tips of refs;\n\n - Traverse the DAG, starting from the commit and notice when you\n   see commits that are at the tips of refs you learned in the first\n   step, arranging to stop when you found the \"closest\" tip.\n\nBut with annotated tags (and \"git describe\" is designed to be\nprimarily used with them; you would need \"--tags\" option to make it\nnotice unannotated tags), the object name you see sitting at the tip\nwill never appear during the DAG traversal.  You will only see\ncommits from the latter, so you would need to parse the tips to\nlearn what commits they refer to.\n\nAnd of course, \"then parse only annotated tags, without parsing\ncommits\" would not work, because you wouldn't know what the object\nis without looking at it ;-)\n"},{"id":"219258","messageId":"7v1u8jgk55.fsf@alter.siamese.dyndns.org","threadId":"33984","inReplyTo":"7vfvwzi29f.fsf@alter.siamese.dyndns.org","subject":"Re: Poor performance of git describe in big repos","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-03T17:48:54Z","receivedAt":"2013-06-03T17:48:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Alex Bennée <kernel-hacker@bennee.com> writes:\n>\n>> Why is git attempting to parse a commit not on the DAG for the branch\n>> I'm attempting to describe?\n>\n> I think that is because you need to parse the objects at the tip of\n> refs to see if they are on the DAG in the first place.\n>\n> If there weren't any annotated tag, conceivably you could do without\n> parsing these objects.  You would:\n>\n>  - First read the refs without parsing anything to learn the object\n>    name of the tips of refs;\n>\n>  - Traverse the DAG, starting from the commit and notice when you\n>    see commits that are at the tips of refs you learned in the first\n>    step, arranging to stop when you found the \"closest\" tip.\n>\n> But with annotated tags (and \"git describe\" is designed to be\n> primarily used with them; you would need \"--tags\" option to make it\n> notice unannotated tags), the object name you see sitting at the tip\n> will never appear during the DAG traversal.  You will only see\n> commits from the latter, so you would need to parse the tips to\n> learn what commits they refer to.\n>\n> And of course, \"then parse only annotated tags, without parsing\n> commits\" would not work, because you wouldn't know what the object\n> is without looking at it ;-)\n\nHaving said all that, with changes by Peff and Michael Haggerty\naround f85354b5c7b8 (pack_one_ref(): use function peel_entry(),\n2013-04-22), recent Git does not \"parse\" as many refs as it used to,\nonly to figure out what commit an annotated tag points at when your\nrefs are packed, so we may be a lot closer to the optimum than I\nhinted by the above description.\n"}]}