{"thread":{"id":"27589","subject":"Git is not scalable with too many refs/*","startedAt":"2011-06-09T03:44:42Z","lastAt":"2011-10-09T05:43:47Z","messageCount":126,"participants":["NAKAMURA Takumi","Sverre Rabbelier","Jakub Narebski","Shawn Pearce","Stephen Bash","A Large Angry SCM","Jeff King","Andreas Ericsson","Junio C Hamano","Johan Herland","Martin Fick","Thomas Rast","Michael Haggerty","Jens Lehmann","Christian Couder","Julian Phillips","David Michael Barr","David Barr","Nguyen Thai Ngoc Duy","René Scharfe"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"169726","messageId":"BANLkTimnCqaEBVreMhnbRBV3r-r1ZzkFcg@mail.gmail.com","threadId":"27589","inReplyTo":null,"subject":"Git is not scalable with too many refs/*","fromName":"NAKAMURA Takumi","fromEmail":"geek4civic@gmail.com","sentAt":"2011-06-09T03:44:42Z","receivedAt":"2011-06-09T03:44:42Z","isPatch":false,"sender":{"key":"geek4civic@gmail.com","avatar":null},"body":"Hello, Git. It is my 1st post here.\n\nI have tried tagging each commit as \"refs/tags/rXXXXXX\" on git-svn\nrepo locally. (over 100k refs/tags.)\nIndeed, it made something extremely slower, even with packed-refs and\npack objects.\nI gave up, then, to push tags to upstream. (it must be terror) :p\n\nI know it might be crazy in the git way, but it would bring me conveniences.\n(eg. git log --oneline --decorate shows me each svn revision)\nI would like to work for Git to live with many tags.\n\n* Issues as far as I have investigated;\n\n  - git show --decorate is always slow.\n    in decorate.c, every commits are inspected.\n  - git rev-tree --quiet --objects $upstream --not --all spends so much time,\n    even if it is expected to return with 0.\n    As you know, it is used in builtin/fetch.c.\n  - git-upload-pack shows \"all\" refs to me if upstream has too many refs.\n\nI would like to work as below if they were valuable.\n\n  - Get rid of inspecting commits in packed-refs on decorate stuff.\n  - Implement sort-by-hash packed-refs, (not sort-by-name)\n  - Implement more effective pruning --not --all on revision.c.\n  - Think about enhancement of protocol to transfer many refs more effectively.\n\nI am happy to consider the issue, thank you.\n\n...Takumi\n"},{"id":"169733","messageId":"BANLkTinfVNxYX3kj4DBm1ra=8Ar5ca9UvQ@mail.gmail.com","threadId":"27589","inReplyTo":"BANLkTimnCqaEBVreMhnbRBV3r-r1ZzkFcg@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-06-09T06:50:27Z","receivedAt":"2011-06-09T06:50:27Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\n[+shawn, who runs into something similar with Gerrit]\n\nOn Thu, Jun 9, 2011 at 05:44, NAKAMURA Takumi <geek4civic@gmail.com> wrote:\n> Hello, Git. It is my 1st post here.\n>\n> I have tried tagging each commit as \"refs/tags/rXXXXXX\" on git-svn\n> repo locally. (over 100k refs/tags.)\n> Indeed, it made something extremely slower, even with packed-refs and\n> pack objects.\n> I gave up, then, to push tags to upstream. (it must be terror) :p\n>\n> I know it might be crazy in the git way, but it would bring me conveniences.\n> (eg. git log --oneline --decorate shows me each svn revision)\n> I would like to work for Git to live with many tags.\n>\n> * Issues as far as I have investigated;\n>\n>  - git show --decorate is always slow.\n>    in decorate.c, every commits are inspected.\n>  - git rev-tree --quiet --objects $upstream --not --all spends so much time,\n>    even if it is expected to return with 0.\n>    As you know, it is used in builtin/fetch.c.\n>  - git-upload-pack shows \"all\" refs to me if upstream has too many refs.\n>\n> I would like to work as below if they were valuable.\n>\n>  - Get rid of inspecting commits in packed-refs on decorate stuff.\n>  - Implement sort-by-hash packed-refs, (not sort-by-name)\n>  - Implement more effective pruning --not --all on revision.c.\n>  - Think about enhancement of protocol to transfer many refs more effectively.\n>\n> I am happy to consider the issue, thank you.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"169754","messageId":"m38vtbdzjq.fsf@localhost.localdomain","threadId":"27589","inReplyTo":"BANLkTimnCqaEBVreMhnbRBV3r-r1ZzkFcg@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-06-09T11:18:09Z","receivedAt":"2011-06-09T11:18:09Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"NAKAMURA Takumi <geek4civic@gmail.com> writes:\n\n> Hello, Git. It is my 1st post here.\n> \n> I have tried tagging each commit as \"refs/tags/rXXXXXX\" on git-svn\n> repo locally. (over 100k refs/tags.)\n[...]\n\nThat's insane.  You would do much better to mark each commit with\nnote.  Notes are designed to be scalable.  See e.g. this thread\n\n  [RFD] Proposal for git-svn: storing SVN metadata (git-svn-id) in notes\n  http://article.gmane.org/gmane.comp.version-control.git/174657\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"169768","messageId":"BANLkTi=PnYmJVXe8tuqdb9UiYnethf1GSw@mail.gmail.com","threadId":"27589","inReplyTo":"BANLkTinfVNxYX3kj4DBm1ra=8Ar5ca9UvQ@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-06-09T15:23:44Z","receivedAt":"2011-06-09T15:23:44Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Jun 8, 2011 at 23:50, Sverre Rabbelier <srabbelier@gmail.com> wrote:\n> [+shawn, who runs into something similar with Gerrit]\n\n> On Thu, Jun 9, 2011 at 05:44, NAKAMURA Takumi <geek4civic@gmail.com> wrote:\n>> Hello, Git. It is my 1st post here.\n>>\n>> I have tried tagging each commit as \"refs/tags/rXXXXXX\" on git-svn\n>> repo locally. (over 100k refs/tags.)\n\nAs Jakub pointed out, use git notes for this. They were designed to\nscale to >100,000 annotations.\n\n>> Indeed, it made something extremely slower, even with packed-refs and\n>> pack objects.\n\nHaving a reference to every commit in the repository is horrifically\nslow. We run into this with Gerrit Code Review and I need to find\nanother solution. Git just wasn't meant to process repositories like\nthis.\n\n-- \nShawn.\n"},{"id":"169769","messageId":"5313676.596.1307634150272.JavaMail.root@mail.hq.genarts.com","threadId":"27589","inReplyTo":"m38vtbdzjq.fsf@localhost.localdomain","subject":"Re: Git is not scalable with too many refs/*","fromName":"Stephen Bash","fromEmail":"bash@genarts.com","sentAt":"2011-06-09T15:42:30Z","receivedAt":"2011-06-09T15:42:30Z","isPatch":false,"sender":{"key":"bash@genarts.com","avatar":null},"body":"----- Original Message -----\n> From: \"Jakub Narebski\" <jnareb@gmail.com>\n> To: \"NAKAMURA Takumi\" <geek4civic@gmail.com>\n> Cc: \"git\" <git@vger.kernel.org>\n> Sent: Thursday, June 9, 2011 7:18:09 AM\n> Subject: Re: Git is not scalable with too many refs/*\n> NAKAMURA Takumi <geek4civic@gmail.com> writes:\n> \n> > Hello, Git. It is my 1st post here.\n> >\n> > I have tried tagging each commit as \"refs/tags/rXXXXXX\" on git-svn\n> > repo locally. (over 100k refs/tags.)\n> [...]\n> \n> That's insane. You would do much better to mark each commit with\n> note. Notes are designed to be scalable. See e.g. this thread\n> \n> [RFD] Proposal for git-svn: storing SVN metadata (git-svn-id) in notes\n> http://article.gmane.org/gmane.comp.version-control.git/174657\n\nAs a reformed SVN user (i.e. not using it anymore ;]) I agree that 100k tags seems crazy, but I was contemplating doing the exact same thing as Takumi.  Skimming that thread, I didn't see the key point (IMO): notes can map from commits to a \"name\" (or other information), tags map from a \"name\" to commits.\n\nI've seen two different workflows develop:\n  1) Hacking on some code in Git the programmer finds something wrong.  Using Git tools he can pickaxe/bisect/etc. and find that the problem traces back to a commit imported from Subversion.\n  2) The programmer finds something wrong, asks coworker, coworker says \"see bug XYZ\", bug XYZ says \"Fixed in r20356\".\n\nI agree notes is the right answer for (1), but for (2) you really want a cross reference table from Subversion rev number to Git commit.\n\nIn our office we created the cross reference table once by walking the Git tree and storing it as a file (we had some degenerate cases where one SVN rev mapped to multiple Git commits, but I don't remember the details), but it's not really usable from Git.  Lightweight tags would be an awesome solution (if they worked).  Perhaps a custom subcommand is a reasonable middle ground.\n\nThanks,\nStephen\n"},{"id":"169773","messageId":"4DF0EC32.40001@gmail.com","threadId":"27589","inReplyTo":"BANLkTi=PnYmJVXe8tuqdb9UiYnethf1GSw@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2011-06-09T15:52:18Z","receivedAt":"2011-06-09T15:52:18Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 06/09/2011 11:23 AM, Shawn Pearce wrote:\n> On Wed, Jun 8, 2011 at 23:50, Sverre Rabbelier<srabbelier@gmail.com>  wrote:\n>> [+shawn, who runs into something similar with Gerrit]\n>\n>> On Thu, Jun 9, 2011 at 05:44, NAKAMURA Takumi<geek4civic@gmail.com>  wrote:\n>>> Hello, Git. It is my 1st post here.\n>>>\n>>> I have tried tagging each commit as \"refs/tags/rXXXXXX\" on git-svn\n>>> repo locally. (over 100k refs/tags.)\n>\n> As Jakub pointed out, use git notes for this. They were designed to\n> scale to>100,000 annotations.\n>\n>>> Indeed, it made something extremely slower, even with packed-refs and\n>>> pack objects.\n>\n> Having a reference to every commit in the repository is horrifically\n> slow. We run into this with Gerrit Code Review and I need to find\n> another solution. Git just wasn't meant to process repositories like\n> this.\n\nAssuming a very large number of refs, what is it that makes git so \nhorrifically slow? Is there a design or implementation lesson here?\n"},{"id":"169781","messageId":"BANLkTimk06eibz99AO_0BwzoL6FWb5pR8A@mail.gmail.com","threadId":"27589","inReplyTo":"4DF0EC32.40001@gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-06-09T15:56:50Z","receivedAt":"2011-06-09T15:56:50Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Thu, Jun 9, 2011 at 08:52, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> On 06/09/2011 11:23 AM, Shawn Pearce wrote:\n>> Having a reference to every commit in the repository is horrifically\n>> slow. We run into this with Gerrit Code Review and I need to find\n>> another solution. Git just wasn't meant to process repositories like\n>> this.\n>\n> Assuming a very large number of refs, what is it that makes git so\n> horrifically slow? Is there a design or implementation lesson here?\n\nA few things.\n\nGit does a sequential scan of all references when it first needs to\naccess references for an operation. This requires reading the entire\npacked-refs file, and the recursive scan of the \"refs/\" subdirectory\nfor any loose refs that might override the packed-refs file.\n\nA lot of operations toss every commit that a reference points at into\nthe revision walker's LRU queue. If you have a tag pointing to every\ncommit, then the entire project history enters the LRU queue at once,\nup front. That queue is managed with O(N^2) insertion time. And the\nentire queue has to be filled before anything can be output.\n\n-- \nShawn.\n"},{"id":"169786","messageId":"20110609162604.GC25885@sigill.intra.peff.net","threadId":"27589","inReplyTo":"BANLkTimk06eibz99AO_0BwzoL6FWb5pR8A@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-06-09T16:26:04Z","receivedAt":"2011-06-09T16:26:04Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 09, 2011 at 08:56:50AM -0700, Shawn O. Pearce wrote:\n\n> A lot of operations toss every commit that a reference points at into\n> the revision walker's LRU queue. If you have a tag pointing to every\n> commit, then the entire project history enters the LRU queue at once,\n> up front. That queue is managed with O(N^2) insertion time. And the\n> entire queue has to be filled before anything can be output.\n\nWe ran into this recently at github. Since our many-refs repos were\nmostly forks, we had a lot of duplicate commits, and were able to solve\nit with ea5f220 (fetch: avoid repeated commits in mark_complete,\n2011-05-19).\n\nHowever, I also worked up a faster priority queue implementation that\nwould work in the general case:\n\n  http://thread.gmane.org/gmane.comp.version-control.git/174003/focus=174005\n\nI suspect it would speed up the original poster's slow fetch. The\nproblem is that a fast priority queue doesn't have quite the same access\npatterns as a linked list, so replacing all of the commit_lists in git\nwith the priority queue would be quite a painful undertaking. So we are\nleft with using the fast queue only in specific hot-spots.\n\n-Peff\n"},{"id":"169832","messageId":"BANLkTimEGjBMrbQpkZfWYPTZ93syiKFHdw@mail.gmail.com","threadId":"27589","inReplyTo":"20110609162604.GC25885@sigill.intra.peff.net","subject":"Re: Git is not scalable with too many refs/*","fromName":"NAKAMURA Takumi","fromEmail":"geek4civic@gmail.com","sentAt":"2011-06-10T03:59:47Z","receivedAt":"2011-06-10T03:59:47Z","isPatch":false,"sender":{"key":"geek4civic@gmail.com","avatar":null},"body":"Good afternoon Git! Thank you guys to give me comments.\n\nJakub and Shawn,\n\nSure, Notes should be used at the case, I agree.\n\n> (eg. git log --oneline --decorate shows me each svn revision)\n\nMy example might misunderstand you. I intended tags could show me\npretty abbrev everywhere on Git. I would be happier if tags might be\navailable bi-directional alias, as Stephen mentions.\n\nIt would be better git-svn could record metadata into notes, I think, too. :D\n\nStephen,\n\n2011/6/10 Stephen Bash <bash@genarts.com>:\n> I've seen two different workflows develop:\n>  1) Hacking on some code in Git the programmer finds something wrong.  Using Git tools he can pickaxe/bisect/etc. and find that the problem traces back to a commit imported from Subversion.\n>  2) The programmer finds something wrong, asks coworker, coworker says \"see bug XYZ\", bug XYZ says \"Fixed in r20356\".\n>\n> I agree notes is the right answer for (1), but for (2) you really want a cross reference table from Subversion rev number to Git commit.\n\nIt is the point I wanted to say, thank you! I am working with svn-men.\nThey often speak svn revision number. (And I have to tell them svn\nrevs then)\n\n> In our office we created the cross reference table once by walking the Git tree and storing it as a file (we had some degenerate cases where one SVN rev mapped to multiple Git commits, but I don't remember the details), but it's not really usable from Git.  Lightweight tags would be an awesome solution (if they worked).  Perhaps a custom subcommand is a reasonable middle ground.\n\nReconstructing svnrev-commits mapping can be done by git-svn itself.\nUnfortunately, git-svn's .rev-map is sorted by revision number. I\nthink it would be useless to make subcommands unless they were\npluggable into Git as \"smart-tag resolver\".\n\nPeff,\n\nAt first, thank you to work for Github! Awesome!\nI didn't know Github has refs issues. (yeah, I should not push 100k of\ntags to Github for now :p )\n\nI am working on linux and windows. Many-refs-repo can make Git awfully\nslow (than linux!) I hope I could work also for windows to improve\nvarious performance issue.\n\nFYI, I have tweaked git-rev-list for commits not to sort by date with\n--quiet. It improves git-fetch (git-rev-list --not --all) performance\nwhen objects is well-packed.\n\n\n...Takumi\n"},{"id":"169837","messageId":"4DF1CAC1.7060705@op5.se","threadId":"27589","inReplyTo":"BANLkTimk06eibz99AO_0BwzoL6FWb5pR8A@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2011-06-10T07:41:53Z","receivedAt":"2011-06-10T07:41:53Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"On 06/09/2011 05:56 PM, Shawn Pearce wrote:\n> On Thu, Jun 9, 2011 at 08:52, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>> On 06/09/2011 11:23 AM, Shawn Pearce wrote:\n>>> Having a reference to every commit in the repository is horrifically\n>>> slow. We run into this with Gerrit Code Review and I need to find\n>>> another solution. Git just wasn't meant to process repositories like\n>>> this.\n>>\n>> Assuming a very large number of refs, what is it that makes git so\n>> horrifically slow? Is there a design or implementation lesson here?\n> \n> A few things.\n> \n> Git does a sequential scan of all references when it first needs to\n> access references for an operation. This requires reading the entire\n> packed-refs file, and the recursive scan of the \"refs/\" subdirectory\n> for any loose refs that might override the packed-refs file.\n> \n> A lot of operations toss every commit that a reference points at into\n> the revision walker's LRU queue. If you have a tag pointing to every\n> commit, then the entire project history enters the LRU queue at once,\n> up front. That queue is managed with O(N^2) insertion time. And the\n> entire queue has to be filled before anything can be output.\n> \n\nHmm. Since we're using pre-hashed data with an obvious lookup method\nwe should be able to do much, much better than O(n^2) for insertion\nand better than O(n) for worst-case lookups. I'm thinking a 1-byte\ntrie, resulting in a depth, lookup and insertion complexity of 20. It\nwould waste some memory but it might be worth it for fixed asymptotic\ncomplexity for both insertion and lookup.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"169851","messageId":"BANLkTi=4zfO5jKKzbncJk7ihcoHX7Rst4Q@mail.gmail.com","threadId":"27589","inReplyTo":"4DF1CAC1.7060705@op5.se","subject":"Re: Git is not scalable with too many refs/*","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-06-10T19:41:39Z","receivedAt":"2011-06-10T19:41:39Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jun 10, 2011 at 00:41, Andreas Ericsson <ae@op5.se> wrote:\n> On 06/09/2011 05:56 PM, Shawn Pearce wrote:\n>>\n>> A lot of operations toss every commit that a reference points at into\n>> the revision walker's LRU queue. If you have a tag pointing to every\n>> commit, then the entire project history enters the LRU queue at once,\n>> up front. That queue is managed with O(N^2) insertion time. And the\n>> entire queue has to be filled before anything can be output.\n>\n> Hmm. Since we're using pre-hashed data with an obvious lookup method\n> we should be able to do much, much better than O(n^2) for insertion\n> and better than O(n) for worst-case lookups. I'm thinking a 1-byte\n> trie, resulting in a depth, lookup and insertion complexity of 20. It\n> would waste some memory but it might be worth it for fixed asymptotic\n> complexity for both insertion and lookup.\n\nNot really.\n\nThe queue isn't sorting by SHA-1. Its sorting by commit timestamp,\ndescending. Those aren't pre-hashed. The O(N^2) insertion is because\nthe code is trying to find where this commit belongs in the list of\ncommits as sorted by commit timestamp.\n\nThere are some priority queue datastructures designed for this sort of\nwork, e.g. a calendar queue might help. But its not as simple as a 1\nbyte trie.\n\n-- \nShawn.\n"},{"id":"169852","messageId":"m3vcwdcuqm.fsf@localhost.localdomain","threadId":"27589","inReplyTo":"BANLkTi=4zfO5jKKzbncJk7ihcoHX7Rst4Q@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-06-10T20:12:12Z","receivedAt":"2011-06-10T20:12:12Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n> On Fri, Jun 10, 2011 at 00:41, Andreas Ericsson <ae@op5.se> wrote:\n>> On 06/09/2011 05:56 PM, Shawn Pearce wrote:\n>>>\n>>> A lot of operations toss every commit that a reference points at into\n>>> the revision walker's LRU queue. If you have a tag pointing to every\n>>> commit, then the entire project history enters the LRU queue at once,\n>>> up front. That queue is managed with O(N^2) insertion time. And the\n>>> entire queue has to be filled before anything can be output.\n>>\n>> Hmm. Since we're using pre-hashed data with an obvious lookup method\n>> we should be able to do much, much better than O(n^2) for insertion\n>> and better than O(n) for worst-case lookups. I'm thinking a 1-byte\n>> trie, resulting in a depth, lookup and insertion complexity of 20. It\n>> would waste some memory but it might be worth it for fixed asymptotic\n>> complexity for both insertion and lookup.\n> \n> Not really.\n> \n> The queue isn't sorting by SHA-1. Its sorting by commit timestamp,\n> descending. Those aren't pre-hashed. The O(N^2) insertion is because\n> the code is trying to find where this commit belongs in the list of\n> commits as sorted by commit timestamp.\n> \n> There are some priority queue datastructures designed for this sort of\n> work, e.g. a calendar queue might help. But its not as simple as a 1\n> byte trie.\n\nIn the case of Subversion numbers (revision number to hash mapping)\nsorted by name (in version order at least) means sorted by date.  I\nwonder if there is data structure for which this is optimum insertion\norder (like for insertion sort almost sorted data is best case).\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"169860","messageId":"20110610203545.GA2564@sigill.intra.peff.net","threadId":"27589","inReplyTo":"BANLkTi=4zfO5jKKzbncJk7ihcoHX7Rst4Q@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-06-10T20:35:45Z","receivedAt":"2011-06-10T20:35:45Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jun 10, 2011 at 12:41:39PM -0700, Shawn O. Pearce wrote:\n\n> Not really.\n> \n> The queue isn't sorting by SHA-1. Its sorting by commit timestamp,\n> descending. Those aren't pre-hashed. The O(N^2) insertion is because\n> the code is trying to find where this commit belongs in the list of\n> commits as sorted by commit timestamp.\n> \n> There are some priority queue datastructures designed for this sort of\n> work, e.g. a calendar queue might help. But its not as simple as a 1\n> byte trie.\n\nAll you really need is a heap-based priority queue, which gives O(lg n)\ninsertion and popping (and O(1) peeking at the top). I even wrote one\nand posted it recently (I won't dig up the reference, but I posted it\nelsewhere in this thread, I think).\n\nThe problem is that many parts of the code assume that commit_list is a\nlinked list and do fast iterations, or even splicing. It's nothing you\ncouldn't get around with some work, but it turns out to involve a lot\nof code changes.\n\n-Peff\n"},{"id":"169921","messageId":"4DF5B759.8090401@op5.se","threadId":"27589","inReplyTo":"BANLkTi=4zfO5jKKzbncJk7ihcoHX7Rst4Q@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2011-06-13T07:08:09Z","receivedAt":"2011-06-13T07:08:09Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"On 06/10/2011 09:41 PM, Shawn Pearce wrote:\n> On Fri, Jun 10, 2011 at 00:41, Andreas Ericsson<ae@op5.se>  wrote:\n>> On 06/09/2011 05:56 PM, Shawn Pearce wrote:\n>>>\n>>> A lot of operations toss every commit that a reference points at into\n>>> the revision walker's LRU queue. If you have a tag pointing to every\n>>> commit, then the entire project history enters the LRU queue at once,\n>>> up front. That queue is managed with O(N^2) insertion time. And the\n>>> entire queue has to be filled before anything can be output.\n>>\n>> Hmm. Since we're using pre-hashed data with an obvious lookup method\n>> we should be able to do much, much better than O(n^2) for insertion\n>> and better than O(n) for worst-case lookups. I'm thinking a 1-byte\n>> trie, resulting in a depth, lookup and insertion complexity of 20. It\n>> would waste some memory but it might be worth it for fixed asymptotic\n>> complexity for both insertion and lookup.\n> \n> Not really.\n> \n> The queue isn't sorting by SHA-1. Its sorting by commit timestamp,\n> descending. Those aren't pre-hashed. The O(N^2) insertion is because\n> the code is trying to find where this commit belongs in the list of\n> commits as sorted by commit timestamp.\n> \n\nHmm. We should still be able to do better than that, and particularly\nfor the \"tag-each-commit\" workflow. Since it's most likely those tags\nare generated using incrementing numbers, we could have a cut-off where\nwe first parse all the refs and make an optimistic assumption that an\nalphabetical sort of the refs provides a map of insertion-points for\nthe commits. Since the best case behaviour is still O(1) for insertion\nsort and it's unlikely that thousands of refs are in random order, that\nshould cause the vast majority of the refs we insert to follow the best\ncase scenario.\n\nThis will fall on its arse when people start doing hg-ref -> git-commit\ntags ofcourse, but that doesn't seem to be happening, or at least not to\nthe same extent as with svn-revisions -> git-gommit mapping.\n\nWe're still not improving the asymptotic complexity, but it's a pretty\nsafe bet that we for a vast majority of cases improve wallclock runtime\nby a hefty amount with a relatively minor effort.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"169969","messageId":"20110613222734.GD21390@sigill.intra.peff.net","threadId":"27589","inReplyTo":"BANLkTimEGjBMrbQpkZfWYPTZ93syiKFHdw@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-06-13T22:27:34Z","receivedAt":"2011-06-13T22:27:34Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jun 10, 2011 at 12:59:47PM +0900, NAKAMURA Takumi wrote:\n\n> 2011/6/10 Stephen Bash <bash@genarts.com>:\n> > I've seen two different workflows develop:\n> >  1) Hacking on some code in Git the programmer finds something wrong.  Using Git tools he can pickaxe/bisect/etc. and find that the problem traces back to a commit imported from Subversion.\n> >  2) The programmer finds something wrong, asks coworker, coworker says \"see bug XYZ\", bug XYZ says \"Fixed in r20356\".\n> >\n> > I agree notes is the right answer for (1), but for (2) you really want a cross reference table from Subversion rev number to Git commit.\n> \n> It is the point I wanted to say, thank you! I am working with svn-men.\n> They often speak svn revision number. (And I have to tell them svn\n> revs then)\n\nYeah, there is no simple way to do the bi-directional mapping in git.\nIf all you want are decorations on commits, notes are definitely the way\nto go. They are optimized for lookup in of commit -> data. But if you\nwant data -> commit, the only mapping we have is refs, and they are not\nwell optimized for the many-refs use case.\n\nPacked-refs are better than loose refs, but I think right now we just\nload them all in to an in-memory linked list. We could load them into a\nmore efficient in-memory data structure, or we could perhaps even mmap\nthe packed-refs file and binary search it in place.\n\nBut lookup is only part of the problem. There are algorithms that want\nto look at all the refs (notably fetching and pushing), which are going\nto be a bit slower. We don't have a way to tell those algorithms that\nthose refs are uninteresting for reachability analysis, because they are\njust pointing to parts of the graph that are already reachable by\nregular refs. Maybe there could be a part of the refs namespace that is\nignored by \"--all\". I dunno. That seems like a weird inconsistency.\n\n> FYI, I have tweaked git-rev-list for commits not to sort by date with\n> --quiet. It improves git-fetch (git-rev-list --not --all) performance\n> when objects is well-packed.\n\nI'm not sure that is a good solution. Even with --quiet, we will be\nwalking the commit graph to find merge bases to see if things are\nconnected. The walking code expects date-sorting; I'm not sure what\nchanging that assumption will do to the code.\n\n-Peff\n"},{"id":"169972","messageId":"4DF6A8B6.9030301@op5.se","threadId":"27589","inReplyTo":"BANLkTimEGjBMrbQpkZfWYPTZ93syiKFHdw@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2011-06-14T00:17:58Z","receivedAt":"2011-06-14T00:17:58Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"On 06/10/2011 05:59 AM, NAKAMURA Takumi wrote:\n> Good afternoon Git! Thank you guys to give me comments.\n> \n> Jakub and Shawn,\n> \n> Sure, Notes should be used at the case, I agree.\n> \n>> (eg. git log --oneline --decorate shows me each svn revision)\n> \n> My example might misunderstand you. I intended tags could show me\n> pretty abbrev everywhere on Git. I would be happier if tags might be\n> available bi-directional alias, as Stephen mentions.\n> \n> It would be better git-svn could record metadata into notes, I think, too. :D\n> \n> Stephen,\n> \n> 2011/6/10 Stephen Bash<bash@genarts.com>:\n>> I've seen two different workflows develop:\n>>   1) Hacking on some code in Git the programmer finds something wrong.  Using Git tools he can pickaxe/bisect/etc. and find that the problem traces back to a commit imported from Subversion.\n>>   2) The programmer finds something wrong, asks coworker, coworker says \"see bug XYZ\", bug XYZ says \"Fixed in r20356\".\n>>\n>> I agree notes is the right answer for (1), but for (2) you really want a cross reference table from Subversion rev number to Git commit.\n> \n\nIf you're using svn metadata in the commit text, you can always do\n\"git log -p --grep=@20356\" to get the commits relevant to that one.\nIt's not as fast as \"git show svn-20356\", but it's not exactly\nglacial either and would avoid the problems you're having now.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"169973","messageId":"20110614003029.GA31447@sigill.intra.peff.net","threadId":"27589","inReplyTo":"4DF6A8B6.9030301@op5.se","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-06-14T00:30:29Z","receivedAt":"2011-06-14T00:30:29Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 14, 2011 at 02:17:58AM +0200, Andreas Ericsson wrote:\n\n> If you're using svn metadata in the commit text, you can always do\n> \"git log -p --grep=@20356\" to get the commits relevant to that one.\n> It's not as fast as \"git show svn-20356\", but it's not exactly\n> glacial either and would avoid the problems you're having now.\n\nIf we do end up putting this data into notes eventually (which I think\nwe _should_ do, because then you aren't locked into having this svn\ncruft in your commit messages for all time, but can rather choose\nwhether or not to display it), it would be nice to have a --grep-notes\nfeature in git-log. Or maybe --grep should look in notes by default,\ntoo, if we are showing them.\n\nI suspect the feature would be really easy to implement, if somebody is\nlooking for a gentle introduction to git, or a fun way to spend an hour.\n:)\n\n-Peff\n"},{"id":"169977","messageId":"7vtybtm3dl.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"20110614003029.GA31447@sigill.intra.peff.net","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-06-14T04:41:10Z","receivedAt":"2011-06-14T04:41:10Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I suspect the feature would be really easy to implement, if somebody is\n> looking for a gentle introduction to git, or a fun way to spend an hour.\n\nI would rather want to see if somebody can come up with a flexible reverse\nmapping feature around notes. It does not have to be completely generic,\njust being flexible enough is fine.\n"},{"id":"169983","messageId":"BANLkTimNoh3-Jde_-arzwBa=aUR+KK3Xhw@mail.gmail.com","threadId":"27589","inReplyTo":"7vtybtm3dl.fsf@alter.siamese.dyndns.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-06-14T07:26:10Z","receivedAt":"2011-06-14T07:26:10Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Jun 14, 2011 at 06:41, Junio C Hamano <gitster@pobox.com> wrote:\n> I would rather want to see if somebody can come up with a flexible reverse\n> mapping feature around notes. It does not have to be completely generic,\n> just being flexible enough is fine.\n\nWouldn't it be enough to simply create a note on 'r651235' with as\ncontents the git ref?\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"169989","messageId":"201106141202.46720.johan@herland.net","threadId":"27589","inReplyTo":"BANLkTimNoh3-Jde_-arzwBa=aUR+KK3Xhw@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2011-06-14T10:02:46Z","receivedAt":"2011-06-14T10:02:46Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Tuesday 14 June 2011, Sverre Rabbelier wrote:\n> Heya,\n> \n> On Tue, Jun 14, 2011 at 06:41, Junio C Hamano <gitster@pobox.com> wrote:\n> > I would rather want to see if somebody can come up with a flexible\n> > reverse mapping feature around notes. It does not have to be\n> > completely generic, just being flexible enough is fine.\n> \n> Wouldn't it be enough to simply create a note on 'r651235' with as\n> contents the git ref?\n\nNot quite sure what you mean by \"create a note on 'r651235'\". You could \ndevise a scheme where you SHA1('r651235'), and then create a note on the \nresulting hash.\n\nNotes are named by the SHA1 of the object they annotate, but there is no \nhard requirement (as long as you stay away from \"git notes prune\") that the \nSHA1 annotated actually exists as a valid Git object in your repo.\n\nHence, you can use notes to annotate _anything_ that can be uniquely reduced \nto a SHA1 hash.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"169993","messageId":"BANLkTikxwUhrfiYkd0ci0cps2=S4TYcRoQ@mail.gmail.com","threadId":"27589","inReplyTo":"201106141202.46720.johan@herland.net","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-06-14T10:34:47Z","receivedAt":"2011-06-14T10:34:47Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Jun 14, 2011 at 12:02, Johan Herland <johan@herland.net> wrote:\n> Not quite sure what you mean by \"create a note on 'r651235'\". You could\n> devise a scheme where you SHA1('r651235'), and then create a note on the\n> resulting hash.\n\nI was thinking they could annotate anything, even non-sha's, but in\nthat case, yes, the sha of the revision would work just as well.\n\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"170016","messageId":"20110614170214.GB26764@sigill.intra.peff.net","threadId":"27589","inReplyTo":"201106141202.46720.johan@herland.net","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-06-14T17:02:14Z","receivedAt":"2011-06-14T17:02:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 14, 2011 at 12:02:46PM +0200, Johan Herland wrote:\n\n> > Wouldn't it be enough to simply create a note on 'r651235' with as\n> > contents the git ref?\n> \n> Not quite sure what you mean by \"create a note on 'r651235'\". You could \n> devise a scheme where you SHA1('r651235'), and then create a note on the \n> resulting hash.\n> \n> Notes are named by the SHA1 of the object they annotate, but there is no \n> hard requirement (as long as you stay away from \"git notes prune\") that the \n> SHA1 annotated actually exists as a valid Git object in your repo.\n> \n> Hence, you can use notes to annotate _anything_ that can be uniquely reduced \n> to a SHA1 hash.\n\nI lean against that as a solution. I think \"git gc\" will probably\neventually learn to do \"git notes prune\", at which point we would start\nlosing people's data. So I think it is better to keep the definition of\nnotes a little tighter now, and say \"the left-hand side of a notes\nmapping must be a referenced object\". We can always loosen it later.\n\nOn top of that, though, the sha1 solution is not all that pleasant. It\nlets you do exact lookups, but you have no way of iterating over the\nlist of svn revisions.\n\nI also think we can do something a little more lightweight. The user has\nalready created and is maintaining a mapping in one direction via the\nnotes. We just need the inverse mapping, which we can generate\nprogramatically. So it can be a straight cache, with the sha1 of the\nnotes tree determining the cache validity (i.e., if the forward mapping\nin the notes tree changes, you regenerate the cache from scratch).\n\nWe would want to store the cache in an on-disk format that could be\nsearched easily. Possibly something like the packed-refs format would be\nsufficient, if we mmap'd and binary searched it. It would be dirt simple\nif we used an existing key/value store like gdbm or tokyocabinet, but we\nusually try to avoid extra dependencies.\n\n-Peff\n"},{"id":"170022","messageId":"BANLkTin0CjnM_hMaEpMroZdDhhavaoKAv00_4xBqeHj9biToVA@mail.gmail.com","threadId":"27589","inReplyTo":"20110614170214.GB26764@sigill.intra.peff.net","subject":"Re: Git is not scalable with too many refs/*","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-06-14T19:20:29Z","receivedAt":"2011-06-14T19:20:29Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Jun 14, 2011 at 10:02, Jeff King <peff@peff.net> wrote:\n> I also think we can do something a little more lightweight. The user has\n> already created and is maintaining a mapping in one direction via the\n> notes. We just need the inverse mapping, which we can generate\n> programatically. So it can be a straight cache, with the sha1 of the\n> notes tree determining the cache validity (i.e., if the forward mapping\n> in the notes tree changes, you regenerate the cache from scratch).\n>\n> We would want to store the cache in an on-disk format that could be\n> searched easily. Possibly something like the packed-refs format would be\n> sufficient, if we mmap'd and binary searched it. It would be dirt simple\n> if we used an existing key/value store like gdbm or tokyocabinet, but we\n> usually try to avoid extra dependencies.\n\nYea, not a bad idea. Use a series of SSTable like things, like Hadoop\nuses. It doesn't need to be as complex as the Hadoop SSTable concept.\nBut a simple sorted string to string mapping file that is immutable,\nwith edits applied by creating an overlay file that contains\nnew/updated entries.\n\nAs you point out, we can use the notes tree to tell us the validity of\nthe cache, and do incremental updates. If the current cache doesn't\nmatch the notes ref, compute the tree diff between the current cache's\nsource tree and the new tree, and create a new SSTable like thing that\nhas the relevant updates as an overlay of the existing tables. After\nsome time you will have many of these little overlay files, and a GC\ncan just merge them down to a single file.\n\nThe only problem is, you probably want this \"reverse notes index\" to\nbe indexing a portion of the note blob text, not all of it. That is,\nwe want the SVN note text to say something like \"SVN Revision: r1828\"\nso `git log --notes=svn` shows us something more useful than just\n\"r1828\". But in the reverse index, we may only want the key to be\n\"r1828\". So you need some sort of small mapping function to decide\nwhat to put into that reverse index.\n\n-- \nShawn.\n"},{"id":"170028","messageId":"20110614194749.GA1567@sigill.intra.peff.net","threadId":"27589","inReplyTo":"BANLkTin0CjnM_hMaEpMroZdDhhavaoKAv00_4xBqeHj9biToVA@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-06-14T19:47:49Z","receivedAt":"2011-06-14T19:47:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jun 14, 2011 at 12:20:29PM -0700, Shawn O. Pearce wrote:\n\n> > We would want to store the cache in an on-disk format that could be\n> > searched easily. Possibly something like the packed-refs format would be\n> > sufficient, if we mmap'd and binary searched it. It would be dirt simple\n> > if we used an existing key/value store like gdbm or tokyocabinet, but we\n> > usually try to avoid extra dependencies.\n> \n> Yea, not a bad idea. Use a series of SSTable like things, like Hadoop\n> uses. It doesn't need to be as complex as the Hadoop SSTable concept.\n> But a simple sorted string to string mapping file that is immutable,\n> with edits applied by creating an overlay file that contains\n> new/updated entries.\n> \n> As you point out, we can use the notes tree to tell us the validity of\n> the cache, and do incremental updates. If the current cache doesn't\n> match the notes ref, compute the tree diff between the current cache's\n> source tree and the new tree, and create a new SSTable like thing that\n> has the relevant updates as an overlay of the existing tables. After\n> some time you will have many of these little overlay files, and a GC\n> can just merge them down to a single file.\n\nI was really hoping that it would be fast enough that we could simply\nblow away the old mapping and recreate it from scratch. That gets us out\nof writing any journaling-type code with overlays. For something like\nsvn revisions, it's probably fine to take an extra second or two to\nbuild the cache after we do a fetch. But it wouldn't scale to something\nthat was getting updated frequently.\n\nIf we're going to start doing clever database-y things, I'd much rather\nuse a proven key/value db solution like tokyocabinet. I'm just not sure\nhow to degrade gracefully when the db library isn't available. Don't\nallow reverse mappings? Fallback to something slow?\n\n> The only problem is, you probably want this \"reverse notes index\" to\n> be indexing a portion of the note blob text, not all of it. That is,\n> we want the SVN note text to say something like \"SVN Revision: r1828\"\n> so `git log --notes=svn` shows us something more useful than just\n> \"r1828\". But in the reverse index, we may only want the key to be\n> \"r1828\". So you need some sort of small mapping function to decide\n> what to put into that reverse index.\n\nI had assumed that we would just be writing r1828 into the note. The\noutput via git log is actually pretty readable:\n\n  $ git notes --ref=svn/revisions add -m r1828\n  $ git show --notes=svn/revisions\n  ...\n  Notes (svn/revisions):\n      r1828\n\nOf course this is just one use case.\n\nFor that matter, we have to figure out how one would actually reference\nthe reverse mapping. If we have a simple, pure-reverse mapping, we can\njust generate and cache them on the fly, and give a special syntax.\nLike:\n\n  $ git log notes/svn/revisions@{revnote:r1828}\n\nwhich would invert the notes/svn/revisions tree, search for r1828, and\nreference the resulting commit.\n\nIf you had something more heavyweight that actually needed to parse\nduring the mapping, you might have something like:\n\n  $ : set up the mapping\n  $ git config revnote.svn.map 'SVN Revision: (r[0-9]+)'\n\n  $ : do the reverse; we should be able to build the cache on the fly\n  $ git notes reverse r1828\n  346ab9aaa1cf7b1ed2dd2c0a67bccc5b8ec23f7c\n\n  $ : so really you could have a similar ref syntax like, though\n  $ : this would require some ref parser updates, as we currently\n  $ : assume anything to the left of @{} is a real ref\n  $ git log r1828@{revnote:svn}\n\nThe syntaxes are not as nice as having a real ref. In the last example,\nwe could probably look for the contents of \"@{}\" as a possible revnote\nmapping (since we've already had to name it via the configuration), to\nmake it \"r1828@{svn}\". Or you could even come up with a default set of\nrevnotes to consider, so that if we lookup \"r1828\" and it isn't a real\nref, we fall back to trying r1828@{revnote:svn}.\n\nI dunno. I'm just throwing ideas out at this point.\n\n-Peff\n"},{"id":"170029","messageId":"BANLkTi=GZDLu-ey1=h8LLDbWssoSpsM_jd7R-oFr+b+82Otb8g@mail.gmail.com","threadId":"27589","inReplyTo":"20110614194749.GA1567@sigill.intra.peff.net","subject":"Re: Git is not scalable with too many refs/*","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-06-14T20:12:07Z","receivedAt":"2011-06-14T20:12:07Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Jun 14, 2011 at 12:47, Jeff King <peff@peff.net> wrote:\n> On Tue, Jun 14, 2011 at 12:20:29PM -0700, Shawn O. Pearce wrote:\n>\n>> > We would want to store the cache in an on-disk format that could be\n>> > searched easily. Possibly something like the packed-refs format would be\n>> > sufficient, if we mmap'd and binary searched it. It would be dirt simple\n>> > if we used an existing key/value store like gdbm or tokyocabinet, but we\n>> > usually try to avoid extra dependencies.\n>>\n>> Yea, not a bad idea. Use a series of SSTable like things, like Hadoop\n>> uses. It doesn't need to be as complex as the Hadoop SSTable concept.\n>> But a simple sorted string to string mapping file that is immutable,\n>> with edits applied by creating an overlay file that contains\n>> new/updated entries.\n>>\n>> As you point out, we can use the notes tree to tell us the validity of\n>> the cache, and do incremental updates. If the current cache doesn't\n>> match the notes ref, compute the tree diff between the current cache's\n>> source tree and the new tree, and create a new SSTable like thing that\n>> has the relevant updates as an overlay of the existing tables. After\n>> some time you will have many of these little overlay files, and a GC\n>> can just merge them down to a single file.\n>\n> I was really hoping that it would be fast enough that we could simply\n> blow away the old mapping and recreate it from scratch. That gets us out\n> of writing any journaling-type code with overlays. For something like\n> svn revisions, it's probably fine to take an extra second or two to\n> build the cache after we do a fetch. But it wouldn't scale to something\n> that was getting updated frequently.\n>\n> If we're going to start doing clever database-y things, I'd much rather\n> use a proven key/value db solution like tokyocabinet. I'm just not sure\n> how to degrade gracefully when the db library isn't available. Don't\n> allow reverse mappings? Fallback to something slow?\n\nThis is why I would prefer to build the solution into Git.\n\nIts not that bad to do a sorted string file. Take something simple\nthat is similar to a pack file:\n\n  GSST | vers | rcnt | srctree | base\n  [ klen key vlen value ]*\n  [ roff ]*\n  SHA1(all_of_above)\n\nWhere vers is a version number, rcnt is the number of records, srctree\nis the SHA-1 of the notes tree this thing indexed from, and base is\nthe SHA-1 of the notes tree this \"applies on top of\". There are then\nrcnt records in the file, each using a variable length key length and\nvalue length field (klen, vlen), with variable length key and values.\nAt the end of the file are rcnt 4 byte offsets to the start of each\nkey.\n\nWhen writing the file, write all of the above to a temporary file,\nthen rename it to $GIT_DIR/cache/db-$SHA1.db, as its a unique name.\nIts easy to prepare the list of entries in memory as an array of\nstructs of key/value pairs, sort them with qsort(), write them out and\nupdate offsets as you go, then dump out the offset table at the end.\nOne could compress the offset table by only storing every N offsets,\nreaders perform binary search until they find the first key that is\nbefore their desired key, then sequentially scan records until they\nlocate the correct entry... but I'm not sure the space savings is\nreally worthwhile here.\n\nWhen reading, scan the directory and read the headers of each file. If\nthe file has your target srctree, your cache is current and you can\nread it. If a key isn't in this file, you open the file named\n$GIT_DIR/cache/db-$base.db and try again there, walking back along\nthat base chain until base is '0'x40. (or some other marker in the\nheader to denote there is no base file).\n\nGC is just a matter of merging the sorted files together. Follow along\nall of the base pointers, open all of them, scan through the records\nand write out the first key that is defined. I guess we need a small\n\"delete\" bit in the record to indicate a particular key/value was\nremoved from the database. Since this is a reverse mapping, duplicates\nare possible, and readers that want all values need to scan back to\nthe base file, but skip base entries that were marked deleted in a\nnewer file.\n\nUpdating is just preparing a new file that uses the current srctree as\nyour base, and only inserting/sorting the paths that were different in\nthe notes.\n\nWe probably need to store these files keyed by their notes ref, so we\ncan find \"svn/revisions\" differently from \"bugzilla\" (another\nhypothetical mapping of Bugzilla bug ids to commit SHA-1s, based on\nnotes that attached bug numbers to commits).\n\nI don't think its that bad. Maybe its a bit too much complexity for\nversion 1 to have these incremental update files be supported, but it\nshouldn't be that hard.\n\n>> The only problem is, you probably want this \"reverse notes index\" to\n>> be indexing a portion of the note blob text, not all of it. That is,\n>> we want the SVN note text to say something like \"SVN Revision: r1828\"\n>> so `git log --notes=svn` shows us something more useful than just\n>> \"r1828\". But in the reverse index, we may only want the key to be\n>> \"r1828\". So you need some sort of small mapping function to decide\n>> what to put into that reverse index.\n>\n> I had assumed that we would just be writing r1828 into the note. The\n> output via git log is actually pretty readable:\n>\n>  $ git notes --ref=svn/revisions add -m r1828\n>  $ git show --notes=svn/revisions\n>  ...\n>  Notes (svn/revisions):\n>      r1828\n>\n> Of course this is just one use case.\n\nThanks, I keep forgetting that the notes prints the note ref name out\nbefore the text, so its already got this annotation present. This\nmakes it much more likely that the bare \"r1828\" text is acceptable in\nthe note, and that the reverse index is just the entire content of the\nblob as the key.  :-)\n\n> For that matter, we have to figure out how one would actually reference\n> the reverse mapping. If we have a simple, pure-reverse mapping, we can\n> just generate and cache them on the fly, and give a special syntax.\n> Like:\n>\n>  $ git log notes/svn/revisions@{revnote:r1828}\n\nUhm. Ick.\n\n> The syntaxes are not as nice as having a real ref. In the last example,\n> we could probably look for the contents of \"@{}\" as a possible revnote\n> mapping (since we've already had to name it via the configuration), to\n> make it \"r1828@{svn}\". Or you could even come up with a default set of\n> revnotes to consider, so that if we lookup \"r1828\" and it isn't a real\n> ref, we fall back to trying r1828@{revnote:svn}.\n\nOr, what about setting up a fake ref namespace:\n\n  git config ref.refs/remotes/svn/*.from refs/notes/svn/revisions\n\nThen `git log svn/r1828` works. But these aren't real references. We\nwould only want to consider them if a request matched the glob, so\n`git for-each-ref` and `git upload-pack` aren't reporting these things\nby default, and neither is `git log --all` or `gitk --all`.\n\nI agree a syntax that works out of the box without a configuration\nfile change would be nicer. But we are running out of operators to do\nthat with. `git log notes/svn/revisions@{revnote:r1828}` as you\npropose above is at least workable...\n\n-- \nShawn.\n"},{"id":"175133","messageId":"1315511619144-6773496.post@n2.nabble.com","threadId":"27589","inReplyTo":"BANLkTi=GZDLu-ey1=h8LLDbWssoSpsM_jd7R-oFr+b+82Otb8g@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-08T19:53:39Z","receivedAt":"2011-09-08T19:53:39Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"Just thought that I should add some numbers to this thread as it seems that\nthe later versions of git are worse off by several orders of magnitude on\nthis one.  \n\nWe have a Gerrit repo with just under 100K refs in refs/changes/*.  When I\nfetch them all with git 1.7.6 it does not seem to complete.  Even after 5\ndays, it is just under half way through the ref #s!   It appears, but I am\nnot sure, that it is getting slower with time also, so it may not even\ncomplete after 10 days, I couldn't wait any longer.  However, the same\ncommand works in under 8 mins with git 1.7.3.3 on the same machine!\n\nSyncing 100K refs:\n\n  git 1.7.6      > 8 days?\n  git 1.7.3.3   ~8mins\n\nThat is quite a difference!  Have there been any obvious changes to git that\nshould cause this?  If needed, I can bisect git to find out where things go\nsour, but I thought that perhaps there would be someone who already\nunderstands the problem and why older versions aren't nearly as bas as\nrecent ones.\n\nSome more things that I have tried:  after syncing the repo locally with all\n100K refs under refs/changes, I cloned it locally again and tried fetching\nlocally with both git 1.7.6 and 1.7.3.3.  I got the same results as\nremotely, so it does not appear to be related to round trips.\n\nThe original git remote syncing takes just a bit of time, and then it\noutputs lines like these:\n ...\n * [new branch]      refs/changes/13/66713/2 -> refs/changes/13/66713/2\n * [new branch]      refs/changes/13/66713/3 -> refs/changes/13/66713/3\n * [new branch]      refs/changes/13/66713/4 -> refs/changes/13/66713/4\n * [new branch]      refs/changes/13/66713/5 -> refs/changes/13/66713/5\n ...\n\nThis is the part that takes forever.  The lines seem to scroll by slower and\nslower (with git 1.7.6).  In the beginning, the lines might be a screens\nworth a minute, after 5 days, about 1  a minute.  My CPU is pegged at 100%\nduring this time (one core).  Since I have some good test data for this, let\nme know if I should test anything specific.\n\nThanks,\n\n-Martin\n\nEmployee of Qualcomm Innovation Center, Inc. which is a member of Code\nAurora Forum\n\n\n--\nView this message in context: http://git.661346.n2.nabble.com/Git-is-not-scalable-with-too-many-refs-tp6456443p6773496.html\nSent from the git mailing list archive at Nabble.com.\n"},{"id":"175134","messageId":"1315529522448-6774328.post@n2.nabble.com","threadId":"27589","inReplyTo":"1315511619144-6773496.post@n2.nabble.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-09T00:52:02Z","receivedAt":"2011-09-09T00:52:02Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"An update, I bisected it down to this commit:\n\n  88a21979c5717e3f37b9691e90b6dbf2b94c751a\n\n   fetch/pull: recurse into submodules when necessary\n\nSince this can be disabled with the --no-recurse-submodules switch, I tried\nthat and indeed, even with the latest 1.7.7rc it becomes fast (~8mins)\nagain. The strange part about this is that the repository does not have any\nsubmodules. Anyway, I hope that this can be useful to others since it is a\nworkaround which speed things up enormously. Let me know if you have any\nother tests that you want me to perform,\n\n-Martin\n\n--\nView this message in context: http://git.661346.n2.nabble.com/Git-is-not-scalable-with-too-many-refs-tp6456443p6774328.html\nSent from the git mailing list archive at Nabble.com.\n"},{"id":"175137","messageId":"201109090305.15896.trast@student.ethz.ch","threadId":"27589","inReplyTo":"1315529522448-6774328.post@n2.nabble.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2011-09-09T01:05:15Z","receivedAt":"2011-09-09T01:05:15Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Martin Fick wrote:\n> An update, I bisected it down to this commit:\n> \n>   88a21979c5717e3f37b9691e90b6dbf2b94c751a\n> \n>    fetch/pull: recurse into submodules when necessary\n> \n> Since this can be disabled with the --no-recurse-submodules switch, I tried\n> that and indeed, even with the latest 1.7.7rc it becomes fast (~8mins)\n> again. The strange part about this is that the repository does not have any\n> submodules. Anyway, I hope that this can be useful to others since it is a\n> workaround which speed things up enormously. Let me know if you have any\n> other tests that you want me to perform,\n\nJens should know about this, so let's Cc him.\n\nI took a quick look and I'm guessing that there's at least one\nquadratic behaviour: in check_for_new_submodule_commits(), I see\n\n+       const char *argv[] = {NULL, NULL, \"--not\", \"--all\", NULL};\n+       int argc = ARRAY_SIZE(argv) - 1;\n+\n+       init_revisions(&rev, NULL);\n\nwhich means that the --all needs to walk all commits reachable from\nall refs and flag them as uninteresting.  But that function is called\nfor every ref update, so IIUC the time spent is on the order of\n#ref updates*#commits.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"175140","messageId":"201109090313.56898.trast@student.ethz.ch","threadId":"27589","inReplyTo":"201109090305.15896.trast@student.ethz.ch","subject":"Re: Git is not scalable with too many refs/*","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2011-09-09T01:13:56Z","receivedAt":"2011-09-09T01:13:56Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Thomas Rast wrote:\n> +       const char *argv[] = {NULL, NULL, \"--not\", \"--all\", NULL};\n> +       int argc = ARRAY_SIZE(argv) - 1;\n> +\n> +       init_revisions(&rev, NULL);\n> \n> which means that the --all needs to walk all commits reachable from\n> all refs and flag them as uninteresting.\n\nScratch that, it \"only\" needs to mark every tip commit and then walk\nthem back to about where the interesting commits end.\n\nIn any case, since the uninteresting set only gets larger, it should\nbe possible to reuse the same revision walker.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"175187","messageId":"4E6A19AD.80100@alum.mit.edu","threadId":"27589","inReplyTo":"1315511619144-6773496.post@n2.nabble.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-09-09T13:50:37Z","receivedAt":"2011-09-09T13:50:37Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 09/08/2011 09:53 PM, Martin Fick wrote:\n> Just thought that I should add some numbers to this thread as it seems that\n> the later versions of git are worse off by several orders of magnitude on\n> this one.  \n> \n> We have a Gerrit repo with just under 100K refs in refs/changes/*.  When I\n> fetch them all with git 1.7.6 it does not seem to complete.  Even after 5\n> days, it is just under half way through the ref #s! [...]\n\nI recently reported very slow performance when doing a \"git\nfilter-branch\" involving only about 1000 tags, with hints of O(N^3)\nscaling [1].  That could certainly explain enormous runtimes for 100k refs.\n\nReferences are cached in git in a single linked list, so it is easy to\nimagine O(N^2) all over the place (which is bad enough for 100k\nreferences).  I am working on improving the situation by reorganizing\nhow the reference cache is stored in memory, but progress is slow.\n\nI'm not sure whether your problem is related.  For example, it is not\nobvious to me why the commit that you cite (88a21979) would make the\nreference problem so dramatically worse.\n\nI suggest the following experiments to characterize the problem:\n\n1. Fetch the references in batches of a few hundred each, and see if\nthat dramatically decreases the total time.\n\n2. Same as (1), except run \"git pack-refs --all --prune\" between the\nbatches.  In my experiments, packing references made a dramatic\ndifference in runtimes.\n\n3. Try using the --no-replace-objects option (I assume that it can be\nused like \"git --no-replace-objects fetch ...\").  In my case this option\nmade a dramatic improvement in the runtimes.\n\n4. Try a test using a repository generated something like the test\nscript that I posted in [1].  If it also gives pathologically bad\nperformance, then it can serve as a test case to use while we debug the\nproblem.\n\nYours,\nMichael\n\n[1] http://comments.gmane.org/gmane.comp.version-control.git/177103\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"175202","messageId":"4E6A35E8.8060502@alum.mit.edu","threadId":"27589","inReplyTo":"4E6A19AD.80100@alum.mit.edu","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-09-09T15:51:04Z","receivedAt":"2011-09-09T15:51:04Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"I have answered some of my own questions:\n\nOn 09/09/2011 03:50 PM, Michael Haggerty wrote:\n> 3. Try using the --no-replace-objects option (I assume that it can be\n> used like \"git --no-replace-objects fetch ...\").  In my case this option\n> made a dramatic improvement in the runtimes.\n\nThis does not seem to help much.\n\n> 4. Try a test using a repository generated something like the test\n> script that I posted in [1].  If it also gives pathologically bad\n> performance, then it can serve as a test case to use while we debug the\n> problem.\n\nYes, a simple test repo like that created by the script is enough to\nreproduce the problem.  The slowdown becomes very obvious after only a\nfew hundred references.\n\nCuriously, \"git clone\" is very fast under the same circumstances that\n\"git fetch\" is excruciatingly slow.\n\nAccording to strace, git seems to be repopulating the ref cache after\neach new ref is created (it walks through the whole refs subdirectory\nand reads every file).  Apparently the ref cache is being discarded\ncompletely whenever a ref is added (which can and should be fixed) and\nthen being reloaded for some reason (though single refs can be inspected\nmuch faster without reading the cache).  This situation should be\nimproved by the hierarchical refcache changes that I'm working on plus\nsmarter updating (rather than discarding) of the cache when a new\nreference is created.\n\nSome earlier speculation in this thread was that that slowdowns might be\ncaused by \"pessimal\" ordering of revisions in the walker queue.  But my\ntest repository shards the references in such a way that the lexical\norder of the refnames does not correspond to the topological order of\nthe commits.  So that can't be the whole story.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"175205","messageId":"4E6A37D8.8050400@web.de","threadId":"27589","inReplyTo":"1315529522448-6774328.post@n2.nabble.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2011-09-09T15:59:20Z","receivedAt":"2011-09-09T15:59:20Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 09.09.2011 02:52, schrieb Martin Fick:\n> An update, I bisected it down to this commit:\n> \n>   88a21979c5717e3f37b9691e90b6dbf2b94c751a\n> \n>    fetch/pull: recurse into submodules when necessary\n> \n> Since this can be disabled with the --no-recurse-submodules switch, I tried\n> that and indeed, even with the latest 1.7.7rc it becomes fast (~8mins)\n> again. The strange part about this is that the repository does not have any\n> submodules. Anyway, I hope that this can be useful to others since it is a\n> workaround which speed things up enormously. Let me know if you have any\n> other tests that you want me to perform,\n\nThanks for nailing that one down. I'm currently looking into bringing back\ndecent performance here.\n"},{"id":"175209","messageId":"4E6A38D3.3010900@web.de","threadId":"27589","inReplyTo":"4E6A19AD.80100@alum.mit.edu","subject":"Re: Git is not scalable with too many refs/*","fromName":"Jens Lehmann","fromEmail":"jens.lehmann@web.de","sentAt":"2011-09-09T16:03:31Z","receivedAt":"2011-09-09T16:03:31Z","isPatch":false,"sender":{"key":"jens.lehmann@web.de","avatar":"https://avatars.githubusercontent.com/u/135220?v=4"},"body":"Am 09.09.2011 15:50, schrieb Michael Haggerty:\n> On 09/08/2011 09:53 PM, Martin Fick wrote:\n>> Just thought that I should add some numbers to this thread as it seems that\n>> the later versions of git are worse off by several orders of magnitude on\n>> this one.  \n>>\n>> We have a Gerrit repo with just under 100K refs in refs/changes/*.  When I\n>> fetch them all with git 1.7.6 it does not seem to complete.  Even after 5\n>> days, it is just under half way through the ref #s! [...]\n> \n> I recently reported very slow performance when doing a \"git\n> filter-branch\" involving only about 1000 tags, with hints of O(N^3)\n> scaling [1].  That could certainly explain enormous runtimes for 100k refs.\n> \n> References are cached in git in a single linked list, so it is easy to\n> imagine O(N^2) all over the place (which is bad enough for 100k\n> references).  I am working on improving the situation by reorganizing\n> how the reference cache is stored in memory, but progress is slow.\n> \n> I'm not sure whether your problem is related.  For example, it is not\n> obvious to me why the commit that you cite (88a21979) would make the\n> reference problem so dramatically worse.\n\n88a21979 is the reason, as since then a \"git rev-list <sha1> --not --all\" is\nrun for *every* updated ref to find out all new commits fetched for that ref.\nAnd if you have 100K of them ...\n"},{"id":"176179","messageId":"201109251443.28243.mfick@codeaurora.org","threadId":"27589","inReplyTo":"1315529522448-6774328.post@n2.nabble.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-25T20:43:27Z","receivedAt":"2011-09-25T20:43:27Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"A coworker of mine pointed out to me that a simple\n\n  git checkout \n\ncan also take rather long periods of time > 3 mins when run \non a repo with ~100K refs.  \n\nWhile this is not massive like the other problem I reported, \nit still seems like it is more than one would expect.  So, I \ntried an older version of git, and to my surprise/delight, \nit was much faster (.2s).  So, I bisected this issue also, \nand it seems that the \"offending\" commit is \n680955702990c1d4bfb3c6feed6ae9c6cb5c3c07:\n\n\ncommit 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07\nAuthor: Christian Couder <chriscool@tuxfamily.org>\nDate:   Fri Jan 23 10:06:53 2009 +0100\n\n    replace_object: add mechanism to replace objects found \nin \"refs/replace/\"\n\n    The code implementing this mechanism has been copied \nmore-or-less\n    from the commit graft code.\n\n    This mechanism is used in \"read_sha1_file\". sha1 passed \nto this\n    function that match a ref name in \"refs/replace/\" are \nreplaced by\n    the sha1 that has been read in the ref.\n\n    We \"die\" if the replacement recursion depth is too high \nor if we\n    can't read the replacement object.\n\n    Signed-off-by: Christian Couder \n<chriscool@tuxfamily.org>\n    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n\n\n\nNow, I suspect this commit is desirable, but I was hoping \nthat perhaps a look at it might inspire someone to find an \nobvious problem with it.  \n\nThanks,\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176200","messageId":"CAP8UFD3TWQHU0wLPuxMDnc3bRSz90Yd+yDMBe03kofeo-nr7yA@mail.gmail.com","threadId":"27589","inReplyTo":"201109251443.28243.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2011-09-26T12:41:04Z","receivedAt":"2011-09-26T12:41:04Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Sun, Sep 25, 2011 at 10:43 PM, Martin Fick <mfick@codeaurora.org> wrote:\n> A coworker of mine pointed out to me that a simple\n>\n>  git checkout\n>\n> can also take rather long periods of time > 3 mins when run\n> on a repo with ~100K refs.\n\nAre all these refs packed?\n\n> While this is not massive like the other problem I reported,\n> it still seems like it is more than one would expect.  So, I\n> tried an older version of git, and to my surprise/delight,\n> it was much faster (.2s).  So, I bisected this issue also,\n> and it seems that the \"offending\" commit is\n> 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07:\n>\n> commit 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07\n> Author: Christian Couder <chriscool@tuxfamily.org>\n> Date:   Fri Jan 23 10:06:53 2009 +0100\n>\n>    replace_object: add mechanism to replace objects found\n> in \"refs/replace/\"\n\n[...]\n\n> Now, I suspect this commit is desirable, but I was hoping\n> that perhaps a look at it might inspire someone to find an\n> obvious problem with it.\n\nI don't think there is an obvious problem with it, but it would be\nnice if you could dig a bit deeper.\n\nThe first thing that could take a lot of time is the call to\nfor_each_replace_ref() in this function:\n\n+static void prepare_replace_object(void)\n+{\n+       static int replace_object_prepared;\n+\n+       if (replace_object_prepared)\n+               return;\n+\n+       for_each_replace_ref(register_replace_ref, NULL);\n+       replace_object_prepared = 1;\n+}\n\nAnother thing is calling replace_object_pos() repeatedly in\nlookup_replace_object().\n\nThanks,\nChristian.\n"},{"id":"176216","messageId":"201109260915.29285.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109251443.28243.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T15:15:29Z","receivedAt":"2011-09-26T15:15:29Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"OK, I have found what I believe is another performance \nregression for large ref counts (~100K).  \n\nWhen I run git br on my repo which only has one branch, but \nhas ~100K refs under ref/changes (a gerrit repo), it takes \nnormally 3-6mins depending on whether my caches are fresh or \nnot.  After bisecting some older changes, I noticed that \nthis ref seems to be where things start to get slow: \nc774aab98ce6c5ef7aaacbef38da0a501eb671d4\n\n\ncommit c774aab98ce6c5ef7aaacbef38da0a501eb671d4\nAuthor: Julian Phillips <julian@quantumfyre.co.uk>\nDate:   Tue Apr 17 02:42:50 2007 +0100\n\n    refs.c: add a function to sort a ref list, rather then \nsorting on add\n\n    Rather than sorting the refs list while building it, \nsort in one\n    go after it is built using a merge sort.  This has a \nlarge\n    performance boost with large numbers of refs.\n\n    It shouldn't happen that we read duplicate entries into \nthe same\n    list, but just in case sort_ref_list drops them if the \nSHA1s are\n    the same, or dies, as we have no way of knowing which \none is the\n    correct one.\n\n    Signed-off-by: Julian Phillips \n<julian@quantumfyre.co.uk>\n    Acked-by: Linus Torvalds <torvalds@linux-foundation.org>\n    Signed-off-by: Junio C Hamano <junkio@cox.net>\n\n\n\nwhich is a bit strange since that commit's purpose was to \nactually speed things up in the case of many refs.  Just to \nverify, I reverted the commit on 1.7.7.rc0.73 and sure \nenough, things speed up down to the 14-20s range depending \non caching.\n\nIf this change does not actually speed things up, should it \nbe reverted?  Or was there a bug in the change that makes it \nnot do what it was supposed to do?\n\nThanks,\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176217","messageId":"CAGdFq_iuY0-PBDOmtz1pRvh6J9EDRiRJHsWkTN_cHjGU20PYTQ@mail.gmail.com","threadId":"27589","inReplyTo":"201109260915.29285.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-09-26T15:21:30Z","receivedAt":"2011-09-26T15:21:30Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Mon, Sep 26, 2011 at 17:15, Martin Fick <mfick@codeaurora.org> wrote:\n> If this change does not actually speed things up, should it\n> be reverted?  Or was there a bug in the change that makes it\n> not do what it was supposed to do?\n\nIt probably looks at the refs in refs/changes while it shouldn't,\nhence worsening your performance compared to not looking at those\nrefs. I assume that it does improve your situation if you have all\nthose refs under say refs/heads.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"176218","messageId":"4E809AFE.6010901@alum.mit.edu","threadId":"27589","inReplyTo":"201109251443.28243.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-09-26T15:32:14Z","receivedAt":"2011-09-26T15:32:14Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 09/25/2011 10:43 PM, Martin Fick wrote:\n> A coworker of mine pointed out to me that a simple\n> \n>   git checkout \n> \n> can also take rather long periods of time > 3 mins when run \n> on a repo with ~100K refs.  \n> \n> While this is not massive like the other problem I reported, \n> it still seems like it is more than one would expect.  So, I \n> tried an older version of git, and to my surprise/delight, \n> it was much faster (.2s).  So, I bisected this issue also, \n> and it seems that the \"offending\" commit is \n> 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07:\n\nI'm still working on changes to store references hierarchically in the\ncache and read them lazily.  I hope that it will help some scaling\nproblems with large number of refs.\n\nUnfortunately I keep getting tangled up in side issues, so it is taking\na lot longer than expected.  But there's still hope.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176219","messageId":"201109260942.08299.mfick@codeaurora.org","threadId":"27589","inReplyTo":"4E809AFE.6010901@alum.mit.edu","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T15:42:08Z","receivedAt":"2011-09-26T15:42:08Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 09:32:14 am Michael Haggerty \nwrote:\n> On 09/25/2011 10:43 PM, Martin Fick wrote:\n> > A coworker of mine pointed out to me that a simple\n> > \n> >   git checkout\n> > \n> > can also take rather long periods of time > 3 mins when\n> > run on a repo with ~100K refs.\n> > \n> > While this is not massive like the other problem I\n> > reported, it still seems like it is more than one\n> > would expect.  So, I tried an older version of git,\n> > and to my surprise/delight, it was much faster (.2s). \n> > So, I bisected this issue also, and it seems that the\n> > \"offending\" commit is\n> \n> > 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07:\n> I'm still working on changes to store references\n> hierarchically in the cache and read them lazily.  I\n> hope that it will help some scaling problems with large\n> number of refs.\n> \n> Unfortunately I keep getting tangled up in side issues,\n> so it is taking a lot longer than expected.  But there's\n> still hope.\n> \n> Michael\n\nThanks Michael, I look forward to those changes.  In the \nmeantime however, I will try to take advantage of the \ncurrent inefficiencies of large ref counts to attempt to \nfind places where there are obvious problems in the code \npaths.  I suspect that there are several commands in git \nwhich inadvertently scan all the refs when they probably \nshouldn't.  Since this is likely very slow now, it should be \neasy to find those, if it were faster, this might get \noverlooked.  I feel like git checkout is one of those cases, \nit does not seem like git checkout should be affected by the \nnumber of refs in a repo?\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176220","messageId":"201109260948.25312.mfick@codeaurora.org","threadId":"27589","inReplyTo":"CAGdFq_iuY0-PBDOmtz1pRvh6J9EDRiRJHsWkTN_cHjGU20PYTQ@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T15:48:25Z","receivedAt":"2011-09-26T15:48:25Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 09:21:30 am Sverre Rabbelier \nwrote:\n> Heya,\n> \n> On Mon, Sep 26, 2011 at 17:15, Martin Fick \n<mfick@codeaurora.org> wrote:\n> > If this change does not actually speed things up,\n> > should it be reverted?  Or was there a bug in the\n> > change that makes it not do what it was supposed to\n> > do?\n> \n> It probably looks at the refs in refs/changes while it\n> shouldn't, hence worsening your performance compared to\n> not looking at those refs. I assume that it does improve\n> your situation if you have all those refs under say\n> refs/heads.\n\nHmm, I was thinking that too, and I just did a test. \n\nInstead of storing the changes under refs/changes, I fetched \nthem under refs/heads/changes and then ran git 1.7.6 and it \ntook about 3 mins.  Then, I ran the 1.7.7.rc0.73 with \nc774aab98ce6c5ef7aaacbef38da0a501eb671d4 reverted and it \nonly took 13s!  So, if this indeed tests what you were \nsuggesting, I think it shows that even in the intended case \nthis change slowed things down?\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176221","messageId":"CAGdFq_hRmSif4=V+9h8=S1fWfPCj+meRY8xGyfgv=SWk+DrBQw@mail.gmail.com","threadId":"27589","inReplyTo":"201109260948.25312.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-09-26T15:56:50Z","receivedAt":"2011-09-26T15:56:50Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Mon, Sep 26, 2011 at 17:48, Martin Fick <mfick@codeaurora.org> wrote:\n> Hmm, I was thinking that too, and I just did a test.\n>\n> Instead of storing the changes under refs/changes, I fetched\n> them under refs/heads/changes and then ran git 1.7.6 and it\n> took about 3 mins.  Then, I ran the 1.7.7.rc0.73 with\n> c774aab98ce6c5ef7aaacbef38da0a501eb671d4 reverted and it\n> only took 13s!  So, if this indeed tests what you were\n> suggesting, I think it shows that even in the intended case\n> this change slowed things down?\n\nAnd if you run 1.7.7 without that commit reverted?\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"176222","messageId":"201109261825.39831.trast@student.ethz.ch","threadId":"27589","inReplyTo":"201109260942.08299.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2011-09-26T16:25:39Z","receivedAt":"2011-09-26T16:25:39Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Martin Fick wrote:\n> \n> I suspect that there are several commands in git \n> which inadvertently scan all the refs when they probably \n> shouldn't. [...] I feel like git checkout is one of those cases, \n> it does not seem like git checkout should be affected by the \n> number of refs in a repo?\n\ngit-checkout checks whether you are leaving any unreferenced\n(orphaned) commits behind when you leave a detached HEAD, which\nrequires that it scan the history of all refs for the commit you just\nleft.\n\nSo unless you disable that warning it'll be pretty expensive\nregardless.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"176224","messageId":"201109261038.34527.mfick@codeaurora.org","threadId":"27589","inReplyTo":"CAGdFq_hRmSif4=V+9h8=S1fWfPCj+meRY8xGyfgv=SWk+DrBQw@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T16:38:34Z","receivedAt":"2011-09-26T16:38:34Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 09:56:50 am Sverre Rabbelier \nwrote:\n> Heya,\n> \n> On Mon, Sep 26, 2011 at 17:48, Martin Fick \n<mfick@codeaurora.org> wrote:\n> > Hmm, I was thinking that too, and I just did a test.\n> > \n> > Instead of storing the changes under refs/changes, I\n> > fetched them under refs/heads/changes and then ran git\n> > 1.7.6 and it took about 3 mins.  Then, I ran the\n> > 1.7.7.rc0.73 with\n> > c774aab98ce6c5ef7aaacbef38da0a501eb671d4 reverted and\n> > it only took 13s!  So, if this indeed tests what you\n> > were suggesting, I think it shows that even in the\n> > intended case this change slowed things down?\n> \n> And if you run 1.7.7 without that commit reverted?\n\nSorry, I probably confused things by mentioning 1.7.6, the \nbad commit was way before that early 1.5 days...  \n\nAs for 1.7.7, I don't think that exists yet, so did you mean \nthe 1.7.7.rc0.73 version that I mentioned above without the \nrevert?  Strangely enough, that ends up being \n1.7.7.rc0.72.g4b5ea.  That is also slow with \nrefs/heads/changes > 3mins.\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176226","messageId":"2b7c307b138bd6061209795056a593ca@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109261038.34527.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-26T16:49:06Z","receivedAt":"2011-09-26T16:49:06Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 10:38:34 -0600, Martin Fick wrote:\n> On Monday, September 26, 2011 09:56:50 am Sverre Rabbelier\n> wrote:\n>> Heya,\n>>\n>> On Mon, Sep 26, 2011 at 17:48, Martin Fick\n> <mfick@codeaurora.org> wrote:\n>> > Hmm, I was thinking that too, and I just did a test.\n>> >\n>> > Instead of storing the changes under refs/changes, I\n>> > fetched them under refs/heads/changes and then ran git\n>> > 1.7.6 and it took about 3 mins.  Then, I ran the\n>> > 1.7.7.rc0.73 with\n>> > c774aab98ce6c5ef7aaacbef38da0a501eb671d4 reverted and\n>> > it only took 13s!  So, if this indeed tests what you\n>> > were suggesting, I think it shows that even in the\n>> > intended case this change slowed things down?\n>>\n>> And if you run 1.7.7 without that commit reverted?\n>\n> Sorry, I probably confused things by mentioning 1.7.6, the\n> bad commit was way before that early 1.5 days...\n>\n> As for 1.7.7, I don't think that exists yet, so did you mean\n> the 1.7.7.rc0.73 version that I mentioned above without the\n> revert?  Strangely enough, that ends up being\n> 1.7.7.rc0.72.g4b5ea.  That is also slow with\n> refs/heads/changes > 3mins.\n\nHmm ... something interesting is going on.\n\nI created a little test repo with ~100k unpacked refs.\n\nI tried \"time git branch\" with three versions of git, and I got (hot \ncache times):\n\ngit version 1.7.6.1: ~1.2s\ngit version 1.7.7.rc3: ~1.2s\ngit version 1.7.7.rc3.1.gbc93f: ~40s\n\nWhere the third was with the commit reverted.  That was almost 40s of \n100% CPU - my poor laptop had to turn the fans up to noisy ...\n\n> -Martin\n\n-- \nJulian\n"},{"id":"176234","messageId":"201109261147.56607.mfick@codeaurora.org","threadId":"27589","inReplyTo":"CAP8UFD3TWQHU0wLPuxMDnc3bRSz90Yd+yDMBe03kofeo-nr7yA@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T17:47:56Z","receivedAt":"2011-09-26T17:47:56Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 06:41:04 am Christian Couder \nwrote:\n> On Sun, Sep 25, 2011 at 10:43 PM, Martin Fick \n<mfick@codeaurora.org> wrote:\n> > A coworker of mine pointed out to me that a simple\n> > \n> >  git checkout\n> > \n> > can also take rather long periods of time > 3 mins when\n> > run on a repo with ~100K refs.\n> \n> Are all these refs packed?\n\nI think so, is there a way to find out for sure?\n\n-Martin\n"},{"id":"176240","messageId":"201109261207.52736.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109260915.29285.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T18:07:52Z","receivedAt":"2011-09-26T18:07:52Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 09:15:29 am Martin Fick wrote:\n> OK, I have found what I believe is another performance\n> regression for large ref counts (~100K).\n> \n> When I run git br on my repo which only has one branch,\n> but has ~100K refs under ref/changes (a gerrit repo), it\n> takes normally 3-6mins depending on whether my caches\n> are fresh or not.  After bisecting some older changes, I\n> noticed that this ref seems to be where things start to\n> get slow: c774aab98ce6c5ef7aaacbef38da0a501eb671d4\n> \n> \n> commit c774aab98ce6c5ef7aaacbef38da0a501eb671d4\n> Author: Julian Phillips <julian@quantumfyre.co.uk>\n> Date:   Tue Apr 17 02:42:50 2007 +0100\n> \n>     refs.c: add a function to sort a ref list, rather\n> then sorting on add\n> \n>     Rather than sorting the refs list while building it,\n> sort in one\n>     go after it is built using a merge sort.  This has a\n> large\n>     performance boost with large numbers of refs.\n> \n>     It shouldn't happen that we read duplicate entries\n> into the same\n>     list, but just in case sort_ref_list drops them if\n> the SHA1s are\n>     the same, or dies, as we have no way of knowing which\n> one is the\n>     correct one.\n> \n>     Signed-off-by: Julian Phillips\n> <julian@quantumfyre.co.uk>\n>     Acked-by: Linus Torvalds\n> <torvalds@linux-foundation.org> Signed-off-by: Junio C\n> Hamano <junkio@cox.net>\n> \n> \n> \n> which is a bit strange since that commit's purpose was to\n> actually speed things up in the case of many refs.  Just\n> to verify, I reverted the commit on 1.7.7.rc0.73 and\n> sure enough, things speed up down to the 14-20s range\n> depending on caching.\n> \n> If this change does not actually speed things up, should\n> it be reverted?  Or was there a bug in the change that\n> makes it not do what it was supposed to do?\n\n\nAhh, I think I have some more clues.  So while this change \ndoes not speed things up for me normally, I found a case \nwhere it does!  I  set my .git/config to have\n\n  [core]\n        compression = 0\n\nand ran git-gc on my repo.  Now, with a modern git with this \noptimization in it (1.7.6, 1.7.7.rc0...), 'git branch' is \nalmost instantaneous (.05s)!  But, if I revert c774aa it \ntakes > ~15s.  \n\nSo, it appears that this optimization is foiled by \ncompression?  In the case when this optimization helps, it \nsave about 15s, when it hurts (with compression), it seems \nto cost > 3mins.  I am not sure this optimization is worth \nit?  Would there be someway for it to adjust to the repo \nconditions?\n\n \nThanks,\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176242","messageId":"88a00eadcbb4a7946dbe8d70dd0e933d@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109261207.52736.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-26T18:37:10Z","receivedAt":"2011-09-26T18:37:10Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 12:07:52 -0600, Martin Fick wrote:\n-- snip --\n> Ahh, I think I have some more clues.  So while this change\n> does not speed things up for me normally, I found a case\n> where it does!  I  set my .git/config to have\n>\n>   [core]\n>         compression = 0\n>\n> and ran git-gc on my repo.  Now, with a modern git with this\n> optimization in it (1.7.6, 1.7.7.rc0...), 'git branch' is\n> almost instantaneous (.05s)!  But, if I revert c774aa it\n> takes > ~15s.\n\nI don't understand this.  I don't see why core.compression should have \nanything to do with refs ...\n\n> So, it appears that this optimization is foiled by\n> compression?  In the case when this optimization helps, it\n> save about 15s, when it hurts (with compression), it seems\n> to cost > 3mins.  I am not sure this optimization is worth\n> it?  Would there be someway for it to adjust to the repo\n> conditions?\n\nWell, in the case I tried it was 1.2s vs 40s.  It would seem that you \nhave managed to find some corner case.  It doesn't seem right to punish \neveryone who has large numbers of refs by making their commands take \norders of magnitude longer to save one person 3m.  Much better to find, \nunderstand and fix the actual cause.\n\nI really can't see what effect core.compression can have on loading the \nref_list.  Certainly the sort doesn't load anything from the object \ndatabase.  It would be really good to profile and find out what is \ntaking all the time - I am assuming that the CPU is at 100% for the 3+ \nminutes?\n\nRandom thought.  What happens to the with compression case if you leave \nthe commit in, but add a sleep(15) to the end of sort_refs_list?\n\n-- \nJulian\n"},{"id":"176245","messageId":"201109262056.04279.chriscool@tuxfamily.org","threadId":"27589","inReplyTo":"201109261147.56607.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Christian Couder","fromEmail":"chriscool@tuxfamily.org","sentAt":"2011-09-26T18:56:04Z","receivedAt":"2011-09-26T18:56:04Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Monday 26 September 2011 19:47:56 Martin Fick wrote:\n> On Monday, September 26, 2011 06:41:04 am Christian Couder\n> wrote:\n> > \n> > Are all these refs packed?\n> \n> I think so, is there a way to find out for sure?\n\nAfter \"git pack-refs --all\" I get:\n\n$ find .git/refs/ -type f\n.git/refs/remotes/origin/HEAD\n.git/refs/stash\n\nSo I suppose that if such a find gives you only a few files all (or most of) \nyour refs are packed.\n\nBest regards,\nChristian.\n"},{"id":"176253","messageId":"201109261401.38624.mfick@codeaurora.org","threadId":"27589","inReplyTo":"88a00eadcbb4a7946dbe8d70dd0e933d@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T20:01:38Z","receivedAt":"2011-09-26T20:01:38Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 12:37:10 pm Julian Phillips \nwrote:\n> On Mon, 26 Sep 2011 12:07:52 -0600, Martin Fick wrote:\n> -- snip --\n> \n> > Ahh, I think I have some more clues.  So while this\n> > change does not speed things up for me normally, I\n> > found a case where it does!  I  set my .git/config to\n> > have\n> > \n> >   [core]\n> >   \n> >         compression = 0\n> > \n> > and ran git-gc on my repo.  Now, with a modern git with\n> > this optimization in it (1.7.6, 1.7.7.rc0...), 'git\n> > branch' is almost instantaneous (.05s)!  But, if I\n> > revert c774aa it takes > ~15s.\n> \n> I don't understand this.  I don't see why\n> core.compression should have anything to do with refs\n> ...\n> \n> > So, it appears that this optimization is foiled by\n> > compression?  In the case when this optimization helps,\n> > it save about 15s, when it hurts (with compression),\n> > it seems to cost > 3mins.  I am not sure this\n> > optimization is worth it?  Would there be someway for\n> > it to adjust to the repo conditions?\n> \n> Well, in the case I tried it was 1.2s vs 40s.  It would\n> seem that you have managed to find some corner case.  It\n> doesn't seem right to punish everyone who has large\n> numbers of refs by making their commands take orders of\n> magnitude longer to save one person 3m.  Much better to\n> find, understand and fix the actual cause.\n\nI am not sure mine is the corner case, it is a real repo \n(albeit a Gerrit repo with strange refs/changes), while it \nsounds like yours is a test repo.  It seems likely that \nwhatever you did to create the test repo makes it perform \nwell?  I am also guessing that it is not the refs which are \nthe problem but the objects since the refs don't get \ncompressed do they?  Does your repo have real data in it \n(not just 100K refs)?  \n\nMy repo compressed is about ~2G and uncompressed is ~1.1G\nYes, the compressed one is larger than the uncompressed one.\nSince the compressed repo above was larger, I thought that I \nshould at lest gc it.  After git gc, it is ~1.1G, so it \nlooks like the size difference was really because of not \nhaving gced it at first after fetching the 100K refs.\n\nAfter a gc, the repo does perform the similar to the \nuncompressed one (which was achieved via gc).  After gc, it \ntakes ~.05s do to a 'git branch' with 1.7.6 and \ngit.1.7.7.rc0.72.g4b5ea.  It also takes a bit more than 15s \nwith the patch reverted.  So it appears that compression is \nnot likely the culprit, but rather the need to be gced.\n\nSo, maybe you are correct, maybe my repo is the corner case?  \nIs a repo which needs to be gced considered a corner case?  \nShould git be able to detect that the repo is so in \ndesperate need of gcing?  Is it normal for git to need to gc \nright after a clone and then fetching ~100K refs?\n\nI am not sure what is right here, if this patch makes a repo \nwhich needs gcing degrade 5 to 10 times worse than the \nbenefit of this patch, it still seems questionable to me.\n\n\n \n> I really can't see what effect core.compression can have\n> on loading the ref_list.  Certainly the sort doesn't\n> load anything from the object database.  It would be\n> really good to profile and find out what is taking all\n> the time - I am assuming that the CPU is at 100% for the\n> 3+ minutes?\n\nYes, 100% CPU (I mostly run the tests at least twice and \nhave 8G of RAM, so I think the entire repo gets cached).\n\n\n> Random thought.  What happens to the with compression\n> case if you leave the commit in, but add a sleep(15) to\n> the end of sort_refs_list?\n\nWhy, what are you thinking?  Hmm, I am trying this on the \nnon gced repo and it doesn't seem to be completing (no cpu \nusage)!  It appears that perhaps it is being called many \ntimes (the sleeping would explain no cpu usage)?!?  This\ncould be a real problem, this should only get called once \nright?\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176255","messageId":"7vvcsf130i.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"201109261401.38624.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-26T20:07:25Z","receivedAt":"2011-09-26T20:07:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Fick <mfick@codeaurora.org> writes:\n\n> After a gc, the repo does perform the similar to the \n> uncompressed one (which was achieved via gc).  After gc, it \n> takes ~.05s do to a 'git branch' with 1.7.6 and \n> git.1.7.7.rc0.72.g4b5ea.  It also takes a bit more than 15s \n> with the patch reverted.  So it appears that compression is \n> not likely the culprit, but rather the need to be gced.\n\nIsn't packing refs part of \"gc\" these days?\n"},{"id":"176256","messageId":"9ae990f15489d7b51a172d08e63ca458@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109261401.38624.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-26T20:28:53Z","receivedAt":"2011-09-26T20:28:53Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 14:01:38 -0600, Martin Fick wrote:\n-- snip --\n> So, maybe you are correct, maybe my repo is the corner case?\n> Is a repo which needs to be gced considered a corner case?\n> Should git be able to detect that the repo is so in\n> desperate need of gcing?  Is it normal for git to need to gc\n> right after a clone and then fetching ~100K refs?\n\nWere you 100k refs packed before the gc?  If not, perhaps your refs are \ncausing a lot of trouble for the merge sort?  They will be written out \nsorted to the packed-refs file, so the merge sort won't have to do any \nreal work when loading them after that...\n\n> I am not sure what is right here, if this patch makes a repo\n> which needs gcing degrade 5 to 10 times worse than the\n> benefit of this patch, it still seems questionable to me.\n\nWell - it does this _for your repo_, that doesn't automatically mean \nthat it does generally, or frequently.  For instance, none of my normal \nrepos that have a lot of refs are Gerrit ones, and I wouldn't be \nsurprised if they benefitted from the merge sort (assuming that I am \nright that the merge sort is taking a long time on your gerrit refs).\n\nBesides, you would be better off running gc, and thus getting the \nbenefit too.\n\n>> Random thought.  What happens to the with compression\n>> case if you leave the commit in, but add a sleep(15) to\n>> the end of sort_refs_list?\n>\n> Why, what are you thinking?  Hmm, I am trying this on the\n> non gced repo and it doesn't seem to be completing (no cpu\n> usage)!  It appears that perhaps it is being called many\n> times (the sleeping would explain no cpu usage)?!?  This\n> could be a real problem, this should only get called once\n> right?\n\nI was just wondering if the time taken to get the refs was changing the \ninteraction with something else.  Not very likely, but ...\n\nI added a print statement, and it was called four times when I had \nunpacked refs, and once with packed.  So, maybe you are hitting some \nnasty case with unpacked refs.  If you use a print statement instead of \na sleep, how many times does sort_refs_lists get called in your unpacked \ncase?  It may well also be worth calculating the time taken to do the \nsort.\n\n-- \nJulian\n"},{"id":"176264","messageId":"201109261539.33437.mfick@codeaurora.org","threadId":"27589","inReplyTo":"9ae990f15489d7b51a172d08e63ca458@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T21:39:33Z","receivedAt":"2011-09-26T21:39:33Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 02:28:53 pm Julian Phillips \nwrote:\n> On Mon, 26 Sep 2011 14:01:38 -0600, Martin Fick wrote:\n> -- snip --\n> \n> > So, maybe you are correct, maybe my repo is the corner\n> > case? Is a repo which needs to be gced considered a\n> > corner case? Should git be able to detect that the\n> > repo is so in desperate need of gcing?  Is it normal\n> > for git to need to gc right after a clone and then\n> > fetching ~100K refs?\n> \n> Were you 100k refs packed before the gc?  If not, perhaps\n> your refs are causing a lot of trouble for the merge\n> sort?  They will be written out sorted to the\n> packed-refs file, so the merge sort won't have to do any\n> real work when loading them after that...\n\nI am not sure how to determine that (?), but I think they \nwere packed.  Under .git/objects/pack there were 2 large \nfiles, both close to 500MB.  Those 2 files constituted most \nof the space in the repo (I was wrong about the repo sizes, \nthat included the working dir, so think about half the \nquoted sizes for all of .git).  So does that mean it is \nmostly packed?  Aside from the pack and idx files, there was \nnothing else under the objects dir.  After gcing, it is down \nto just one ~500MB pack file.\n\n\n> > I am not sure what is right here, if this patch makes a\n> > repo which needs gcing degrade 5 to 10 times worse\n> > than the benefit of this patch, it still seems\n> > questionable to me.\n> \n> Well - it does this _for your repo_, that doesn't\n> automatically mean that it does generally, or\n> frequently.  \n\nOh, def agreed! I just didn't want to discount it so quickly \nas being a corner case.\n\n\n> For instance, none of my normal repos that\n> have a lot of refs are Gerrit ones, and I wouldn't be\n> surprised if they benefitted from the merge sort\n> (assuming that I am right that the merge sort is taking\n> a long time on your gerrit refs).\n> \n> Besides, you would be better off running gc, and thus\n> getting the benefit too.\n\nAgreed, which is why I was asking if git should have noticed \nmy \"degenerate\" case and auto gced?  But hopefully, there is \nan actual bug here somewhere and we both will get to eat our \ncake. :)\n\n\n\n> >> Random thought.  What happens to the with compression\n> >> case if you leave the commit in, but add a sleep(15)\n> >> to the end of sort_refs_list?\n> > \n> > Why, what are you thinking?  Hmm, I am trying this on\n> > the non gced repo and it doesn't seem to be completing\n> > (no cpu usage)!  It appears that perhaps it is being\n> > called many times (the sleeping would explain no cpu\n> > usage)?!?  This could be a real problem, this should\n> > only get called once right?\n> \n> I was just wondering if the time taken to get the refs\n> was changing the interaction with something else.  Not\n> very likely, but ...\n> \n> I added a print statement, and it was called four times\n> when I had unpacked refs, and once with packed.  So,\n> maybe you are hitting some nasty case with unpacked\n> refs.  If you use a print statement instead of a sleep,\n> how many times does sort_refs_lists get called in your\n> unpacked case?  It may well also be worth calculating\n> the time taken to do the sort.\n\nIn my case it was called 18785 times!  Any other tests I \nshould run?\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176267","messageId":"201109261552.04946.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109261539.33437.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-26T21:52:04Z","receivedAt":"2011-09-26T21:52:04Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 03:39:33 pm Martin Fick wrote:\n> On Monday, September 26, 2011 02:28:53 pm Julian Phillips\n> wrote:\n> > >> Random thought.  What happens to the with\n> > >> compression case if you leave the commit in, but\n> > >> add a sleep(15) to the end of sort_refs_list?\n> > > \n> > > Why, what are you thinking?  Hmm, I am trying this on\n> > > the non gced repo and it doesn't seem to be\n> > > completing (no cpu usage)!  It appears that perhaps\n> > > it is being called many times (the sleeping would\n> > > explain no cpu usage)?!?  This could be a real\n> > > problem, this should only get called once right?\n> > \n> > I was just wondering if the time taken to get the refs\n> > was changing the interaction with something else.  Not\n> > very likely, but ...\n> > \n> > I added a print statement, and it was called four times\n> > when I had unpacked refs, and once with packed.  So,\n> > maybe you are hitting some nasty case with unpacked\n> > refs.  If you use a print statement instead of a sleep,\n> > how many times does sort_refs_lists get called in your\n> > unpacked case?  It may well also be worth calculating\n> > the time taken to do the sort.\n> \n> In my case it was called 18785 times!  Any other tests I\n> should run?\n\nGerrit stores the changes in directories under refs/changes \nnamed after the last 2 digits of the change.  Then under \neach change it stores each patchset.  So it looks like this:   \nrefs/changes/dd/change_num/ps_num\n\nI noticed that:\n\n ls refs/changes/* | wc -l \n -> 18876\n\nsomewhat close, but not super close to 18785,  I am not sure \nif that is a clue.  It's almost like each change is causing \na re-sort,\n\n \n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176273","messageId":"981f42c890e69e8c4ac1958df52b2214@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109261539.33437.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-26T22:30:46Z","receivedAt":"2011-09-26T22:30:46Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 15:39:33 -0600, Martin Fick wrote:\n> On Monday, September 26, 2011 02:28:53 pm Julian Phillips\n> wrote:\n>> On Mon, 26 Sep 2011 14:01:38 -0600, Martin Fick wrote:\n>> -- snip --\n>>\n>> > So, maybe you are correct, maybe my repo is the corner\n>> > case? Is a repo which needs to be gced considered a\n>> > corner case? Should git be able to detect that the\n>> > repo is so in desperate need of gcing?  Is it normal\n>> > for git to need to gc right after a clone and then\n>> > fetching ~100K refs?\n>>\n>> Were you 100k refs packed before the gc?  If not, perhaps\n>> your refs are causing a lot of trouble for the merge\n>> sort?  They will be written out sorted to the\n>> packed-refs file, so the merge sort won't have to do any\n>> real work when loading them after that...\n>\n> I am not sure how to determine that (?), but I think they\n> were packed.  Under .git/objects/pack there were 2 large\n> files, both close to 500MB.  Those 2 files constituted most\n> of the space in the repo (I was wrong about the repo sizes,\n> that included the working dir, so think about half the\n> quoted sizes for all of .git).  So does that mean it is\n> mostly packed?  Aside from the pack and idx files, there was\n> nothing else under the objects dir.  After gcing, it is down\n> to just one ~500MB pack file.\n\nIf refs are listed under .git/refs/... they are unpacked, if they are \nlisted in .git/packed-refs they are packed.\nThey can be in both if updated since the last pack.\n\n>> > I am not sure what is right here, if this patch makes a\n>> > repo which needs gcing degrade 5 to 10 times worse\n>> > than the benefit of this patch, it still seems\n>> > questionable to me.\n>>\n>> Well - it does this _for your repo_, that doesn't\n>> automatically mean that it does generally, or\n>> frequently.\n>\n> Oh, def agreed! I just didn't want to discount it so quickly\n> as being a corner case.\n>\n>\n>> For instance, none of my normal repos that\n>> have a lot of refs are Gerrit ones, and I wouldn't be\n>> surprised if they benefitted from the merge sort\n>> (assuming that I am right that the merge sort is taking\n>> a long time on your gerrit refs).\n>>\n>> Besides, you would be better off running gc, and thus\n>> getting the benefit too.\n>\n> Agreed, which is why I was asking if git should have noticed\n> my \"degenerate\" case and auto gced?  But hopefully, there is\n> an actual bug here somewhere and we both will get to eat our\n> cake. :)\n\nI think automatic gc is currently only triggered by unpacked objects, \nnot unpacked refs ... perhaps the auto-gc should cover refs too?\n\n>> >> Random thought.  What happens to the with compression\n>> >> case if you leave the commit in, but add a sleep(15)\n>> >> to the end of sort_refs_list?\n>> >\n>> > Why, what are you thinking?  Hmm, I am trying this on\n>> > the non gced repo and it doesn't seem to be completing\n>> > (no cpu usage)!  It appears that perhaps it is being\n>> > called many times (the sleeping would explain no cpu\n>> > usage)?!?  This could be a real problem, this should\n>> > only get called once right?\n>>\n>> I was just wondering if the time taken to get the refs\n>> was changing the interaction with something else.  Not\n>> very likely, but ...\n>>\n>> I added a print statement, and it was called four times\n>> when I had unpacked refs, and once with packed.  So,\n>> maybe you are hitting some nasty case with unpacked\n>> refs.  If you use a print statement instead of a sleep,\n>> how many times does sort_refs_lists get called in your\n>> unpacked case?  It may well also be worth calculating\n>> the time taken to do the sort.\n>\n> In my case it was called 18785 times!  Any other tests I\n> should run?\n\nThat's a lot of sorts.  I really can't see why there would need to be \nmore than one ...\n\nI've created a new test repo, using a more complicated method to \nconstruct the 100k refs, and it took ~40m to run \"git branch\" instead of \nthe 1.2s for the previous repo.  So, I think the ref naming pattern used \nby Gerrit is definitely triggering something odd.  However, progress is \na bit slow - now that it takes over 1/2 an hour to try things out ...\n\n-- \nJulian\n"},{"id":"176284","messageId":"ece30e6a1b74bcddde5634003408f61f@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109261552.04946.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-26T23:26:55Z","receivedAt":"2011-09-26T23:26:55Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 15:52:04 -0600, Martin Fick wrote:\n> On Monday, September 26, 2011 03:39:33 pm Martin Fick wrote:\n>> On Monday, September 26, 2011 02:28:53 pm Julian Phillips\n>> wrote:\n>> > >> Random thought.  What happens to the with\n>> > >> compression case if you leave the commit in, but\n>> > >> add a sleep(15) to the end of sort_refs_list?\n>> > >\n>> > > Why, what are you thinking?  Hmm, I am trying this on\n>> > > the non gced repo and it doesn't seem to be\n>> > > completing (no cpu usage)!  It appears that perhaps\n>> > > it is being called many times (the sleeping would\n>> > > explain no cpu usage)?!?  This could be a real\n>> > > problem, this should only get called once right?\n>> >\n>> > I was just wondering if the time taken to get the refs\n>> > was changing the interaction with something else.  Not\n>> > very likely, but ...\n>> >\n>> > I added a print statement, and it was called four times\n>> > when I had unpacked refs, and once with packed.  So,\n>> > maybe you are hitting some nasty case with unpacked\n>> > refs.  If you use a print statement instead of a sleep,\n>> > how many times does sort_refs_lists get called in your\n>> > unpacked case?  It may well also be worth calculating\n>> > the time taken to do the sort.\n>>\n>> In my case it was called 18785 times!  Any other tests I\n>> should run?\n>\n> Gerrit stores the changes in directories under refs/changes\n> named after the last 2 digits of the change.  Then under\n> each change it stores each patchset.  So it looks like this:\n> refs/changes/dd/change_num/ps_num\n>\n> I noticed that:\n>\n>  ls refs/changes/* | wc -l\n>  -> 18876\n>\n> somewhat close, but not super close to 18785,  I am not sure\n> if that is a clue.  It's almost like each change is causing\n> a re-sort,\n\nbasically, it is ...\n\nBack when I made that change, I failed to notice that get_ref_dir was \nrecursive for subdirectories ... sorry ...\n\nHopefully this should speed things up.  My test repo went from ~17m \nuser time, to ~2.5s.\nPacking still make things much faster of course.\n\ndiff --git a/refs.c b/refs.c\nindex a615043..212e7ec 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -319,7 +319,7 @@ static struct ref_list *get_ref_dir(const char \n*submodule, c\n                 free(ref);\n                 closedir(dir);\n         }\n-       return sort_ref_list(list);\n+       return list;\n  }\n\n  struct warn_if_dangling_data {\n@@ -361,11 +361,13 @@ static struct ref_list *get_loose_refs(const char \n*submodu\n         if (submodule) {\n                 free_ref_list(submodule_refs.loose);\n                 submodule_refs.loose = get_ref_dir(submodule, \"refs\", \nNULL);\n+               submodule_refs.loose = \nsort_refs_list(submodule_refs.loose);\n                 return submodule_refs.loose;\n         }\n\n         if (!cached_refs.did_loose) {\n                 cached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n+               cached_refs.loose = sort_refs_list(cached_refs.loose);\n                 cached_refs.did_loose = 1;\n         }\n         return cached_refs.loose;\n\n\n\n>\n>\n> -Martin\n\n-- \nJulian\n"},{"id":"176286","messageId":"CAFfmPPNCCCo=40CVvjRebXvkR7H_wh9+cz=tGxHZ1LtarE+w+A@mail.gmail.com","threadId":"27589","inReplyTo":"ece30e6a1b74bcddde5634003408f61f@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"David Michael Barr","fromEmail":"davidbarr@google.com","sentAt":"2011-09-26T23:37:59Z","receivedAt":"2011-09-26T23:37:59Z","isPatch":false,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"On Tue, Sep 27, 2011 at 9:26 AM, Julian Phillips\n<julian@quantumfyre.co.uk> wrote:\n>\n> On Mon, 26 Sep 2011 15:52:04 -0600, Martin Fick wrote:\n>>\n>> On Monday, September 26, 2011 03:39:33 pm Martin Fick wrote:\n>>>\n>>> On Monday, September 26, 2011 02:28:53 pm Julian Phillips\n>>> wrote:\n>>> > >> Random thought.  What happens to the with\n>>> > >> compression case if you leave the commit in, but\n>>> > >> add a sleep(15) to the end of sort_refs_list?\n>>> > >\n>>> > > Why, what are you thinking?  Hmm, I am trying this on\n>>> > > the non gced repo and it doesn't seem to be\n>>> > > completing (no cpu usage)!  It appears that perhaps\n>>> > > it is being called many times (the sleeping would\n>>> > > explain no cpu usage)?!?  This could be a real\n>>> > > problem, this should only get called once right?\n>>> >\n>>> > I was just wondering if the time taken to get the refs\n>>> > was changing the interaction with something else.  Not\n>>> > very likely, but ...\n>>> >\n>>> > I added a print statement, and it was called four times\n>>> > when I had unpacked refs, and once with packed.  So,\n>>> > maybe you are hitting some nasty case with unpacked\n>>> > refs.  If you use a print statement instead of a sleep,\n>>> > how many times does sort_refs_lists get called in your\n>>> > unpacked case?  It may well also be worth calculating\n>>> > the time taken to do the sort.\n>>>\n>>> In my case it was called 18785 times!  Any other tests I\n>>> should run?\n>>\n>> Gerrit stores the changes in directories under refs/changes\n>> named after the last 2 digits of the change.  Then under\n>> each change it stores each patchset.  So it looks like this:\n>> refs/changes/dd/change_num/ps_num\n>>\n>> I noticed that:\n>>\n>>  ls refs/changes/* | wc -l\n>>  -> 18876\n>>\n>> somewhat close, but not super close to 18785,  I am not sure\n>> if that is a clue.  It's almost like each change is causing\n>> a re-sort,\n>\n> basically, it is ...\n>\n> Back when I made that change, I failed to notice that get_ref_dir was recursive for subdirectories ... sorry ...\n>\n> Hopefully this should speed things up.  My test repo went from ~17m user time, to ~2.5s.\n> Packing still make things much faster of course.\n>\n> diff --git a/refs.c b/refs.c\n> index a615043..212e7ec 100644\n> --- a/refs.c\n> +++ b/refs.c\n> @@ -319,7 +319,7 @@ static struct ref_list *get_ref_dir(const char *submodule, c\n>                free(ref);\n>                closedir(dir);\n>        }\n> -       return sort_ref_list(list);\n> +       return list;\n>  }\n>\n>  struct warn_if_dangling_data {\n> @@ -361,11 +361,13 @@ static struct ref_list *get_loose_refs(const char *submodu\n>        if (submodule) {\n>                free_ref_list(submodule_refs.loose);\n>                submodule_refs.loose = get_ref_dir(submodule, \"refs\", NULL);\n> +               submodule_refs.loose = sort_refs_list(submodule_refs.loose);\n>                return submodule_refs.loose;\n>        }\n>\n>        if (!cached_refs.did_loose) {\n>                cached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n> +               cached_refs.loose = sort_refs_list(cached_refs.loose);\n>                cached_refs.did_loose = 1;\n>        }\n>        return cached_refs.loose;\n>\n>\n>\n>>\n>>\n>> -Martin\n>\n> --\n> Julian\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\nWell done! I'll try to compose a patch attributed to Julian with the\ninformation from this thread.\n\n--\nDavid Barr\n"},{"id":"176287","messageId":"7vsjnizxf5.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"ece30e6a1b74bcddde5634003408f61f@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-26T23:38:54Z","receivedAt":"2011-09-26T23:38:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Julian Phillips <julian@quantumfyre.co.uk> writes:\n\n> Back when I made that change, I failed to notice that get_ref_dir was\n> recursive for subdirectories ... sorry ...\n\nAha, I also was blind while I was watching this discussion from the\nsideline, and I thought I re-read the codepath involved X-<. Indeed\nwe were sorting the list way too early and the patch looks correct.\n\nThanks.\n\n> Hopefully this should speed things up.  My test repo went from ~17m\n> user time, to ~2.5s.\n> Packing still make things much faster of course.\n>\n> diff --git a/refs.c b/refs.c\n> index a615043..212e7ec 100644\n> --- a/refs.c\n> +++ b/refs.c\n> @@ -319,7 +319,7 @@ static struct ref_list *get_ref_dir(const char\n> *submodule, c\n>                 free(ref);\n>                 closedir(dir);\n>         }\n> -       return sort_ref_list(list);\n> +       return list;\n>  }\n>\n>  struct warn_if_dangling_data {\n> @@ -361,11 +361,13 @@ static struct ref_list *get_loose_refs(const\n> char *submodu\n>         if (submodule) {\n>                 free_ref_list(submodule_refs.loose);\n>                 submodule_refs.loose = get_ref_dir(submodule, \"refs\",\n> NULL);\n> +               submodule_refs.loose =\n> sort_refs_list(submodule_refs.loose);\n>                 return submodule_refs.loose;\n>         }\n>\n>         if (!cached_refs.did_loose) {\n>                 cached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n> +               cached_refs.loose = sort_refs_list(cached_refs.loose);\n>                 cached_refs.did_loose = 1;\n>         }\n>         return cached_refs.loose;\n>\n>\n>\n>>\n>>\n>> -Martin\n"},{"id":"176291","messageId":"20110927000010.79913.71464.julian@quantumfyre.co.uk","threadId":"27589","inReplyTo":"7vsjnizxf5.fsf@alter.siamese.dyndns.org","subject":"[PATCH] Don't sort ref_list too early","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-27T00:00:09Z","receivedAt":"2011-09-27T00:00:09Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"get_ref_dir is called recursively for subdirectories, which means that\nwe were calling sort_ref_list for each directory of refs instead of\nonce for all the refs.  This is a massive wast of processing, so now\njust call sort_ref_list on the result of the top-level get_ref_dir, so\nthat the sort is only done once.\n\nIn the common case of only a few different directories of refs the\ndifference isn't very noticable, but it becomes very noticeable when\nyou have a large number of direcotries containing refs (e.g. as\ncreated by Gerrit).\n\nReported by Martin Fick.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\n\nThis time the typos are fixed too ... perhaps I wrote the original commit at 1am\ntoo ... :$\n\n refs.c |    4 +++-\n 1 files changed, 3 insertions(+), 1 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex a615043..a49ff74 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -319,7 +319,7 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\tfree(ref);\n \t\tclosedir(dir);\n \t}\n-\treturn sort_ref_list(list);\n+\treturn list;\n }\n \n struct warn_if_dangling_data {\n@@ -361,11 +361,13 @@ static struct ref_list *get_loose_refs(const char *submodule)\n \tif (submodule) {\n \t\tfree_ref_list(submodule_refs.loose);\n \t\tsubmodule_refs.loose = get_ref_dir(submodule, \"refs\", NULL);\n+\t\tsubmodule_refs.loose = sort_ref_list(submodule_refs.loose);\n \t\treturn submodule_refs.loose;\n \t}\n \n \tif (!cached_refs.did_loose) {\n \t\tcached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n+\t\tcached_refs.loose = sort_ref_list(cached_refs.loose);\n \t\tcached_refs.did_loose = 1;\n \t}\n \treturn cached_refs.loose;\n-- \n1.7.6.1\n"},{"id":"176293","messageId":"201109261812.31738.mfick@codeaurora.org","threadId":"27589","inReplyTo":"ece30e6a1b74bcddde5634003408f61f@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-27T00:12:31Z","receivedAt":"2011-09-27T00:12:31Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 05:26:55 pm Julian Phillips \nwrote:\n> On Mon, 26 Sep 2011 15:52:04 -0600, Martin Fick wrote:\n> > On Monday, September 26, 2011 03:39:33 pm Martin Fick \nwrote:\n> >> On Monday, September 26, 2011 02:28:53 pm Julian\n> >> Phillips\n> >> In my case it was called 18785 times!  Any other tests\n> >> I should run?\n> > \n> > Gerrit stores the changes in directories under\n> > refs/changes named after the last 2 digits of the\n> > change.  Then under each change it stores each\n> > patchset.  So it looks like this:\n> > refs/changes/dd/change_num/ps_num\n> > \n> > I noticed that:\n> >  ls refs/changes/* | wc -l\n> >  -> 18876\n> > \n> > somewhat close, but not super close to 18785,  I am not\n> > sure if that is a clue.  It's almost like each change\n> > is causing a re-sort,\n> \n> basically, it is ...\n> \n> Back when I made that change, I failed to notice that\n> get_ref_dir was recursive for subdirectories ... sorry\n> ...\n> \n> Hopefully this should speed things up.  My test repo went\n> from ~17m user time, to ~2.5s.\n> Packing still make things much faster of course.\n\nExcellent!  This works (almost, in my refs.c it is called \nsort_ref_list, not sort_refs_list).  So, on the non garbage \ncollected repo, git branch now takes ~.5s, and in the \ngarbage collected one it takes only ~.05s!\n\nThanks way much!!!\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176294","messageId":"97c45128ddeb8269273a4431b3941478@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109261812.31738.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-27T00:22:32Z","receivedAt":"2011-09-27T00:22:32Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 18:12:31 -0600, Martin Fick wrote:\n> On Monday, September 26, 2011 05:26:55 pm Julian Phillips\n> wrote:\n-- snip --\n>> Back when I made that change, I failed to notice that\n>> get_ref_dir was recursive for subdirectories ... sorry\n>> ...\n>>\n>> Hopefully this should speed things up.  My test repo went\n>> from ~17m user time, to ~2.5s.\n>> Packing still make things much faster of course.\n>\n> Excellent!  This works (almost, in my refs.c it is called\n> sort_ref_list, not sort_refs_list).\n\nYeah, in mine too ;)  It's late and I got the compile/send mail \nsequence backwards. :$\nIt's fixed in the proper patch email.\n\n>  So, on the non garbage\n> collected repo, git branch now takes ~.5s, and in the\n> garbage collected one it takes only ~.05s!\n\nThat sounds a lot better.  Hopefully other commands should be faster \nnow too.\n\n> Thanks way much!!!\n\nNo problem.  Thank you for all the time you've put in to help chase \nthis down.  Makes it so much easier when the person with original \nproblem mucks in with the investigation.\nJust think how much time you've saved for anyone with a large number of \nthose Gerrit change refs ;)\n\n-- \nJulian\n"},{"id":"176295","messageId":"1317085283-33943-1-git-send-email-davidbarr@google.com","threadId":"27589","inReplyTo":"CAFfmPPNCCCo=40CVvjRebXvkR7H_wh9+cz=tGxHZ1LtarE+w+A@mail.gmail.com","subject":"[PATCH] refs.c: Fix slowness with numerous loose refs","fromName":"David Barr","fromEmail":"davidbarr@google.com","sentAt":"2011-09-27T01:01:23Z","receivedAt":"2011-09-27T01:01:23Z","isPatch":true,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"Martin Fick reported:\n OK, I have found what I believe is another performance\n regression for large ref counts (~100K).\n\n When I run git br on my repo which only has one branch, but\n has ~100K refs under ref/changes (a gerrit repo), it takes\n normally 3-6mins depending on whether my caches are fresh or\n not.  After bisecting some older changes, I noticed that\n this ref seems to be where things start to get slow:\n v1.5.2-rc0~21^2 (refs.c: add a function to sort a ref list,\n rather then sorting on add) (Julian Phillips, Apr 17, 2007)\n\nMartin Fick observed that sort_refs_lists() was called almost\nas many times as there were loose refs.\n\nJulian Phillips commented:\n Back when I made that change, I failed to notice that get_ref_dir\n was recursive for subdirectories ... sorry ...\n\n Hopefully this should speed things up. My test repo went from\n ~17m user time, to ~2.5s.\n Packing still make things much faster of course.\n\nMartin Fick acked:\n Excellent!  This works (almost, in my refs.c it is called\n sort_ref_list, not sort_refs_list).  So, on the non garbage\n collected repo, git branch now takes ~.5s, and in the\n garbage collected one it takes only ~.05s!\n\n[db: summarised transcript, rewrote patch to fix callee not callers]\n\n[attn jch: patch applies to maint]\n\nAnalyzed-by: Martin Fick <mfick@codeaurora.org>\nInspired-by: Julian Phillips <julian@quantumfyre.co.uk>\nAcked-by: Martin Fick <mfick@codeaurora.org>\nSigned-off-by: David Barr <davidbarr@google.com>\n---\n refs.c |   14 ++++++++++----\n 1 files changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex 4c1fd47..e40a09c 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -255,8 +255,8 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \treturn refs->packed;\n }\n \n-static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n-\t\t\t\t    struct ref_list *list)\n+static struct ref_list *walk_ref_dir(const char *submodule, const char *base,\n+\t\t\t\t     struct ref_list *list)\n {\n \tDIR *dir;\n \tconst char *path;\n@@ -299,7 +299,7 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\tif (stat(refdir, &st) < 0)\n \t\t\t\tcontinue;\n \t\t\tif (S_ISDIR(st.st_mode)) {\n-\t\t\t\tlist = get_ref_dir(submodule, ref, list);\n+\t\t\t\tlist = walk_ref_dir(submodule, ref, list);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (submodule) {\n@@ -319,7 +319,13 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\tfree(ref);\n \t\tclosedir(dir);\n \t}\n-\treturn sort_ref_list(list);\n+\treturn list;\n+}\n+\n+static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n+\t\t\t\t    struct ref_list *list)\n+{\n+\treturn sort_ref_list(walk_ref_dir(submodule, base, list));\n }\n \n struct warn_if_dangling_data {\n-- \n1.7.5.75.g69330\n"},{"id":"176296","messageId":"CAFfmPPMx9_nRE2Zfg2g0hwzybWDPJARc6LCHbSK8y-uZWQCZqQ@mail.gmail.com","threadId":"27589","inReplyTo":"1317085283-33943-1-git-send-email-davidbarr@google.com","subject":"Re: [PATCH] refs.c: Fix slowness with numerous loose refs","fromName":"David Michael Barr","fromEmail":"davidbarr@google.com","sentAt":"2011-09-27T02:04:43Z","receivedAt":"2011-09-27T02:04:43Z","isPatch":true,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"+cc Shawn O. Pearce\n\nI used the following to generate a test repo shaped like\na gerrit mirror with unpacked refs (10k, because life is too short for\n100k tests):\n\ncd test.git\ngit init\ntouch empty\ngit add empty\ngit commit -m 'empty'\nREV=`git rev-parse HEAD`\nfor ((d=0;d<100;++d)); do\n for ((n=0;n<100;++n)); do\n  let r=n*100+d\n  mkdir -p .git/refs/changes/$d/$r\n  echo $REV > .git/refs/changes/$d/$r/1\n done\ndone\ntime git branch xyz\n\nWith warm caches...\n\nGit 1.7.6.4:\nreal\t0m8.232s\nuser\t0m7.842s\nsys\t0m0.385s\n\nGit 1.7.6.4, with patch below:\nreal\t0m0.394s\nuser\t0m0.069s\nsys\t0m0.324s\n\nOn Tue, Sep 27, 2011 at 11:01 AM, David Barr <davidbarr@google.com> wrote:\n> Martin Fick reported:\n>  OK, I have found what I believe is another performance\n>  regression for large ref counts (~100K).\n>\n>  When I run git br on my repo which only has one branch, but\n>  has ~100K refs under ref/changes (a gerrit repo), it takes\n>  normally 3-6mins depending on whether my caches are fresh or\n>  not.  After bisecting some older changes, I noticed that\n>  this ref seems to be where things start to get slow:\n>  v1.5.2-rc0~21^2 (refs.c: add a function to sort a ref list,\n>  rather then sorting on add) (Julian Phillips, Apr 17, 2007)\n>\n> Martin Fick observed that sort_refs_lists() was called almost\n> as many times as there were loose refs.\n>\n> Julian Phillips commented:\n>  Back when I made that change, I failed to notice that get_ref_dir\n>  was recursive for subdirectories ... sorry ...\n>\n>  Hopefully this should speed things up. My test repo went from\n>  ~17m user time, to ~2.5s.\n>  Packing still make things much faster of course.\n>\n> Martin Fick acked:\n>  Excellent!  This works (almost, in my refs.c it is called\n>  sort_ref_list, not sort_refs_list).  So, on the non garbage\n>  collected repo, git branch now takes ~.5s, and in the\n>  garbage collected one it takes only ~.05s!\n>\n> [db: summarised transcript, rewrote patch to fix callee not callers]\n>\n> [attn jch: patch applies to maint]\n>\n> Analyzed-by: Martin Fick <mfick@codeaurora.org>\n> Inspired-by: Julian Phillips <julian@quantumfyre.co.uk>\n> Acked-by: Martin Fick <mfick@codeaurora.org>\n> Signed-off-by: David Barr <davidbarr@google.com>\n> ---\n>  refs.c |   14 ++++++++++----\n>  1 files changed, 10 insertions(+), 4 deletions(-)\n>\n> diff --git a/refs.c b/refs.c\n> index 4c1fd47..e40a09c 100644\n> --- a/refs.c\n> +++ b/refs.c\n> @@ -255,8 +255,8 @@ static struct ref_list *get_packed_refs(const char *submodule)\n>        return refs->packed;\n>  }\n>\n> -static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n> -                                   struct ref_list *list)\n> +static struct ref_list *walk_ref_dir(const char *submodule, const char *base,\n> +                                    struct ref_list *list)\n>  {\n>        DIR *dir;\n>        const char *path;\n> @@ -299,7 +299,7 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n>                        if (stat(refdir, &st) < 0)\n>                                continue;\n>                        if (S_ISDIR(st.st_mode)) {\n> -                               list = get_ref_dir(submodule, ref, list);\n> +                               list = walk_ref_dir(submodule, ref, list);\n>                                continue;\n>                        }\n>                        if (submodule) {\n> @@ -319,7 +319,13 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n>                free(ref);\n>                closedir(dir);\n>        }\n> -       return sort_ref_list(list);\n> +       return list;\n> +}\n> +\n> +static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n> +                                   struct ref_list *list)\n> +{\n> +       return sort_ref_list(walk_ref_dir(submodule, base, list));\n>  }\n>\n>  struct warn_if_dangling_data {\n> --\n> 1.7.5.75.g69330\n>\n>\n\n\n\n-- \n\nDavid Barr | Software Engineer | davidbarr@google.com | 614-3438-8348\n"},{"id":"176301","messageId":"3539dab7-0fb2-4759-baaf-8e22efab2904@email.android.com","threadId":"27589","inReplyTo":"97c45128ddeb8269273a4431b3941478@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-27T02:34:02Z","receivedAt":"2011-09-27T02:34:02Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"\n\n> Julian Phillips <julian@quantumfyre.co.uk> wrote:\n>On Mon, 26 Sep 2011 18:12:31 -0600, Martin Fick wrote:\n>That sounds a lot better.  Hopefully other commands should be faster \n>now too.\n\nYeah, I will try this in a few other places to see.\n\n>> Thanks way much!!!\n>\n>No problem.  Thank you for all the time you've put in to help chase \n>this down.  Makes it so much easier when the person with original \n>problem mucks in with the investigation.\n>Just think how much time you've saved for anyone with a large number of\n>\n>those Gerrit change refs ;)\n\n Perhaps this is a naive question, but why are all these refs being put into a list to be sorted, only to be discarded soon thereafter anyway?  After all, git branch knows that it isn't going to print these, and the refs are stored precategorized, so why not only grab the refs which matter upfront?\n\n-Martin \n"},{"id":"176313","messageId":"2ad152520590530dd1bdcfae9941ef9d@quantumfyre.co.uk","threadId":"27589","inReplyTo":"3539dab7-0fb2-4759-baaf-8e22efab2904@email.android.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-27T07:59:24Z","receivedAt":"2011-09-27T07:59:24Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 26 Sep 2011 20:34:02 -0600, Martin Fick wrote:\n>> Julian Phillips <julian@quantumfyre.co.uk> wrote:\n>>On Mon, 26 Sep 2011 18:12:31 -0600, Martin Fick wrote:\n>>That sounds a lot better.  Hopefully other commands should be faster\n>>now too.\n>\n> Yeah, I will try this in a few other places to see.\n>\n>>> Thanks way much!!!\n>>\n>>No problem.  Thank you for all the time you've put in to help chase\n>>this down.  Makes it so much easier when the person with original\n>>problem mucks in with the investigation.\n>>Just think how much time you've saved for anyone with a large number \n>> of\n>>\n>>those Gerrit change refs ;)\n>\n>  Perhaps this is a naive question, but why are all these refs being\n> put into a list to be sorted, only to be discarded soon thereafter\n> anyway?  After all, git branch knows that it isn't going to print\n> these, and the refs are stored precategorized, so why not only grab\n> the refs which matter upfront?\n\nI can't say that I am aware of a specific decision having been taken on \nthe subject, but I'll have a guess at the reason:\n\nThe extra code it would take to have an API for getting a list of only \na subset of the refs has never been considered worth the cost.  It would \ntake effort to implement, test and maintain - and it would have to be \ndone separately for packed and unpacked cases to avoid still loading and \ndiscarding unwanted refs.  All that to not do something that no-one has \nnoticed taking any time?  Until now, I doubt anyone has considered it \nsomething that was a problem - and now that even with 100k refs it takes \nless than a second, I doubt anyone will feel all that inclined to have a \ncrack at it now either.\n\n-- \nJulian\n"},{"id":"176314","messageId":"CAGdFq_hvR1MPF33YFcjDCzCM0iOO2zpiiePFFS4dBabu84cwTg@mail.gmail.com","threadId":"27589","inReplyTo":"ece30e6a1b74bcddde5634003408f61f@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-09-27T08:20:29Z","receivedAt":"2011-09-27T08:20:29Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Sep 27, 2011 at 01:26, Julian Phillips <julian@quantumfyre.co.uk> wrote:\n> Back when I made that change, I failed to notice that get_ref_dir was\n> recursive for subdirectories ... sorry ...\n>\n> Hopefully this should speed things up.  My test repo went from ~17m user\n> time, to ~2.5s.\n> Packing still make things much faster of course.\n\nCan we perhaps also have some tests to prevent this from happening again?\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"176318","messageId":"22f055b34840e3c64f3339f7b3dc6920@quantumfyre.co.uk","threadId":"27589","inReplyTo":"CAGdFq_hvR1MPF33YFcjDCzCM0iOO2zpiiePFFS4dBabu84cwTg@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-27T09:01:56Z","receivedAt":"2011-09-27T09:01:56Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Tue, 27 Sep 2011 10:20:29 +0200, Sverre Rabbelier wrote:\n> Heya,\n>\n> On Tue, Sep 27, 2011 at 01:26, Julian Phillips\n> <julian@quantumfyre.co.uk> wrote:\n>> Back when I made that change, I failed to notice that get_ref_dir \n>> was\n>> recursive for subdirectories ... sorry ...\n>>\n>> Hopefully this should speed things up.  My test repo went from ~17m \n>> user\n>> time, to ~2.5s.\n>> Packing still make things much faster of course.\n>\n> Can we perhaps also have some tests to prevent this from happening \n> again?\n\nUm ... any suggestion what to test?\n\nIt has to be hot-cache, otherwise time taken to read the refs from disk \nwill mean that it is always slow.  On my Mac it seems to _always_ be \nslow reading the refs from disk, so even the \"fast\" case still takes \n~17m.\n\nAlso, what counts as ok, and what as broken?\n\n-- \nJulian\n"},{"id":"176325","messageId":"CAGdFq_j2aa8bwxWuJvEsgA_1zkR4mMzoKjGs9TQVqw+0XYr98A@mail.gmail.com","threadId":"27589","inReplyTo":"22f055b34840e3c64f3339f7b3dc6920@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-09-27T10:01:24Z","receivedAt":"2011-09-27T10:01:24Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Tue, Sep 27, 2011 at 11:01, Julian Phillips <julian@quantumfyre.co.uk> wrote:\n> It has to be hot-cache, otherwise time taken to read the refs from disk will\n> mean that it is always slow.  On my Mac it seems to _always_ be slow reading\n> the refs from disk, so even the \"fast\" case still takes ~17m.\n\nAh, that seems unfortunate. Not sure how to test it then.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"176326","messageId":"CACsJy8Dxj72U+wiu6MUV3sQYso948Gi6DhTLSO7=ddt6BxdBYQ@mail.gmail.com","threadId":"27589","inReplyTo":"CAGdFq_j2aa8bwxWuJvEsgA_1zkR4mMzoKjGs9TQVqw+0XYr98A@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-09-27T10:25:08Z","receivedAt":"2011-09-27T10:25:08Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Sep 27, 2011 at 8:01 PM, Sverre Rabbelier <srabbelier@gmail.com> wrote:\n> Heya,\n>\n> On Tue, Sep 27, 2011 at 11:01, Julian Phillips <julian@quantumfyre.co.uk> wrote:\n>> It has to be hot-cache, otherwise time taken to read the refs from disk will\n>> mean that it is always slow.  On my Mac it seems to _always_ be slow reading\n>> the refs from disk, so even the \"fast\" case still takes ~17m.\n>\n> Ah, that seems unfortunate. Not sure how to test it then.\n\nIf you care about performance, a perf test suite could be made,\nperhaps as a separate project. The output would be charts or\nspreadsheets, that interesting parties can look at and point out\nregressions. We may start with a set of common used operations.\n-- \nDuy\n"},{"id":"176327","messageId":"4E81AE63.8040008@alum.mit.edu","threadId":"27589","inReplyTo":"22f055b34840e3c64f3339f7b3dc6920@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-09-27T11:07:15Z","receivedAt":"2011-09-27T11:07:15Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 09/27/2011 11:01 AM, Julian Phillips wrote:\n> It has to be hot-cache, otherwise time taken to read the refs from disk\n> will mean that it is always slow.  On my Mac it seems to _always_ be\n> slow reading the refs from disk, so even the \"fast\" case still takes ~17m.\n\nThis case should be helped by lazy-loading of loose references, which I\nam working on.  So if you develop some benchmarking code, it would help\nme with my work.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176329","messageId":"e44cba8038381a722e127fb47b8ddfe4@quantumfyre.co.uk","threadId":"27589","inReplyTo":"4E81AE63.8040008@alum.mit.edu","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-27T12:10:25Z","receivedAt":"2011-09-27T12:10:25Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Tue, 27 Sep 2011 13:07:15 +0200, Michael Haggerty wrote:\n> On 09/27/2011 11:01 AM, Julian Phillips wrote:\n>> It has to be hot-cache, otherwise time taken to read the refs from \n>> disk\n>> will mean that it is always slow.  On my Mac it seems to _always_ be\n>> slow reading the refs from disk, so even the \"fast\" case still takes \n>> ~17m.\n>\n> This case should be helped by lazy-loading of loose references, which \n> I\n> am working on.  So if you develop some benchmarking code, it would \n> help\n> me with my work.\n\nThe attached script creates the repo structure I was testing with ...\n\nIf you create a repo with 100k refs it takes quite a while to read the \nrefs from disk.  If you are lazy-loading then it should take practically \nno time, since the only interesting ref is refs/heads/master.\n\nThe following is the hot-cache timing for \"./refs-stress c 40000\", with \nthe sorting patch applied (wasn't prepared to wait for numbers with 100k \nrefs).\n\njp3@rayne: refs>(cd c; time ~/misc/git/git/git branch)\n* master\n\nreal    0m0.885s\nuser    0m0.161s\nsys     0m0.722s\n\nAfter doing \"rm -rf c/.git/refs/changes/*\", I get:\n\njp3@rayne: refs>(cd c; time ~/misc/git/git/git branch)\n* master\n\nreal    0m0.004s\nuser    0m0.001s\nsys     0m0.002s\n\n-- \nJulian\n\n#!/usr/bin/env python\n\nimport os\nimport random\nimport subprocess\nimport sys\n\ndef die(msg):\n    print >> sys.stderr, msg\n    sys.exit(1)\n\ndef new_ref(a, b, commit):\n    d = \".git/refs/changes/%d/%d\" % (a, b)\n    if not os.path.exists(d):\n        os.makedirs(d)\n    e = 1\n    p = \"%s/%d\" % (d, e)\n    while os.path.exists(p):\n        e += 1\n        p = \"%s/%d\" % (d, e)\n    f = open(p, \"w\")\n    f.write(commit)\n    f.close()\n\ndef make_refs(count, commit):\n    while count > 0:\n        sys.stdout.write(\"left: %d%s\\r\" % (count, \" \" * 30))\n        a = random.randrange(10, 30)\n        b = random.randrange(10000, 50000)\n        new_ref(a, b, commit)\n        count -= 1\n    print \"refs complete\"\n\ndef main():\n    if len(sys.argv) != 3:\n        die(\"usage: %s <name> <ref count>\" % sys.argv[0])\n\n    _, name, refs = sys.argv\n\n    os.mkdir(name)\n    os.chdir(name)\n\n    if subprocess.call([\"git\", \"init\"]) != 0:\n        die(\"failed to init repo\")\n\n    f = open(\"foobar.txt\", \"w\")\n    f.write(\"%s: %s refs\\n\" % (name, refs))\n    f.close()\n\n    if subprocess.call([\"git\", \"add\", \"foobar.txt\"]) != 0:\n        die(\"failed to add foobar.txt\")\n\n    if subprocess.call([\"git\", \"commit\", \"-m\", \"inital commit\"]) != 0:\n        die(\"failed to create initial commit\")\n\n    commit = subprocess.check_output([\"git\", \"show-ref\", \"-s\", \"master\"]).strip()\n\n    make_refs(int(refs), commit)\n\nif __name__ == \"__main__\":\n    main()\n"},{"id":"176451","messageId":"201109281338.04378.mfick@codeaurora.org","threadId":"27589","inReplyTo":"CAP8UFD3TWQHU0wLPuxMDnc3bRSz90Yd+yDMBe03kofeo-nr7yA@mail.gmail.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-28T19:38:04Z","receivedAt":"2011-09-28T19:38:04Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, September 26, 2011 06:41:04 am Christian Couder \nwrote:\n> On Sun, Sep 25, 2011 at 10:43 PM, Martin Fick \n<mfick@codeaurora.org> wrote:\n...\n> >  git checkout\n> > \n> > can also take rather long periods of time > 3 mins when\n> > run on a repo with ~100K refs.\n...\n> >  So, I bisected this issue also, and it seems that the\n> > \"offending\" commit is\n...\n> > commit 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07\n> > Author: Christian Couder <chriscool@tuxfamily.org>\n> > \n> >    replace_object: add mechanism to replace objects\n> > found in \"refs/replace/\"\n...\n\n> I don't think there is an obvious problem with it, but it\n> would be nice if you could dig a bit deeper.\n> \n> The first thing that could take a lot of time is the call\n> to for_each_replace_ref() in this function:\n> \n> +static void prepare_replace_object(void)\n> +{\n> +       static int replace_object_prepared;\n> +\n> +       if (replace_object_prepared)\n> +               return;\n> +\n> +       for_each_replace_ref(register_replace_ref, NULL);\n> +       replace_object_prepared = 1;\n> +}\n\nThe time was actually spent in for_each_replace_ref()\nwhich calls get_loose_refs() which has the recursive bug \nthat Julian Phillips fixed 2 days ago.  Good to see that \nthis fix helps other use cases too.\n\nSo with that bug fixed, the thing taking the most time now \nfor a git checkout with ~100K refs seems to be the orphan \ncheck as Thomas predicted.  The strange part with this, is \nthat the orphan check seems to take only about ~20s in the \nrepo where the refs aren't packed.  However, in the repo \nwhere they are packed, this check takes at least 5min!  This \nseems a bit unusual, doesn't it?  Is the filesystem that \nmuch better at indexing refs than git's pack mechanism?  \nSeems unlikely, the unpacked refs take 312M in the FS, the \npacked ones only take about 4.3M.  I suspect their is \nsomething else unexpected going on here in the packed ref \ncase.  \n\nAny thoughts?  I will dig deeper...\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176456","messageId":"201109281610.49322.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109281338.04378.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-28T22:10:48Z","receivedAt":"2011-09-28T22:10:48Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, September 28, 2011 01:38:04 pm Martin Fick \nwrote:\n> On Monday, September 26, 2011 06:41:04 am Christian\n> Couder\n> \n> wrote:\n> > On Sun, Sep 25, 2011 at 10:43 PM, Martin Fick\n> \n> <mfick@codeaurora.org> wrote:\n> ...\n> \n> > >  git checkout\n> > > \n> > > can also take rather long periods of time > 3 mins\n> > > when run on a repo with ~100K refs.\n> \n> ...\n> \n> > >  So, I bisected this issue also, and it seems that\n> > >  the\n> > > \n> > > \"offending\" commit is\n> \n> ...\n> \n> > > commit 680955702990c1d4bfb3c6feed6ae9c6cb5c3c07\n> > > Author: Christian Couder <chriscool@tuxfamily.org>\n> > > \n> > >    replace_object: add mechanism to replace objects\n> > > \n> > > found in \"refs/replace/\"\n> \n> ...\n> \n> > I don't think there is an obvious problem with it, but\n> > it would be nice if you could dig a bit deeper.\n> > \n> > The first thing that could take a lot of time is the\n> > call to for_each_replace_ref() in this function:\n> > \n> > +static void prepare_replace_object(void)\n> > +{\n> > +       static int replace_object_prepared;\n> > +\n> > +       if (replace_object_prepared)\n> > +               return;\n> > +\n> > +       for_each_replace_ref(register_replace_ref,\n> > NULL); +       replace_object_prepared = 1;\n> > +}\n> \n> The time was actually spent in for_each_replace_ref()\n> which calls get_loose_refs() which has the recursive bug\n> that Julian Phillips fixed 2 days ago.  Good to see that\n> this fix helps other use cases too.\n> \n> So with that bug fixed, the thing taking the most time\n> now for a git checkout with ~100K refs seems to be the\n> orphan check as Thomas predicted.  The strange part with\n> this, is that the orphan check seems to take only about\n> ~20s in the repo where the refs aren't packed.  However,\n> in the repo where they are packed, this check takes at\n> least 5min!  This seems a bit unusual, doesn't it?  Is\n> the filesystem that much better at indexing refs than\n> git's pack mechanism? Seems unlikely, the unpacked refs\n> take 312M in the FS, the packed ones only take about\n> 4.3M.  I suspect their is something else unexpected\n> going on here in the packed ref case.\n> \n> Any thoughts?  I will dig deeper...\n\nI think the problem is that resolve_ref() walks a linked \nlist of searching for the packed ref.  Does this mean that \npacked refs are not indexed at all?\n\n> \n> -Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176466","messageId":"c76d7f65203c0fc2c6e4e14fe2f33274@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109281610.49322.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T00:54:00Z","receivedAt":"2011-09-29T00:54:00Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Wed, 28 Sep 2011 16:10:48 -0600, Martin Fick wrote:\n> On Wednesday, September 28, 2011 01:38:04 pm Martin Fick\n> wrote:\n-- snip --\n>> So with that bug fixed, the thing taking the most time\n>> now for a git checkout with ~100K refs seems to be the\n>> orphan check as Thomas predicted.  The strange part with\n>> this, is that the orphan check seems to take only about\n>> ~20s in the repo where the refs aren't packed.  However,\n>> in the repo where they are packed, this check takes at\n>> least 5min!  This seems a bit unusual, doesn't it?  Is\n>> the filesystem that much better at indexing refs than\n>> git's pack mechanism? Seems unlikely, the unpacked refs\n>> take 312M in the FS, the packed ones only take about\n>> 4.3M.  I suspect their is something else unexpected\n>> going on here in the packed ref case.\n>>\n>> Any thoughts?  I will dig deeper...\n>\n> I think the problem is that resolve_ref() walks a linked\n> list of searching for the packed ref.  Does this mean that\n> packed refs are not indexed at all?\n\nAre you sure that it is walking the linked list that is the problem?  \nI've created a test repo with ~100k refs/changes/... style refs, and \n~40000 refs/heads/... style refs, and checkout can walk the list of \n~140k refs seven times in 85ms user time including doing whatever other \nprocessing is needed for checkout.  The real time is only 114ms - but \nthen my test repo has no real data in.\n\nIf resolve_ref() walking the linked list of refs was the problem, then \nI would expect my test repo to show the same problem.  It doesn't, a pre \nref-packing checkout took minutes (~0.5s user time), whereas a \nref-packed checkout takes ~0.1s.  So, I would suggest that the problem \nlies elsewhere.\n\nHave you tried running a checkout whilst profiling?\n\n-- \nJulian\n"},{"id":"176469","messageId":"960aacbf-8d4d-4b2a-8902-f6380ff9febd@email.android.com","threadId":"27589","inReplyTo":"c76d7f65203c0fc2c6e4e14fe2f33274@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-29T01:37:18Z","receivedAt":"2011-09-29T01:37:18Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday 28 September 2011 18:59:09 Martin Fick wrote: \n> Julian Phillips <julian@quantumfyre.co.uk> wrote: \n> > On Wed, 28 Sep 2011 16:10:48 -0600, Martin Fick wrote: \n> >> So with that bug fixed, the thing taking the most time \n> >> now for a git checkout with ~100K refs seems to be the \n> >> orphan check as Thomas predicted. The strange part with \n> >> this, is that the orphan check seems to take only about \n> >> ~20s in the repo where the refs aren't packed. However, \n> >> in the repo where they are packed, this check takes at \n> >> least 5min! This seems a bit unusual, doesn't it? Is \n> >> the filesystem that much better at indexing refs than \n> >> git's pack mechanism? Seems unlikely, the unpacked refs \n> >> take 312M in the FS, the packed ones only take about \n> >> 4.3M. I suspect their is something else unexpected \n> >> going on here in the packed ref case. \n> >> \n> >> Any thoughts? I will dig deeper... \n> > \n> > I think the problem is that resolve_ref() walks a linked \n> > list of searching for the packed ref. Does this mean that \n> > packed refs are not indexed at all? \n> > Are you sure that it is walking the linked list that is the problem?\n\nIt sure seems like it.\n\n> I've created a test repo with ~100k refs/changes/... style refs, and \n> ~40000 refs/heads/... style refs, and checkout can walk the list of \n> ~140k refs seven times in 85ms user time including doing whatever other \n> processing is needed for checkout. The real time is only 114ms - but \n> then my test repo has no real data in.\n\nIf I understand what you are saying, it sounds like you do not have a very good test case. The amount of time it takes for checkout depends on how long it takes to find a ref with the sha1 that you are on. If that sha1 is so early in the list of refs that it only took you 7 traversals to find it, then that is not a very good testcase. I think that you should probably try making an orphaned ref (checkout a detached head, commit to it), that is probably the worst testcase since it should then have to search all 140K refs to eventually give up.\n\nAgain, if I understand what you are saying, if it took 85ms for 7 traversals, then it takes approximately 10ms per traversal, that's only 100/s! If you have to traverse it 140K times, that should work out to 1400s ~ 23mins.\n\n> If resolve_ref() walking the linked list of refs was the problem, then > I would expect my test repo to show the same problem. It doesn't, a pre \n> ref-packing checkout took minutes (~0.5s user time), whereas a \n> ref-packed checkout takes ~0.1s. So, I would suggest that the problem > lies elsewhere. \n> \n> Have you tried running a checkout whilst profiling?\n\nNo, to be honest, I am not familiar with any profilling tools.\n\n-Martin\n\nEmployee of Qualcomm Innovation Center,Inc. which is a member of Code Aurora Forum\n"},{"id":"176472","messageId":"7c0105c6cca7dd0aa336522f90617fe4@quantumfyre.co.uk","threadId":"27589","inReplyTo":"960aacbf-8d4d-4b2a-8902-f6380ff9febd@email.android.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T02:19:16Z","receivedAt":"2011-09-29T02:19:16Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Wed, 28 Sep 2011 19:37:18 -0600, Martin Fick wrote:\n> On Wednesday 28 September 2011 18:59:09 Martin Fick wrote:\n>> Julian Phillips <julian@quantumfyre.co.uk> wrote:\n-- snip --\n>> I've created a test repo with ~100k refs/changes/... style refs, and\n>> ~40000 refs/heads/... style refs, and checkout can walk the list of\n>> ~140k refs seven times in 85ms user time including doing whatever \n>> other\n>> processing is needed for checkout. The real time is only 114ms - but\n>> then my test repo has no real data in.\n>\n> If I understand what you are saying, it sounds like you do not have a\n> very good test case. The amount of time it takes for checkout depends\n> on how long it takes to find a ref with the sha1 that you are on. If\n> that sha1 is so early in the list of refs that it only took you 7\n> traversals to find it, then that is not a very good testcase. I think\n> that you should probably try making an orphaned ref (checkout a\n> detached head, commit to it), that is probably the worst testcase\n> since it should then have to search all 140K refs to eventually give\n> up.\n>\n> Again, if I understand what you are saying, if it took 85ms for 7\n> traversals, then it takes approximately 10ms per traversal, that's\n> only 100/s! If you have to traverse it 140K times, that should work\n> out to 1400s ~ 23mins.\n\nWell, it's no more than 10ms per traversal - since the rest of the work \npresumably takes some time too ...\n\nHowever, I had forgotten to make the orphaned commit as you suggest - \nand then _bang_ 7N^2, it tries seven different variants of each ref \n(which is silly as they are all fully qualified), and with packed refs \nit has to search for them each time, all to turn names into hashes that \nwe already know to start with.\n\nSo, yes - it is that list traversal.\n\nDoes the following help?\n\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 5e356a6..f0f4ca1 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -605,7 +605,7 @@ static int add_one_ref_to_rev_list_arg(const char \n*refname,\n                                        int flags,\n                                        void *cb_data)\n  {\n-       add_one_rev_list_arg(cb_data, refname);\n+       add_one_rev_list_arg(cb_data, strdup(sha1_to_hex(sha1)));\n         return 0;\n  }\n\n-- \nJulian\n"},{"id":"176530","messageId":"20110929041811.5363.33396.julian@quantumfyre.co.uk","threadId":"27589","inReplyTo":"7vy5x7rwq9.fsf@alter.siamese.dyndns.org","subject":"[PATCH] refs: Use binary search to lookup refs faster","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T04:18:10Z","receivedAt":"2011-09-29T04:18:10Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"Currently we linearly search through lists of refs when we need to\nfind a specific ref.  This can be very slow if we need to lookup a\nlarge number of refs.  By changing to a binary search we can make this\nfaster.\n\nIn order to be able to use a binary search we need to change from\nusing linked lists to arrays, which we can manage using ALLOC_GROW.\n\nWe can now also use the standard library qsort function to sort the\nrefs arrays.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\n\nSomething like this?\n\n refs.c |  328 ++++++++++++++++++++++++++--------------------------------------\n 1 files changed, 131 insertions(+), 197 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex a49ff74..e411bea 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -8,14 +8,18 @@\n #define REF_KNOWS_PEELED 04\n #define REF_BROKEN 010\n \n-struct ref_list {\n-\tstruct ref_list *next;\n+struct ref_entry {\n \tunsigned char flag; /* ISSYMREF? ISPACKED? */\n \tunsigned char sha1[20];\n \tunsigned char peeled[20];\n \tchar name[FLEX_ARRAY];\n };\n \n+struct ref_array {\n+\tint nr, alloc;\n+\tstruct ref_entry **refs;\n+};\n+\n static const char *parse_ref_line(char *line, unsigned char *sha1)\n {\n \t/*\n@@ -44,108 +48,55 @@ static const char *parse_ref_line(char *line, unsigned char *sha1)\n \treturn line;\n }\n \n-static struct ref_list *add_ref(const char *name, const unsigned char *sha1,\n-\t\t\t\tint flag, struct ref_list *list,\n-\t\t\t\tstruct ref_list **new_entry)\n+static void add_ref(const char *name, const unsigned char *sha1,\n+\t\t    int flag, struct ref_array *refs,\n+\t\t    struct ref_entry **new_entry)\n {\n \tint len;\n-\tstruct ref_list *entry;\n+\tstruct ref_entry *entry;\n \n \t/* Allocate it and add it in.. */\n \tlen = strlen(name) + 1;\n-\tentry = xmalloc(sizeof(struct ref_list) + len);\n+\tentry = xmalloc(sizeof(struct ref) + len);\n \thashcpy(entry->sha1, sha1);\n \thashclr(entry->peeled);\n \tmemcpy(entry->name, name, len);\n \tentry->flag = flag;\n-\tentry->next = list;\n \tif (new_entry)\n \t\t*new_entry = entry;\n-\treturn entry;\n+\tALLOC_GROW(refs->refs, refs->nr + 1, refs->alloc);\n+\trefs->refs[refs->nr++] = entry;\n }\n \n-/* merge sort the ref list */\n-static struct ref_list *sort_ref_list(struct ref_list *list)\n+static int ref_entry_cmp(const void *a, const void *b)\n {\n-\tint psize, qsize, last_merge_count, cmp;\n-\tstruct ref_list *p, *q, *l, *e;\n-\tstruct ref_list *new_list = list;\n-\tint k = 1;\n-\tint merge_count = 0;\n-\n-\tif (!list)\n-\t\treturn list;\n-\n-\tdo {\n-\t\tlast_merge_count = merge_count;\n-\t\tmerge_count = 0;\n-\n-\t\tpsize = 0;\n-\n-\t\tp = new_list;\n-\t\tq = new_list;\n-\t\tnew_list = NULL;\n-\t\tl = NULL;\n+\tstruct ref_entry *one = *(struct ref_entry **)a;\n+\tstruct ref_entry *two = *(struct ref_entry **)b;\n+\treturn strcmp(one->name, two->name);\n+}\n \n-\t\twhile (p) {\n-\t\t\tmerge_count++;\n+static void sort_ref_array(struct ref_array *array)\n+{\n+\tqsort(array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n+}\n \n-\t\t\twhile (psize < k && q->next) {\n-\t\t\t\tq = q->next;\n-\t\t\t\tpsize++;\n-\t\t\t}\n-\t\t\tqsize = k;\n-\n-\t\t\twhile ((psize > 0) || (qsize > 0 && q)) {\n-\t\t\t\tif (qsize == 0 || !q) {\n-\t\t\t\t\te = p;\n-\t\t\t\t\tp = p->next;\n-\t\t\t\t\tpsize--;\n-\t\t\t\t} else if (psize == 0) {\n-\t\t\t\t\te = q;\n-\t\t\t\t\tq = q->next;\n-\t\t\t\t\tqsize--;\n-\t\t\t\t} else {\n-\t\t\t\t\tcmp = strcmp(q->name, p->name);\n-\t\t\t\t\tif (cmp < 0) {\n-\t\t\t\t\t\te = q;\n-\t\t\t\t\t\tq = q->next;\n-\t\t\t\t\t\tqsize--;\n-\t\t\t\t\t} else if (cmp > 0) {\n-\t\t\t\t\t\te = p;\n-\t\t\t\t\t\tp = p->next;\n-\t\t\t\t\t\tpsize--;\n-\t\t\t\t\t} else {\n-\t\t\t\t\t\tif (hashcmp(q->sha1, p->sha1))\n-\t\t\t\t\t\t\tdie(\"Duplicated ref, and SHA1s don't match: %s\",\n-\t\t\t\t\t\t\t    q->name);\n-\t\t\t\t\t\twarning(\"Duplicated ref: %s\", q->name);\n-\t\t\t\t\t\te = q;\n-\t\t\t\t\t\tq = q->next;\n-\t\t\t\t\t\tqsize--;\n-\t\t\t\t\t\tfree(e);\n-\t\t\t\t\t\te = p;\n-\t\t\t\t\t\tp = p->next;\n-\t\t\t\t\t\tpsize--;\n-\t\t\t\t\t}\n-\t\t\t\t}\n+static struct ref_entry *search_ref_array(struct ref_array *array, const char *name)\n+{\n+\tstruct ref_entry *e, **r;\n+\tint len;\n \n-\t\t\t\te->next = NULL;\n+\tlen = strlen(name) + 1;\n+\te = xmalloc(sizeof(struct ref) + len);\n+\tmemcpy(e->name, name, len);\n \n-\t\t\t\tif (l)\n-\t\t\t\t\tl->next = e;\n-\t\t\t\tif (!new_list)\n-\t\t\t\t\tnew_list = e;\n-\t\t\t\tl = e;\n-\t\t\t}\n+\tr = bsearch(&e, array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n \n-\t\t\tp = q;\n-\t\t};\n+\tfree(e);\n \n-\t\tk = k * 2;\n-\t} while ((last_merge_count != merge_count) || (last_merge_count != 1));\n+\tif (r == NULL)\n+\t\treturn NULL;\n \n-\treturn new_list;\n+\treturn *r;\n }\n \n /*\n@@ -155,38 +106,37 @@ static struct ref_list *sort_ref_list(struct ref_list *list)\n static struct cached_refs {\n \tchar did_loose;\n \tchar did_packed;\n-\tstruct ref_list *loose;\n-\tstruct ref_list *packed;\n+\tstruct ref_array loose;\n+\tstruct ref_array packed;\n } cached_refs, submodule_refs;\n-static struct ref_list *current_ref;\n+static struct ref_entry *current_ref;\n \n-static struct ref_list *extra_refs;\n+static struct ref_array extra_refs;\n \n-static void free_ref_list(struct ref_list *list)\n+static void free_ref_array(struct ref_array *array)\n {\n-\tstruct ref_list *next;\n-\tfor ( ; list; list = next) {\n-\t\tnext = list->next;\n-\t\tfree(list);\n-\t}\n+\tint i;\n+\tfor (i = 0; i < array->nr; i++)\n+\t\tfree(array->refs[i]);\n+\tfree(array->refs);\n+\tarray->nr = array->alloc = 0;\n+\tarray->refs = NULL;\n }\n \n static void invalidate_cached_refs(void)\n {\n \tstruct cached_refs *ca = &cached_refs;\n \n-\tif (ca->did_loose && ca->loose)\n-\t\tfree_ref_list(ca->loose);\n-\tif (ca->did_packed && ca->packed)\n-\t\tfree_ref_list(ca->packed);\n-\tca->loose = ca->packed = NULL;\n+\tif (ca->did_loose)\n+\t\tfree_ref_array(&ca->loose);\n+\tif (ca->did_packed)\n+\t\tfree_ref_array(&ca->packed);\n \tca->did_loose = ca->did_packed = 0;\n }\n \n static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n {\n-\tstruct ref_list *list = NULL;\n-\tstruct ref_list *last = NULL;\n+\tstruct ref_entry *last = NULL;\n \tchar refline[PATH_MAX];\n \tint flag = REF_ISPACKED;\n \n@@ -205,7 +155,7 @@ static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n \n \t\tname = parse_ref_line(refline, sha1);\n \t\tif (name) {\n-\t\t\tlist = add_ref(name, sha1, flag, list, &last);\n+\t\t\tadd_ref(name, sha1, flag, &cached_refs->packed, &last);\n \t\t\tcontinue;\n \t\t}\n \t\tif (last &&\n@@ -215,21 +165,20 @@ static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n \t\t    !get_sha1_hex(refline + 1, sha1))\n \t\t\thashcpy(last->peeled, sha1);\n \t}\n-\tcached_refs->packed = sort_ref_list(list);\n+\tsort_ref_array(&cached_refs->packed);\n }\n \n void add_extra_ref(const char *name, const unsigned char *sha1, int flag)\n {\n-\textra_refs = add_ref(name, sha1, flag, extra_refs, NULL);\n+\tadd_ref(name, sha1, flag, &extra_refs, NULL);\n }\n \n void clear_extra_refs(void)\n {\n-\tfree_ref_list(extra_refs);\n-\textra_refs = NULL;\n+\tfree_ref_array(&extra_refs);\n }\n \n-static struct ref_list *get_packed_refs(const char *submodule)\n+static struct ref_array *get_packed_refs(const char *submodule)\n {\n \tconst char *packed_refs_file;\n \tstruct cached_refs *refs;\n@@ -237,7 +186,7 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \tif (submodule) {\n \t\tpacked_refs_file = git_path_submodule(submodule, \"packed-refs\");\n \t\trefs = &submodule_refs;\n-\t\tfree_ref_list(refs->packed);\n+\t\tfree_ref_array(&refs->packed);\n \t} else {\n \t\tpacked_refs_file = git_path(\"packed-refs\");\n \t\trefs = &cached_refs;\n@@ -245,18 +194,17 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \n \tif (!refs->did_packed || submodule) {\n \t\tFILE *f = fopen(packed_refs_file, \"r\");\n-\t\trefs->packed = NULL;\n \t\tif (f) {\n \t\t\tread_packed_refs(f, refs);\n \t\t\tfclose(f);\n \t\t}\n \t\trefs->did_packed = 1;\n \t}\n-\treturn refs->packed;\n+\treturn &refs->packed;\n }\n \n-static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n-\t\t\t\t    struct ref_list *list)\n+static void get_ref_dir(const char *submodule, const char *base,\n+\t\t\tstruct ref_array *array)\n {\n \tDIR *dir;\n \tconst char *path;\n@@ -299,7 +247,7 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\tif (stat(refdir, &st) < 0)\n \t\t\t\tcontinue;\n \t\t\tif (S_ISDIR(st.st_mode)) {\n-\t\t\t\tlist = get_ref_dir(submodule, ref, list);\n+\t\t\t\tget_ref_dir(submodule, ref, array);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (submodule) {\n@@ -314,12 +262,11 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\t\t\thashclr(sha1);\n \t\t\t\t\tflag |= REF_BROKEN;\n \t\t\t\t}\n-\t\t\tlist = add_ref(ref, sha1, flag, list, NULL);\n+\t\t\tadd_ref(ref, sha1, flag, array, NULL);\n \t\t}\n \t\tfree(ref);\n \t\tclosedir(dir);\n \t}\n-\treturn list;\n }\n \n struct warn_if_dangling_data {\n@@ -356,21 +303,21 @@ void warn_dangling_symref(FILE *fp, const char *msg_fmt, const char *refname)\n \tfor_each_rawref(warn_if_dangling_symref, &data);\n }\n \n-static struct ref_list *get_loose_refs(const char *submodule)\n+static struct ref_array *get_loose_refs(const char *submodule)\n {\n \tif (submodule) {\n-\t\tfree_ref_list(submodule_refs.loose);\n-\t\tsubmodule_refs.loose = get_ref_dir(submodule, \"refs\", NULL);\n-\t\tsubmodule_refs.loose = sort_ref_list(submodule_refs.loose);\n-\t\treturn submodule_refs.loose;\n+\t\tfree_ref_array(&submodule_refs.loose);\n+\t\tget_ref_dir(submodule, \"refs\", &submodule_refs.loose);\n+\t\tsort_ref_array(&submodule_refs.loose);\n+\t\treturn &submodule_refs.loose;\n \t}\n \n \tif (!cached_refs.did_loose) {\n-\t\tcached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n-\t\tcached_refs.loose = sort_ref_list(cached_refs.loose);\n+\t\tget_ref_dir(NULL, \"refs\", &cached_refs.loose);\n+\t\tsort_ref_array(&cached_refs.loose);\n \t\tcached_refs.did_loose = 1;\n \t}\n-\treturn cached_refs.loose;\n+\treturn &cached_refs.loose;\n }\n \n /* We allow \"recursive\" symbolic refs. Only within reason, though */\n@@ -381,8 +328,8 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n {\n \tFILE *f;\n \tstruct cached_refs refs;\n-\tstruct ref_list *ref;\n-\tint retval;\n+\tstruct ref_entry *ref;\n+\tint retval = -1;\n \n \tstrcpy(name + pathlen, \"packed-refs\");\n \tf = fopen(name, \"r\");\n@@ -390,17 +337,12 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n \t\treturn -1;\n \tread_packed_refs(f, &refs);\n \tfclose(f);\n-\tref = refs.packed;\n-\tretval = -1;\n-\twhile (ref) {\n-\t\tif (!strcmp(ref->name, refname)) {\n-\t\t\tretval = 0;\n-\t\t\tmemcpy(result, ref->sha1, 20);\n-\t\t\tbreak;\n-\t\t}\n-\t\tref = ref->next;\n+\tref = search_ref_array(&refs.packed, refname);\n+\tif (ref != NULL) {\n+\t\tmemcpy(result, ref->sha1, 20);\n+\t\tretval = 0;\n \t}\n-\tfree_ref_list(refs.packed);\n+\tfree_ref_array(&refs.packed);\n \treturn retval;\n }\n \n@@ -501,15 +443,13 @@ const char *resolve_ref(const char *ref, unsigned char *sha1, int reading, int *\n \t\tgit_snpath(path, sizeof(path), \"%s\", ref);\n \t\t/* Special case: non-existing file. */\n \t\tif (lstat(path, &st) < 0) {\n-\t\t\tstruct ref_list *list = get_packed_refs(NULL);\n-\t\t\twhile (list) {\n-\t\t\t\tif (!strcmp(ref, list->name)) {\n-\t\t\t\t\thashcpy(sha1, list->sha1);\n-\t\t\t\t\tif (flag)\n-\t\t\t\t\t\t*flag |= REF_ISPACKED;\n-\t\t\t\t\treturn ref;\n-\t\t\t\t}\n-\t\t\t\tlist = list->next;\n+\t\t\tstruct ref_array *packed = get_packed_refs(NULL);\n+\t\t\tstruct ref_entry *r = search_ref_array(packed, ref);\n+\t\t\tif (r != NULL) {\n+\t\t\t\thashcpy(sha1, r->sha1);\n+\t\t\t\tif (flag)\n+\t\t\t\t\t*flag |= REF_ISPACKED;\n+\t\t\t\treturn ref;\n \t\t\t}\n \t\t\tif (reading || errno != ENOENT)\n \t\t\t\treturn NULL;\n@@ -584,7 +524,7 @@ int read_ref(const char *ref, unsigned char *sha1)\n \n #define DO_FOR_EACH_INCLUDE_BROKEN 01\n static int do_one_ref(const char *base, each_ref_fn fn, int trim,\n-\t\t      int flags, void *cb_data, struct ref_list *entry)\n+\t\t      int flags, void *cb_data, struct ref_entry *entry)\n {\n \tif (prefixcmp(entry->name, base))\n \t\treturn 0;\n@@ -630,18 +570,12 @@ int peel_ref(const char *ref, unsigned char *sha1)\n \t\treturn -1;\n \n \tif ((flag & REF_ISPACKED)) {\n-\t\tstruct ref_list *list = get_packed_refs(NULL);\n+\t\tstruct ref_array *array = get_packed_refs(NULL);\n+\t\tstruct ref_entry *r = search_ref_array(array, ref);\n \n-\t\twhile (list) {\n-\t\t\tif (!strcmp(list->name, ref)) {\n-\t\t\t\tif (list->flag & REF_KNOWS_PEELED) {\n-\t\t\t\t\thashcpy(sha1, list->peeled);\n-\t\t\t\t\treturn 0;\n-\t\t\t\t}\n-\t\t\t\t/* older pack-refs did not leave peeled ones */\n-\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tlist = list->next;\n+\t\tif (r != NULL && r->flag & REF_KNOWS_PEELED) {\n+\t\t\thashcpy(sha1, r->peeled);\n+\t\t\treturn 0;\n \t\t}\n \t}\n \n@@ -660,36 +594,39 @@ fallback:\n static int do_for_each_ref(const char *submodule, const char *base, each_ref_fn fn,\n \t\t\t   int trim, int flags, void *cb_data)\n {\n-\tint retval = 0;\n-\tstruct ref_list *packed = get_packed_refs(submodule);\n-\tstruct ref_list *loose = get_loose_refs(submodule);\n+\tint retval = 0, i, p = 0, l = 0;\n+\tstruct ref_array *packed = get_packed_refs(submodule);\n+\tstruct ref_array *loose = get_loose_refs(submodule);\n \n-\tstruct ref_list *extra;\n+\tstruct ref_array *extra = &extra_refs;\n \n-\tfor (extra = extra_refs; extra; extra = extra->next)\n-\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, extra);\n+\tfor (i = 0; i < extra->nr; i++)\n+\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, extra->refs[i]);\n \n-\twhile (packed && loose) {\n-\t\tstruct ref_list *entry;\n-\t\tint cmp = strcmp(packed->name, loose->name);\n+\twhile (p < packed->nr && l < loose->nr) {\n+\t\tstruct ref_entry *entry;\n+\t\tint cmp = strcmp(packed->refs[p]->name, loose->refs[l]->name);\n \t\tif (!cmp) {\n-\t\t\tpacked = packed->next;\n+\t\t\tp++;\n \t\t\tcontinue;\n \t\t}\n \t\tif (cmp > 0) {\n-\t\t\tentry = loose;\n-\t\t\tloose = loose->next;\n+\t\t\tentry = loose->refs[l++];\n \t\t} else {\n-\t\t\tentry = packed;\n-\t\t\tpacked = packed->next;\n+\t\t\tentry = packed->refs[p++];\n \t\t}\n \t\tretval = do_one_ref(base, fn, trim, flags, cb_data, entry);\n \t\tif (retval)\n \t\t\tgoto end_each;\n \t}\n \n-\tfor (packed = packed ? packed : loose; packed; packed = packed->next) {\n-\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, packed);\n+\tif (l < loose->nr) {\n+\t\tp = l;\n+\t\tpacked = loose;\n+\t}\n+\n+\tfor (; p < packed->nr; p++) {\n+\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, packed->refs[p]);\n \t\tif (retval)\n \t\t\tgoto end_each;\n \t}\n@@ -1005,24 +942,24 @@ static int remove_empty_directories(const char *file)\n }\n \n static int is_refname_available(const char *ref, const char *oldref,\n-\t\t\t\tstruct ref_list *list, int quiet)\n-{\n-\tint namlen = strlen(ref); /* e.g. 'foo/bar' */\n-\twhile (list) {\n-\t\t/* list->name could be 'foo' or 'foo/bar/baz' */\n-\t\tif (!oldref || strcmp(oldref, list->name)) {\n-\t\t\tint len = strlen(list->name);\n+\t\t\t\tstruct ref_array *array, int quiet)\n+{\n+\tint i, namlen = strlen(ref); /* e.g. 'foo/bar' */\n+\tfor (i = 0; i < array->nr; i++ ) {\n+\t\tstruct ref_entry *entry = array->refs[i];\n+\t\t/* entry->name could be 'foo' or 'foo/bar/baz' */\n+\t\tif (!oldref || strcmp(oldref, entry->name)) {\n+\t\t\tint len = strlen(entry->name);\n \t\t\tint cmplen = (namlen < len) ? namlen : len;\n-\t\t\tconst char *lead = (namlen < len) ? list->name : ref;\n-\t\t\tif (!strncmp(ref, list->name, cmplen) &&\n+\t\t\tconst char *lead = (namlen < len) ? entry->name : ref;\n+\t\t\tif (!strncmp(ref, entry->name, cmplen) &&\n \t\t\t    lead[cmplen] == '/') {\n \t\t\t\tif (!quiet)\n \t\t\t\t\terror(\"'%s' exists; cannot create '%s'\",\n-\t\t\t\t\t      list->name, ref);\n+\t\t\t\t\t      entry->name, ref);\n \t\t\t\treturn 0;\n \t\t\t}\n \t\t}\n-\t\tlist = list->next;\n \t}\n \treturn 1;\n }\n@@ -1129,18 +1066,13 @@ static struct lock_file packlock;\n \n static int repack_without_ref(const char *refname)\n {\n-\tstruct ref_list *list, *packed_ref_list;\n-\tint fd;\n-\tint found = 0;\n+\tstruct ref_array *packed;\n+\tstruct ref_entry *ref;\n+\tint fd, i;\n \n-\tpacked_ref_list = get_packed_refs(NULL);\n-\tfor (list = packed_ref_list; list; list = list->next) {\n-\t\tif (!strcmp(refname, list->name)) {\n-\t\t\tfound = 1;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n-\tif (!found)\n+\tpacked = get_packed_refs(NULL);\n+\tref = search_ref_array(packed, refname);\n+\tif (ref == NULL)\n \t\treturn 0;\n \tfd = hold_lock_file_for_update(&packlock, git_path(\"packed-refs\"), 0);\n \tif (fd < 0) {\n@@ -1148,17 +1080,19 @@ static int repack_without_ref(const char *refname)\n \t\treturn error(\"cannot delete '%s' from packed refs\", refname);\n \t}\n \n-\tfor (list = packed_ref_list; list; list = list->next) {\n+\tfor (i = 0; i < packed->nr; i++) {\n \t\tchar line[PATH_MAX + 100];\n \t\tint len;\n \n-\t\tif (!strcmp(refname, list->name))\n+\t\tref = packed->refs[i];\n+\n+\t\tif (!strcmp(refname, ref->name))\n \t\t\tcontinue;\n \t\tlen = snprintf(line, sizeof(line), \"%s %s\\n\",\n-\t\t\t       sha1_to_hex(list->sha1), list->name);\n+\t\t\t       sha1_to_hex(ref->sha1), ref->name);\n \t\t/* this should not happen but just being defensive */\n \t\tif (len > sizeof(line))\n-\t\t\tdie(\"too long a refname '%s'\", list->name);\n+\t\t\tdie(\"too long a refname '%s'\", ref->name);\n \t\twrite_or_die(fd, line, len);\n \t}\n \treturn commit_lock_file(&packlock);\n-- \n1.7.6.1\n"},{"id":"176495","messageId":"201109291038.45290.mfick@codeaurora.org","threadId":"27589","inReplyTo":"7c0105c6cca7dd0aa336522f90617fe4@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-29T16:38:44Z","receivedAt":"2011-09-29T16:38:44Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, September 28, 2011 08:19:16 pm Julian Phillips \nwrote:\n> On Wed, 28 Sep 2011 19:37:18 -0600, Martin Fick wrote:\n> > On Wednesday 28 September 2011 18:59:09 Martin Fick \nwrote:\n> >> Julian Phillips <julian@quantumfyre.co.uk> wrote:\n> -- snip --\n> \n> >> I've created a test repo with ~100k refs/changes/...\n> >> style refs, and ~40000 refs/heads/... style refs, and\n> >> checkout can walk the list of ~140k refs seven times\n> >> in 85ms user time including doing whatever other\n> >> processing is needed for checkout. The real time is\n> >> only 114ms - but then my test repo has no real data\n> >> in.\n> > \n> > If I understand what you are saying, it sounds like you\n> > do not have a very good test case. The amount of time\n> > it takes for checkout depends on how long it takes to\n> > find a ref with the sha1 that you are on. If that sha1\n> > is so early in the list of refs that it only took you\n> > 7 traversals to find it, then that is not a very good\n> > testcase. I think that you should probably try making\n> > an orphaned ref (checkout a detached head, commit to\n> > it), that is probably the worst testcase since it\n> > should then have to search all 140K refs to eventually\n> > give up.\n> > \n> > Again, if I understand what you are saying, if it took\n> > 85ms for 7 traversals, then it takes approximately\n> > 10ms per traversal, that's only 100/s! If you have to\n> > traverse it 140K times, that should work out to 1400s\n> > ~ 23mins.\n> \n> Well, it's no more than 10ms per traversal - since the\n> rest of the work presumably takes some time too ...\n> \n> However, I had forgotten to make the orphaned commit as\n> you suggest - and then _bang_ 7N^2, it tries seven\n> different variants of each ref (which is silly as they\n> are all fully qualified), and with packed refs it has to\n> search for them each time, all to turn names into hashes\n> that we already know to start with.\n> \n> So, yes - it is that list traversal.\n> \n> Does the following help?\n> \n> diff --git a/builtin/checkout.c b/builtin/checkout.c\n> index 5e356a6..f0f4ca1 100644\n> --- a/builtin/checkout.c\n> +++ b/builtin/checkout.c\n> @@ -605,7 +605,7 @@ static int\n> add_one_ref_to_rev_list_arg(const char *refname,\n>                                         int flags,\n>                                         void *cb_data)\n>   {\n> -       add_one_rev_list_arg(cb_data, refname);\n> +       add_one_rev_list_arg(cb_data,\n> strdup(sha1_to_hex(sha1))); return 0;\n>   }\n\n\nYes, but in some strange ways. :)\n\nFirst, let me clarify that all the tests here involve your \n\"sort fix\" from 2 days ago applied first.\n\nIn the packed ref repo, it brings the time down to about \n~10s (from > 5 mins).  In the unpacked ref repo, it brings \nit down to about the same thing ~10s, but it was only \nstarting at about ~20s.\n\nSo, I have to ask, what does that change do, I don't quite \nunderstand it?  Does it just do only one lookup per ref by \nnormalizing it?  Is the list still being traversed, just \nabout 7 time less now?  Should the packed_ref list simply be \nput in an array which could be binary searched instead, it \nis a fixed list once loaded, no?  \n\nI prototyped a packed_ref implementation using the hash.c \nprovided in the git sources and it seemed to speed a \ncheckout up to almost instantaneous, but I was getting a few \ncollisions so the implementation was not good enough.  That \nis when I started to wonder if an array wouldn't be better \nin this case?\n\n\n\nNow I also decided to go back and test a noop fetch (a \nrefetch) of all the changes (since this use case is still \ntaking way longer than I think it should, even with the \nsubmodule fix posted earlier).  Up until this point, even \nthe sorting fix did not help.  So I tried it with this fix.  \nIn the unpackref case, it did not seem to change (2~4mins).  \nHowever, in the packed ref change (which was previously also \nabout 2-4mins), this now only takes about 10-15s!\n\nAny clues as to why the unpacked refs would still be so slow \non noop fetches and not be sped up by this?\n\n\n-Martin\n\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176508","messageId":"942277b4139229d6d3e921999859250f@quantumfyre.co.uk","threadId":"27589","inReplyTo":"201109291038.45290.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T18:26:24Z","receivedAt":"2011-09-29T18:26:24Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Thu, 29 Sep 2011 10:38:44 -0600, Martin Fick wrote:\n> On Wednesday, September 28, 2011 08:19:16 pm Julian Phillips\n> wrote:\n-- snip --\n>> However, I had forgotten to make the orphaned commit as\n>> you suggest - and then _bang_ 7N^2, it tries seven\n>> different variants of each ref (which is silly as they\n>> are all fully qualified), and with packed refs it has to\n>> search for them each time, all to turn names into hashes\n>> that we already know to start with.\n>>\n>> So, yes - it is that list traversal.\n>>\n>> Does the following help?\n>>\n>> diff --git a/builtin/checkout.c b/builtin/checkout.c\n>> index 5e356a6..f0f4ca1 100644\n>> --- a/builtin/checkout.c\n>> +++ b/builtin/checkout.c\n>> @@ -605,7 +605,7 @@ static int\n>> add_one_ref_to_rev_list_arg(const char *refname,\n>>                                         int flags,\n>>                                         void *cb_data)\n>>   {\n>> -       add_one_rev_list_arg(cb_data, refname);\n>> +       add_one_rev_list_arg(cb_data,\n>> strdup(sha1_to_hex(sha1))); return 0;\n>>   }\n>\n>\n> Yes, but in some strange ways. :)\n>\n> First, let me clarify that all the tests here involve your\n> \"sort fix\" from 2 days ago applied first.\n>\n> In the packed ref repo, it brings the time down to about\n> ~10s (from > 5 mins).  In the unpacked ref repo, it brings\n> it down to about the same thing ~10s, but it was only\n> starting at about ~20s.\n>\n> So, I have to ask, what does that change do, I don't quite\n> understand it?  Does it just do only one lookup per ref by\n> normalizing it?  Is the list still being traversed, just\n> about 7 time less now?\n\nIn order to check for orphaned commits, checkout effectively calls \nrev-list passing it a list of the names of all the refs as input.  The \nrev-list code then has to go through this list and convert each entry \ninto an actual hash that it can look up in the object database.  This is \nwhere the N^2 comes in for packed refs, as it calles resolve_ref() for \neach ref in the list (N), which then loops through the list of all refs \n(N) to find a match.  However, the code that creates the list of refs to \npass to the rev-list code already knows the hash for each ref.  So the \nchange above passes the hashes to rev-list, which then doesn't need to \nlookup the ref - it just converts the string form hash back to binary \nform, avoiding the N^2 work altogether.  This is why packed and unpacked \nare about the same speed, as they are now doing the same amount of work.\n\n> Should the packed_ref list simply be\n> put in an array which could be binary searched instead, it\n> is a fixed list once loaded, no?\n\nA quick look at the code suggests that probably both the list of loose \nrefs, and the list of packed refs could both be stored as binary \nsearchable arrays, or in an ordered hash table.  Though whether it is \nactually necessary I don't know.  So far, it seems to have been possible \nto fix performance issues whilst keeping the simple lists ...\n\n> I prototyped a packed_ref implementation using the hash.c\n> provided in the git sources and it seemed to speed a\n> checkout up to almost instantaneous, but I was getting a few\n> collisions so the implementation was not good enough.  That\n> is when I started to wonder if an array wouldn't be better\n> in this case?\n>\n>\n>\n> Now I also decided to go back and test a noop fetch (a\n> refetch) of all the changes (since this use case is still\n> taking way longer than I think it should, even with the\n> submodule fix posted earlier).  Up until this point, even\n> the sorting fix did not help.  So I tried it with this fix.\n> In the unpackref case, it did not seem to change (2~4mins).\n> However, in the packed ref change (which was previously also\n> about 2-4mins), this now only takes about 10-15s!\n>\n> Any clues as to why the unpacked refs would still be so slow\n> on noop fetches and not be sped up by this?\n\nNot really.  I wouldn't expect this change to have any effect on fetch, \nbut I haven't actually looked into it.\n\n-- \nJulian\n"},{"id":"176509","messageId":"4E84B89F.4060304@lsrfire.ath.cx","threadId":"27589","inReplyTo":"7c0105c6cca7dd0aa336522f90617fe4@quantumfyre.co.uk","subject":"Re: Git is not scalable with too many refs/*","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-09-29T18:27:43Z","receivedAt":"2011-09-29T18:27:43Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 29.09.2011 04:19, schrieb Julian Phillips:\n> Does the following help?\n> \n> diff --git a/builtin/checkout.c b/builtin/checkout.c\n> index 5e356a6..f0f4ca1 100644\n> --- a/builtin/checkout.c\n> +++ b/builtin/checkout.c\n> @@ -605,7 +605,7 @@ static int add_one_ref_to_rev_list_arg(const char\n> *refname,\n>                                        int flags,\n>                                        void *cb_data)\n>  {\n> -       add_one_rev_list_arg(cb_data, refname);\n> +       add_one_rev_list_arg(cb_data, strdup(sha1_to_hex(sha1)));\n>         return 0;\n>  }\n\nHmm.  Can we get rid of the multiple ref lookups fixed by the above\n*and* the overhead of dealing with a textual argument list at the same\ntime by calling add_pending_object directly, like this?  (Factoring\nout add_pending_sha1 should be a separate patch..)\n\nRené\n\n---\n builtin/checkout.c |   39 ++++++++++++---------------------------\n revision.c         |   11 ++++++++---\n revision.h         |    1 +\n 3 files changed, 21 insertions(+), 30 deletions(-)\n\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 5e356a6..84e0cdc 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -588,24 +588,11 @@ static void update_refs_for_switch(struct checkout_opts *opts,\n \t\treport_tracking(new);\n }\n \n-struct rev_list_args {\n-\tint argc;\n-\tint alloc;\n-\tconst char **argv;\n-};\n-\n-static void add_one_rev_list_arg(struct rev_list_args *args, const char *s)\n-{\n-\tALLOC_GROW(args->argv, args->argc + 1, args->alloc);\n-\targs->argv[args->argc++] = s;\n-}\n-\n-static int add_one_ref_to_rev_list_arg(const char *refname,\n-\t\t\t\t       const unsigned char *sha1,\n-\t\t\t\t       int flags,\n-\t\t\t\t       void *cb_data)\n+static int add_pending_uninteresting_ref(const char *refname,\n+\t\t\t\t\t const unsigned char *sha1,\n+\t\t\t\t\t int flags, void *cb_data)\n {\n-\tadd_one_rev_list_arg(cb_data, refname);\n+\tadd_pending_sha1(cb_data, refname, sha1, flags | UNINTERESTING);\n \treturn 0;\n }\n \n@@ -685,19 +672,17 @@ static void suggest_reattach(struct commit *commit, struct rev_info *revs)\n  */\n static void orphaned_commit_warning(struct commit *commit)\n {\n-\tstruct rev_list_args args = { 0, 0, NULL };\n \tstruct rev_info revs;\n-\n-\tadd_one_rev_list_arg(&args, \"(internal)\");\n-\tadd_one_rev_list_arg(&args, sha1_to_hex(commit->object.sha1));\n-\tadd_one_rev_list_arg(&args, \"--not\");\n-\tfor_each_ref(add_one_ref_to_rev_list_arg, &args);\n-\tadd_one_rev_list_arg(&args, \"--\");\n-\tadd_one_rev_list_arg(&args, NULL);\n+\tstruct object *object = &commit->object;\n \n \tinit_revisions(&revs, NULL);\n-\tif (setup_revisions(args.argc - 1, args.argv, &revs, NULL) != 1)\n-\t\tdie(_(\"internal error: only -- alone should have been left\"));\n+\tsetup_revisions(0, NULL, &revs, NULL);\n+\n+\tobject->flags &= ~UNINTERESTING;\n+\tadd_pending_object(&revs, object, sha1_to_hex(object->sha1));\n+\n+\tfor_each_ref(add_pending_uninteresting_ref, &revs);\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(_(\"internal error in revision walk\"));\n \tif (!(commit->object.flags & UNINTERESTING))\ndiff --git a/revision.c b/revision.c\nindex c46cfaa..2e8aa33 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -185,6 +185,13 @@ static struct object *get_reference(struct rev_info *revs, const char *name, con\n \treturn object;\n }\n \n+void add_pending_sha1(struct rev_info *revs, const char *name,\n+\t\t      const unsigned char *sha1, unsigned int flags)\n+{\n+\tstruct object *object = get_reference(revs, name, sha1, flags);\n+\tadd_pending_object(revs, object, name);\n+}\n+\n static struct commit *handle_commit(struct rev_info *revs, struct object *object, const char *name)\n {\n \tunsigned long flags = object->flags;\n@@ -832,9 +839,7 @@ struct all_refs_cb {\n static int handle_one_ref(const char *path, const unsigned char *sha1, int flag, void *cb_data)\n {\n \tstruct all_refs_cb *cb = cb_data;\n-\tstruct object *object = get_reference(cb->all_revs, path, sha1,\n-\t\t\t\t\t      cb->all_flags);\n-\tadd_pending_object(cb->all_revs, object, path);\n+\tadd_pending_sha1(cb->all_revs, path, sha1, cb->all_flags);\n \treturn 0;\n }\n \ndiff --git a/revision.h b/revision.h\nindex 3d64ada..4541265 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -191,6 +191,7 @@ extern void add_object(struct object *obj,\n \t\t       const char *name);\n \n extern void add_pending_object(struct rev_info *revs, struct object *obj, const char *name);\n+extern void add_pending_sha1(struct rev_info *revs, const char *name, const unsigned char *sha1, unsigned int flags);\n \n extern void add_head_to_pending(struct rev_info *);\n \n-- \n1.7.7.rc1\n"},{"id":"176513","messageId":"7vy5x7rwq9.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"4E84B89F.4060304@lsrfire.ath.cx","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-29T19:10:06Z","receivedAt":"2011-09-29T19:10:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <rene.scharfe@lsrfire.ath.cx> writes:\n\n> Hmm.  Can we get rid of the multiple ref lookups fixed by the above\n> *and* the overhead of dealing with a textual argument list at the same\n> time by calling add_pending_object directly, like this?  (Factoring\n> out add_pending_sha1 should be a separate patch..)\n\nI haven't tested it or thought about it through, but it smells right ;-)\n\nAlso we would probably want to drop \"next\" field from \"struct ref_list\"\n(i.e. making it not a linear list), introduce a new \"struct ref_array\"\nthat is a ALLOC_GROW() managed array of pointers to \"struct ref_list\",\nmake get_packed_refs() and get_loose_refs() return a pointer to \"struct\nref_array\" after sorting the array contents by \"name\". Then resolve_ref()\ncan do a bisection search in the packed refs array when it does not find a\nloose ref.\n"},{"id":"176514","messageId":"67aa39596ee72e1abbecf488a2916923@quantumfyre.co.uk","threadId":"27589","inReplyTo":"4E84B89F.4060304@lsrfire.ath.cx","subject":"Re: Git is not scalable with too many refs/*","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T19:10:57Z","receivedAt":"2011-09-29T19:10:57Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Thu, 29 Sep 2011 20:27:43 +0200, René Scharfe wrote:\n> Am 29.09.2011 04:19, schrieb Julian Phillips:\n>> Does the following help?\n>>\n>> diff --git a/builtin/checkout.c b/builtin/checkout.c\n>> index 5e356a6..f0f4ca1 100644\n>> --- a/builtin/checkout.c\n>> +++ b/builtin/checkout.c\n>> @@ -605,7 +605,7 @@ static int add_one_ref_to_rev_list_arg(const \n>> char\n>> *refname,\n>>                                        int flags,\n>>                                        void *cb_data)\n>>  {\n>> -       add_one_rev_list_arg(cb_data, refname);\n>> +       add_one_rev_list_arg(cb_data, strdup(sha1_to_hex(sha1)));\n>>         return 0;\n>>  }\n>\n> Hmm.  Can we get rid of the multiple ref lookups fixed by the above\n> *and* the overhead of dealing with a textual argument list at the \n> same\n> time by calling add_pending_object directly, like this?  (Factoring\n> out add_pending_sha1 should be a separate patch..)\n\nSeems like a good idea.  I get the same sort of times as with my patch, \nbut it makes the code _feel_ much nicer (and slightly smaller).  Mine \nwas definitely more of a \"it's 2am, but I think the problem is here\" \ntype of patch ;)\n\n-- \nJulian\n"},{"id":"176517","messageId":"201109291411.06733.mfick@codeaurora.org","threadId":"27589","inReplyTo":"4E84B89F.4060304@lsrfire.ath.cx","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-29T20:11:06Z","receivedAt":"2011-09-29T20:11:06Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Thursday, September 29, 2011 12:27:43 pm René Scharfe \nwrote:\n> Hmm.  Can we get rid of the multiple ref lookups fixed by\n> the above *and* the overhead of dealing with a textual\n> argument list at the same time by calling\n> add_pending_object directly, like this?  (Factoring out\n> add_pending_sha1 should be a separate patch..)\n \nRené,\n\nYour patch works well for me.  It achieves about the same \ngains as Julian's patch. Thanks!\n\nAfter all the performance fixes get merged for large ref \ncounts, it sure should help the Gerrit community.  I wonder \nhow it might impact Gerrit mirroring...\n\n-Martin\n\n\nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176521","messageId":"201109291444.33076.mfick@codeaurora.org","threadId":"27589","inReplyTo":"7vy5x7rwq9.fsf@alter.siamese.dyndns.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-29T20:44:32Z","receivedAt":"2011-09-29T20:44:32Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Thursday, September 29, 2011 01:10:06 pm Junio C Hamano \nwrote:\n> Also we would probably want to drop \"next\" field from\n> \"struct ref_list\" (i.e. making it not a linear list),\n> introduce a new \"struct ref_array\" that is a\n> ALLOC_GROW() managed array of pointers to \"struct\n> ref_list\", make get_packed_refs() and get_loose_refs()\n> return a pointer to \"struct ref_array\" after sorting the\n> array contents by \"name\". Then resolve_ref() can do a\n> bisection search in the packed refs array when it does\n> not find a loose ref.\n\nThat would be nice, and I suspect it would shave a bit more \nof the orphan check and possibly even a fetch.  If I \nunderstood all that, I might try.  But I might need some \nhand holding, my C is pretty rusty... Is there a bisection \nsearch library in git already to use?  Is there a git \nsorting library for the array also?\n\n-Martin\n\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176533","messageId":"7vzkhnqae6.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"20110929041811.5363.33396.julian@quantumfyre.co.uk","subject":"Re: [PATCH] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-29T21:57:53Z","receivedAt":"2011-09-29T21:57:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Julian Phillips <julian@quantumfyre.co.uk> writes:\n\n> Currently we linearly search through lists of refs when we need to\n> find a specific ref.  This can be very slow if we need to lookup a\n> large number of refs.  By changing to a binary search we can make this\n> faster.\n>\n> In order to be able to use a binary search we need to change from\n> using linked lists to arrays, which we can manage using ALLOC_GROW.\n>\n> We can now also use the standard library qsort function to sort the\n> refs arrays.\n>\n> Signed-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n> ---\n>\n> Something like this?\n>\n>  refs.c |  328 ++++++++++++++++++++++++++--------------------------------------\n>  1 files changed, 131 insertions(+), 197 deletions(-)\n>\n> diff --git a/refs.c b/refs.c\n> index a49ff74..e411bea 100644\n> --- a/refs.c\n> +++ b/refs.c\n> @@ -8,14 +8,18 @@\n>  #define REF_KNOWS_PEELED 04\n>  #define REF_BROKEN 010\n>  \n> -struct ref_list {\n> -\tstruct ref_list *next;\n> +struct ref_entry {\n>  \tunsigned char flag; /* ISSYMREF? ISPACKED? */\n>  \tunsigned char sha1[20];\n>  \tunsigned char peeled[20];\n>  \tchar name[FLEX_ARRAY];\n>  };\n>  \n> +struct ref_array {\n> +\tint nr, alloc;\n> +\tstruct ref_entry **refs;\n> +};\n> +\n\nYeah, I can say \"something like that\" without looking at the rest of the\npatch ;-) The rest should naturally follow from the above data structures.\n"},{"id":"176536","messageId":"20110929220435.22811.56434.julian@quantumfyre.co.uk","threadId":"27589","inReplyTo":"20110929041811.5363.33396.julian@quantumfyre.co.uk","subject":"[PATCH v2] refs: Use binary search to lookup refs faster","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T22:04:34Z","receivedAt":"2011-09-29T22:04:34Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"Currently we linearly search through lists of refs when we need to\nfind a specific ref.  This can be very slow if we need to lookup a\nlarge number of refs.  By changing to a binary search we can make this\nfaster.\n\nIn order to be able to use a binary search we need to change from\nusing linked lists to arrays, which we can manage using ALLOC_GROW.\n\nWe can now also use the standard library qsort function to sort the\nrefs arrays.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\n\nPrevious version caused a regression in the test suite ... :$\n\n refs.c |  329 ++++++++++++++++++++++++++--------------------------------------\n 1 files changed, 133 insertions(+), 196 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex a49ff74..35bba97 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -8,14 +8,18 @@\n #define REF_KNOWS_PEELED 04\n #define REF_BROKEN 010\n \n-struct ref_list {\n-\tstruct ref_list *next;\n+struct ref_entry {\n \tunsigned char flag; /* ISSYMREF? ISPACKED? */\n \tunsigned char sha1[20];\n \tunsigned char peeled[20];\n \tchar name[FLEX_ARRAY];\n };\n \n+struct ref_array {\n+\tint nr, alloc;\n+\tstruct ref_entry **refs;\n+};\n+\n static const char *parse_ref_line(char *line, unsigned char *sha1)\n {\n \t/*\n@@ -44,108 +48,58 @@ static const char *parse_ref_line(char *line, unsigned char *sha1)\n \treturn line;\n }\n \n-static struct ref_list *add_ref(const char *name, const unsigned char *sha1,\n-\t\t\t\tint flag, struct ref_list *list,\n-\t\t\t\tstruct ref_list **new_entry)\n+static void add_ref(const char *name, const unsigned char *sha1,\n+\t\t    int flag, struct ref_array *refs,\n+\t\t    struct ref_entry **new_entry)\n {\n \tint len;\n-\tstruct ref_list *entry;\n+\tstruct ref_entry *entry;\n \n \t/* Allocate it and add it in.. */\n \tlen = strlen(name) + 1;\n-\tentry = xmalloc(sizeof(struct ref_list) + len);\n+\tentry = xmalloc(sizeof(struct ref) + len);\n \thashcpy(entry->sha1, sha1);\n \thashclr(entry->peeled);\n \tmemcpy(entry->name, name, len);\n \tentry->flag = flag;\n-\tentry->next = list;\n \tif (new_entry)\n \t\t*new_entry = entry;\n-\treturn entry;\n+\tALLOC_GROW(refs->refs, refs->nr + 1, refs->alloc);\n+\trefs->refs[refs->nr++] = entry;\n }\n \n-/* merge sort the ref list */\n-static struct ref_list *sort_ref_list(struct ref_list *list)\n+static int ref_entry_cmp(const void *a, const void *b)\n {\n-\tint psize, qsize, last_merge_count, cmp;\n-\tstruct ref_list *p, *q, *l, *e;\n-\tstruct ref_list *new_list = list;\n-\tint k = 1;\n-\tint merge_count = 0;\n-\n-\tif (!list)\n-\t\treturn list;\n-\n-\tdo {\n-\t\tlast_merge_count = merge_count;\n-\t\tmerge_count = 0;\n-\n-\t\tpsize = 0;\n+\tstruct ref_entry *one = *(struct ref_entry **)a;\n+\tstruct ref_entry *two = *(struct ref_entry **)b;\n+\treturn strcmp(one->name, two->name);\n+}\n \n-\t\tp = new_list;\n-\t\tq = new_list;\n-\t\tnew_list = NULL;\n-\t\tl = NULL;\n+static void sort_ref_array(struct ref_array *array)\n+{\n+\tqsort(array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n+}\n \n-\t\twhile (p) {\n-\t\t\tmerge_count++;\n+static struct ref_entry *search_ref_array(struct ref_array *array, const char *name)\n+{\n+\tstruct ref_entry *e, **r;\n+\tint len;\n \n-\t\t\twhile (psize < k && q->next) {\n-\t\t\t\tq = q->next;\n-\t\t\t\tpsize++;\n-\t\t\t}\n-\t\t\tqsize = k;\n-\n-\t\t\twhile ((psize > 0) || (qsize > 0 && q)) {\n-\t\t\t\tif (qsize == 0 || !q) {\n-\t\t\t\t\te = p;\n-\t\t\t\t\tp = p->next;\n-\t\t\t\t\tpsize--;\n-\t\t\t\t} else if (psize == 0) {\n-\t\t\t\t\te = q;\n-\t\t\t\t\tq = q->next;\n-\t\t\t\t\tqsize--;\n-\t\t\t\t} else {\n-\t\t\t\t\tcmp = strcmp(q->name, p->name);\n-\t\t\t\t\tif (cmp < 0) {\n-\t\t\t\t\t\te = q;\n-\t\t\t\t\t\tq = q->next;\n-\t\t\t\t\t\tqsize--;\n-\t\t\t\t\t} else if (cmp > 0) {\n-\t\t\t\t\t\te = p;\n-\t\t\t\t\t\tp = p->next;\n-\t\t\t\t\t\tpsize--;\n-\t\t\t\t\t} else {\n-\t\t\t\t\t\tif (hashcmp(q->sha1, p->sha1))\n-\t\t\t\t\t\t\tdie(\"Duplicated ref, and SHA1s don't match: %s\",\n-\t\t\t\t\t\t\t    q->name);\n-\t\t\t\t\t\twarning(\"Duplicated ref: %s\", q->name);\n-\t\t\t\t\t\te = q;\n-\t\t\t\t\t\tq = q->next;\n-\t\t\t\t\t\tqsize--;\n-\t\t\t\t\t\tfree(e);\n-\t\t\t\t\t\te = p;\n-\t\t\t\t\t\tp = p->next;\n-\t\t\t\t\t\tpsize--;\n-\t\t\t\t\t}\n-\t\t\t\t}\n+\tif (name == NULL)\n+\t\treturn NULL;\n \n-\t\t\t\te->next = NULL;\n+\tlen = strlen(name) + 1;\n+\te = xmalloc(sizeof(struct ref) + len);\n+\tmemcpy(e->name, name, len);\n \n-\t\t\t\tif (l)\n-\t\t\t\t\tl->next = e;\n-\t\t\t\tif (!new_list)\n-\t\t\t\t\tnew_list = e;\n-\t\t\t\tl = e;\n-\t\t\t}\n+\tr = bsearch(&e, array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n \n-\t\t\tp = q;\n-\t\t};\n+\tfree(e);\n \n-\t\tk = k * 2;\n-\t} while ((last_merge_count != merge_count) || (last_merge_count != 1));\n+\tif (r == NULL)\n+\t\treturn NULL;\n \n-\treturn new_list;\n+\treturn *r;\n }\n \n /*\n@@ -155,38 +109,37 @@ static struct ref_list *sort_ref_list(struct ref_list *list)\n static struct cached_refs {\n \tchar did_loose;\n \tchar did_packed;\n-\tstruct ref_list *loose;\n-\tstruct ref_list *packed;\n+\tstruct ref_array loose;\n+\tstruct ref_array packed;\n } cached_refs, submodule_refs;\n-static struct ref_list *current_ref;\n+static struct ref_entry *current_ref;\n \n-static struct ref_list *extra_refs;\n+static struct ref_array extra_refs;\n \n-static void free_ref_list(struct ref_list *list)\n+static void free_ref_array(struct ref_array *array)\n {\n-\tstruct ref_list *next;\n-\tfor ( ; list; list = next) {\n-\t\tnext = list->next;\n-\t\tfree(list);\n-\t}\n+\tint i;\n+\tfor (i = 0; i < array->nr; i++)\n+\t\tfree(array->refs[i]);\n+\tfree(array->refs);\n+\tarray->nr = array->alloc = 0;\n+\tarray->refs = NULL;\n }\n \n static void invalidate_cached_refs(void)\n {\n \tstruct cached_refs *ca = &cached_refs;\n \n-\tif (ca->did_loose && ca->loose)\n-\t\tfree_ref_list(ca->loose);\n-\tif (ca->did_packed && ca->packed)\n-\t\tfree_ref_list(ca->packed);\n-\tca->loose = ca->packed = NULL;\n+\tif (ca->did_loose)\n+\t\tfree_ref_array(&ca->loose);\n+\tif (ca->did_packed)\n+\t\tfree_ref_array(&ca->packed);\n \tca->did_loose = ca->did_packed = 0;\n }\n \n static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n {\n-\tstruct ref_list *list = NULL;\n-\tstruct ref_list *last = NULL;\n+\tstruct ref_entry *last = NULL;\n \tchar refline[PATH_MAX];\n \tint flag = REF_ISPACKED;\n \n@@ -205,7 +158,7 @@ static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n \n \t\tname = parse_ref_line(refline, sha1);\n \t\tif (name) {\n-\t\t\tlist = add_ref(name, sha1, flag, list, &last);\n+\t\t\tadd_ref(name, sha1, flag, &cached_refs->packed, &last);\n \t\t\tcontinue;\n \t\t}\n \t\tif (last &&\n@@ -215,21 +168,20 @@ static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n \t\t    !get_sha1_hex(refline + 1, sha1))\n \t\t\thashcpy(last->peeled, sha1);\n \t}\n-\tcached_refs->packed = sort_ref_list(list);\n+\tsort_ref_array(&cached_refs->packed);\n }\n \n void add_extra_ref(const char *name, const unsigned char *sha1, int flag)\n {\n-\textra_refs = add_ref(name, sha1, flag, extra_refs, NULL);\n+\tadd_ref(name, sha1, flag, &extra_refs, NULL);\n }\n \n void clear_extra_refs(void)\n {\n-\tfree_ref_list(extra_refs);\n-\textra_refs = NULL;\n+\tfree_ref_array(&extra_refs);\n }\n \n-static struct ref_list *get_packed_refs(const char *submodule)\n+static struct ref_array *get_packed_refs(const char *submodule)\n {\n \tconst char *packed_refs_file;\n \tstruct cached_refs *refs;\n@@ -237,7 +189,7 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \tif (submodule) {\n \t\tpacked_refs_file = git_path_submodule(submodule, \"packed-refs\");\n \t\trefs = &submodule_refs;\n-\t\tfree_ref_list(refs->packed);\n+\t\tfree_ref_array(&refs->packed);\n \t} else {\n \t\tpacked_refs_file = git_path(\"packed-refs\");\n \t\trefs = &cached_refs;\n@@ -245,18 +197,17 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \n \tif (!refs->did_packed || submodule) {\n \t\tFILE *f = fopen(packed_refs_file, \"r\");\n-\t\trefs->packed = NULL;\n \t\tif (f) {\n \t\t\tread_packed_refs(f, refs);\n \t\t\tfclose(f);\n \t\t}\n \t\trefs->did_packed = 1;\n \t}\n-\treturn refs->packed;\n+\treturn &refs->packed;\n }\n \n-static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n-\t\t\t\t    struct ref_list *list)\n+static void get_ref_dir(const char *submodule, const char *base,\n+\t\t\tstruct ref_array *array)\n {\n \tDIR *dir;\n \tconst char *path;\n@@ -299,7 +250,7 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\tif (stat(refdir, &st) < 0)\n \t\t\t\tcontinue;\n \t\t\tif (S_ISDIR(st.st_mode)) {\n-\t\t\t\tlist = get_ref_dir(submodule, ref, list);\n+\t\t\t\tget_ref_dir(submodule, ref, array);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (submodule) {\n@@ -314,12 +265,11 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\t\t\thashclr(sha1);\n \t\t\t\t\tflag |= REF_BROKEN;\n \t\t\t\t}\n-\t\t\tlist = add_ref(ref, sha1, flag, list, NULL);\n+\t\t\tadd_ref(ref, sha1, flag, array, NULL);\n \t\t}\n \t\tfree(ref);\n \t\tclosedir(dir);\n \t}\n-\treturn list;\n }\n \n struct warn_if_dangling_data {\n@@ -356,21 +306,21 @@ void warn_dangling_symref(FILE *fp, const char *msg_fmt, const char *refname)\n \tfor_each_rawref(warn_if_dangling_symref, &data);\n }\n \n-static struct ref_list *get_loose_refs(const char *submodule)\n+static struct ref_array *get_loose_refs(const char *submodule)\n {\n \tif (submodule) {\n-\t\tfree_ref_list(submodule_refs.loose);\n-\t\tsubmodule_refs.loose = get_ref_dir(submodule, \"refs\", NULL);\n-\t\tsubmodule_refs.loose = sort_ref_list(submodule_refs.loose);\n-\t\treturn submodule_refs.loose;\n+\t\tfree_ref_array(&submodule_refs.loose);\n+\t\tget_ref_dir(submodule, \"refs\", &submodule_refs.loose);\n+\t\tsort_ref_array(&submodule_refs.loose);\n+\t\treturn &submodule_refs.loose;\n \t}\n \n \tif (!cached_refs.did_loose) {\n-\t\tcached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n-\t\tcached_refs.loose = sort_ref_list(cached_refs.loose);\n+\t\tget_ref_dir(NULL, \"refs\", &cached_refs.loose);\n+\t\tsort_ref_array(&cached_refs.loose);\n \t\tcached_refs.did_loose = 1;\n \t}\n-\treturn cached_refs.loose;\n+\treturn &cached_refs.loose;\n }\n \n /* We allow \"recursive\" symbolic refs. Only within reason, though */\n@@ -381,8 +331,8 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n {\n \tFILE *f;\n \tstruct cached_refs refs;\n-\tstruct ref_list *ref;\n-\tint retval;\n+\tstruct ref_entry *ref;\n+\tint retval = -1;\n \n \tstrcpy(name + pathlen, \"packed-refs\");\n \tf = fopen(name, \"r\");\n@@ -390,17 +340,12 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n \t\treturn -1;\n \tread_packed_refs(f, &refs);\n \tfclose(f);\n-\tref = refs.packed;\n-\tretval = -1;\n-\twhile (ref) {\n-\t\tif (!strcmp(ref->name, refname)) {\n-\t\t\tretval = 0;\n-\t\t\tmemcpy(result, ref->sha1, 20);\n-\t\t\tbreak;\n-\t\t}\n-\t\tref = ref->next;\n+\tref = search_ref_array(&refs.packed, refname);\n+\tif (ref != NULL) {\n+\t\tmemcpy(result, ref->sha1, 20);\n+\t\tretval = 0;\n \t}\n-\tfree_ref_list(refs.packed);\n+\tfree_ref_array(&refs.packed);\n \treturn retval;\n }\n \n@@ -501,15 +446,13 @@ const char *resolve_ref(const char *ref, unsigned char *sha1, int reading, int *\n \t\tgit_snpath(path, sizeof(path), \"%s\", ref);\n \t\t/* Special case: non-existing file. */\n \t\tif (lstat(path, &st) < 0) {\n-\t\t\tstruct ref_list *list = get_packed_refs(NULL);\n-\t\t\twhile (list) {\n-\t\t\t\tif (!strcmp(ref, list->name)) {\n-\t\t\t\t\thashcpy(sha1, list->sha1);\n-\t\t\t\t\tif (flag)\n-\t\t\t\t\t\t*flag |= REF_ISPACKED;\n-\t\t\t\t\treturn ref;\n-\t\t\t\t}\n-\t\t\t\tlist = list->next;\n+\t\t\tstruct ref_array *packed = get_packed_refs(NULL);\n+\t\t\tstruct ref_entry *r = search_ref_array(packed, ref);\n+\t\t\tif (r != NULL) {\n+\t\t\t\thashcpy(sha1, r->sha1);\n+\t\t\t\tif (flag)\n+\t\t\t\t\t*flag |= REF_ISPACKED;\n+\t\t\t\treturn ref;\n \t\t\t}\n \t\t\tif (reading || errno != ENOENT)\n \t\t\t\treturn NULL;\n@@ -584,7 +527,7 @@ int read_ref(const char *ref, unsigned char *sha1)\n \n #define DO_FOR_EACH_INCLUDE_BROKEN 01\n static int do_one_ref(const char *base, each_ref_fn fn, int trim,\n-\t\t      int flags, void *cb_data, struct ref_list *entry)\n+\t\t      int flags, void *cb_data, struct ref_entry *entry)\n {\n \tif (prefixcmp(entry->name, base))\n \t\treturn 0;\n@@ -630,18 +573,12 @@ int peel_ref(const char *ref, unsigned char *sha1)\n \t\treturn -1;\n \n \tif ((flag & REF_ISPACKED)) {\n-\t\tstruct ref_list *list = get_packed_refs(NULL);\n+\t\tstruct ref_array *array = get_packed_refs(NULL);\n+\t\tstruct ref_entry *r = search_ref_array(array, ref);\n \n-\t\twhile (list) {\n-\t\t\tif (!strcmp(list->name, ref)) {\n-\t\t\t\tif (list->flag & REF_KNOWS_PEELED) {\n-\t\t\t\t\thashcpy(sha1, list->peeled);\n-\t\t\t\t\treturn 0;\n-\t\t\t\t}\n-\t\t\t\t/* older pack-refs did not leave peeled ones */\n-\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tlist = list->next;\n+\t\tif (r != NULL && r->flag & REF_KNOWS_PEELED) {\n+\t\t\thashcpy(sha1, r->peeled);\n+\t\t\treturn 0;\n \t\t}\n \t}\n \n@@ -660,36 +597,39 @@ fallback:\n static int do_for_each_ref(const char *submodule, const char *base, each_ref_fn fn,\n \t\t\t   int trim, int flags, void *cb_data)\n {\n-\tint retval = 0;\n-\tstruct ref_list *packed = get_packed_refs(submodule);\n-\tstruct ref_list *loose = get_loose_refs(submodule);\n+\tint retval = 0, i, p = 0, l = 0;\n+\tstruct ref_array *packed = get_packed_refs(submodule);\n+\tstruct ref_array *loose = get_loose_refs(submodule);\n \n-\tstruct ref_list *extra;\n+\tstruct ref_array *extra = &extra_refs;\n \n-\tfor (extra = extra_refs; extra; extra = extra->next)\n-\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, extra);\n+\tfor (i = 0; i < extra->nr; i++)\n+\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, extra->refs[i]);\n \n-\twhile (packed && loose) {\n-\t\tstruct ref_list *entry;\n-\t\tint cmp = strcmp(packed->name, loose->name);\n+\twhile (p < packed->nr && l < loose->nr) {\n+\t\tstruct ref_entry *entry;\n+\t\tint cmp = strcmp(packed->refs[p]->name, loose->refs[l]->name);\n \t\tif (!cmp) {\n-\t\t\tpacked = packed->next;\n+\t\t\tp++;\n \t\t\tcontinue;\n \t\t}\n \t\tif (cmp > 0) {\n-\t\t\tentry = loose;\n-\t\t\tloose = loose->next;\n+\t\t\tentry = loose->refs[l++];\n \t\t} else {\n-\t\t\tentry = packed;\n-\t\t\tpacked = packed->next;\n+\t\t\tentry = packed->refs[p++];\n \t\t}\n \t\tretval = do_one_ref(base, fn, trim, flags, cb_data, entry);\n \t\tif (retval)\n \t\t\tgoto end_each;\n \t}\n \n-\tfor (packed = packed ? packed : loose; packed; packed = packed->next) {\n-\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, packed);\n+\tif (l < loose->nr) {\n+\t\tp = l;\n+\t\tpacked = loose;\n+\t}\n+\n+\tfor (; p < packed->nr; p++) {\n+\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, packed->refs[p]);\n \t\tif (retval)\n \t\t\tgoto end_each;\n \t}\n@@ -1005,24 +945,24 @@ static int remove_empty_directories(const char *file)\n }\n \n static int is_refname_available(const char *ref, const char *oldref,\n-\t\t\t\tstruct ref_list *list, int quiet)\n-{\n-\tint namlen = strlen(ref); /* e.g. 'foo/bar' */\n-\twhile (list) {\n-\t\t/* list->name could be 'foo' or 'foo/bar/baz' */\n-\t\tif (!oldref || strcmp(oldref, list->name)) {\n-\t\t\tint len = strlen(list->name);\n+\t\t\t\tstruct ref_array *array, int quiet)\n+{\n+\tint i, namlen = strlen(ref); /* e.g. 'foo/bar' */\n+\tfor (i = 0; i < array->nr; i++ ) {\n+\t\tstruct ref_entry *entry = array->refs[i];\n+\t\t/* entry->name could be 'foo' or 'foo/bar/baz' */\n+\t\tif (!oldref || strcmp(oldref, entry->name)) {\n+\t\t\tint len = strlen(entry->name);\n \t\t\tint cmplen = (namlen < len) ? namlen : len;\n-\t\t\tconst char *lead = (namlen < len) ? list->name : ref;\n-\t\t\tif (!strncmp(ref, list->name, cmplen) &&\n+\t\t\tconst char *lead = (namlen < len) ? entry->name : ref;\n+\t\t\tif (!strncmp(ref, entry->name, cmplen) &&\n \t\t\t    lead[cmplen] == '/') {\n \t\t\t\tif (!quiet)\n \t\t\t\t\terror(\"'%s' exists; cannot create '%s'\",\n-\t\t\t\t\t      list->name, ref);\n+\t\t\t\t\t      entry->name, ref);\n \t\t\t\treturn 0;\n \t\t\t}\n \t\t}\n-\t\tlist = list->next;\n \t}\n \treturn 1;\n }\n@@ -1129,18 +1069,13 @@ static struct lock_file packlock;\n \n static int repack_without_ref(const char *refname)\n {\n-\tstruct ref_list *list, *packed_ref_list;\n-\tint fd;\n-\tint found = 0;\n+\tstruct ref_array *packed;\n+\tstruct ref_entry *ref;\n+\tint fd, i;\n \n-\tpacked_ref_list = get_packed_refs(NULL);\n-\tfor (list = packed_ref_list; list; list = list->next) {\n-\t\tif (!strcmp(refname, list->name)) {\n-\t\t\tfound = 1;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n-\tif (!found)\n+\tpacked = get_packed_refs(NULL);\n+\tref = search_ref_array(packed, refname);\n+\tif (ref == NULL)\n \t\treturn 0;\n \tfd = hold_lock_file_for_update(&packlock, git_path(\"packed-refs\"), 0);\n \tif (fd < 0) {\n@@ -1148,17 +1083,19 @@ static int repack_without_ref(const char *refname)\n \t\treturn error(\"cannot delete '%s' from packed refs\", refname);\n \t}\n \n-\tfor (list = packed_ref_list; list; list = list->next) {\n+\tfor (i = 0; i < packed->nr; i++) {\n \t\tchar line[PATH_MAX + 100];\n \t\tint len;\n \n-\t\tif (!strcmp(refname, list->name))\n+\t\tref = packed->refs[i];\n+\n+\t\tif (!strcmp(refname, ref->name))\n \t\t\tcontinue;\n \t\tlen = snprintf(line, sizeof(line), \"%s %s\\n\",\n-\t\t\t       sha1_to_hex(list->sha1), list->name);\n+\t\t\t       sha1_to_hex(ref->sha1), ref->name);\n \t\t/* this should not happen but just being defensive */\n \t\tif (len > sizeof(line))\n-\t\t\tdie(\"too long a refname '%s'\", list->name);\n+\t\t\tdie(\"too long a refname '%s'\", ref->name);\n \t\twrite_or_die(fd, line, len);\n \t}\n \treturn commit_lock_file(&packlock);\n-- \n1.7.6.1\n"},{"id":"176535","messageId":"7vvcsbqa0k.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"20110929041811.5363.33396.julian@quantumfyre.co.uk","subject":"Re: [PATCH] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-29T22:06:03Z","receivedAt":"2011-09-29T22:06:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Julian Phillips <julian@quantumfyre.co.uk> writes:\n\n> +static void add_ref(const char *name, const unsigned char *sha1,\n> +\t\t    int flag, struct ref_array *refs,\n> +\t\t    struct ref_entry **new_entry)\n>  {\n>  \tint len;\n> -\tstruct ref_list *entry;\n> +\tstruct ref_entry *entry;\n>  \n>  \t/* Allocate it and add it in.. */\n>  \tlen = strlen(name) + 1;\n> -\tentry = xmalloc(sizeof(struct ref_list) + len);\n> +\tentry = xmalloc(sizeof(struct ref) + len);\n\nThis should be sizeof(struct ref_entry), no?  There is another such\nmisallocation in search_ref_array() where it prepares a temporary.\n"},{"id":"176539","messageId":"20110929221143.23806.25666.julian@quantumfyre.co.uk","threadId":"27589","inReplyTo":"7vvcsbqa0k.fsf@alter.siamese.dyndns.org","subject":"[PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-29T22:11:42Z","receivedAt":"2011-09-29T22:11:42Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"Currently we linearly search through lists of refs when we need to\nfind a specific ref.  This can be very slow if we need to lookup a\nlarge number of refs.  By changing to a binary search we can make this\nfaster.\n\nIn order to be able to use a binary search we need to change from\nusing linked lists to arrays, which we can manage using ALLOC_GROW.\n\nWe can now also use the standard library qsort function to sort the\nrefs arrays.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\n\nOn Thu, 29 Sep 2011 15:06:03 -0700, Junio C Hamano wrote:\n> Julian Phillips <julian@quantumfyre.co.uk> writes:\n>\n>> +static void add_ref(const char *name, const unsigned char *sha1,\n>> +\t\t    int flag, struct ref_array *refs,\n>> +\t\t    struct ref_entry **new_entry)\n>>  {\n>>  \tint len;\n>> -\tstruct ref_list *entry;\n>> +\tstruct ref_entry *entry;\n>>\n>>  \t/* Allocate it and add it in.. */\n>>  \tlen = strlen(name) + 1;\n>> -\tentry = xmalloc(sizeof(struct ref_list) + len);\n>> +\tentry = xmalloc(sizeof(struct ref) + len);\n>\n> This should be sizeof(struct ref_entry), no?  There is another such\n> misallocation in search_ref_array() where it prepares a temporary.\n\nIndeed, thanks.\n\nLooks like two instances of not noticing that \"struct ref\" already existed\nmanaged to survive.  Drat.  Of course since \"struct ref\" is bigger than \"struct\nref_entry\", everthing worked fine ... so no failed tests to tip me off.\n\n refs.c |  329 ++++++++++++++++++++++++++--------------------------------------\n 1 files changed, 133 insertions(+), 196 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex a49ff74..4c01d79 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -8,14 +8,18 @@\n #define REF_KNOWS_PEELED 04\n #define REF_BROKEN 010\n \n-struct ref_list {\n-\tstruct ref_list *next;\n+struct ref_entry {\n \tunsigned char flag; /* ISSYMREF? ISPACKED? */\n \tunsigned char sha1[20];\n \tunsigned char peeled[20];\n \tchar name[FLEX_ARRAY];\n };\n \n+struct ref_array {\n+\tint nr, alloc;\n+\tstruct ref_entry **refs;\n+};\n+\n static const char *parse_ref_line(char *line, unsigned char *sha1)\n {\n \t/*\n@@ -44,108 +48,58 @@ static const char *parse_ref_line(char *line, unsigned char *sha1)\n \treturn line;\n }\n \n-static struct ref_list *add_ref(const char *name, const unsigned char *sha1,\n-\t\t\t\tint flag, struct ref_list *list,\n-\t\t\t\tstruct ref_list **new_entry)\n+static void add_ref(const char *name, const unsigned char *sha1,\n+\t\t    int flag, struct ref_array *refs,\n+\t\t    struct ref_entry **new_entry)\n {\n \tint len;\n-\tstruct ref_list *entry;\n+\tstruct ref_entry *entry;\n \n \t/* Allocate it and add it in.. */\n \tlen = strlen(name) + 1;\n-\tentry = xmalloc(sizeof(struct ref_list) + len);\n+\tentry = xmalloc(sizeof(struct ref_entry) + len);\n \thashcpy(entry->sha1, sha1);\n \thashclr(entry->peeled);\n \tmemcpy(entry->name, name, len);\n \tentry->flag = flag;\n-\tentry->next = list;\n \tif (new_entry)\n \t\t*new_entry = entry;\n-\treturn entry;\n+\tALLOC_GROW(refs->refs, refs->nr + 1, refs->alloc);\n+\trefs->refs[refs->nr++] = entry;\n }\n \n-/* merge sort the ref list */\n-static struct ref_list *sort_ref_list(struct ref_list *list)\n+static int ref_entry_cmp(const void *a, const void *b)\n {\n-\tint psize, qsize, last_merge_count, cmp;\n-\tstruct ref_list *p, *q, *l, *e;\n-\tstruct ref_list *new_list = list;\n-\tint k = 1;\n-\tint merge_count = 0;\n-\n-\tif (!list)\n-\t\treturn list;\n-\n-\tdo {\n-\t\tlast_merge_count = merge_count;\n-\t\tmerge_count = 0;\n-\n-\t\tpsize = 0;\n+\tstruct ref_entry *one = *(struct ref_entry **)a;\n+\tstruct ref_entry *two = *(struct ref_entry **)b;\n+\treturn strcmp(one->name, two->name);\n+}\n \n-\t\tp = new_list;\n-\t\tq = new_list;\n-\t\tnew_list = NULL;\n-\t\tl = NULL;\n+static void sort_ref_array(struct ref_array *array)\n+{\n+\tqsort(array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n+}\n \n-\t\twhile (p) {\n-\t\t\tmerge_count++;\n+static struct ref_entry *search_ref_array(struct ref_array *array, const char *name)\n+{\n+\tstruct ref_entry *e, **r;\n+\tint len;\n \n-\t\t\twhile (psize < k && q->next) {\n-\t\t\t\tq = q->next;\n-\t\t\t\tpsize++;\n-\t\t\t}\n-\t\t\tqsize = k;\n-\n-\t\t\twhile ((psize > 0) || (qsize > 0 && q)) {\n-\t\t\t\tif (qsize == 0 || !q) {\n-\t\t\t\t\te = p;\n-\t\t\t\t\tp = p->next;\n-\t\t\t\t\tpsize--;\n-\t\t\t\t} else if (psize == 0) {\n-\t\t\t\t\te = q;\n-\t\t\t\t\tq = q->next;\n-\t\t\t\t\tqsize--;\n-\t\t\t\t} else {\n-\t\t\t\t\tcmp = strcmp(q->name, p->name);\n-\t\t\t\t\tif (cmp < 0) {\n-\t\t\t\t\t\te = q;\n-\t\t\t\t\t\tq = q->next;\n-\t\t\t\t\t\tqsize--;\n-\t\t\t\t\t} else if (cmp > 0) {\n-\t\t\t\t\t\te = p;\n-\t\t\t\t\t\tp = p->next;\n-\t\t\t\t\t\tpsize--;\n-\t\t\t\t\t} else {\n-\t\t\t\t\t\tif (hashcmp(q->sha1, p->sha1))\n-\t\t\t\t\t\t\tdie(\"Duplicated ref, and SHA1s don't match: %s\",\n-\t\t\t\t\t\t\t    q->name);\n-\t\t\t\t\t\twarning(\"Duplicated ref: %s\", q->name);\n-\t\t\t\t\t\te = q;\n-\t\t\t\t\t\tq = q->next;\n-\t\t\t\t\t\tqsize--;\n-\t\t\t\t\t\tfree(e);\n-\t\t\t\t\t\te = p;\n-\t\t\t\t\t\tp = p->next;\n-\t\t\t\t\t\tpsize--;\n-\t\t\t\t\t}\n-\t\t\t\t}\n+\tif (name == NULL)\n+\t\treturn NULL;\n \n-\t\t\t\te->next = NULL;\n+\tlen = strlen(name) + 1;\n+\te = xmalloc(sizeof(struct ref_entry) + len);\n+\tmemcpy(e->name, name, len);\n \n-\t\t\t\tif (l)\n-\t\t\t\t\tl->next = e;\n-\t\t\t\tif (!new_list)\n-\t\t\t\t\tnew_list = e;\n-\t\t\t\tl = e;\n-\t\t\t}\n+\tr = bsearch(&e, array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n \n-\t\t\tp = q;\n-\t\t};\n+\tfree(e);\n \n-\t\tk = k * 2;\n-\t} while ((last_merge_count != merge_count) || (last_merge_count != 1));\n+\tif (r == NULL)\n+\t\treturn NULL;\n \n-\treturn new_list;\n+\treturn *r;\n }\n \n /*\n@@ -155,38 +109,37 @@ static struct ref_list *sort_ref_list(struct ref_list *list)\n static struct cached_refs {\n \tchar did_loose;\n \tchar did_packed;\n-\tstruct ref_list *loose;\n-\tstruct ref_list *packed;\n+\tstruct ref_array loose;\n+\tstruct ref_array packed;\n } cached_refs, submodule_refs;\n-static struct ref_list *current_ref;\n+static struct ref_entry *current_ref;\n \n-static struct ref_list *extra_refs;\n+static struct ref_array extra_refs;\n \n-static void free_ref_list(struct ref_list *list)\n+static void free_ref_array(struct ref_array *array)\n {\n-\tstruct ref_list *next;\n-\tfor ( ; list; list = next) {\n-\t\tnext = list->next;\n-\t\tfree(list);\n-\t}\n+\tint i;\n+\tfor (i = 0; i < array->nr; i++)\n+\t\tfree(array->refs[i]);\n+\tfree(array->refs);\n+\tarray->nr = array->alloc = 0;\n+\tarray->refs = NULL;\n }\n \n static void invalidate_cached_refs(void)\n {\n \tstruct cached_refs *ca = &cached_refs;\n \n-\tif (ca->did_loose && ca->loose)\n-\t\tfree_ref_list(ca->loose);\n-\tif (ca->did_packed && ca->packed)\n-\t\tfree_ref_list(ca->packed);\n-\tca->loose = ca->packed = NULL;\n+\tif (ca->did_loose)\n+\t\tfree_ref_array(&ca->loose);\n+\tif (ca->did_packed)\n+\t\tfree_ref_array(&ca->packed);\n \tca->did_loose = ca->did_packed = 0;\n }\n \n static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n {\n-\tstruct ref_list *list = NULL;\n-\tstruct ref_list *last = NULL;\n+\tstruct ref_entry *last = NULL;\n \tchar refline[PATH_MAX];\n \tint flag = REF_ISPACKED;\n \n@@ -205,7 +158,7 @@ static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n \n \t\tname = parse_ref_line(refline, sha1);\n \t\tif (name) {\n-\t\t\tlist = add_ref(name, sha1, flag, list, &last);\n+\t\t\tadd_ref(name, sha1, flag, &cached_refs->packed, &last);\n \t\t\tcontinue;\n \t\t}\n \t\tif (last &&\n@@ -215,21 +168,20 @@ static void read_packed_refs(FILE *f, struct cached_refs *cached_refs)\n \t\t    !get_sha1_hex(refline + 1, sha1))\n \t\t\thashcpy(last->peeled, sha1);\n \t}\n-\tcached_refs->packed = sort_ref_list(list);\n+\tsort_ref_array(&cached_refs->packed);\n }\n \n void add_extra_ref(const char *name, const unsigned char *sha1, int flag)\n {\n-\textra_refs = add_ref(name, sha1, flag, extra_refs, NULL);\n+\tadd_ref(name, sha1, flag, &extra_refs, NULL);\n }\n \n void clear_extra_refs(void)\n {\n-\tfree_ref_list(extra_refs);\n-\textra_refs = NULL;\n+\tfree_ref_array(&extra_refs);\n }\n \n-static struct ref_list *get_packed_refs(const char *submodule)\n+static struct ref_array *get_packed_refs(const char *submodule)\n {\n \tconst char *packed_refs_file;\n \tstruct cached_refs *refs;\n@@ -237,7 +189,7 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \tif (submodule) {\n \t\tpacked_refs_file = git_path_submodule(submodule, \"packed-refs\");\n \t\trefs = &submodule_refs;\n-\t\tfree_ref_list(refs->packed);\n+\t\tfree_ref_array(&refs->packed);\n \t} else {\n \t\tpacked_refs_file = git_path(\"packed-refs\");\n \t\trefs = &cached_refs;\n@@ -245,18 +197,17 @@ static struct ref_list *get_packed_refs(const char *submodule)\n \n \tif (!refs->did_packed || submodule) {\n \t\tFILE *f = fopen(packed_refs_file, \"r\");\n-\t\trefs->packed = NULL;\n \t\tif (f) {\n \t\t\tread_packed_refs(f, refs);\n \t\t\tfclose(f);\n \t\t}\n \t\trefs->did_packed = 1;\n \t}\n-\treturn refs->packed;\n+\treturn &refs->packed;\n }\n \n-static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n-\t\t\t\t    struct ref_list *list)\n+static void get_ref_dir(const char *submodule, const char *base,\n+\t\t\tstruct ref_array *array)\n {\n \tDIR *dir;\n \tconst char *path;\n@@ -299,7 +250,7 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\tif (stat(refdir, &st) < 0)\n \t\t\t\tcontinue;\n \t\t\tif (S_ISDIR(st.st_mode)) {\n-\t\t\t\tlist = get_ref_dir(submodule, ref, list);\n+\t\t\t\tget_ref_dir(submodule, ref, array);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (submodule) {\n@@ -314,12 +265,11 @@ static struct ref_list *get_ref_dir(const char *submodule, const char *base,\n \t\t\t\t\thashclr(sha1);\n \t\t\t\t\tflag |= REF_BROKEN;\n \t\t\t\t}\n-\t\t\tlist = add_ref(ref, sha1, flag, list, NULL);\n+\t\t\tadd_ref(ref, sha1, flag, array, NULL);\n \t\t}\n \t\tfree(ref);\n \t\tclosedir(dir);\n \t}\n-\treturn list;\n }\n \n struct warn_if_dangling_data {\n@@ -356,21 +306,21 @@ void warn_dangling_symref(FILE *fp, const char *msg_fmt, const char *refname)\n \tfor_each_rawref(warn_if_dangling_symref, &data);\n }\n \n-static struct ref_list *get_loose_refs(const char *submodule)\n+static struct ref_array *get_loose_refs(const char *submodule)\n {\n \tif (submodule) {\n-\t\tfree_ref_list(submodule_refs.loose);\n-\t\tsubmodule_refs.loose = get_ref_dir(submodule, \"refs\", NULL);\n-\t\tsubmodule_refs.loose = sort_ref_list(submodule_refs.loose);\n-\t\treturn submodule_refs.loose;\n+\t\tfree_ref_array(&submodule_refs.loose);\n+\t\tget_ref_dir(submodule, \"refs\", &submodule_refs.loose);\n+\t\tsort_ref_array(&submodule_refs.loose);\n+\t\treturn &submodule_refs.loose;\n \t}\n \n \tif (!cached_refs.did_loose) {\n-\t\tcached_refs.loose = get_ref_dir(NULL, \"refs\", NULL);\n-\t\tcached_refs.loose = sort_ref_list(cached_refs.loose);\n+\t\tget_ref_dir(NULL, \"refs\", &cached_refs.loose);\n+\t\tsort_ref_array(&cached_refs.loose);\n \t\tcached_refs.did_loose = 1;\n \t}\n-\treturn cached_refs.loose;\n+\treturn &cached_refs.loose;\n }\n \n /* We allow \"recursive\" symbolic refs. Only within reason, though */\n@@ -381,8 +331,8 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n {\n \tFILE *f;\n \tstruct cached_refs refs;\n-\tstruct ref_list *ref;\n-\tint retval;\n+\tstruct ref_entry *ref;\n+\tint retval = -1;\n \n \tstrcpy(name + pathlen, \"packed-refs\");\n \tf = fopen(name, \"r\");\n@@ -390,17 +340,12 @@ static int resolve_gitlink_packed_ref(char *name, int pathlen, const char *refna\n \t\treturn -1;\n \tread_packed_refs(f, &refs);\n \tfclose(f);\n-\tref = refs.packed;\n-\tretval = -1;\n-\twhile (ref) {\n-\t\tif (!strcmp(ref->name, refname)) {\n-\t\t\tretval = 0;\n-\t\t\tmemcpy(result, ref->sha1, 20);\n-\t\t\tbreak;\n-\t\t}\n-\t\tref = ref->next;\n+\tref = search_ref_array(&refs.packed, refname);\n+\tif (ref != NULL) {\n+\t\tmemcpy(result, ref->sha1, 20);\n+\t\tretval = 0;\n \t}\n-\tfree_ref_list(refs.packed);\n+\tfree_ref_array(&refs.packed);\n \treturn retval;\n }\n \n@@ -501,15 +446,13 @@ const char *resolve_ref(const char *ref, unsigned char *sha1, int reading, int *\n \t\tgit_snpath(path, sizeof(path), \"%s\", ref);\n \t\t/* Special case: non-existing file. */\n \t\tif (lstat(path, &st) < 0) {\n-\t\t\tstruct ref_list *list = get_packed_refs(NULL);\n-\t\t\twhile (list) {\n-\t\t\t\tif (!strcmp(ref, list->name)) {\n-\t\t\t\t\thashcpy(sha1, list->sha1);\n-\t\t\t\t\tif (flag)\n-\t\t\t\t\t\t*flag |= REF_ISPACKED;\n-\t\t\t\t\treturn ref;\n-\t\t\t\t}\n-\t\t\t\tlist = list->next;\n+\t\t\tstruct ref_array *packed = get_packed_refs(NULL);\n+\t\t\tstruct ref_entry *r = search_ref_array(packed, ref);\n+\t\t\tif (r != NULL) {\n+\t\t\t\thashcpy(sha1, r->sha1);\n+\t\t\t\tif (flag)\n+\t\t\t\t\t*flag |= REF_ISPACKED;\n+\t\t\t\treturn ref;\n \t\t\t}\n \t\t\tif (reading || errno != ENOENT)\n \t\t\t\treturn NULL;\n@@ -584,7 +527,7 @@ int read_ref(const char *ref, unsigned char *sha1)\n \n #define DO_FOR_EACH_INCLUDE_BROKEN 01\n static int do_one_ref(const char *base, each_ref_fn fn, int trim,\n-\t\t      int flags, void *cb_data, struct ref_list *entry)\n+\t\t      int flags, void *cb_data, struct ref_entry *entry)\n {\n \tif (prefixcmp(entry->name, base))\n \t\treturn 0;\n@@ -630,18 +573,12 @@ int peel_ref(const char *ref, unsigned char *sha1)\n \t\treturn -1;\n \n \tif ((flag & REF_ISPACKED)) {\n-\t\tstruct ref_list *list = get_packed_refs(NULL);\n+\t\tstruct ref_array *array = get_packed_refs(NULL);\n+\t\tstruct ref_entry *r = search_ref_array(array, ref);\n \n-\t\twhile (list) {\n-\t\t\tif (!strcmp(list->name, ref)) {\n-\t\t\t\tif (list->flag & REF_KNOWS_PEELED) {\n-\t\t\t\t\thashcpy(sha1, list->peeled);\n-\t\t\t\t\treturn 0;\n-\t\t\t\t}\n-\t\t\t\t/* older pack-refs did not leave peeled ones */\n-\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tlist = list->next;\n+\t\tif (r != NULL && r->flag & REF_KNOWS_PEELED) {\n+\t\t\thashcpy(sha1, r->peeled);\n+\t\t\treturn 0;\n \t\t}\n \t}\n \n@@ -660,36 +597,39 @@ fallback:\n static int do_for_each_ref(const char *submodule, const char *base, each_ref_fn fn,\n \t\t\t   int trim, int flags, void *cb_data)\n {\n-\tint retval = 0;\n-\tstruct ref_list *packed = get_packed_refs(submodule);\n-\tstruct ref_list *loose = get_loose_refs(submodule);\n+\tint retval = 0, i, p = 0, l = 0;\n+\tstruct ref_array *packed = get_packed_refs(submodule);\n+\tstruct ref_array *loose = get_loose_refs(submodule);\n \n-\tstruct ref_list *extra;\n+\tstruct ref_array *extra = &extra_refs;\n \n-\tfor (extra = extra_refs; extra; extra = extra->next)\n-\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, extra);\n+\tfor (i = 0; i < extra->nr; i++)\n+\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, extra->refs[i]);\n \n-\twhile (packed && loose) {\n-\t\tstruct ref_list *entry;\n-\t\tint cmp = strcmp(packed->name, loose->name);\n+\twhile (p < packed->nr && l < loose->nr) {\n+\t\tstruct ref_entry *entry;\n+\t\tint cmp = strcmp(packed->refs[p]->name, loose->refs[l]->name);\n \t\tif (!cmp) {\n-\t\t\tpacked = packed->next;\n+\t\t\tp++;\n \t\t\tcontinue;\n \t\t}\n \t\tif (cmp > 0) {\n-\t\t\tentry = loose;\n-\t\t\tloose = loose->next;\n+\t\t\tentry = loose->refs[l++];\n \t\t} else {\n-\t\t\tentry = packed;\n-\t\t\tpacked = packed->next;\n+\t\t\tentry = packed->refs[p++];\n \t\t}\n \t\tretval = do_one_ref(base, fn, trim, flags, cb_data, entry);\n \t\tif (retval)\n \t\t\tgoto end_each;\n \t}\n \n-\tfor (packed = packed ? packed : loose; packed; packed = packed->next) {\n-\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, packed);\n+\tif (l < loose->nr) {\n+\t\tp = l;\n+\t\tpacked = loose;\n+\t}\n+\n+\tfor (; p < packed->nr; p++) {\n+\t\tretval = do_one_ref(base, fn, trim, flags, cb_data, packed->refs[p]);\n \t\tif (retval)\n \t\t\tgoto end_each;\n \t}\n@@ -1005,24 +945,24 @@ static int remove_empty_directories(const char *file)\n }\n \n static int is_refname_available(const char *ref, const char *oldref,\n-\t\t\t\tstruct ref_list *list, int quiet)\n-{\n-\tint namlen = strlen(ref); /* e.g. 'foo/bar' */\n-\twhile (list) {\n-\t\t/* list->name could be 'foo' or 'foo/bar/baz' */\n-\t\tif (!oldref || strcmp(oldref, list->name)) {\n-\t\t\tint len = strlen(list->name);\n+\t\t\t\tstruct ref_array *array, int quiet)\n+{\n+\tint i, namlen = strlen(ref); /* e.g. 'foo/bar' */\n+\tfor (i = 0; i < array->nr; i++ ) {\n+\t\tstruct ref_entry *entry = array->refs[i];\n+\t\t/* entry->name could be 'foo' or 'foo/bar/baz' */\n+\t\tif (!oldref || strcmp(oldref, entry->name)) {\n+\t\t\tint len = strlen(entry->name);\n \t\t\tint cmplen = (namlen < len) ? namlen : len;\n-\t\t\tconst char *lead = (namlen < len) ? list->name : ref;\n-\t\t\tif (!strncmp(ref, list->name, cmplen) &&\n+\t\t\tconst char *lead = (namlen < len) ? entry->name : ref;\n+\t\t\tif (!strncmp(ref, entry->name, cmplen) &&\n \t\t\t    lead[cmplen] == '/') {\n \t\t\t\tif (!quiet)\n \t\t\t\t\terror(\"'%s' exists; cannot create '%s'\",\n-\t\t\t\t\t      list->name, ref);\n+\t\t\t\t\t      entry->name, ref);\n \t\t\t\treturn 0;\n \t\t\t}\n \t\t}\n-\t\tlist = list->next;\n \t}\n \treturn 1;\n }\n@@ -1129,18 +1069,13 @@ static struct lock_file packlock;\n \n static int repack_without_ref(const char *refname)\n {\n-\tstruct ref_list *list, *packed_ref_list;\n-\tint fd;\n-\tint found = 0;\n+\tstruct ref_array *packed;\n+\tstruct ref_entry *ref;\n+\tint fd, i;\n \n-\tpacked_ref_list = get_packed_refs(NULL);\n-\tfor (list = packed_ref_list; list; list = list->next) {\n-\t\tif (!strcmp(refname, list->name)) {\n-\t\t\tfound = 1;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n-\tif (!found)\n+\tpacked = get_packed_refs(NULL);\n+\tref = search_ref_array(packed, refname);\n+\tif (ref == NULL)\n \t\treturn 0;\n \tfd = hold_lock_file_for_update(&packlock, git_path(\"packed-refs\"), 0);\n \tif (fd < 0) {\n@@ -1148,17 +1083,19 @@ static int repack_without_ref(const char *refname)\n \t\treturn error(\"cannot delete '%s' from packed refs\", refname);\n \t}\n \n-\tfor (list = packed_ref_list; list; list = list->next) {\n+\tfor (i = 0; i < packed->nr; i++) {\n \t\tchar line[PATH_MAX + 100];\n \t\tint len;\n \n-\t\tif (!strcmp(refname, list->name))\n+\t\tref = packed->refs[i];\n+\n+\t\tif (!strcmp(refname, ref->name))\n \t\t\tcontinue;\n \t\tlen = snprintf(line, sizeof(line), \"%s %s\\n\",\n-\t\t\t       sha1_to_hex(list->sha1), list->name);\n+\t\t\t       sha1_to_hex(ref->sha1), ref->name);\n \t\t/* this should not happen but just being defensive */\n \t\tif (len > sizeof(line))\n-\t\t\tdie(\"too long a refname '%s'\", list->name);\n+\t\t\tdie(\"too long a refname '%s'\", ref->name);\n \t\twrite_or_die(fd, line, len);\n \t}\n \treturn commit_lock_file(&packlock);\n-- \n1.7.6.1\n"},{"id":"176544","messageId":"7v62karjv3.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"20110929221143.23806.25666.julian@quantumfyre.co.uk","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-29T23:48:00Z","receivedAt":"2011-09-29T23:48:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This version looks sane, although I have a suspicion that it may have\nsome interaction with what Michael may be working on.\n\nThanks.\n"},{"id":"176547","messageId":"201109291913.34196.mfick@codeaurora.org","threadId":"27589","inReplyTo":"20110929221143.23806.25666.julian@quantumfyre.co.uk","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T01:13:08Z","receivedAt":"2011-09-30T01:13:08Z","isPatch":true,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Thursday, September 29, 2011 04:11:42 pm Julian Phillips \nwrote:\n> Currently we linearly search through lists of refs when\n> we need to find a specific ref.  This can be very slow\n> if we need to lookup a large number of refs.  By\n> changing to a binary search we can make this faster.\n> \n> In order to be able to use a binary search we need to\n> change from using linked lists to arrays, which we can\n> manage using ALLOC_GROW.\n> \n> We can now also use the standard library qsort function\n> to sort the refs arrays.\n> \n\nThis works for me, however unfortunately, I cannot find any \nscenarios where it improves anything over the previous fix \nby René.  :(\n\nI tested many things, clones, fetches, fetch noops, \ncheckouts, garbage collection.  I am a bit surprised, \nbecause I thought that my hack of a hash map did improve \nstill on checkouts on packed refs, but it could just be that \nmy hack was buggy and did not actually do a full orphan \ncheck.\n\nThanks,\n\n-Martin\n"},{"id":"176551","messageId":"7vwrcqpuc7.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"201109291913.34196.mfick@codeaurora.org","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-30T03:44:40Z","receivedAt":"2011-09-30T03:44:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Fick <mfick@codeaurora.org> writes:\n\n> This works for me, however unfortunately, I cannot find any \n> scenarios where it improves anything over the previous fix \n> by René.  :(\n\nNevertheless, I would appreciate it if you can try this _without_ René's\npatch. This attempts to make resolve_ref() cheap for _any_ caller. René's\npatch avoids calling it in one specific callchain.\n\nThey address different issues. René's patch is probably an independently\ngood change (I haven't thought about the interactions with the topics in\nflight and its implications on the future direction), but would not help\nother/new callers that make many calls to resolve_ref().\n"},{"id":"176556","messageId":"a9f3dba5f48adfa603d76b7d49111e3d@quantumfyre.co.uk","threadId":"27589","inReplyTo":"7vwrcqpuc7.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-30T08:04:02Z","receivedAt":"2011-09-30T08:04:02Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Thu, 29 Sep 2011 20:44:40 -0700, Junio C Hamano wrote:\n> Martin Fick <mfick@codeaurora.org> writes:\n>\n>> This works for me, however unfortunately, I cannot find any\n>> scenarios where it improves anything over the previous fix\n>> by René.  :(\n>\n> Nevertheless, I would appreciate it if you can try this _without_ \n> René's\n> patch. This attempts to make resolve_ref() cheap for _any_ caller. \n> René's\n> patch avoids calling it in one specific callchain.\n>\n> They address different issues. René's patch is probably an \n> independently\n> good change (I haven't thought about the interactions with the topics \n> in\n> flight and its implications on the future direction), but would not \n> help\n> other/new callers that make many calls to resolve_ref().\n\nIt certainly helps with my test repo (~140k refs, of which ~40k are \nbranches).  User times for checkout starting from an orphaned commit \nare:\n\nNo fix          : ~16m8s\n+ Binary Search : ~4s\n+ René's patch  : ~2s\n\n(The 2s includes both patches, though the timing is the same for René's \npatch alone)\n\n-- \nJulian\n"},{"id":"176561","messageId":"4E8587E8.9070606@lsrfire.ath.cx","threadId":"27589","inReplyTo":"201109291411.06733.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-09-30T09:12:08Z","receivedAt":"2011-09-30T09:12:08Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Hi Martin,\n\nAm 29.09.2011 22:11, schrieb Martin Fick:\n> Your patch works well for me.  It achieves about the same \n> gains as Julian's patch. Thanks!\n\nOK, and what happens if you apply the following patch on top of my first\none?  It avoids going through all the refs a second time during cleanup,\nat the cost of going through the list of all known objects.  I wonder if\nthat's any faster in your case.\n\nThanks,\nRené\n\n\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 84e0cdc..a4b1003 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -596,15 +596,14 @@ static int add_pending_uninteresting_ref(const char *refname,\n \treturn 0;\n }\n \n-static int clear_commit_marks_from_one_ref(const char *refname,\n-\t\t\t\t      const unsigned char *sha1,\n-\t\t\t\t      int flags,\n-\t\t\t\t      void *cb_data)\n+static void clear_commit_marks_for_all(unsigned int mark)\n {\n-\tstruct commit *commit = lookup_commit_reference_gently(sha1, 1);\n-\tif (commit)\n-\t\tclear_commit_marks(commit, -1);\n-\treturn 0;\n+\tunsigned int i, max = get_max_object_index();\n+\tfor (i = 0; i < max; i++) {\n+\t\tstruct object *object = get_indexed_object(i);\n+\t\tif (object && object->type == OBJ_COMMIT)\n+\t\t\tobject->flags &= ~mark;\n+\t}\n }\n \n static void describe_one_orphan(struct strbuf *sb, struct commit *commit)\n@@ -690,8 +689,7 @@ static void orphaned_commit_warning(struct commit *commit)\n \telse\n \t\tdescribe_detached_head(_(\"Previous HEAD position was\"), commit);\n \n-\tclear_commit_marks(commit, -1);\n-\tfor_each_ref(clear_commit_marks_from_one_ref, NULL);\n+\tclear_commit_marks_for_all(ALL_REV_FLAGS);\n }\n \n static int switch_branches(struct checkout_opts *opts, struct branch_info *new)\n"},{"id":"176573","messageId":"4E85E07C.5070402@alum.mit.edu","threadId":"27589","inReplyTo":"7v62karjv3.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-09-30T15:30:04Z","receivedAt":"2011-09-30T15:30:04Z","isPatch":true,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 09/30/2011 01:48 AM, Junio C Hamano wrote:\n> This version looks sane, although I have a suspicion that it may have\n> some interaction with what Michael may be working on.\n\nIndeed, I have almost equivalent changes in the giant patch series that\nI am working on [1].  The branch is very experimental.  The tip\ncurrently passes all the tests, but it has a known performance\nregression in connection if \"git fetch\" is used to fetch many commits.\n\n\nBut before comparing ref-related optimizations, we have an *urgent* need\nfor a decent performance test suite.  There are many slightly different\nscenarios that have very different performance characteristics, and we\nhave to be sure that we are optimizing for the whole palette of\nmany-reference use cases.  So I made an attempt at a kludgey but\nsomewhat flexible performance-testing script [2].  I don't know whether\nsomething like this should be integrated into the git project, and if so\nwhere; suggestions are welcome.\n\n\nTo run the tests, from the root of the git source tree:\n\n    make # make sure git is up-to-date\n    t/make-refperf-repo --help\n    t/make-refperf-repo [OPTIONS]\n    t/refperf\n    cat refperf.times # See the results\n\nThe default repo has 5k commits in a linear series with one reference on\neach commit.  (These numbers can both be adjusted.)\n\nThe reference namespace can be laid out a few ways:\n\n* Many references in a single \"directory\" vs. sharded over many\n\"directories\"\n\n* In lexicographic order by commit, in reverse order, or \"shuffled\".\n\nBy default, the repo is written to \"refperf-repo\".\n\nThe time it takes to create the test repository is itself also an\ninteresting benchmark.  For example, on the maint branch it is terribly\nslow unless it is passed either the --pack-refs-interval=N (with N, say\n100) or --no-replace-object option.  I also noticed that if it is run like\n\n    t/make-refperf-repo --refs=5000 --commits=5000 \\\n            --pack-refs-interval=100\n\n(one ref per commit), git-pack-refs becomes precipitously and\ndramatically slower after the 2000th commit.\n\nI haven't had time yet for systematic benchmarks of other git versions.\n\nSee the refperf script to see what sorts of benchmarks that I have built\ninto it so far.  The refperf test is non-destructive; it always copies\nfrom \"refperf-repo\" to \"refperf-repo-copy\" and does its tests in the\ncopy; therefore a test repo can be reused.  The timing data are written\nto \"refperf.times\" and other output to \"refperf.log\".\n\nHere are my refperf results for the \"maint\" branch on my notebook with\nthe default \"make-refperf-repo\" arguments (times in seconds):\n\n3.36 git branch (cold)\n0.01 git branch (warm)\n0.04 git for-each-ref\n3.08 git checkout (cold)\n0.01 git checkout (warm)\n0.00 git checkout --orphan (warm)\n0.15 git checkout from detached orphan\n0.12 git pack-refs\n1.17 git branch (cold)\n0.00 git branch (warm)\n0.17 git for-each-ref\n0.95 git checkout (cold)\n0.00 git checkout (warm)\n0.00 git checkout --orphan (warm)\n0.21 git checkout from detached orphan\n0.18 git branch -a --contains\n7.67 git clone\n0.06 git fetch (nothing)\n0.01 git pack-refs\n0.05 git fetch (nothing, packed)\n0.10 git clone of a ref-packed repo\n0.63 git fetch (everything)\n\nProbably we should test with even more references than this, but this\ntest already shows that some commands are quite sluggish.\n\nThere are some more things that could be added, like:\n\n* Branches vs. annotated tags\n\n* References on the tips of branches in a more typical \"branchy\" repository.\n\n* git describe --all\n\n* git log --decorate\n\n* git gc\n\n* git filter-branch\n  (This has very different performance characteristics because it is a\nscript that invokes git many times.)\n\nI suggest that we try to do systematic benchmarking of any changes that\nwe claim are performance optimizations and share before/after results in\nthe cover letter for the patch series.\n\nMichael\n\n[1] branch hierarchical-refs at git://github.com/mhagger/git.git\n[2] branch refperf at git://github.com/mhagger/git.git\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176576","messageId":"201109300945.34728.mfick@codeaurora.org","threadId":"27589","inReplyTo":"7vwrcqpuc7.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T15:45:34Z","receivedAt":"2011-09-30T15:45:34Z","isPatch":true,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Thursday, September 29, 2011 09:44:40 pm Junio C Hamano \nwrote:\n> Martin Fick <mfick@codeaurora.org> writes:\n> > This works for me, however unfortunately, I cannot find\n> > any scenarios where it improves anything over the\n> > previous fix by René.  :(\n> \n> Nevertheless, I would appreciate it if you can try this\n> _without_ René's patch. This attempts to make\n> resolve_ref() cheap for _any_ caller. René's patch\n> avoids calling it in one specific callchain.\n> \n> They address different issues. René's patch is probably\n> an independently good change (I haven't thought about\n> the interactions with the topics in flight and its\n> implications on the future direction), but would not\n> help other/new callers that make many calls to\n> resolve_ref().\n\nAgreed.  Here is what I am seeing without René's patch.\n\nCheckout in NON packed ref repo takes about 20s, with patch \nv3 of binary search, it takes about 11s (1s slower than \nRené's patch).\n\nCheckout in packed ref repo takes about 5:30min, with patch \nv3 of binary search, it takes about 10s (also 1s slower than \nRené's patch).\n\nI'd say that's not bad, it seems like the 1s difference is \ndoing the search 60K+times (my tests don't quite scan the \nfull list), so the search seems to scale well with patch v3.\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176577","messageId":"201109301009.05583.mfick@codeaurora.org","threadId":"27589","inReplyTo":"4E8587E8.9070606@lsrfire.ath.cx","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T16:09:05Z","receivedAt":"2011-09-30T16:09:05Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Friday, September 30, 2011 03:12:08 am René Scharfe \nwrote:\n> OK, and what happens if you apply the following patch on\n> top of my first one?  It avoids going through all the\n> refs a second time during cleanup, at the cost of going\n> through the list of all known objects.  I wonder if\n> that's any faster in your case.\n\n\nThis patch helps a bit more.  It seems to shave about \nanother .5s off in packed and non packed case w or w/o \nbinary search.\n\n-Martin\n\n\n\n> diff --git a/builtin/checkout.c b/builtin/checkout.c\n> index 84e0cdc..a4b1003 100644\n> --- a/builtin/checkout.c\n> +++ b/builtin/checkout.c\n> @@ -596,15 +596,14 @@ static int\n> add_pending_uninteresting_ref(const char *refname,\n> return 0;\n>  }\n> \n> -static int clear_commit_marks_from_one_ref(const char\n> *refname, -\t\t\t\t      const unsigned \nchar *sha1,\n> -\t\t\t\t      int flags,\n> -\t\t\t\t      void *cb_data)\n> +static void clear_commit_marks_for_all(unsigned int\n> mark) {\n> -\tstruct commit *commit =\n> lookup_commit_reference_gently(sha1, 1); -\tif (commit)\n> -\t\tclear_commit_marks(commit, -1);\n> -\treturn 0;\n> +\tunsigned int i, max = get_max_object_index();\n> +\tfor (i = 0; i < max; i++) {\n> +\t\tstruct object *object = \nget_indexed_object(i);\n> +\t\tif (object && object->type == OBJ_COMMIT)\n> +\t\t\tobject->flags &= ~mark;\n> +\t}\n>  }\n> \n>  static void describe_one_orphan(struct strbuf *sb,\n> struct commit *commit) @@ -690,8 +689,7 @@ static void\n> orphaned_commit_warning(struct commit *commit) else\n>  \t\tdescribe_detached_head(_(\"Previous HEAD \nposition\n> was\"), commit);\n> \n> -\tclear_commit_marks(commit, -1);\n> -\tfor_each_ref(clear_commit_marks_from_one_ref, NULL);\n> +\tclear_commit_marks_for_all(ALL_REV_FLAGS);\n>  }\n> \n>  static int switch_branches(struct checkout_opts *opts,\n> struct branch_info *new) --\n> To unsubscribe from this list: send the line \"unsubscribe\n> git\" in the body of a message to\n> majordomo@vger.kernel.org More majordomo info at \n> http://vger.kernel.org/majordomo-info.html\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176578","messageId":"7vk48qouht.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"4E85E07C.5070402@alum.mit.edu","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-30T16:38:54Z","receivedAt":"2011-09-30T16:38:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu> writes:\n\n> On 09/30/2011 01:48 AM, Junio C Hamano wrote:\n>> This version looks sane, although I have a suspicion that it may have\n>> some interaction with what Michael may be working on.\n>\n> Indeed, I have almost equivalent changes in the giant patch series that\n> I am working on [1].\n\nGood; that was the primary thing I wanted to know.  I want to take\nJulian's patch early but if the approach and data structures were\ndrastically different from what you are cooking, that would force\nunnecessary reroll on your part, which I wanted to avoid.\n\nThanks.\n"},{"id":"176579","messageId":"201109301041.13848.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109262056.04279.chriscool@tuxfamily.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T16:41:13Z","receivedAt":"2011-09-30T16:41:13Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"> On Monday, September 26, 2011 12:56:04 pm Christian Couder \nwrote:\n> After \"git pack-refs --all\" I get:\n\nOK.   So many great improvements in ref scalability, thanks \neveryone!\n\nIt is getting so good, that I had to take a step back and \nre-evaluate what we consider good/bad.  On doing so, I can't \nhelp but think that fetches still need some improvement.\n\nFetches had the worst regression of all > 8days, so the \nmassive fix to bring it down to 7.5mins was awesome.  \n7-8mins sounded pretty good 2 weeks ago, especially when a \ncheckout took 5+ mins!  but now that almost every other \noperation has been sped up, that is starting to feel a bit \non the slow side still.  My spidey sense tells me something \nis still not quite right in the fetch path.\n\nHere is some more data to backup my spidey sense: after all \nthe improvements, a noop fetch of all the changes (noop \nmeaning they are all already uptodate) takes around \n3mins with a non gced (non packed refs) case.  That same \nnoop only takes ~12s in the gced (packed ref case)!\n\nI dug into this a bit further.  I took a non gced and non \npacked refs repo and this time instead of gcing it to get \npackedrefs, I only ran the above git pack-refs --all so that\nobjects did not get gced.  With this, the noop fetch was \nalso only around 12s.  This confirmed that the non gced \nobjects are not interfering with the noop fetch, the problem \nreally is just the unpacked refs.  Just to confirm that the \nFS is not horribly slow, I did a \"find .git/refs\" and it \nonly takes about .4s for about 80Kresults!\n\nSo, while I understand that a full fetch will actually have \nto transfer quite a bit of data, the noop fetch seems like \nit is still suffering in the non gced (non packed ref case).  \nIf that time were improved, I suspect that the full fetch \nwill improve at least by an equivalent amount, if not more.\n\nAny thoughts?\n\n-Martin\n\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176580","messageId":"7vfwjeotv1.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"4E8587E8.9070606@lsrfire.ath.cx","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-09-30T16:52:34Z","receivedAt":"2011-09-30T16:52:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <rene.scharfe@lsrfire.ath.cx> writes:\n\n> Hi Martin,\n>\n> Am 29.09.2011 22:11, schrieb Martin Fick:\n>> Your patch works well for me.  It achieves about the same \n>> gains as Julian's patch. Thanks!\n>\n> OK, and what happens if you apply the following patch on top of my first\n> one?  It avoids going through all the refs a second time during cleanup,\n> at the cost of going through the list of all known objects.  I wonder if\n> that's any faster in your case.\n> ...\n>  static void describe_one_orphan(struct strbuf *sb, struct commit *commit)\n> @@ -690,8 +689,7 @@ static void orphaned_commit_warning(struct commit *commit)\n>  \telse\n>  \t\tdescribe_detached_head(_(\"Previous HEAD position was\"), commit);\n>  \n> -\tclear_commit_marks(commit, -1);\n> -\tfor_each_ref(clear_commit_marks_from_one_ref, NULL);\n> +\tclear_commit_marks_for_all(ALL_REV_FLAGS);\n>  }\n\nThe function already clears all the flag bits from commits near the tip of\nall the refs (i.e. whatever commit it traverses until it gets to the fork\npoint), so it cannot be reused in other contexts where the caller\n\n - first marks commit objects with some flag bits for its own purpose,\n   unrelated to the \"orphaned\"-ness check;\n - calls this function to issue a warning; and then\n - use the flag it earlier set to do something useful.\n\nwhich requires \"cleaning after yourself, by clearing only the bits you\nused without disturbing other bits that you do not use\" pattern.\n\nIt might be a better solution to not bother to clear the marks at all;\nwould it break anything in this codepath?\n"},{"id":"176582","messageId":"20110930175658.1220.11552.julian@quantumfyre.co.uk","threadId":"27589","inReplyTo":"7vk48qouht.fsf@alter.siamese.dyndns.org","subject":"[PATCH] refs: Remove duplicates after sorting with qsort","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2011-09-30T17:56:57Z","receivedAt":"2011-09-30T17:56:57Z","isPatch":true,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"The previous custom merge sort would drop duplicate entries as part of\nthe sort.  It would also die if the duplicate entries had different\nsha1 values.  The standard library qsort doesn't do this, so we have\nto do it manually afterwards.\n\nSigned-off-by: Julian Phillips <julian@quantumfyre.co.uk>\n---\n\nOn Fri, 30 Sep 2011 09:38:54 -0700, Junio C Hamano wrote:\n> Michael Haggerty <mhagger@alum.mit.edu> writes:\n>\n>> On 09/30/2011 01:48 AM, Junio C Hamano wrote:\n>>> This version looks sane, although I have a suspicion that it may \n>>> have\n>>> some interaction with what Michael may be working on.\n>>\n>> Indeed, I have almost equivalent changes in the giant patch series \n>> that\n>> I am working on [1].\n>\n> Good; that was the primary thing I wanted to know.  I want to take\n> Julian's patch early but if the approach and data structures were\n> drastically different from what you are cooking, that would force\n> unnecessary reroll on your part, which I wanted to avoid.\n>\n> Thanks.\n\nI had a quick look at Michael's code, and it reminded me that I had missed one\nthing out.  If we want to keep the duplicate detection & removal from the\noriginal merge sort then this patch is needed on top of v3 of the binary search.\n\nThough I never could figure out how duplicate refs were supposed to appear ... I\ntested by editing packed-refs, but I assume that isn't \"supported\".\n\n refs.c |   22 ++++++++++++++++++++++\n 1 files changed, 22 insertions(+), 0 deletions(-)\n\ndiff --git a/refs.c b/refs.c\nindex 4c01d79..cf080ee 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -77,7 +77,29 @@ static int ref_entry_cmp(const void *a, const void *b)\n \n static void sort_ref_array(struct ref_array *array)\n {\n+\tint i = 0, j = 1;\n+\n+\t/* Nothing to sort unless there are at least two entries */\n+\tif (array->nr < 2)\n+\t\treturn;\n+\n \tqsort(array->refs, array->nr, sizeof(*array->refs), ref_entry_cmp);\n+\n+\t/* Remove any duplicates from the ref_array */\n+\tfor (; j < array->nr; j++) {\n+\t\tstruct ref_entry *a = array->refs[i];\n+\t\tstruct ref_entry *b = array->refs[j];\n+\t\tif (!strcmp(a->name, b->name)) {\n+\t\t\tif (hashcmp(a->sha1, b->sha1))\n+\t\t\t\tdie(\"Duplicated ref, and SHA1s don't match: %s\",\n+\t\t\t\t    a->name);\n+\t\t\twarning(\"Duplicated ref: %s\", a->name);\n+\t\t\tcontinue;\n+\t\t}\n+\t\ti++;\n+\t\tarray->refs[i] = array->refs[j];\n+\t}\n+\tarray->nr = i + 1;\n }\n \n static struct ref_entry *search_ref_array(struct ref_array *array, const char *name)\n-- \n1.7.6.1\n"},{"id":"176584","messageId":"4E8607B6.2040800@lsrfire.ath.cx","threadId":"27589","inReplyTo":"7vfwjeotv1.fsf@alter.siamese.dyndns.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-09-30T18:17:26Z","receivedAt":"2011-09-30T18:17:26Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 30.09.2011 18:52, schrieb Junio C Hamano:\n> René Scharfe <rene.scharfe@lsrfire.ath.cx> writes:\n> \n>> Hi Martin,\n>>\n>> Am 29.09.2011 22:11, schrieb Martin Fick:\n>>> Your patch works well for me.  It achieves about the same \n>>> gains as Julian's patch. Thanks!\n>>\n>> OK, and what happens if you apply the following patch on top of my first\n>> one?  It avoids going through all the refs a second time during cleanup,\n>> at the cost of going through the list of all known objects.  I wonder if\n>> that's any faster in your case.\n>> ...\n>>  static void describe_one_orphan(struct strbuf *sb, struct commit *commit)\n>> @@ -690,8 +689,7 @@ static void orphaned_commit_warning(struct commit *commit)\n>>  \telse\n>>  \t\tdescribe_detached_head(_(\"Previous HEAD position was\"), commit);\n>>  \n>> -\tclear_commit_marks(commit, -1);\n>> -\tfor_each_ref(clear_commit_marks_from_one_ref, NULL);\n>> +\tclear_commit_marks_for_all(ALL_REV_FLAGS);\n>>  }\n> \n> The function already clears all the flag bits from commits near the tip of\n> all the refs (i.e. whatever commit it traverses until it gets to the fork\n> point), so it cannot be reused in other contexts where the caller\n> \n>  - first marks commit objects with some flag bits for its own purpose,\n>    unrelated to the \"orphaned\"-ness check;\n>  - calls this function to issue a warning; and then\n>  - use the flag it earlier set to do something useful.\n> \n> which requires \"cleaning after yourself, by clearing only the bits you\n> used without disturbing other bits that you do not use\" pattern.\n\nYes, clear_commit_marks_for_all is a bit brutal.  Callers could clear\nspecfic bits (e.g. SEEN|UNINTERESTING) instead of ALL_REV_FLAGS, though.\n\n> It might be a better solution to not bother to clear the marks at all;\n> would it break anything in this codepath?\n\nUnfortunately, yes; the cleanup part was added by 5c08dc48 later, when\nit become apparent that it's really needed.\n\nHowever, since the patch only buys us a 5% speedup I'm not sure it's\nworth it in its current form.\n\nRené\n"},{"id":"176591","messageId":"201109301326.03711.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109301041.13848.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T19:26:03Z","receivedAt":"2011-09-30T19:26:03Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Friday, September 30, 2011 10:41:13 am Martin Fick wrote:\n> I dug into this a bit further.  I took a non gced and non\n> packed refs repo and this time instead of gcing it to get\n> packedrefs, I only ran the above git pack-refs --all so\n> that objects did not get gced.  With this, the noop\n> fetch was also only around 12s.  This confirmed that the\n> non gced objects are not interfering with the noop\n> fetch, the problem really is just the unpacked refs. \n> Just to confirm that the FS is not horribly slow, I did\n> a \"find .git/refs\" and it only takes about .4s for about\n> 80Kresults!\n\nIs there a way I can for refs to always be packed?  I didn't \nsee a config option for this.  I would like to try a fetch \nthis way even if I have to make a small code tweak.  \n\nI tried simulating on the fly ref packing every now and then \nby running the pack from another repo during the fetch, it \nactually slowed things down (by more than the time it took \nto do the packs).\n\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176594","messageId":"201109301502.30617.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109301041.13848.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T21:02:30Z","receivedAt":"2011-09-30T21:02:30Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Friday, September 30, 2011 10:41:13 am Martin Fick wrote:\n> massive fix to bring it down to 7.5mins was awesome.\n> 7-8mins sounded pretty good 2 weeks ago, especially when\n> a checkout took 5+ mins!  but now that almost every\n> other operation has been sped up, that is starting to\n> feel a bit on the slow side still.  My spidey sense\n> tells me something is still not quite right in the fetch\n> path.\n\nI guess I overlooked that there were 2 sides to this \nequation.  Even though I have been doing my fetches locally, \nI was using the file:// protocol and it appears that the \nremote was running git 1.7.6 which was in my path the whole \ntime.  So eliminating that from my path and pointing to the \nthe \"best\" binary with all the fixes for both remote and \nlocal, the full fetch does indeed speed up quite a bit, it \ngoes from about 7.5mins down to ~5m!  Previously the remote \nseemed to primarily spend the extra time after:\n\n remote: Counting objects: 316961\n\nyet before:\n\n remote: Compressing objects\n\n\n> Here is some more data to backup my spidey sense: after\n> all the improvements, a noop fetch of all the changes\n> (noop meaning they are all already uptodate) takes\n> around 3mins with a non gced (non packed refs) case. \n> That same noop only takes ~12s in the gced (packed ref\n> case)!\n\nI believe (it is hard to be go back and be sure) that this \nmeans that the timings above which gave me 3mins were \nbecause the remote was using git 1.7.6.  Now, with the good \nbinary, in both repos (packed and unpacked), I get great \nwarm cache times of about 11-13s for a noop fetch.  It is \ninteresting to note that cold cache times are 20s for packed \nrefs and 1m30s for unpacked refs.  I guess that makes some \nsense.  \n\nBut, this does leave me thinking that packed refs should \nbecome the default and that there should be a config option \nto disable it?  This still might help a fetch?\n\nSince a full sync is now done to about 5mins, I broke down \nthe output a bit.  It appears that the longest part (2:45m) \nis now the time spent scrolling though each change still.  \nEach one of these takes about 2ms:\n * [new branch]      refs/changes/99/71199/1 -> \nrefs/changes/99/71199/1\n\nSeems fast, but at about 80K... So, are there any obvious N \nloops over the refs happening inside each of of the [new \nbranch] iterations?\n\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176600","messageId":"201109301606.31748.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109301502.30617.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-09-30T22:06:31Z","receivedAt":"2011-09-30T22:06:31Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Friday, September 30, 2011 03:02:30 pm Martin Fick wrote:\n> On Friday, September 30, 2011 10:41:13 am Martin Fick \nwrote:\n> Since a full sync is now done to about 5mins, I broke\n> down the output a bit.  It appears that the longest part\n> (2:45m) is now the time spent scrolling though each\n> change still. Each one of these takes about 2ms:\n>  * [new branch]      refs/changes/99/71199/1 ->\n> refs/changes/99/71199/1\n> \n> Seems fast, but at about 80K... So, are there any obvious\n> N loops over the refs happening inside each of of the\n> [new branch] iterations?\n\nOK, I narrowed it down I believe.  If I comment out the \ninvalidate_cached_refs() line in write_ref_sha1(), it speeds \nthrough this section.  \n\nI guess this makes sense, we invalidate the cache and have \nto rebuild it after every new ref is added?  Perhaps a \nsimple fix would be to move the invalidation right after all \nthe refs are updated?  Maybe write_ref_sha1 could take in a \nflag to tell it to not invalidate the cache so that during \niterative updates it could be disabled and then run manually \nafter the update?\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176638","messageId":"4E8731AF.2040305@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8607B6.2040800@lsrfire.ath.cx","subject":"Re: Git is not scalable with too many refs/*","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T15:28:47Z","receivedAt":"2011-10-01T15:28:47Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 30.09.2011 20:17, schrieb René Scharfe:\n> Am 30.09.2011 18:52, schrieb Junio C Hamano:\n>> It might be a better solution to not bother to clear the marks at\n>> all; would it break anything in this codepath?\n> \n> Unfortunately, yes; the cleanup part was added by 5c08dc48 later,\n> when it become apparent that it's really needed.\n> \n> However, since the patch only buys us a 5% speedup I'm not sure it's \n> worth it in its current form.\n\nI found something better: A trick used by bisect and bundle.  They copy\nthe list of pending objects from rev_info before calling\nprepare_revision_walk and then go through it to clean up the commit\nmarks without going through the refs again.  And I think we can even\nimprove it a little.\n\nThe following patches tighten some orphan/detached head tests a little,\nthen comes a resend of my first patch on this topic, only split up into\ntwo, then four patches that introduce the trick mentioned above (which\ncould be squashed together perhaps) and the last one is a bonus\nrefactoring patch.\n\n bisect.c                   |   20 +++++++-------\n builtin/checkout.c         |   58 +++++++++++++------------------------------\n bundle.c                   |   11 +++-----\n commit.c                   |   14 ++++++++++\n commit.h                   |    1 +\n revision.c                 |   14 +++++++---\n revision.h                 |    2 +\n t/t2020-checkout-detach.sh |    7 ++++-\n 8 files changed, 64 insertions(+), 63 deletions(-)\n\nRené\n"},{"id":"176639","messageId":"4E8733FA.6070201@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 1/8] checkout: check for \"Previous HEAD\" notice in t2020","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T15:38:34Z","receivedAt":"2011-10-01T15:38:34Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"If we leave a detached head, exactly one of two things happens: either\ncheckout warns about it being an orphan or describes it as a courtesy.\nTest t2020 already checked that the warning is shown as needed.  This\npatch also checks for the description.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\n t/t2020-checkout-detach.sh |    7 +++++--\n 1 files changed, 5 insertions(+), 2 deletions(-)\n\ndiff --git a/t/t2020-checkout-detach.sh b/t/t2020-checkout-detach.sh\nindex 2366f0f..068fba4 100755\n--- a/t/t2020-checkout-detach.sh\n+++ b/t/t2020-checkout-detach.sh\n@@ -12,11 +12,14 @@ check_not_detached () {\n }\n \n ORPHAN_WARNING='you are leaving .* commit.*behind'\n+PREV_HEAD_DESC='Previous HEAD position was'\n check_orphan_warning() {\n-\ttest_i18ngrep \"$ORPHAN_WARNING\" \"$1\"\n+\ttest_i18ngrep \"$ORPHAN_WARNING\" \"$1\" &&\n+\ttest_i18ngrep ! \"$PREV_HEAD_DESC\" \"$1\"\n }\n check_no_orphan_warning() {\n-\ttest_i18ngrep ! \"$ORPHAN_WARNING\" \"$1\"\n+\ttest_i18ngrep ! \"$ORPHAN_WARNING\" \"$1\" &&\n+\ttest_i18ngrep \"$PREV_HEAD_DESC\" \"$1\"\n }\n \n reset () {\n-- \n1.7.7\n"},{"id":"176640","messageId":"4E873538.8080401@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 2/8] revision: factor out add_pending_sha1","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T15:43:52Z","receivedAt":"2011-10-01T15:43:52Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"This function is a combination of the static get_reference and\nadd_pending_object.  It can be used to easily queue objects by hash.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\nThe next patch is going to use it in checkout.\n\n revision.c |   11 ++++++++---\n revision.h |    1 +\n 2 files changed, 9 insertions(+), 3 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex c46cfaa..2e8aa33 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -185,6 +185,13 @@ static struct object *get_reference(struct rev_info *revs, const char *name, con\n \treturn object;\n }\n \n+void add_pending_sha1(struct rev_info *revs, const char *name,\n+\t\t      const unsigned char *sha1, unsigned int flags)\n+{\n+\tstruct object *object = get_reference(revs, name, sha1, flags);\n+\tadd_pending_object(revs, object, name);\n+}\n+\n static struct commit *handle_commit(struct rev_info *revs, struct object *object, const char *name)\n {\n \tunsigned long flags = object->flags;\n@@ -832,9 +839,7 @@ struct all_refs_cb {\n static int handle_one_ref(const char *path, const unsigned char *sha1, int flag, void *cb_data)\n {\n \tstruct all_refs_cb *cb = cb_data;\n-\tstruct object *object = get_reference(cb->all_revs, path, sha1,\n-\t\t\t\t\t      cb->all_flags);\n-\tadd_pending_object(cb->all_revs, object, path);\n+\tadd_pending_sha1(cb->all_revs, path, sha1, cb->all_flags);\n \treturn 0;\n }\n \ndiff --git a/revision.h b/revision.h\nindex 3d64ada..4541265 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -191,6 +191,7 @@ extern void add_object(struct object *obj,\n \t\t       const char *name);\n \n extern void add_pending_object(struct rev_info *revs, struct object *obj, const char *name);\n+extern void add_pending_sha1(struct rev_info *revs, const char *name, const unsigned char *sha1, unsigned int flags);\n \n extern void add_head_to_pending(struct rev_info *);\n \n-- \n1.7.7\n"},{"id":"176641","messageId":"4E87370B.4060908@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 3/8] checkout: use add_pending_{object,sha1} in orphan check","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T15:51:39Z","receivedAt":"2011-10-01T15:51:39Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Instead of building a list of textual arguments for setup_revisions, use\nadd_pending_object and add_pending_sha1 to queue the objects directly.\nThis is both faster and simpler.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\n builtin/checkout.c |   39 ++++++++++++---------------------------\n 1 files changed, 12 insertions(+), 27 deletions(-)\n\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 5e356a6..84e0cdc 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -588,24 +588,11 @@ static void update_refs_for_switch(struct checkout_opts *opts,\n \t\treport_tracking(new);\n }\n \n-struct rev_list_args {\n-\tint argc;\n-\tint alloc;\n-\tconst char **argv;\n-};\n-\n-static void add_one_rev_list_arg(struct rev_list_args *args, const char *s)\n-{\n-\tALLOC_GROW(args->argv, args->argc + 1, args->alloc);\n-\targs->argv[args->argc++] = s;\n-}\n-\n-static int add_one_ref_to_rev_list_arg(const char *refname,\n-\t\t\t\t       const unsigned char *sha1,\n-\t\t\t\t       int flags,\n-\t\t\t\t       void *cb_data)\n+static int add_pending_uninteresting_ref(const char *refname,\n+\t\t\t\t\t const unsigned char *sha1,\n+\t\t\t\t\t int flags, void *cb_data)\n {\n-\tadd_one_rev_list_arg(cb_data, refname);\n+\tadd_pending_sha1(cb_data, refname, sha1, flags | UNINTERESTING);\n \treturn 0;\n }\n \n@@ -685,19 +672,17 @@ static void suggest_reattach(struct commit *commit, struct rev_info *revs)\n  */\n static void orphaned_commit_warning(struct commit *commit)\n {\n-\tstruct rev_list_args args = { 0, 0, NULL };\n \tstruct rev_info revs;\n-\n-\tadd_one_rev_list_arg(&args, \"(internal)\");\n-\tadd_one_rev_list_arg(&args, sha1_to_hex(commit->object.sha1));\n-\tadd_one_rev_list_arg(&args, \"--not\");\n-\tfor_each_ref(add_one_ref_to_rev_list_arg, &args);\n-\tadd_one_rev_list_arg(&args, \"--\");\n-\tadd_one_rev_list_arg(&args, NULL);\n+\tstruct object *object = &commit->object;\n \n \tinit_revisions(&revs, NULL);\n-\tif (setup_revisions(args.argc - 1, args.argv, &revs, NULL) != 1)\n-\t\tdie(_(\"internal error: only -- alone should have been left\"));\n+\tsetup_revisions(0, NULL, &revs, NULL);\n+\n+\tobject->flags &= ~UNINTERESTING;\n+\tadd_pending_object(&revs, object, sha1_to_hex(object->sha1));\n+\n+\tfor_each_ref(add_pending_uninteresting_ref, &revs);\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(_(\"internal error in revision walk\"));\n \tif (!(commit->object.flags & UNINTERESTING))\n-- \n1.7.7\n"},{"id":"176642","messageId":"4E873818.6080006@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 4/8] revision: add leak_pending flag","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T15:56:08Z","receivedAt":"2011-10-01T15:56:08Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"The new flag leak_pending in struct rev_info can be used to prevent\nprepare_revision_walk from freeing the list of pending objects.  It\nwill still forget about them, so it really is leaked.  This behaviour\nmay look weird at first, but it can be useful if the pointer to the\nlist is saved before calling prepare_revision_walk.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\nThe next three patches are going to use this flag.\n\n revision.c |    3 ++-\n revision.h |    1 +\n 2 files changed, 3 insertions(+), 1 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 2e8aa33..6d329b4 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -1974,7 +1974,8 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\t}\n \t\te++;\n \t}\n-\tfree(list);\n+\tif (!revs->leak_pending)\n+\t\tfree(list);\n \n \tif (revs->no_walk)\n \t\treturn 0;\ndiff --git a/revision.h b/revision.h\nindex 4541265..366a9b4 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -97,6 +97,7 @@ struct rev_info {\n \t\t\tdate_mode_explicit:1,\n \t\t\tpreserve_subject:1;\n \tunsigned int\tdisable_stdin:1;\n+\tunsigned int\tleak_pending:1;\n \n \tenum date_mode date_mode;\n \n-- \n1.7.7\n"},{"id":"176643","messageId":"4E873948.6030404@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 5/8] bisect: use leak_pending flag","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T16:01:12Z","receivedAt":"2011-10-01T16:01:12Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Instead of creating a copy of the list of pending objects, copy the\nstruct object_array that points to it, turn on leak_pending, and thus\ncause prepare_revision_walk to leave it to us.  And free it once\nwe're done.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\n bisect.c |   13 ++++++++-----\n 1 files changed, 8 insertions(+), 5 deletions(-)\n\ndiff --git a/bisect.c b/bisect.c\nindex c7b7d79..a05504f 100644\n--- a/bisect.c\n+++ b/bisect.c\n@@ -831,12 +831,14 @@ static int check_ancestors(const char *prefix)\n \tbisect_rev_setup(&revs, prefix, \"^%s\", \"%s\", 0);\n \n \t/* Save pending objects, so they can be cleaned up later. */\n-\tmemset(&pending_copy, 0, sizeof(pending_copy));\n-\tfor (i = 0; i < revs.pending.nr; i++)\n-\t\tadd_object_array(revs.pending.objects[i].item,\n-\t\t\t\t revs.pending.objects[i].name,\n-\t\t\t\t &pending_copy);\n+\tpending_copy = revs.pending;\n+\trevs.leak_pending = 1;\n \n+\t/*\n+\t * bisect_common calls prepare_revision_walk right away, which\n+\t * (together with .leak_pending = 1) makes us the sole owner of\n+\t * the list of pending objects.\n+\t */\n \tbisect_common(&revs);\n \tres = (revs.commits != NULL);\n \n@@ -845,6 +847,7 @@ static int check_ancestors(const char *prefix)\n \t\tstruct object *o = pending_copy.objects[i].item;\n \t\tclear_commit_marks((struct commit *)o, ALL_REV_FLAGS);\n \t}\n+\tfree(pending_copy.objects);\n \n \treturn res;\n }\n-- \n1.7.7\n"},{"id":"176644","messageId":"4E87399C.5010700@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 6/8] bundle: use leak_pending flag","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T16:02:36Z","receivedAt":"2011-10-01T16:02:36Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Instead of creating a copy of the list of pending objects, copy the\nstruct object_array that points to it, turn on leak_pending, and thus\ncause prepare_revision_walk to leave it to us.  And free it once\nwe're done.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\n bundle.c |    8 +++-----\n 1 files changed, 3 insertions(+), 5 deletions(-)\n\ndiff --git a/bundle.c b/bundle.c\nindex f48fd7d..26cc9ab 100644\n--- a/bundle.c\n+++ b/bundle.c\n@@ -122,11 +122,8 @@ int verify_bundle(struct bundle_header *header, int verbose)\n \treq_nr = revs.pending.nr;\n \tsetup_revisions(2, argv, &revs, NULL);\n \n-\tmemset(&refs, 0, sizeof(struct object_array));\n-\tfor (i = 0; i < revs.pending.nr; i++) {\n-\t\tstruct object_array_entry *e = revs.pending.objects + i;\n-\t\tadd_object_array(e->item, e->name, &refs);\n-\t}\n+\trefs = revs.pending;\n+\trevs.leak_pending = 1;\n \n \tif (prepare_revision_walk(&revs))\n \t\tdie(\"revision walk setup failed\");\n@@ -146,6 +143,7 @@ int verify_bundle(struct bundle_header *header, int verbose)\n \n \tfor (i = 0; i < refs.nr; i++)\n \t\tclear_commit_marks((struct commit *)refs.objects[i].item, -1);\n+\tfree(refs.objects);\n \n \tif (verbose) {\n \t\tstruct ref_list *r;\n-- \n1.7.7\n"},{"id":"176645","messageId":"4E873B40.7030409@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 7/8] checkout: use leak_pending flag","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T16:09:36Z","receivedAt":"2011-10-01T16:09:36Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Instead of going through all the references again when we clear the\ncommit marks, do it like bisect and bundle and gain ownership of the\nlist of pending objects which we constructed from those references.\n\nWe simply copy the struct object_array that points to the list, set\nthe flag leak_pending and then prepare_revision_walk won't destroy\nit and it's ours.  We use it to clear the marks and  free it at the\nend.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\n builtin/checkout.c |   25 ++++++++++++-------------\n 1 files changed, 12 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 84e0cdc..cfd7e59 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -596,17 +596,6 @@ static int add_pending_uninteresting_ref(const char *refname,\n \treturn 0;\n }\n \n-static int clear_commit_marks_from_one_ref(const char *refname,\n-\t\t\t\t      const unsigned char *sha1,\n-\t\t\t\t      int flags,\n-\t\t\t\t      void *cb_data)\n-{\n-\tstruct commit *commit = lookup_commit_reference_gently(sha1, 1);\n-\tif (commit)\n-\t\tclear_commit_marks(commit, -1);\n-\treturn 0;\n-}\n-\n static void describe_one_orphan(struct strbuf *sb, struct commit *commit)\n {\n \tparse_commit(commit);\n@@ -674,6 +663,8 @@ static void orphaned_commit_warning(struct commit *commit)\n {\n \tstruct rev_info revs;\n \tstruct object *object = &commit->object;\n+\tstruct object_array refs;\n+\tunsigned int i;\n \n \tinit_revisions(&revs, NULL);\n \tsetup_revisions(0, NULL, &revs, NULL);\n@@ -683,6 +674,9 @@ static void orphaned_commit_warning(struct commit *commit)\n \n \tfor_each_ref(add_pending_uninteresting_ref, &revs);\n \n+\trefs = revs.pending;\n+\trevs.leak_pending = 1;\n+\n \tif (prepare_revision_walk(&revs))\n \t\tdie(_(\"internal error in revision walk\"));\n \tif (!(commit->object.flags & UNINTERESTING))\n@@ -690,8 +684,13 @@ static void orphaned_commit_warning(struct commit *commit)\n \telse\n \t\tdescribe_detached_head(_(\"Previous HEAD position was\"), commit);\n \n-\tclear_commit_marks(commit, -1);\n-\tfor_each_ref(clear_commit_marks_from_one_ref, NULL);\n+\tfor (i = 0; i < refs.nr; i++) {\n+\t\tstruct object *o = refs.objects[i].item;\n+\t\tstruct commit *c = lookup_commit_reference_gently(o->sha1, 1);\n+\t\tif (c)\n+\t\t\tclear_commit_marks(c, ALL_REV_FLAGS);\n+\t}\n+\tfree(refs.objects);\n }\n \n static int switch_branches(struct checkout_opts *opts, struct branch_info *new)\n-- \n1.7.7\n"},{"id":"176646","messageId":"4E873CC8.6060100@lsrfire.ath.cx","threadId":"27589","inReplyTo":"4E8731AF.2040305@lsrfire.ath.cx","subject":"[PATCH 8/8] commit: factor out clear_commit_marks_for_object_array","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2011-10-01T16:16:08Z","receivedAt":"2011-10-01T16:16:08Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Factor out the code to clear the commit marks for a whole struct\nobject_array from builtin/checkout.c into its own exported function\nclear_commit_marks_for_object_array and use it in bisect and bundle\nas well.  It handles tags and commits and ignores objects of any\nother type.\n\nSigned-off-by: Rene Scharfe <rene.scharfe@lsrfire.ath.cx>\n---\n bisect.c           |    7 ++-----\n builtin/checkout.c |    8 +-------\n bundle.c           |    3 +--\n commit.c           |   14 ++++++++++++++\n commit.h           |    1 +\n 5 files changed, 19 insertions(+), 14 deletions(-)\n\ndiff --git a/bisect.c b/bisect.c\nindex a05504f..b4547b9 100644\n--- a/bisect.c\n+++ b/bisect.c\n@@ -826,7 +826,7 @@ static int check_ancestors(const char *prefix)\n {\n \tstruct rev_info revs;\n \tstruct object_array pending_copy;\n-\tint i, res;\n+\tint res;\n \n \tbisect_rev_setup(&revs, prefix, \"^%s\", \"%s\", 0);\n \n@@ -843,10 +843,7 @@ static int check_ancestors(const char *prefix)\n \tres = (revs.commits != NULL);\n \n \t/* Clean up objects used, as they will be reused. */\n-\tfor (i = 0; i < pending_copy.nr; i++) {\n-\t\tstruct object *o = pending_copy.objects[i].item;\n-\t\tclear_commit_marks((struct commit *)o, ALL_REV_FLAGS);\n-\t}\n+\tclear_commit_marks_for_object_array(&pending_copy, ALL_REV_FLAGS);\n \tfree(pending_copy.objects);\n \n \treturn res;\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex cfd7e59..683819b 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -664,7 +664,6 @@ static void orphaned_commit_warning(struct commit *commit)\n \tstruct rev_info revs;\n \tstruct object *object = &commit->object;\n \tstruct object_array refs;\n-\tunsigned int i;\n \n \tinit_revisions(&revs, NULL);\n \tsetup_revisions(0, NULL, &revs, NULL);\n@@ -684,12 +683,7 @@ static void orphaned_commit_warning(struct commit *commit)\n \telse\n \t\tdescribe_detached_head(_(\"Previous HEAD position was\"), commit);\n \n-\tfor (i = 0; i < refs.nr; i++) {\n-\t\tstruct object *o = refs.objects[i].item;\n-\t\tstruct commit *c = lookup_commit_reference_gently(o->sha1, 1);\n-\t\tif (c)\n-\t\t\tclear_commit_marks(c, ALL_REV_FLAGS);\n-\t}\n+\tclear_commit_marks_for_object_array(&refs, ALL_REV_FLAGS);\n \tfree(refs.objects);\n }\n \ndiff --git a/bundle.c b/bundle.c\nindex 26cc9ab..a8ea918 100644\n--- a/bundle.c\n+++ b/bundle.c\n@@ -141,8 +141,7 @@ int verify_bundle(struct bundle_header *header, int verbose)\n \t\t\t\trefs.objects[i].name);\n \t\t}\n \n-\tfor (i = 0; i < refs.nr; i++)\n-\t\tclear_commit_marks((struct commit *)refs.objects[i].item, -1);\n+\tclear_commit_marks_for_object_array(&refs, ALL_REV_FLAGS);\n \tfree(refs.objects);\n \n \tif (verbose) {\ndiff --git a/commit.c b/commit.c\nindex 97b4327..50af007 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -430,6 +430,20 @@ void clear_commit_marks(struct commit *commit, unsigned int mark)\n \t}\n }\n \n+void clear_commit_marks_for_object_array(struct object_array *a, unsigned mark)\n+{\n+\tstruct object *object;\n+\tstruct commit *commit;\n+\tunsigned int i;\n+\n+\tfor (i = 0; i < a->nr; i++) {\n+\t\tobject = a->objects[i].item;\n+\t\tcommit = lookup_commit_reference_gently(object->sha1, 1);\n+\t\tif (commit)\n+\t\t\tclear_commit_marks(commit, mark);\n+\t}\n+}\n+\n struct commit *pop_commit(struct commit_list **stack)\n {\n \tstruct commit_list *top = *stack;\ndiff --git a/commit.h b/commit.h\nindex 12d100b..0a4c730 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -126,6 +126,7 @@ struct commit *pop_most_recent_commit(struct commit_list **list,\n struct commit *pop_commit(struct commit_list **stack);\n \n void clear_commit_marks(struct commit *commit, unsigned int mark);\n+void clear_commit_marks_for_object_array(struct object_array *a, unsigned mark);\n \n /*\n  * Performs an in-place topological sort of list supplied.\n-- \n1.7.7\n"},{"id":"176653","messageId":"CAGdFq_gYBg21mh7xiPdqqKLG8ZDM_Eebj5+E0=U4duQG2BJUDw@mail.gmail.com","threadId":"27589","inReplyTo":"4E8733FA.6070201@lsrfire.ath.cx","subject":"Re: [PATCH 1/8] checkout: check for \"Previous HEAD\" notice in t2020","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-10-01T19:02:06Z","receivedAt":"2011-10-01T19:02:06Z","isPatch":true,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Sat, Oct 1, 2011 at 17:38, René Scharfe <rene.scharfe@lsrfire.ath.cx> wrote:\n> If we leave a detached head, exactly one of two things happens: either\n> checkout warns about it being an orphan or describes it as a courtesy.\n> Test t2020 already checked that the warning is shown as needed.  This\n> patch also checks for the description.\n\nA cover letter would have been nice for such a long series :).\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"176658","messageId":"7vwrcola0m.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"201109301606.31748.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-10-01T20:41:45Z","receivedAt":"2011-10-01T20:41:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Fick <mfick@codeaurora.org> writes:\n\n> I guess this makes sense, we invalidate the cache and have \n> to rebuild it after every new ref is added?  Perhaps a \n> simple fix would be to move the invalidation right after all \n> the refs are updated?  Maybe write_ref_sha1 could take in a \n> flag to tell it to not invalidate the cache so that during \n> iterative updates it could be disabled and then run manually \n> after the update?\n\nIt might make sense, on top of Julian's patch, to add a bit that says \"the\ncontents of this ref-array is current but the array is not sorted\", and\nwhenever somebody runs add_ref(), append it also to the ref-array (so that\nthe contents do not have to be re-read from the filesystem) but flip the\n\"unsorted\" bit on. Then update look-up and iteration to sort the array\nwhen \"unsorted\" bit is on without re-reading the contents from the\nfilesystem.\n"},{"id":"176674","messageId":"4E87EF8D.8020801@alum.mit.edu","threadId":"27589","inReplyTo":"20110927000010.79913.71464.julian@quantumfyre.co.uk","subject":"Re: [PATCH] Don't sort ref_list too early","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-10-02T04:58:53Z","receivedAt":"2011-10-02T04:58:53Z","isPatch":true,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 09/27/2011 02:00 AM, Julian Phillips wrote:\n> get_ref_dir is called recursively for subdirectories, which means that\n> we were calling sort_ref_list for each directory of refs instead of\n> once for all the refs.  This is a massive wast of processing, so now\n> just call sort_ref_list on the result of the top-level get_ref_dir, so\n> that the sort is only done once.\n\n+1\n\nI think this patch should also be considered for maint, since it is\nnoninvasive and fixes a bad performance regression.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176676","messageId":"4E87F383.1050403@alum.mit.edu","threadId":"27589","inReplyTo":"7vk48qouht.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-10-02T05:15:47Z","receivedAt":"2011-10-02T05:15:47Z","isPatch":true,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 09/30/2011 06:38 PM, Junio C Hamano wrote:\n> Michael Haggerty <mhagger@alum.mit.edu> writes:\n> \n>> On 09/30/2011 01:48 AM, Junio C Hamano wrote:\n>>> This version looks sane, although I have a suspicion that it may have\n>>> some interaction with what Michael may be working on.\n>>\n>> Indeed, I have almost equivalent changes in the giant patch series that\n>> I am working on [1].\n> \n> Good; that was the primary thing I wanted to know.  I want to take\n> Julian's patch early but if the approach and data structures were\n> drastically different from what you are cooking, that would force\n> unnecessary reroll on your part, which I wanted to avoid.\n\nUm, well, my patch series includes the same changes that Julian's wants\nto introduce, but following lots of other changes, cleanups,\ndocumentation improvements, etc.  Moreover, my patch series builds on\nmh/iterate-refs, with which Julian's patch conflicts.  In other words,\nit would be a real mess to reroll my series on top of Julian's patch.\n(That is of course not to imply that I hold a mutex on refs.c.)  Because\nit changes a data structure that is used throughout refs.c, changes a\nlot of lines of code.\n\nI think that the switch from linked list + linear sort to array plus\nbinary sort is a pretty obvious win in terms of code complexity and\n*potential* performance improvement, but empirically I haven't seen any\nclaims that it brings performance improvements beyond \"René's patch\".\n(Though, honestly, I've lost track of which \"René's patch\" is being\ndiscussed and I don't see anything relevant in Junio's tree.)\n\nIntuitively, given that populating the reference cache involves O(N)\nI/O, speeding up lookups can only help if there are very many ref\nlookups within a single git invocation.  I think we will get a better\nimprovement by avoiding the reading of unneeded loose refs by reading\nthem one subdirectory at a time instead of always reading them en masse.\n I wanted to reach that milestone before submitting my changes.\n\nMy preference would be:\n\n1. Merge jp/get-ref-dir-unsorted, perhaps even into maint.  It is a\nsimple, noninvasive, and obvious improvement and helps performance a lot\nin an important use case.\n\n2. Hold off on merging jp/get-ref-dir-unsorted for a while to give me a\nchance to avoid conflict hell.\n\n3. Evaluate René's patch on its own merits; if it makes sense regardless\nof the binary search speedups, then it can be accepted independently to\ngive most of the performance benefit already.\n\nAre there any other patches in this area that I've forgotten?\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176677","messageId":"4E87F462.3000308@alum.mit.edu","threadId":"27589","inReplyTo":"7vwrcola0m.fsf@alter.siamese.dyndns.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-10-02T05:19:30Z","receivedAt":"2011-10-02T05:19:30Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 10/01/2011 10:41 PM, Junio C Hamano wrote:\n> Martin Fick <mfick@codeaurora.org> writes:\n>> I guess this makes sense, we invalidate the cache and have \n>> to rebuild it after every new ref is added?  Perhaps a \n>> simple fix would be to move the invalidation right after all \n>> the refs are updated?  Maybe write_ref_sha1 could take in a \n>> flag to tell it to not invalidate the cache so that during \n>> iterative updates it could be disabled and then run manually \n>> after the update?\n> \n> It might make sense, on top of Julian's patch, to add a bit that says \"the\n> contents of this ref-array is current but the array is not sorted\", and\n> whenever somebody runs add_ref(), append it also to the ref-array (so that\n> the contents do not have to be re-read from the filesystem) but flip the\n> \"unsorted\" bit on. Then update look-up and iteration to sort the array\n> when \"unsorted\" bit is on without re-reading the contents from the\n> filesystem.\n\nMy WIP patch series does one better than this; it keeps track of what\npart of the array is already sorted so that a reference can be found in\nthe sorted part of the array using binary search, and if it is not found\nthere a linear search is done through the unsorted part of the array.  I\nalso have some code (not pushed) that adds some intelligence to make the\nuse case\n\n    repeat many times:\n        check if reference exists\n        add reference\n\nefficient by picking optimal intervals to re-sort the array.  (This sort\ncan also be faster if most of the array is already sorted: sort the new\nentries using qsort then merge sort them into the already-sorted part of\nthe list.)\n\nBut there is another reason that we cannot currently update the\nreference cache on the fly rather than invalidating it after each\nchange: symbolic references are stored *resolved* in the reference\ncache, and no record is kept of the reference that they refer to.\nTherefore it is possible that the addition or modification of an\narbitrary reference can affect how a symbolic reference is resolved, but\nthere is not enough information in the cache to track this.\n\nIMO the correct solution is to store symbolic references un-resolved.\nGiven that lookup is going to become much faster, the slowdown in\nreference resolution should not be a big performance penalty, whereas\nreference updating could become *much* faster.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176678","messageId":"7vd3egj6aa.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"4E87F383.1050403@alum.mit.edu","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-10-02T05:45:17Z","receivedAt":"2011-10-02T05:45:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu> writes:\n\n> Um, well, my patch series includes the same changes that Julian's wants\n> to introduce, but following lots of other changes, cleanups,\n> documentation improvements, etc.  Moreover, my patch series builds on\n> mh/iterate-refs, with which Julian's patch conflicts.  In other words,\n> it would be a real mess to reroll my series on top of Julian's patch.\n\nConflicts during re-rolling was not something I was worried too much\nabout---that is just the fact of life. We cannot easily resolve two topics\nthat want to go in totally different direction, but we should be able to\nconverge two topics that want to take the same approach in the end,\nespecially one is a subset of the other.\n"},{"id":"176698","messageId":"3e4aa1b3-5b14-4446-ac83-cef41c18a11f@email.android.com","threadId":"27589","inReplyTo":"4E87F462.3000308@alum.mit.edu","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-10-03T00:46:05Z","receivedAt":"2011-10-03T00:46:05Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"\n\nMichael Haggerty <mhagger@alum.mit.edu> wrote:\n\n>On 10/01/2011 10:41 PM, Junio C Hamano wrote:\n>> Martin Fick <mfick@codeaurora.org> writes:\n>>> I guess this makes sense, we invalidate the cache and have \n>>> to rebuild it after every new ref is added?  Perhaps a \n>>> simple fix would be to move the invalidation right after all \n>>> the refs are updated?  Maybe write_ref_sha1 could take in a \n>>> flag to tell it to not invalidate the cache so that during \n>>> iterative updates it could be disabled and then run manually \n>>> after the update?\n>> \n>I\n>also have some code (not pushed) that adds some intelligence to make\n>the use case\n>\n>    repeat many times:\n>        check if reference exists\n>        add reference\n\nWould it be possible to separate the two steps into separate loops somehow?  Could it instead look like this:\n \n>    repeat many times:\n>        check if reference exists\n \n>    repeat many times:\n>        add reference\n\nIt might be difficult with the current functions to achive this, but it would allow the cache to be invalidated over and over in loop two without impacting performance since all the lookups could be done in the first loop.  Of course, this would likely require checking for dups before running the first loop.\n\n-Martin\nEmployee of Qualcomm Innovation Center,Inc. which is a member of Code Aurora Forum\n"},{"id":"176759","messageId":"201110031212.13900.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109301606.31748.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-10-03T18:12:13Z","receivedAt":"2011-10-03T18:12:13Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Friday, September 30, 2011 04:06:31 pm Martin Fick wrote:\n> \n> OK, I narrowed it down I believe.  If I comment out the\n> invalidate_cached_refs() line in write_ref_sha1(), it\n> speeds through this section.\n> \n> I guess this makes sense, we invalidate the cache and\n> have to rebuild it after every new ref is added? \n> Perhaps a simple fix would be to move the invalidation\n> right after all the refs are updated?  Maybe\n> write_ref_sha1 could take in a flag to tell it to not\n> invalidate the cache so that during iterative updates it\n> could be disabled and then run manually after the\n> update?\n\nWould this solution be acceptable if I submitted a patch to \ndo it?  My test shows that this will make a full fetch of \n~80K changes go from 4:50min to 1:50min,\n\n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"176783","messageId":"7vwrcleua9.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"201110031212.13900.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-10-03T19:42:38Z","receivedAt":"2011-10-03T19:42:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Fick <mfick@codeaurora.org> writes:\n\n>> I guess this makes sense, we invalidate the cache and\n>> have to rebuild it after every new ref is added? \n>> Perhaps a simple fix would be to move the invalidation\n>> right after all the refs are updated?  Maybe\n>> write_ref_sha1 could take in a flag to tell it to not\n>> invalidate the cache so that during iterative updates it\n>> could be disabled and then run manually after the\n>> update?\n>\n> Would this solution be acceptable if I submitted a patch to \n> do it?  My test shows that this will make a full fetch of \n> ~80K changes go from 4:50min to 1:50min,\n\nAs long as the resulting code does not introduce new races with another\nprocess updating refs while the bulk update is running, I wouldn't have an\nissue with it.\n"},{"id":"176823","messageId":"4E8ABEF7.7010003@alum.mit.edu","threadId":"27589","inReplyTo":"3e4aa1b3-5b14-4446-ac83-cef41c18a11f@email.android.com","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-10-04T08:08:23Z","receivedAt":"2011-10-04T08:08:23Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 10/03/2011 02:46 AM, Martin Fick wrote:\n> Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> I\n>> also have some code (not pushed) that adds some intelligence to make\n>> the use case\n>>\n>>    repeat many times:\n>>        check if reference exists\n>>        add reference\n> \n> Would it be possible to separate the two steps into separate loops somehow?  Could it instead look like this:\n>  \n>>    repeat many times:\n>>        check if reference exists\n>  \n>>    repeat many times:\n>>        add reference\n\nUndoubtedly this would be possible.  But I'd rather make the refs code\nefficient and general enough that its users don't need to worry about\nsuch things.\n\n> [...] Of course, this would likely require checking for dups\n> before running the first loop.\n\nYes, and this \"checking for dups before running the first loop\" is\napproximately the same work that would have to be done within a smarter\nversion of the refs code.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176824","messageId":"4E8AC0E0.4090000@alum.mit.edu","threadId":"27589","inReplyTo":"201110031212.13900.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-10-04T08:16:32Z","receivedAt":"2011-10-04T08:16:32Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 10/03/2011 08:12 PM, Martin Fick wrote:\n> On Friday, September 30, 2011 04:06:31 pm Martin Fick wrote:\n>> OK, I narrowed it down I believe.  If I comment out the\n>> invalidate_cached_refs() line in write_ref_sha1(), it\n>> speeds through this section.\n>>\n>> I guess this makes sense, we invalidate the cache and\n>> have to rebuild it after every new ref is added? \n>> Perhaps a simple fix would be to move the invalidation\n>> right after all the refs are updated?  Maybe\n>> write_ref_sha1 could take in a flag to tell it to not\n>> invalidate the cache so that during iterative updates it\n>> could be disabled and then run manually after the\n>> update?\n> \n> Would this solution be acceptable if I submitted a patch to \n> do it?  My test shows that this will make a full fetch of \n> ~80K changes go from 4:50min to 1:50min,\n\nNo, no, no.  Let's fix up the refs cache once and for all and avoid\nadding special case code all over the place.\n\n* With minor changes, we can make it possible to invalidate single refs\ninstead of the whole the refs cache.  And we can teach the refs code to\ninvalidate refs by itself when necessary, so that other code can become\nstupider and more decoupled from the refs code.\n\n* With other minor changes (mostly implemented), we can support a\npartly-sorted refs list that decides intelligently when to resort\nitself.  This will give most of the performance benefit of circumventing\nthe refs cache API with none of the chaos.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"176880","messageId":"7vehysbhje.fsf@alter.siamese.dyndns.org","threadId":"27589","inReplyTo":"7vd3egj6aa.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v3] refs: Use binary search to lookup refs faster","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-10-04T20:58:29Z","receivedAt":"2011-10-04T20:58:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Michael Haggerty <mhagger@alum.mit.edu> writes:\n>\n>> Um, well, my patch series includes the same changes that Julian's wants\n>> to introduce, but following lots of other changes, cleanups,\n>> documentation improvements, etc.  Moreover, my patch series builds on\n>> mh/iterate-refs, with which Julian's patch conflicts.  In other words,\n>> it would be a real mess to reroll my series on top of Julian's patch.\n>\n> Conflicts during re-rolling was not something I was worried too much\n> about---that is just the fact of life. We cannot easily resolve two topics\n> that want to go in totally different direction, but we should be able to\n> converge two topics that want to take the same approach in the end,\n> especially one is a subset of the other.\n\nAh, also I should have noted that I have a fix-up between mh/iterate-refs\nand Julian's patch already queued on 'pu'.\n\nI am planning to make mh/iterate-refs graduate to 'master' soonish, so\nhopefully things will become simpler.\n\nThanks.\n"},{"id":"177250","messageId":"201110081459.52174.mfick@codeaurora.org","threadId":"27589","inReplyTo":"201109301606.31748.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2011-10-08T20:59:51Z","receivedAt":"2011-10-08T20:59:51Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Friday, September 30, 2011 04:06:31 pm Martin Fick wrote:\n> On Friday, September 30, 2011 03:02:30 pm Martin Fick \nwrote:\n> > On Friday, September 30, 2011 10:41:13 am Martin Fick\n> \n> wrote:\n> > Since a full sync is now done to about 5mins, I broke\n> > down the output a bit.  It appears that the longest\n> > part (2:45m) is now the time spent scrolling though\n> > each\n> > \n> > change still. Each one of these takes about 2ms:\n> >  * [new branch]      refs/changes/99/71199/1 ->\n> > \n> > refs/changes/99/71199/1\n> > \n> > Seems fast, but at about 80K... So, are there any\n> > obvious N loops over the refs happening inside each of\n> > of the [new branch] iterations?\n> \n> OK, I narrowed it down I believe.  If I comment out the\n> invalidate_cached_refs() line in write_ref_sha1(), it\n> speeds through this section.\n> \n> I guess this makes sense, we invalidate the cache and\n> have to rebuild it after every new ref is added? \n> Perhaps a simple fix would be to move the invalidation\n> right after all the refs are updated?  Maybe\n> write_ref_sha1 could take in a flag to tell it to not\n> invalidate the cache so that during iterative updates it\n> could be disabled and then run manually after the\n> update?\n\nOK, this thing has been bugging me...\n\nI found some more surprising results, I hope you can follow \nbecause there are corner cases here which have surprising \nimpacts.\n\n\n** Important fact: \n** ---------------\n** When I clone my repo, it has about 4K tags which\n** come in packed to the clone.\n**\n\nThis fact has a heavy impact on how I test things. If I \nchoose to delete these packed-refs from the cloned repo and \nthen do a fetch of the changes, all of the tags are also \nfetched along with these changes.  This means that if I want \nto test the impact of having packed-refs vs no packed refs, \non my change fetches, I need to first delete the packed-refs \nfile, and second fetch all the tags again, so that when I \nfetch the changes, the repo only actually fetches changes, \nnot all the tags!\n\nSo, with this in mind, I have discovered, that the fetch \nperformance degradation by invalidating the caches in \nwrite_ref_sha1() is actually due to the packed-refs being \nreloaded and resorted again on each ref insertion (not the \nloose refs)!!!\n\nRemember the important fact above?  Yeah, those silly 4K \nrefs (not a huge number, not 61K!) take a while to reread \nfrom the file and sort.  When this is done for 61K changes, \nit adds a lot of time to a fetch.  The sad part is that, of \ncourse, the packed-refs don't really need to be invalidated \nsince we never add new refs as packed refs during a fetch \n(but apparently we do during a clone)!  Also noteworthy is \nthat invalidating the loose refs, does not cause a big \ndelay.\n\n\nSome data:\n\n1) A fetch of the changes in my series with all good \nexternal patches applied takes about 7:30min.\n\n\n2) A fetch of the changes with #1 invalidate_cache_refs() \ncommented out in write_ref_sha1() takes about 1:50min.\n\n\n3) A fetch of the changes with #1 with \ninvalidate_cache_refs() in write_ref_sha1() replaced with a \ncall to my custom invalidate_loose_cache_refs() takes about \n1:50min.\n\n\n4) A fetch with #1 on a repo with packed-refs deleted after \nthe clone, takes about ~5min.  \n\n** This is a strange regression which threw me off.  In this \ncase, all the tags are refetched in addition to the changes, \nthis seems to cause some weird interaction that makes things \ntake longer than they should (#5 + #6 = 2:10m  <<  #4 5min).\n\n\n5) A fetch with #1 on a repo with packed-refs deleted after \nthe clone, and then a fetch done to get all the tags (see \n#6), takes only 1:30m!!!!\n\n\n6) A fetch to get all the **TAGS** with packed-refs deleted \nafter the clone, takes about 40s.\n\n\n\n---Additional side data/tests:\n\n7) A fetch of the changes with #1 and a special flag causing \nthe packed-refs to be read from the file, but not parsed or \nsorted, takes 2:34min.  So just the repeated reads add at \nleast 40s.\n\n\n8) A fetch of the changes with #1 and a special flag causing \nthe packed-refs to be read from the file, parsed, but NOT \nsorted, takes 3:40min.  So the parsing appears to take an \nadditional minute at least.\n\n\n\n\nI think that all of this might explain why no matter how \ngood Michael's intentions are with his patch series, his \nseries isn't likely to fix this problem unless he does not \ninvalidate the packed-refs after each insertion.  I tried \npreventing this invalidation in his series to prove this, \nbut unfortunately, it appears that in his series it is no \nlonger possible to only invalidate just the packed-refs? :(\nMichael, I hope I am completely wrong about that...\n\n\nAre there any good consistency reasons to invalidate the \npacked refs in write_ref_sha1()?  If not, would you accept a \npatch to simply skip this invalidation (to only invalidate \nthe loose refs)?\n\nThanks,\n \n-Martin\n\n-- \nEmployee of Qualcomm Innovation Center, Inc. which is a \nmember of Code Aurora Forum\n"},{"id":"177260","messageId":"4E913493.7010101@alum.mit.edu","threadId":"27589","inReplyTo":"201110081459.52174.mfick@codeaurora.org","subject":"Re: Git is not scalable with too many refs/*","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2011-10-09T05:43:47Z","receivedAt":"2011-10-09T05:43:47Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 10/08/2011 10:59 PM, Martin Fick wrote:\n> [...]\n> So, with this in mind, I have discovered, that the fetch \n> performance degradation by invalidating the caches in \n> write_ref_sha1() is actually due to the packed-refs being \n> reloaded and resorted again on each ref insertion (not the \n> loose refs)!!!\n\nGood point.\n\n> I think that all of this might explain why no matter how \n> good Michael's intentions are with his patch series, his \n> series isn't likely to fix this problem\n\nI never claimed that my patch fixes all use cases, or cures cancer\neither :-)  One step at a time.\n\n>                                         unless he does not\n> invalidate the packed-refs after each insertion.  I tried \n> preventing this invalidation in his series to prove this, \n> but unfortunately, it appears that in his series it is no \n> longer possible to only invalidate just the packed-refs? :(\n> Michael, I hope I am completely wrong about that...\n\nYes, you are completely wrong.  I just implemented more selective cache\ninvalidation on top of the patch series.\n\nI think your suggestion is safe because only non-symbolic references can\nbe stored in the packed refs; therefore the modification of a loose ref\ncan never affect the value of a packed ref.  Of course a loose ref can\n*hide* the value of a packed ref, but in such cases the packed ref is\nnever read anyway.  And the *deletion* of a loose ref can expose a\npreviously-hidden packed ref, but this case is handled by delete_ref(),\nwhich explicitly invalidates the packed-ref cache.\n\nWhile I was at it, I also:\n\n* In delete_ref(), only invalidate the packed reference cache if the\nreference that is being deleted actually *is* among the packed references.\n\n* Changed the code to stop invalidating the ref caches for submodules.\nIn the code paths where the cache invalidation was being done, only\nmain-module references were being changed.  However, I'm not familiar\nenough with submodules to know if/when submodule references *can* be\nchanged.  It could be that the submodule reference caches have to be\ninvalidated under some circumstances; the current code might be buggy in\nthis area.\n\nThe changes are pushed to github.  They don't make any significant\ndifference to my \"refperf\" results (attached), so perhaps a new\nbenchmark should be added.  But I'm curious to see how they affect your\ntimings.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n\n\n===================================  =======  =======  =======  =======  =======  =======  =======  =======  =======  =======\nTest name                                [0]      [1]      [2]      [3]      [4]      [5]      [6]      [7]      [8]      [9]\n===================================  =======  =======  =======  =======  =======  =======  =======  =======  =======  =======\nbranch-loose-cold                       3.19     3.15     3.10     3.19     3.25     0.70     0.61     0.74     0.66     0.56\nbranch-loose-warm                       0.19     0.19     0.20     0.19     0.19     0.00     0.00     0.00     0.00     0.00\nfor-each-ref-loose-cold                 3.73     3.45     3.55     3.39     3.44     3.40     3.50     3.52     3.70     3.51\nfor-each-ref-loose-warm                 0.44     0.44     0.44     0.43     0.43     0.43     0.43     0.43     0.43     0.43\ncheckout-loose-cold                     3.35     3.23     3.23     3.15     3.29     0.65     0.71     0.76     0.66     0.69\ncheckout-loose-warm                     0.19     0.19     0.20     0.18     0.19     0.01     0.01     0.01     0.01     0.00\ncheckout-orphan-loose                   0.19     0.19     0.19     0.18     0.19     0.00     0.00     0.00     0.00     0.00\ncheckout-from-detached-loose-cold       7.80     4.17     4.17     4.05     4.09     4.07     4.26     4.23     4.18     4.08\ncheckout-from-detached-loose-warm       1.01     1.01     1.02     1.02     1.04     1.03     1.04     1.04     1.02     1.04\nbranch-contains-loose-cold             35.76    35.80    36.15    36.67    35.13    36.29    36.37    36.03    36.70    36.01\nbranch-contains-loose-warm             33.01    33.62    33.52    33.51    32.41    33.51    33.71    32.10    33.70    31.99\npack-refs-loose                         4.19     4.20     4.25     4.21     4.20     4.21     4.20     4.19     4.24     4.21\nbranch-packed-cold                      0.79     0.62     0.60     0.66     0.65     0.58     0.68     0.72     0.60     0.61\nbranch-packed-warm                      0.02     0.02     0.02     0.02     0.02     0.02     0.02     0.02     0.02     0.02\nfor-each-ref-packed-cold                0.96     0.97     0.97     0.93     0.89     0.92     0.98     0.96     0.92     0.96\nfor-each-ref-packed-warm                0.26     0.26     0.26     0.26     0.26     0.26     0.26     0.27     0.27     0.27\ncheckout-packed-cold                   16.14    16.16    16.74     2.04     2.03     2.09     2.06     2.13     2.03     2.00\ncheckout-packed-warm                    0.17     0.17     0.18     0.19     0.18     0.17     0.27     0.18     0.19     0.18\ncheckout-orphan-packed                  0.02     0.01     0.02     0.02     0.02     0.02     0.02     0.02     0.02     0.02\ncheckout-from-detached-packed-cold     16.24    15.96    16.80     1.99     2.06     2.01     2.08     2.10     1.97     1.96\ncheckout-from-detached-packed-warm     15.04    14.96    15.76     0.77     0.81     0.79     0.83     0.80     0.79     0.80\nbranch-contains-packed-cold            36.18    36.98    36.92    35.19    34.97    35.09    33.34    33.87    34.27    34.51\nbranch-contains-packed-warm            35.27    35.12    36.20    33.52    32.76    33.49    33.65    32.96    33.68    32.34\nclone-loose-cold                        9.09     9.22     9.15     9.10     9.19     9.03     9.09     9.25     8.96     9.03\nclone-loose-warm                        5.57     5.85     5.65     5.55     5.61     5.64     5.65     5.61     5.74     5.59\nfetch-nothing-loose                     1.43     1.43     1.44     1.44     1.45     1.45     1.46     1.44     1.44     1.44\npack-refs                               0.08     0.08     0.08     0.08     0.09     0.08     0.09     0.08     0.08     0.08\nfetch-nothing-packed                    1.44     1.43     1.44     1.44     1.44     1.44     1.44     1.44     1.44     1.44\nclone-packed-cold                       1.35     1.26     1.30     1.32     1.28     1.35     1.38     1.35     1.29     1.21\nclone-packed-warm                       0.36     0.35     0.35     0.36     0.36     0.36     0.35     0.36     0.37     0.35\nfetch-everything-cold                  30.29    30.01    29.79    29.04    29.84    29.25    29.30    29.26    29.76    29.30\nfetch-everything-warm                  26.20    26.04    26.40    25.60    26.22    25.83    25.82    25.85    26.68    25.73\n===================================  =======  =======  =======  =======  =======  =======  =======  =======  =======  =======\n\n\n[0] f696543 (tag: v1.7.6) Git 1.7.6\n[1] 703f05a (tag: v1.7.7) Git 1.7.7\n[2] 27897d2 (origin/master) Merge remote-tracking branch 'gitster/mh/iterate-refs'\n[3] 558b49c is_refname_available(): reimplement using do_for_each_ref_in_list()\n[4] 1658397 Store references hierarchically\n[5] 5f5a126 get_ref_dir(): add a recursive option\n[6] a306af1 get_ref_dir(): read one whole directory before descending into subdirs\n[7] fd53cf7 add_ref(): change to take a (struct ref_entry *) as second argument\n[8] 9944c7f (origin/testing) read_packed_refs(): keep track of the directory being worked in\n[9] cb75c57 (origin/ok, origin/hierarchical-refs, origin/HEAD) refs.c: call clear_cached_ref_cache() from repack_without_ref()\n\n"}]}