{"thread":{"id":"48057","subject":"[ANNOUNCE] git-sizer: compute various size-related metrics for your Git repository","startedAt":"2018-03-16T17:27:33Z","lastAt":"2018-03-21T16:03:00Z","messageCount":5,"participants":["Michael Haggerty","Ævar Arnfjörð Bjarmason","Jeff King","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"341879","messageId":"CAMy9T_FaOdLP482YZcMX16mpy_EgM0ok1GKg45rE=X+HTGxSiQ@mail.gmail.com","threadId":"48057","inReplyTo":null,"subject":"[ANNOUNCE] git-sizer: compute various size-related metrics for your Git repository","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2018-03-16T15:28:22Z","receivedAt":"2018-03-16T17:27:33Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"What makes a Git repository unwieldy to work with and host? It turns\nout that the respository's on-disk size in gigabytes is only part of\nthe story. From our experience at GitHub, repositories cause problems\nbecause of poor internal layout at least as often as because of their\noverall size. For example,\n\n* blobs or trees that are too large\n* large blobs that are modified frequently (e.g., database dumps)\n* large trees that are modified frequently\n* trees that expand to unreasonable size when checked out (e.g., \"Git\nbombs\" [2])\n* too many tiny Git objects\n* too many references\n* other oddities, such as giant octopus merges, super long reference\nnames or file paths, huge commit messages, etc.\n\n`git-sizer` [1] is a new open-source tool that computes various\nsize-related statistics for a Git repository and points out those that\nare likely to cause problems or inconvenience to its users.\n\nI tried to make the output of `git-sizer` \"opinionated\" and easy to\ninterpret. Example output for the Linux kernel is appended below. I\nalso made it memory-efficient and resistant against git bombs.\n\nI've written a blog post [3] about `git-sizer` with more explanation\nand examples, and the main project page [1] has a long README with\nsome information about what the individual metrics mean and tips for\nfixing problems.\n\nI also put quite a bit of effort into making `git-sizer` fast. It does\nits work (including figuring out path names for large objects) based\non a single traversal of the repository history using `git rev-list\n--objects --reverse [...]`, followed by using the output of `git\ncat-file --batch` or `git cat-file --batch-check` to get information\nabout individual objects.\n\nOn that subject, let me share some more technical details. `git-sizer`\nis written in Go. I prototyped several ways of extracting object\ninformation, which is critical to the performance because `git-sizer`\nhas to read all of the reachable non-blob objects in the repository.\nThe results surprised me:\n\n| Mechanism for accessing Git data                    | Time   |\n| --------------------------------------------------- | -----: |\n| `libgit2/git2go`                                    | 25.5 s |\n| `libgit2/git2go` with `ManagedTree` optimization    | 18.9 s |\n| `src-d/go-git`                                      | 63.0 s |\n| Git command line client                             |  6.6 s |\n\nIt was almost a factor of four faster to read and parse the output of\nGit plumbing commands (mainly `git for-each-ref`, `git rev-list\n--objects`, `git cat-file --batch-check`, and `git cat-file --batch`)\nthan it was to use the Go bindings to libgit2. (I expect that part of\nthe reason is that Go's peculiar stack layout makes it quite expensive\nto call out to C.) Even after Carlos Martin implemented an\nexperimental `ManagedTree` optimization that removed the need to call\nC for every entry in a tree, it was still not competitive with the Git\nCLI. `go-git`, which is a Git implementation in pure Go, was even\nslower. So the final version of `git-sizer` calls `git` for accessing\nthe repository.\n\nFeedback is welcome, including about the weightings [4] that I use to\ncompute the \"level of concern\" of the various metrics.\n\nHave fun,\nMichael\n\n[1] https://github.com/github/git-sizer\n[2] https://kate.io/blog/git-bomb/\n[3] https://blog.github.com/2018-03-05-measuring-the-many-sizes-of-a-git-repository/\n[4] https://github.com/github/git-sizer/blob/2e9a30f241ac357f2af01d42f0dd51fbbbae4b0b/sizes/output.go#L330-L401\n\n$ git-sizer --verbose\nProcessing blobs: 1652370\nProcessing trees: 3396199\nProcessing commits: 722647\nMatching commits to trees: 722647\nProcessing annotated tags: 534\nProcessing references: 539\n| Name                         | Value     | Level of concern               |\n| ---------------------------- | --------- | ------------------------------ |\n| Overall repository size      |           |                                |\n| * Commits                    |           |                                |\n|   * Count                    |   723 k   | *                              |\n|   * Total size               |   525 MiB | **                             |\n| * Trees                      |           |                                |\n|   * Count                    |  3.40 M   | **                             |\n|   * Total size               |  9.00 GiB | ****                           |\n|   * Total tree entries       |   264 M   | *****                          |\n| * Blobs                      |           |                                |\n|   * Count                    |  1.65 M   | *                              |\n|   * Total size               |  55.8 GiB | *****                          |\n| * Annotated tags             |           |                                |\n|   * Count                    |   534     |                                |\n| * References                 |           |                                |\n|   * Count                    |   539     |                                |\n|                              |           |                                |\n| Biggest objects              |           |                                |\n| * Commits                    |           |                                |\n|   * Maximum size         [1] |  72.7 KiB | *                              |\n|   * Maximum parents      [2] |    66     | ******                         |\n| * Trees                      |           |                                |\n|   * Maximum entries      [3] |  1.68 k   |                                |\n| * Blobs                      |           |                                |\n|   * Maximum size         [4] |  13.5 MiB | *                              |\n|                              |           |                                |\n| History structure            |           |                                |\n| * Maximum history depth      |   136 k   |                                |\n| * Maximum tag depth      [5] |     1     | *                              |\n|                              |           |                                |\n| Biggest checkouts            |           |                                |\n| * Number of directories  [6] |  4.38 k   | **                             |\n| * Maximum path depth     [7] |    13     | *                              |\n| * Maximum path length    [8] |   134 B   | *                              |\n| * Number of files        [9] |  62.3 k   | *                              |\n| * Total size of files    [9] |   747 MiB |                                |\n| * Number of symlinks    [10] |    40     |                                |\n| * Number of submodules       |     0     |                                |\n\n[1]  91cc53b0c78596a73fa708cceb7313e7168bb146\n[2]  2cde51fbd0f310c8a2c5f977e665c0ac3945b46d\n[3]  4f86eed5893207aca2c2da86b35b38f2e1ec1fc8\n(refs/heads/master:arch/arm/boot/dts)\n[4]  a02b6794337286bc12c907c33d5d75537c240bd0\n(refs/heads/master:drivers/gpu/drm/amd/include/asic_reg/vega10/NBIO/nbio_6_1_sh_mask.h)\n[5]  5dc01c595e6c6ec9ccda4f6f69c131c0dd945f8c (refs/tags/v2.6.11)\n[6]  1459754b9d9acc2ffac8525bed6691e15913c6e2\n(589b754df3f37ca0a1f96fccde7f91c59266f38a^{tree})\n[7]  78a269635e76ed927e17d7883f2d90313570fdbc\n(dae09011115133666e47c35673c0564b0a702db7^{tree})\n[8]  ce5f2e31d3bdc1186041fdfd27a5ac96e728f2c5 (refs/heads/master^{tree})\n[9]  532bdadc08402b7a72a4b45a2e02e5c710b7d626\n(e9ef1fe312b533592e39cddc1327463c30b0ed8d^{tree})\n[10] f29a5ea76884ac37e1197bef1941f62fda3f7b99\n(f5308d1b83eba20e69df5e0926ba7257c8dd9074^{tree})\n"},{"id":"341917","messageId":"87370zeqmx.fsf@evledraar.gmail.com","threadId":"48057","inReplyTo":"CAMy9T_FaOdLP482YZcMX16mpy_EgM0ok1GKg45rE=X+HTGxSiQ@mail.gmail.com","subject":"Re: [ANNOUNCE] git-sizer: compute various size-related metrics for your Git repository","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-03-16T20:01:42Z","receivedAt":"2018-03-16T20:02:03Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Mar 16 2018, Michael Haggerty jotted:\n\n> What makes a Git repository unwieldy to work with and host? It turns\n> out that the respository's on-disk size in gigabytes is only part of\n> the story. From our experience at GitHub, repositories cause problems\n> because of poor internal layout at least as often as because of their\n> overall size. For example,\n>\n> * blobs or trees that are too large\n> * large blobs that are modified frequently (e.g., database dumps)\n> * large trees that are modified frequently\n> * trees that expand to unreasonable size when checked out (e.g., \"Git\n> bombs\" [2])\n> * too many tiny Git objects\n> * too many references\n> * other oddities, such as giant octopus merges, super long reference\n> names or file paths, huge commit messages, etc.\n>\n> `git-sizer` [1] is a new open-source tool that computes various\n> size-related statistics for a Git repository and points out those that\n> are likely to cause problems or inconvenience to its users.\n\nThis is a very useful tool. I've been using it to get insight into some\nbad repositories.\n\nSuggestion for a thing to add to it, I don't have the time on the Go\ntuits:\n\nOne thing that can make repositories very pathological is if the ratio\nof trees to commits is too low.\n\nI was dealing with a repo the other day that had several thousand files\nall in the same root directory, and no subdirectories.\n\nThis meant that doing `git log -- <file>` was very expensive. I wrote a\nbit about this on this related ticket the other day:\nhttps://gitlab.com/gitlab-org/gitlab-ce/issues/42104#note_54933512\n\nBut it's not something where you can just say having more trees is\nbetter, because on the other end of the spectrume we can imagine a repo\nlike linux.git where each file like COPYING instead exists at\nC/O/P/Y/I/N/G, that would also be pathological.\n\nIt would be very interesting to do some tests to see what the optimal\nvalue would be.\n\nI also suspect it's not really about the commit / tree ratio, but that\nyou have some reasonable amount of nested trees per file, *and* that\nchanges to them are reasonably spread out. I.e. it doesn't help if you\nhave a doc/ and a src/ directory if 99% of your commits change src/, and\nif you're doing 'git log -- src/something.c'.\n\nWhich is all a very long-winded way of saying that I don't know what the\ngeneral rule is, but I have some suspicions, but having all your files\nin the root is definitely bad.\n"},{"id":"341947","messageId":"20180316212920.GD12333@sigill.intra.peff.net","threadId":"48057","inReplyTo":"87370zeqmx.fsf@evledraar.gmail.com","subject":"Re: [ANNOUNCE] git-sizer: compute various size-related metrics for your Git repository","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-03-16T21:29:21Z","receivedAt":"2018-03-16T21:29:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Mar 16, 2018 at 09:01:42PM +0100, Ævar Arnfjörð Bjarmason wrote:\n\n> Suggestion for a thing to add to it, I don't have the time on the Go\n> tuits:\n> \n> One thing that can make repositories very pathological is if the ratio\n> of trees to commits is too low.\n> \n> I was dealing with a repo the other day that had several thousand files\n> all in the same root directory, and no subdirectories.\n\nWe've definitely run into this problem before (CocoaPods/Specs, for\nexample). The metric that would hopefully show this off is \"what is the\ntree object with the most entries\". Or possibly \"what is the average\nnumber of entries in a tree object\".\n\nThat's not the _whole_ story, because the really pathological case is\nwhen you then touch that giant tree a lot. But if you assume the paths\ntouched by commits are reasonably distributed over the tree, then having\na huge number of entries in one tree will also mean that more commits\nwill touch that tree. Sort of a vaguely quadratic problem.\n\nDoing it at the root is obviously the worst case, but the same thing can\nhappen if you have \"foo/bar\" as a huge tree, and every single commit\nneeds to touch some variant of \"foo/bar/baz\".\n\nThat's why I suspect some \"average per tree object\" may show this type\nof problem, because you'd have a lot of near-identical copies of that\ngiant tree if it's being modified a lot.\n\n> But it's not something where you can just say having more trees is\n> better, because on the other end of the spectrume we can imagine a repo\n> like linux.git where each file like COPYING instead exists at\n> C/O/P/Y/I/N/G, that would also be pathological.\n> \n> It would be very interesting to do some tests to see what the optimal\n> value would be.\n\nI suspect there's some math that could give us the solution. You want\napproximately equal-sized trees, so maybe log(N) levels?\n\n-Peff\n"},{"id":"342111","messageId":"CAMy9T_FNW5ksx-zLJRb48A-Dt4KNikQ9zXmxDshbby40OSLuyw@mail.gmail.com","threadId":"48057","inReplyTo":"20180316212920.GD12333@sigill.intra.peff.net","subject":"Re: [ANNOUNCE] git-sizer: compute various size-related metrics for your Git repository","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2018-03-18T19:06:04Z","receivedAt":"2018-03-18T19:06:16Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On Fri, Mar 16, 2018 at 10:29 PM, Jeff King <peff@peff.net> wrote:\n> On Fri, Mar 16, 2018 at 09:01:42PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>> One thing that can make repositories very pathological is if the ratio\n>> of trees to commits is too low.\n>>\n>> I was dealing with a repo the other day that had several thousand files\n>> all in the same root directory, and no subdirectories.\n>\n> We've definitely run into this problem before (CocoaPods/Specs, for\n> example). The metric that would hopefully show this off is \"what is the\n> tree object with the most entries\". Or possibly \"what is the average\n> number of entries in a tree object\".\n\nI find that the best metric for determining this sort of problem is\n\"Overall repository size -> Trees -> Total tree entries\". If you have\na big directory that is being changed frequently, the *real* problem\nis that every commit has to rewrite the whole tree, with all of its\nmany entries. So \"Total tree entries\" (or equivalently, the total tree\nsize) skyrockets. And this means that a history traversal has to\n*expand* all of those trees again. So a repository that is problematic\nfor this reason will have a very large number of tree entries.\n\nIf you want to detect a bad repository layout like this *before* it\nbecomes a problem, then probably something like \"average tree entries\nper commit\" might be a good leading indicator of a problem.\n\nMichael\n"},{"id":"342412","messageId":"nycvar.QRO.7.76.6.1803211659390.77@ZVAVAG-6OXH6DA.rhebcr.pbec.zvpebfbsg.pbz","threadId":"48057","inReplyTo":"CAMy9T_FaOdLP482YZcMX16mpy_EgM0ok1GKg45rE=X+HTGxSiQ@mail.gmail.com","subject":"Re: [ANNOUNCE] git-sizer: compute various size-related metrics for your Git repository","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2018-03-21T16:02:31Z","receivedAt":"2018-03-21T16:03:00Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Michael,\n\nOn Fri, 16 Mar 2018, Michael Haggerty wrote:\n\n> What makes a Git repository unwieldy to work with and host? It turns\n> out that the respository's on-disk size in gigabytes is only part of\n> the story. From our experience at GitHub, repositories cause problems\n> because of poor internal layout at least as often as because of their\n> overall size. For example,\n> \n> * blobs or trees that are too large\n> * large blobs that are modified frequently (e.g., database dumps)\n> * large trees that are modified frequently\n> * trees that expand to unreasonable size when checked out (e.g., \"Git\n> bombs\" [2])\n> * too many tiny Git objects\n> * too many references\n> * other oddities, such as giant octopus merges, super long reference\n> names or file paths, huge commit messages, etc.\n> \n> `git-sizer` [1] is a new open-source tool that computes various\n> size-related statistics for a Git repository and points out those that\n> are likely to cause problems or inconvenience to its users.\n\nThank you very much for sharing this tool.\n\nI packaged this as a MSYS2 package for use in Git for Windows' SDKs. You\ncan install it via\n\n\tpacman -Sy mingw-w64-x86_64-git-sizer\n\n(obviously, if you are in a 32-bit SDK you want to replace x86_64 by i686)\n\nNote: I am simply re-bundling the binaries you post to the GitHub\nreleases; The main purpose is to make it easier for users to include this\nin their custom installers.\n\nSecond note: I briefly considered including this tool in Git for Windows,\nbut it does increase the size of the installer by a full megabyte, and\ntherefore I decided to keep it as SDK-only, optional package.\n\nThanks!\nDscho\n"}]}