{"thread":{"id":"45256","subject":"Delta compression not so effective","startedAt":"2017-03-01T13:59:33Z","lastAt":"2017-03-07T09:07:52Z","messageCount":15,"participants":["Marius Storm-Olsen","Junio C Hamano","Linus Torvalds","Martin Langhoff","Thomas Braun"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"312947","messageId":"4d2a1852-8c84-2869-78ad-3c863f6dcaf7@gmail.com","threadId":"45256","inReplyTo":null,"subject":"Delta compression not so effective","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2017-03-01T13:51:18Z","receivedAt":"2017-03-01T13:59:33Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"I have just converted an SVN repo to Git (using SubGit), where I feel \ndelta compression has let me down :)\n\nSuffice it to say, this is a \"traditional\" SVN repo, with an extern/ \nblown out of proportion with many binary check-ins. BUT, even still, I \nwould expect Git's delta compression to be quite effective, compared to \nthe compression present in SVN. In this case however, the Git repo ends \nup being 46% larger than the SVN DB.\n\nDetails - SVN:\n     Commits: 32988\n     DB (server) size: 139GB\n     Branches: 103\n     Tags: 1088\n\nDetails - Git:\n     $ git count-objects -v\n       count: 0\n       size: 0\n       in-pack: 666515\n       packs: 1\n       size-pack: 211933109\n       prune-packable: 0\n       garbage: 0\n       size-garbage: 0\n     $ du -sh .\n       203G    .\n\n     $ java -jar ~/sources/bfg/bfg.jar --delete-folders extern \n--no-blob-protection && \\\n       git reflog expire --expire=now --all && \\\n       git gc --prune=now --aggressive\n     $ git count-objects -v\n       count: 0\n       size: 0\n       in-pack: 495070\n       packs: 1\n       size-pack: 5765365\n       prune-packable: 0\n       garbage: 0\n       size-garbage: 0\n     $ du -sh .\n       5.6G    .\n\nWhen first importing, I disabled gc to avoid any repacking until \ncompleted. When done importing, there was 209GB of all loose objects \n(~670k files). With the hopes of quick consolidation, I did a\n     git -c gc.autoDetach=0 -c gc.reflogExpire=0 \\\n           -c gc.reflogExpireUnreachable=0 -c gc.rerereresolved=0 \\\n           -c gc.rerereunresolved=0 -c gc.pruneExpire=now \\\n           gc --prune\nwhich brought it down to 206GB in a single pack. I then ran\n     git repack -a -d -F --window=350 --depth=250\nwhich took it down to 203GB, where I'm at right now.\n\nHowever, this is still miles away from the 139GB in SVN's DB.\n\nAny ideas what's going on, and why my results are so terrible, compared \nto SVN?\n\nThanks!\n\n-- \n.marius\n"},{"id":"312949","messageId":"CAPc5daU0dKXPxOjoM4Er_QxbHgLXqSPDeo_Sxc+d0aBMe837gQ@mail.gmail.com","threadId":"45256","inReplyTo":"CAPc5daWXafzN0dpyd+kdHcLU_YSmZpiNz_i2rXn_2hbMN-9Xww@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-01T16:17:14Z","receivedAt":"2017-03-01T16:17:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Wed, Mar 1, 2017 at 8:06 AM, Junio C Hamano <gitster@pobox.com> wrote:\n\n> Just a hunch. s/F/f/ perhaps?  \"-F\" does not allow Git to recover from poor\n\nNah, sorry for the noise. Between -F and -f there shouldn't be any difference.\n"},{"id":"312950","messageId":"CAPc5daWXafzN0dpyd+kdHcLU_YSmZpiNz_i2rXn_2hbMN-9Xww@mail.gmail.com","threadId":"45256","inReplyTo":"4d2a1852-8c84-2869-78ad-3c863f6dcaf7@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-01T16:06:40Z","receivedAt":"2017-03-01T17:16:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Wed, Mar 1, 2017 at 5:51 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n> ... which brought it down to 206GB in a single pack. I then ran\n>     git repack -a -d -F --window=350 --depth=250\n> which took it down to 203GB, where I'm at right now.\n\nJust a hunch. s/F/f/ perhaps?  \"-F\" does not allow Git to recover from poor\ndelta-base choice the original importer may have made (and if the\noriginal importer\nused fast-import, it is known that its choice of the delta-base is suboptimal).\n"},{"id":"312954","messageId":"CA+55aFzQ0o2R2kShS=AuKu0TLnfPV-0JCkViqx5J_afCK0Yt5g@mail.gmail.com","threadId":"45256","inReplyTo":"4d2a1852-8c84-2869-78ad-3c863f6dcaf7@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T17:36:27Z","receivedAt":"2017-03-01T17:38:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 5:51 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>\n> When first importing, I disabled gc to avoid any repacking until completed.\n> When done importing, there was 209GB of all loose objects (~670k files).\n> With the hopes of quick consolidation, I did a\n>     git -c gc.autoDetach=0 -c gc.reflogExpire=0 \\\n>           -c gc.reflogExpireUnreachable=0 -c gc.rerereresolved=0 \\\n>           -c gc.rerereunresolved=0 -c gc.pruneExpire=now \\\n>           gc --prune\n> which brought it down to 206GB in a single pack. I then ran\n>     git repack -a -d -F --window=350 --depth=250\n> which took it down to 203GB, where I'm at right now.\n\nConsidering that it was 209GB in loose objects, I don't think it\ndelta-packed the big objects at all.\n\nI wonder if the big objects end up hitting some size limit that causes\nthe delta creation to fail.\n\nFor example, we have that HASH_LIMIT  that limits how many hashes\nwe'll create for the same hash bucket, because there's some quadratic\nbehavior in the delta algorithm. It triggered with things like big\nfiles that have lots of repeated content.\n\nWe also have various memory limits, in particular\n'window_memory_limit'. That one should default to 0, but maybe you\nlimited it at some point in a config file and forgot about it?\n\n                                     Linus\n"},{"id":"312957","messageId":"eba83461-34cf-6d64-4013-873b04af9b82@gmail.com","threadId":"45256","inReplyTo":"CA+55aFzQ0o2R2kShS=AuKu0TLnfPV-0JCkViqx5J_afCK0Yt5g@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2017-03-01T17:57:27Z","receivedAt":"2017-03-01T17:59:06Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"On 3/1/2017 11:36, Linus Torvalds wrote:\n> On Wed, Mar 1, 2017 at 5:51 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>>\n>> When first importing, I disabled gc to avoid any repacking until completed.\n>> When done importing, there was 209GB of all loose objects (~670k files).\n>> With the hopes of quick consolidation, I did a\n>>     git -c gc.autoDetach=0 -c gc.reflogExpire=0 \\\n>>           -c gc.reflogExpireUnreachable=0 -c gc.rerereresolved=0 \\\n>>           -c gc.rerereunresolved=0 -c gc.pruneExpire=now \\\n>>           gc --prune\n>> which brought it down to 206GB in a single pack. I then ran\n>>     git repack -a -d -F --window=350 --depth=250\n>> which took it down to 203GB, where I'm at right now.\n>\n> Considering that it was 209GB in loose objects, I don't think it\n> delta-packed the big objects at all.\n>\n> I wonder if the big objects end up hitting some size limit that causes\n> the delta creation to fail.\n\nYou're likely on to something here.\nI just ran\n     git verify-pack --verbose \nobjects/pack/pack-9473815bc36d20fbcd38021d7454fbe09f791931.idx | sort \n-k3n | tail -n15\nand got no blobs with deltas in them.\n   feb35d6dc7af8463e038c71cc3893d163d47c31c blob   36841958 36461935 \n3259424358\n   007b65e603cdcec6644ddc25c2a729a394534927 blob   36845345 36462120 \n3341677889\n   0727a97f68197c99c63fcdf7254e5867f8512f14 blob   37368646 36983862 \n3677338718\n   576ce2e0e7045ee36d0370c2365dc730cb435f40 blob   37399203 37014740 \n3639613780\n   7f6e8b22eed5d8348467d9b0180fc4ae01129052 blob   125296632 83609223 \n5045853543\n   014b9318d2d969c56d46034a70223554589b3dc4 blob   170113524 6124878 \n1118227958\n   22d83cb5240872006c01651eb1166c8db62c62d8 blob   170113524 65941491 \n1257435955\n   292ac84f48a3d5c4de8d12bfb2905e055f9a33b1 blob   170113524 67770601 \n1323377446\n   2b9329277e379dfbdcd0b452b39c6b0bf3549005 blob   170113524 7656690 \n1110571268\n   37517efb4818a15ad7bba79b515170b3ee18063b blob   170113524 133083119 \n1124352836\n   55a4a70500eb3b99735677d0025f33b1bb78624a blob   170113524 6592386 \n1398975989\n   e669421ea5bf2e733d5bf10cf505904d168de749 blob   170113524 7827942 \n1391148047\n   e9916da851962265a9d5b099e72f60659a74c144 blob   170113524 73514361 \n966299538\n   f7bf1313752deb1bae592cc7fc54289aea87ff19 blob   170113524 70756581 \n1039814687\n   8afc6f2a51f0fa1cc4b03b8d10c70599866804ad blob   248959314 237612609 \n606692699\n\nIn fact, I don't see a single \"deltified\" blob until 6355th last line!\n\n\n> For example, we have that HASH_LIMIT  that limits how many hashes\n> we'll create for the same hash bucket, because there's some quadratic\n> behavior in the delta algorithm. It triggered with things like big\n> files that have lots of repeated content.\n>\n> We also have various memory limits, in particular\n> 'window_memory_limit'. That one should default to 0, but maybe you\n> limited it at some point in a config file and forgot about it?\n\nIndeed, I did do a\n     -c pack.threads=20 --window-memory=6g\nto 'git repack', since the machine is a 20-core (40 threads) machine \nwith 126GB of RAM.\n\nSo I guess with these sized objects, even at 6GB per thread, it's not \nenough to get a big enough Window for proper delta-packing?\n\nThis repo took >14hr to repack on 20 threads though (\"compression\" step \nwas very fast, but stuck 95% of the time in \"writing objects\"), so I can \nonly imagine how long a pack.threads=1 will take :)\n\nBut arent't the blobs sorted by some metric for reasonable delta-pack \nlocality, so even with a 6GB window it should have seen ~25 similar \nobjects to deltify against?\n\n\n-- \n.marius\n"},{"id":"312974","messageId":"CA+55aFx7QFqrHw4e72vOdM5z0rw1CCkL2-UX8ej5CLSBWjLNLA@mail.gmail.com","threadId":"45256","inReplyTo":"eba83461-34cf-6d64-4013-873b04af9b82@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T18:30:57Z","receivedAt":"2017-03-01T20:22:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 9:57 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>\n> Indeed, I did do a\n>     -c pack.threads=20 --window-memory=6g\n> to 'git repack', since the machine is a 20-core (40 threads) machine with\n> 126GB of RAM.\n>\n> So I guess with these sized objects, even at 6GB per thread, it's not enough\n> to get a big enough Window for proper delta-packing?\n\nHmm. The 6GB window should be plenty good enough, unless your blobs\nare in the gigabyte range too.\n\n> This repo took >14hr to repack on 20 threads though (\"compression\" step was\n> very fast, but stuck 95% of the time in \"writing objects\"), so I can only\n> imagine how long a pack.threads=1 will take :)\n\nActually, it's usually the compression phase that should be slow - but\nif something is limiting finding deltas (so that we abort early), then\nthat would certainly tend to speed up compression.\n\nThe \"writing objects\" phase should be mainly about the actual IO.\nWhich should be much faster *if* you actually find deltas.\n\n> But arent't the blobs sorted by some metric for reasonable delta-pack\n> locality, so even with a 6GB window it should have seen ~25 similar objects\n> to deltify against?\n\nYes they are. The sorting for delta packing tries to make sure that\nthe window is effective. However, the sorting is also just a\nheuristic, and it may well be that your repository layout ends up\nscrewing up the sorting, so that the windows just work very badly.\n\nFor example, the sorting code thinks that objects with the same name\nacross the history are good sources of deltas. But it may be that for\nyour case, the binary blobs that you have don't tend to actually\nchange in the history, so that heuristic doesn't end up doing\nanything.\n\nThe sorting does use the size and the type too, but the \"filename\nhash\" (which isn't really a hash, it's something nasty to give\nreasonable results for the case where files get renamed) is the main\nsort key.\n\nSo you might well want to look at the sorting code too. If filenames\n(particularly the end of filenames) for the blobs aren't good hints\nfor the sorting code, that sort might end up spreading all the blobs\nout rather than sort them by size.\n\nAnd again, if that happens, the \"can I delta these two objects\" code\nwill notice that the size of the objects are wildly different and\nwon't even bother trying. Which speeds up the \"compressing\" phase, of\ncourse, but then because you don't get any good deltas, the \"writing\nout\" phase sucks donkey balls because it does zlib compression on big\nobjects and writes them out to disk.\n\nSo there are certainly multiple possible reasons for the deltification\nto not work well for you.\n\nHos sensitive is your material? Could you make a smaller repo with\nsome of the blobs that still show the symptoms? I don't think I want\nto download 206GB of data even if my internet access is good.\n\n                    Linus\n"},{"id":"312976","messageId":"CACPiFCKG7K4jGLHRBA5VrdhLwjqj6KPWS2NYtHjmSdfcRhm+2A@mail.gmail.com","threadId":"45256","inReplyTo":"CA+55aFx7QFqrHw4e72vOdM5z0rw1CCkL2-UX8ej5CLSBWjLNLA@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2017-03-01T21:08:04Z","receivedAt":"2017-03-01T21:09:26Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Mar 1, 2017 at 1:30 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> For example, the sorting code thinks that objects with the same name\n> across the history are good sources of deltas.\n\nMarius has indicated he is working with jar files. IME jar and war\nfiles, which are zipfiles containing Java bytecode, range from not\ndelta-ing in a useful fashion, to pretty good deltas.\n\nDepending on the build process (hi Maven!) there can be enough\nvariance in the build metadata to throw all the compression machinery\noff.\n\nOn a simple Maven-driven project I have at hand, two .war files\ncompiled from the same codebase compressed really well in git. I've\nalso seen projects where storage space is ~101% of the \"uncompressed\"\nsize.\n\nmy 2c,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n - ask interesting questions  ~  http://linkedin.com/in/martinlanghoff\n - don't be distracted        ~  http://github.com/martin-langhoff\n   by shiny stuff\n"},{"id":"312977","messageId":"CACPiFC+=ZpHT=xh7Y8f68BcXxNYx8EFJfzqqG2ub4NL=uREu7g@mail.gmail.com","threadId":"45256","inReplyTo":"4d2a1852-8c84-2869-78ad-3c863f6dcaf7@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2017-03-01T20:19:29Z","receivedAt":"2017-03-01T21:18:02Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Mar 1, 2017 at 8:51 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n> BUT, even still, I would expect Git's delta compression to be quite effective, compared to the compression present in SVN.\n\njar files are zipfiles. They don't delta in any useful form, and in\nfact they differ even if they contain identical binary files inside.\n\n>     Commits: 32988\n>     DB (server) size: 139GB\n\nAre you certain of the on-disk storage at the SVN server? Ideally,\nyou've taken the size with a low-level tool like `du -sh\n/path/to/SVNRoot`.\n\nEven with no delta compression (as per Junio and Linus' discussion),\nbased on past experience importing jar/wars/binaries from SVN into\ngit... I'd expect git's worst case to be on-par with SVN, perhaps ~5%\nlarger due to compression headers on uncompressible data.\n\ncheers,\n\n\nm\n-- \n martin.langhoff@gmail.com\n - ask interesting questions  ~  http://linkedin.com/in/martinlanghoff\n - don't be distracted        ~  http://github.com/martin-langhoff\n   by shiny stuff\n"},{"id":"313009","messageId":"6a72bfd4-5032-5e40-5c6d-8b77ca5ae775@gmail.com","threadId":"45256","inReplyTo":"CACPiFC+=ZpHT=xh7Y8f68BcXxNYx8EFJfzqqG2ub4NL=uREu7g@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2017-03-01T23:59:30Z","receivedAt":"2017-03-02T00:06:08Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"On 3/1/2017 14:19, Martin Langhoff wrote:\n> On Wed, Mar 1, 2017 at 8:51 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>> BUT, even still, I would expect Git's delta compression to be quite effective, compared to the compression present in SVN.\n>\n> jar files are zipfiles. They don't delta in any useful form, and in\n> fact they differ even if they contain identical binary files inside.\n\nIf you look through the initial post, you'll see that the jar in \nquestion is in fact a tool (BFG) by Roberto Tyley, which is basically \ngit filter-branch on steroids. I used it to quickly filter out the \nextern/ folder, just to prove most of the original size stems from that \nparticular folder. That's all.\n\nThe repo does not contain zip or jar files. A few images and other \ncompressed formats (except a few 100MBs of proprietary files, which \nnever change), but nothing unusual.\n\n\n>>     Commits: 32988\n>>     DB (server) size: 139GB\n>\n> Are you certain of the on-disk storage at the SVN server? Ideally,\n> you've taken the size with a low-level tool like `du -sh\n> /path/to/SVNRoot`.\n\n139GB is from 'du -sh' on the SVN server. I imported (via SubGit) \ndirectly from the (hotcopied) SVN folder on the server. So true SVN size.\n\n\n> Even with no delta compression (as per Junio and Linus' discussion),\n> based on past experience importing jar/wars/binaries from SVN into\n> git... I'd expect git's worst case to be on-par with SVN, perhaps ~5%\n> larger due to compression headers on uncompressible data.\n\nYes, I was expecting a Git repo <139GB, but like Linus mentioned, \nsomething must be knocking the delta search off its feet, so it bails \nout. Loose object -> 'hard' repack didn't show that much difference.\n\n\nThanks!\n\n-- \n.marius\n"},{"id":"313011","messageId":"603afdf2-159c-6bed-0e85-2824391185d1@gmail.com","threadId":"45256","inReplyTo":"CA+55aFx7QFqrHw4e72vOdM5z0rw1CCkL2-UX8ej5CLSBWjLNLA@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2017-03-02T00:12:10Z","receivedAt":"2017-03-02T00:19:02Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"On 3/1/2017 12:30, Linus Torvalds wrote:\n> On Wed, Mar 1, 2017 at 9:57 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>>\n>> Indeed, I did do a\n>>     -c pack.threads=20 --window-memory=6g\n>> to 'git repack', since the machine is a 20-core (40 threads) machine with\n>> 126GB of RAM.\n>>\n>> So I guess with these sized objects, even at 6GB per thread, it's not enough\n>> to get a big enough Window for proper delta-packing?\n>\n> Hmm. The 6GB window should be plenty good enough, unless your blobs\n> are in the gigabyte range too.\n\nNo, the list of git verify-objects in the previous post was from the \nbottom of the sorted list, so those are the largest blobs, ~249MB..\n\n\n>> This repo took >14hr to repack on 20 threads though (\"compression\" step was\n>> very fast, but stuck 95% of the time in \"writing objects\"), so I can only\n>> imagine how long a pack.threads=1 will take :)\n>\n> Actually, it's usually the compression phase that should be slow - but\n> if something is limiting finding deltas (so that we abort early), then\n> that would certainly tend to speed up compression.\n>\n> The \"writing objects\" phase should be mainly about the actual IO.\n> Which should be much faster *if* you actually find deltas.\n\nSo, this repo must be knocking several parts of Git's insides. I was \ncurious about why it was so slow on the writing objects part, since the \nwhole repo is on a 4x RAID 5, 7k spindels. Now, they are not SSDs sure, \nbut the thing has ~400MB/s continuous throughput available.\n\niostat -m 5 showed trickle read/write to the process, and 80-100% CPU \nsingle thread (since the \"write objects\" stage is single threaded, \nobviously).\n\nThe failing delta must be triggering other negative behavior.\n\n\n> For example, the sorting code thinks that objects with the same name\n> across the history are good sources of deltas. But it may be that for\n> your case, the binary blobs that you have don't tend to actually\n> change in the history, so that heuristic doesn't end up doing\n> anything.\n\nThese are generally just DLLs (debug & release), which content is \nupdated due to upstream project updates. So, filenames/paths tend to \nstay identical, while content changes throughout history.\n\n\n> The sorting does use the size and the type too, but the \"filename\n> hash\" (which isn't really a hash, it's something nasty to give\n> reasonable results for the case where files get renamed) is the main\n> sort key.\n>\n> So you might well want to look at the sorting code too. If filenames\n> (particularly the end of filenames) for the blobs aren't good hints\n> for the sorting code, that sort might end up spreading all the blobs\n> out rather than sort them by size.\n\nFilenames are fairly static, and the bulk of the 6000 biggest \nnon-delta'ed blobs are the same DLLs (multiple of them)\n\n\n> And again, if that happens, the \"can I delta these two objects\" code\n> will notice that the size of the objects are wildly different and\n> won't even bother trying. Which speeds up the \"compressing\" phase, of\n> course, but then because you don't get any good deltas, the \"writing\n> out\" phase sucks donkey balls because it does zlib compression on big\n> objects and writes them out to disk.\n\nRight, now on this machine, I really didn't notice much difference \nbetween standard zlib level and doing -9. The 203GB version was actually \nwith zlib=9.\n\n\n> So there are certainly multiple possible reasons for the deltification\n> to not work well for you.\n>\n> Hos sensitive is your material? Could you make a smaller repo with\n> some of the blobs that still show the symptoms? I don't think I want\n> to download 206GB of data even if my internet access is good.\n\nPretty sensitive, and not sure how I can reproduce this reasonable well. \nHowever, I can easily recompile git with any recommended \ninstrumentation/printfs, if you have any suggestions of good places to \nstart? If anyone have good file/line numbers, I'll give that a go, and \nreport back?\n\nThanks!\n\n-- \n.marius\n"},{"id":"313026","messageId":"CA+55aFxxQUixAJWXkUgVvDNCHD4LuYYuQRTE7dJ_OZTo9Gxqew@mail.gmail.com","threadId":"45256","inReplyTo":"603afdf2-159c-6bed-0e85-2824391185d1@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-02T00:43:24Z","receivedAt":"2017-03-02T01:45:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 4:12 PM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>\n> No, the list of git verify-objects in the previous post was from the bottom\n> of the sorted list, so those are the largest blobs, ~249MB..\n\n.. so with a 6GB window, you should easily sill have 20+ objects. Not\na huge window, but it should find some deltas.\n\nBut a smaller window - _together_ with a suboptimal sorting choice -\ncould then result in lack of successful delta matches.\n\n> So, this repo must be knocking several parts of Git's insides. I was curious\n> about why it was so slow on the writing objects part, since the whole repo\n> is on a 4x RAID 5, 7k spindels. Now, they are not SSDs sure, but the thing\n> has ~400MB/s continuous throughput available.\n>\n> iostat -m 5 showed trickle read/write to the process, and 80-100% CPU single\n> thread (since the \"write objects\" stage is single threaded, obviously).\n\nSo the writing phase isn't multi-threaded because it's not expected to\nmatter. But if you can't even generate deltas, you aren't just\n*writing* much more data, you're compressing all that data with zlib\ntoo.\n\nSo even with a fast disk subsystem, you won't even be able to saturate\nthe disk, simply because the compression will be slower (and\nsingle-threaded).\n\n> Filenames are fairly static, and the bulk of the 6000 biggest non-delta'ed\n> blobs are the same DLLs (multiple of them)\n\nI think the first thing you should test is to repack with fewer\nthreads, and a bigger pack window. Do somethinig like\n\n  -c pack.threads=4 --window-memory=30g\n\ninstead. Just to see if that starts finding deltas.\n\n> Right, now on this machine, I really didn't notice much difference between\n> standard zlib level and doing -9. The 203GB version was actually with\n> zlib=9.\n\nDon't. zlib has *horrible* scaling with higher compressions. It\ndoesn't actually improve the end result very much, and it makes things\n*much* slower.\n\nzlib was a reasonable choice when git started - well-known, stable, easy to use.\n\nBut realistically it's a relatively horrible choice today, just\nbecause there are better alternatives now.\n\n>> Hos sensitive is your material? Could you make a smaller repo with\n>> some of the blobs that still show the symptoms? I don't think I want\n>> to download 206GB of data even if my internet access is good.\n>\n> Pretty sensitive, and not sure how I can reproduce this reasonable well.\n> However, I can easily recompile git with any recommended\n> instrumentation/printfs, if you have any suggestions of good places to\n> start? If anyone have good file/line numbers, I'll give that a go, and\n> report back?\n\nSo the first thing you might want to do is to just print out the\nobjects after sorting them, and before it starts trying to finsd\ndeltas.\n\nSee prepare_pack() in builtin/pack-objects.c, where it does something like this:\n\n        if (nr_deltas && n > 1) {\n                unsigned nr_done = 0;\n                if (progress)\n                        progress_state = start_progress(_(\"Compressing\nobjects\"),\n                                                        nr_deltas);\n                QSORT(delta_list, n, type_size_sort);\n                ll_find_deltas(delta_list, n, window+1, depth, &nr_done);\n                stop_progress(&progress_state);\n\n\nand notice that QSORT() line: that's what sorts the objects. You can\ndo something like\n\n                for (i = 0; i < n; i++)\n                        show_object_entry_details(delta_list[i]);\n\nright after that QSORT(), and make that print out the object hash,\nfilename hash, and size (we don't have the filename that the object\nwas associated with any more at that stage - they take too much\nspace).\n\nSave off that array for off-line processing: when you have the object\nhash, you can see what the contents are, and match it up wuith the\nfile in the git history using something like\n\n   git log --oneline --raw -R --abbrev=40\n\nwhich shows you the log, but also the \"diff\" in the form of \"this\nfilename changed from SHA1 to SHA1\", so you can match up the object\nhashes with where they are in the tree (and where they are in\nhistory).\n\nSo then you could try to figure out if that type_size_sort() heuristic\nis just particularly horrible for you.\n\nIn fact, if your data is not *so* sensitive, and you're ok with making\nthe one-line commit logs and the filenames public, you could make just\nthose things available, and maybe I'll have time to look at it.\n\nI'm in the middle of the kernel merge window, but I'm in the last\nstretch, and because of the SHA1 thing I've been looking at git\nlately. No promises, though.\n\n                   Linus\n"},{"id":"313266","messageId":"9961a973-0d5d-5ff9-ab78-eea07bdb5dbf@gmail.com","threadId":"45256","inReplyTo":"CA+55aFxxQUixAJWXkUgVvDNCHD4LuYYuQRTE7dJ_OZTo9Gxqew@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2017-03-04T08:27:00Z","receivedAt":"2017-03-04T08:55:55Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"On 3/1/2017 18:43, Linus Torvalds wrote:\n>> So, this repo must be knocking several parts of Git's insides. I was curious\n>> about why it was so slow on the writing objects part, since the whole repo\n>> is on a 4x RAID 5, 7k spindels. Now, they are not SSDs sure, but the thing\n>> has ~400MB/s continuous throughput available.\n>>\n>> iostat -m 5 showed trickle read/write to the process, and 80-100% CPU single\n>> thread (since the \"write objects\" stage is single threaded, obviously).\n>\n> So the writing phase isn't multi-threaded because it's not expected to\n> matter. But if you can't even generate deltas, you aren't just\n> *writing* much more data, you're compressing all that data with zlib\n> too.\n>\n> So even with a fast disk subsystem, you won't even be able to saturate\n> the disk, simply because the compression will be slower (and\n> single-threaded).\n\nI did a simple\n     $ time zip -r repo.zip repo/\n...\n     total bytes=219353596620, compressed=214310715074 -> 2% savings\n\n     real    154m6.323s\n     user    133m5.209s\n     sys     5m5.338s\n\nalso using a single thread + same disk, as git repack. But if you \ncompare it to the numbers below, it's 2.6hrs with zip vs 14.2hrs \n(1:5.5). So it can't just be the overhead of having to compress the full \nblobs, due to lacking delta..\n\n\n>> Filenames are fairly static, and the bulk of the 6000 biggest non-delta'ed\n>> blobs are the same DLLs (multiple of them)\n>\n> I think the first thing you should test is to repack with fewer\n> threads, and a bigger pack window. Do somethinig like\n>\n>   -c pack.threads=4 --window-memory=30g\n>\n> instead. Just to see if that starts finding deltas.\n\nI reran the repack with the options above (dropping the zlib=9, as you \nsuggested)\n\n     $ time git -c pack.threads=4 repack -a -d -F \\\n                --window=350 --depth=250 --window-memory=30g\n\n     Delta compression using up to 4 threads.\n     Compressing objects:   100% (609413/609413)\n     Writing objects: 100% (666515/666515), done.\n     Total 666515 (delta 499585), reused 0 (delta 0)\n\n     real\t850m3.473s\n     user\t897m36.280s\n     sys \t10m8.824s\n\nand ended up with\n     $ du -sh .\n     205G\t.\n\nIn other words, going from 6G to 30G window didn't help a lick on \nfinding deltas for those binaries. (205G was what I had with the \nnon-aggressive 'git gc', before zlib=9 repack.)\n\nBUT, oddly enough, even if the new size if almost identical to the \nprevious version without zlib=9,\n     git verify-pack --verbose \nobjects/pack/pack-29b06ae4d458ac03efd98b330702d30e851b2933.idx | sort \n-k3n | tail -n15\ngives me a VERY different list than before\n\n   17e5b2146311256dc8317d6e0ed1291363c31a76 blob   673399562 110248747 \n190398904084\n   04c881d9069eab3bd0d50dd48a047a60f79cc415 blob   673863358 111710559 \n188818868865\n   fdcabd75aeda86ce234d6e43b54d27d993acddcd blob   674523614 111956017 \n185706433825\n   d8815033d1b00b151ae762be8a69ffa35f55c4b4 blob   675286758 112099638 \n185153570292\n   997e0b9d3bcf440af10c7bbe535a597ca46c492c blob   678274978 112654668 \n184041692883\n   dfed141679e5c33caaa921cbe1595a24967a3c2c blob   681692132 113121410 \n186753502634\n   76a4000e71cd5b85f2265e02eb876acf1f33cc55 blob   682673430 112743915 \n184563542298\n   81e7292c4d2da2d2d236fbfaa572b6c4e8d787f4 blob   684543130 112797325 \n181805773038\n   991184c60e1fc6b2721bf40f181012b72b10d02d blob   684543130 112796892 \n182344388066\n   0e9269f4abd1440addd05d4f964c96d74d11cd89 blob   684547270 112809074 \n181070719237\n   6019b6d09759cf5adeac678c8b56d177803a0486 blob   684547270 112809336 \n180517242193\n   70a5f70bd205329472d6f9c660eb3f7d207a596e blob   686852038 112873611 \n183520467528\n   e86a0064d9652be9f5e3a877b11a665f64198ecd blob   686852038 112874133 \n182893219377\n   bae8de0555be5b1ffa0988cbc6cba698f6745c26 blob   894041802 137223252 \n2355250324\n   94dc773600e03ac1e6f3ab077b70b8297325ad77 blob   945197364 145219485 \n16560137220\n\ncompared to the last 3 entries of the previous pack\n   e9916da851962265a9d5b099e72f60659a74c144 blob   170113524 73514361 \n966299538\n   f7bf1313752deb1bae592cc7fc54289aea87ff19 blob   170113524 70756581 \n1039814687\n   8afc6f2a51f0fa1cc4b03b8d10c70599866804ad blob   248959314 237612609 \n606692699\n\n\n> So the first thing you might want to do is to just print out the\n> objects after sorting them, and before it starts trying to finsd\n> deltas.\n...\n> and notice that QSORT() line: that's what sorts the objects. You can\n> do something like\n>\n>                 for (i = 0; i < n; i++)\n>                         show_object_entry_details(delta_list[i]);\n\nI did\n     fprintf(stderr, \"%s %u %lu\\n\",\n             sha1_to_hex(delta_list[i]->idx.sha1),\n             delta_list[i]->hash,\n             delta_list[i]->size);\n\nI assume that's correct?\n\n\n> In fact, if your data is not *so* sensitive, and you're ok with making\n> the one-line commit logs and the filenames public, you could make just\n> those things available, and maybe I'll have time to look at it.\n\nI've removed all commit messages, and \"sanitized\" some filepaths etc, so \nname hashes won't match what's reported, but that should be fine. (the \nobject_entry->hash seems to be just a trivial uint32 hash for sorting \nanyways)\n\nI really don't want the files on the mailinglist, so I'll send you a \nlink directly. However, small snippets for public discussions about \npotential issues would be fine, obviously.\n\nBUT, if I look at the last 3 entries of the sorted git verify-pack \noutput, and look for them in the 'git log --oneline --raw -R \n--abbrev=40' output, I get:\n  :100644 100644 991184c60e1fc6b2721bf40f181012b72b10d02d \ne86a0064d9652be9f5e3a877b11a665f64198ecd M \nextern/win/FlammableV3/x64/lib/FlameProxyLibD.lib\n  :100644 000000 bae8de0555be5b1ffa0988cbc6cba698f6745c26 \n0000000000000000000000000000000000000000 D \nextern/win/gdal-2.0.0/lib/x64/Debug/libgdal.lib\n  :000000 100644 0000000000000000000000000000000000000000 \n94dc773600e03ac1e6f3ab077b70b8297325ad77 A \nextern/win/gdal-2.0.0/lib/x64/Debug/gdal.lib\n\nwhile I cannot find ANY of them in the delta_list output?? Shouldn't \ndelta_list contain all objects, sorted by some heuristics? Or is the \ndelta_list already here limited by some other metric, before the QSORT?\n\nAlso note that the 'git log --oneline --raw -R --abbrev=40' only gave me \nthe log for trunk, so for the second last object, must have been added \nin a branch, and deleted on trunk; so I could only see the deletion of \nthat object in the output.\n\n\nYou might get an idea for how to easily create a repo which reproduces \nthe issue, and which would highlight it more easily for the ML.\n\nI was thinking of maybe scripting up\n     make install prefix=extern\nfor each Git release, and rewrite trunk history with extern/ binary \ncommits at the time of each tag; maybe that would show the same \nbehavior? But then again, most of the binaries are just copies of each \nother, and only ~10M, so probably not a big win.\n\n\nThanks!\n\n-- \n.marius\n"},{"id":"313306","messageId":"CA+55aFw=U4PbfvVzeyuWk2VOsgicZRRKZRkrGp7jr_ppvgP3ng@mail.gmail.com","threadId":"45256","inReplyTo":"9961a973-0d5d-5ff9-ab78-eea07bdb5dbf@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-06T01:14:14Z","receivedAt":"2017-03-06T01:23:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Sat, Mar 4, 2017 at 12:27 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n>\n> I reran the repack with the options above (dropping the zlib=9, as you\n> suggested)\n>\n>     $ time git -c pack.threads=4 repack -a -d -F \\\n>                --window=350 --depth=250 --window-memory=30g\n>\n> and ended up with\n>     $ du -sh .\n>     205G        .\n>\n> In other words, going from 6G to 30G window didn't help a lick on finding\n> deltas for those binaries.\n\nOk.\n\n> I did\n>     fprintf(stderr, \"%s %u %lu\\n\",\n>             sha1_to_hex(delta_list[i]->idx.sha1),\n>             delta_list[i]->hash,\n>             delta_list[i]->size);\n>\n> I assume that's correct?\n\nLooks good.\n\n> I've removed all commit messages, and \"sanitized\" some filepaths etc, so\n> name hashes won't match what's reported, but that should be fine. (the\n> object_entry->hash seems to be just a trivial uint32 hash for sorting\n> anyways)\n\nYes. I see your name list and your pack-file index.\n\n> BUT, if I look at the last 3 entries of the sorted git verify-pack output,\n> and look for them in the 'git log --oneline --raw -R --abbrev=40' output, I\n> get:\n...\n> while I cannot find ANY of them in the delta_list output?? \\\n\nYes. You have a lot of of object names in that log file you sent in\nprivate that aren't in the delta list.\n\nNow, objects smaller than 50 bytes we don't ever try to even delta. I\ncan't see the object sizes when they don't show up in the delta list,\nbut looking at some of those filenames I'd expect them to not fall in\nthat category.\n\nI guess you could do the printout a bit earlier (on the\n\"to_pack.objects[]\" array - to_pack.nr_objects is the count there).\nThat should show all of them. But the small objects shouldn't matter.\n\nBut if you have a file like\n\n   extern/win/FlammableV3/x64/lib/FlameProxyLibD.lib\n\nI would have assumed that it has a size that is > 50. Unless those\n\"extern\" things are placeholders?\n\n> You might get an idea for how to easily create a repo which reproduces the\n> issue, and which would highlight it more easily for the ML.\n\nLooking at your sorted object list ready for packing, it doesn't look\nhorrible. When sorting for size, it still shows a lot of those large\nfiles with the same name hash, so they sorted together in that form\ntoo.\n\nI do wonder if your dll data just simply is absolutely horrible for\nxdelta. We've also limited the delta finding a bit, simply because it\nhad some O(m*n) behavior that gets very expensive on some patterns.\nMaybe your blobs trigger some of those case.\n\nThe diff-delta work all goes back to 2005 and 2006, so it's a long time ago.\n\nWhat I'd ask you to do is try to find if you could make a reposity of\njust one of the bigger DLL's with its history, particularly if you can\nfind some that you don't think is _that_ sensitive.\n\nLooking at it, for example, I see that you have that file\n\n   extern/redhat-5/FlammableV3/x64/plugins/libFlameCUDA-3.0.703.so\n\nthat seems to have changed several times, and is a largish blob. Could\nyou try creating a repository with git fast-import that *only*\ncontains that file (or pick another one), and see if that delta's\nwell?\n\nAnd if you find some case that doesn't xdelta well, and that you feel\nyou could make available outside, we could have a test-case...\n\n                 Linus\n"},{"id":"313339","messageId":"ef0b90a6-4386-7479-b56b-acfe34989bda@gmail.com","threadId":"45256","inReplyTo":"CA+55aFw=U4PbfvVzeyuWk2VOsgicZRRKZRkrGp7jr_ppvgP3ng@mail.gmail.com","subject":"Re: Delta compression not so effective","fromName":"Marius Storm-Olsen","fromEmail":"mstormo@gmail.com","sentAt":"2017-03-06T13:36:45Z","receivedAt":"2017-03-06T13:47:30Z","isPatch":false,"sender":{"key":"mstormo@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1500?v=4"},"body":"On 3/5/2017 19:14, Linus Torvalds wrote:\n> On Sat, Mar 4, 2017 at 12:27 AM, Marius Storm-Olsen <mstormo@gmail.com> wrote:\n> I guess you could do the printout a bit earlier (on the\n> \"to_pack.objects[]\" array - to_pack.nr_objects is the count there).\n> That should show all of them. But the small objects shouldn't matter.\n>\n> But if you have a file like\n>\n>    extern/win/FlammableV3/x64/lib/FlameProxyLibD.lib\n>\n> I would have assumed that it has a size that is > 50. Unless those\n> \"extern\" things are placeholders?\n\nNo placeholders, the FlameProxyLibD.lib is a debug lib, and probably the \nlargest in the whole repo (with a replace count > 5).\n\n\n> I do wonder if your dll data just simply is absolutely horrible for\n> xdelta. We've also limited the delta finding a bit, simply because it\n> had some O(m*n) behavior that gets very expensive on some patterns.\n> Maybe your blobs trigger some of those case.\n\nOk, but given that the SVN delta compression, which forward-linear only, \nis ~45% better, perhaps that particular search could be done fairly \ncheap? Although, I bet time(stamps) are out of the loop at that point, \nso it's not a factor anymore. Even if it where, I'm not sure it would \nsolve anything, if there's other factors also limiting deltafication.\n\n\n> The diff-delta work all goes back to 2005 and 2006, so it's a long time ago.\n>\n> What I'd ask you to do is try to find if you could make a reposity of\n> just one of the bigger DLL's with its history, particularly if you can\n> find some that you don't think is _that_ sensitive.\n>\n> Looking at it, for example, I see that you have that file\n>\n>    extern/redhat-5/FlammableV3/x64/plugins/libFlameCUDA-3.0.703.so\n>\n> that seems to have changed several times, and is a largish blob. Could\n> you try creating a repository with git fast-import that *only*\n> contains that file (or pick another one), and see if that delta's\n> well?\n\nI'll filter-branch to extern/ only, however the whole FlammableV3 needs \nto go too, I'm afaid (extern for that project, but internal to $WORK).\nI'll do some rewrites and see what comes up.\n\n> And if you find some case that doesn't xdelta well, and that you feel\n> you could make available outside, we could have a test-case...\n\nI'll try with this repo first, if not, I'll see if I can construct one.\n\nThanks!\n\n\n-- \n.marius\n"},{"id":"313402","messageId":"1831082528.2738206.1488877652752.JavaMail.open-xchange@app04.ox.hosteurope.de","threadId":"45256","inReplyTo":"9961a973-0d5d-5ff9-ab78-eea07bdb5dbf@gmail.com","subject":"Re: Delta compression not so effective","fromName":"Thomas Braun","fromEmail":"thomas.braun@virtuell-zuhause.de","sentAt":"2017-03-07T09:07:32Z","receivedAt":"2017-03-07T09:07:52Z","isPatch":false,"sender":{"key":"thomas.braun@virtuell-zuhause.de","avatar":"https://avatars.githubusercontent.com/u/1185677?v=4"},"body":"\n\n> Marius Storm-Olsen <mstormo@gmail.com> hat am 4. März 2017 um 09:27\n> geschrieben:\n\n[...]\n\n> I really don't want the files on the mailinglist, so I'll send you a \n> link directly. However, small snippets for public discussions about \n> potential issues would be fine, obviously.\n\ngit fast-export can anonymize a repository [1]. Maybe an anonymized repository\nstill shows the issue you are seeing.\n\n[1]: https://www.git-scm.com/docs/git-fast-export#_anonymizing\n"}]}