{"thread":{"id":"9631","subject":"git-daemon on NSLU2","startedAt":"2007-08-24T05:54:48Z","lastAt":"2007-08-27T16:26:23Z","messageCount":30,"participants":["Jon Smirl","Shawn O. Pearce","Nicolas Pitre","Jakub Narebski","Junio C Hamano","Linus Torvalds","David Kastrup","Salikh Zakirov","Jeff King","Daniel Hulme","Theodore Tso"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"51418","messageId":"9e4733910708232254w4e74ca72o917c7cadae4ee0f4@mail.gmail.com","threadId":"9631","inReplyTo":null,"subject":"git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-24T05:54:48Z","receivedAt":"2007-08-24T05:54:48Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Any ideas on why git protocol clone is failing?\n\n2007-08-24_20:51:33.85649 [9758] Connection from 72.74.92.181:19367\n2007-08-24_20:51:33.85828 [9758] Extended attributes (33 bytes) exist\n<host=git.jonsmirl.is-a-geek.net>\n2007-08-24_20:51:33.96990 [9758] Request upload-pack for\n'/home/git/mpc5200b.git'\n2007-08-24_20:51:45.00789 fatal: Out of memory? mmap failed: Cannot\nallocate memory\n2007-08-24_20:51:45.08746 error: git-upload-pack: git-rev-list died with error.\n2007-08-24_20:51:45.08771 fatal: git-upload-pack: aborting due to\npossible repository corruption on the remote side.\n\nNSLU2 ($70) is 266Mhz ARM with 32MB memory.\nIt's running Debian on a 250GB disk with 180MB swap.\n\nWatching top the process runs up to about 60MB in virtual size and exits.\nSetting the window down made no difference  packedGitWindowSize = 4194304\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51421","messageId":"20070824062106.GV27913@spearce.org","threadId":"9631","inReplyTo":"9e4733910708232254w4e74ca72o917c7cadae4ee0f4@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-08-24T06:21:06Z","receivedAt":"2007-08-24T06:21:06Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jon Smirl <jonsmirl@gmail.com> wrote:\n> Any ideas on why git protocol clone is failing?\n> \n> 2007-08-24_20:51:33.85649 [9758] Connection from 72.74.92.181:19367\n> 2007-08-24_20:51:33.85828 [9758] Extended attributes (33 bytes) exist\n> <host=git.jonsmirl.is-a-geek.net>\n> 2007-08-24_20:51:33.96990 [9758] Request upload-pack for\n> '/home/git/mpc5200b.git'\n> 2007-08-24_20:51:45.00789 fatal: Out of memory? mmap failed: Cannot\n> allocate memory\n> 2007-08-24_20:51:45.08746 error: git-upload-pack: git-rev-list died with error.\n> 2007-08-24_20:51:45.08771 fatal: git-upload-pack: aborting due to\n> possible repository corruption on the remote side.\n> \n> NSLU2 ($70) is 266Mhz ARM with 32MB memory.\n> It's running Debian on a 250GB disk with 180MB swap.\n> \n> Watching top the process runs up to about 60MB in virtual size and exits.\n> Setting the window down made no difference  packedGitWindowSize = 4194304\n\nulimits?  packedGitLimit may also need to be decreased?  Though we\nalways try to free unused windows before we declare we are out\nof memory...\n\n-- \nShawn.\n"},{"id":"51466","messageId":"9e4733910708241238n1899f332j4fafbd6d7ccc48b9@mail.gmail.com","threadId":"9631","inReplyTo":"20070824062106.GV27913@spearce.org","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-24T19:38:32Z","receivedAt":"2007-08-24T19:38:32Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I'm still trying to debug git-daemon\n\nI do find it surprising that git-index-pack can't be happy with in\n20MB of RAM and it has to continuously swap it's 30MB of virtual. My\ndisk is chattering itself to death. It stayed that way for 40 minutes.\n\nI'm practicing on the kernel tree.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51468","messageId":"alpine.LFD.0.999.0708241618070.16727@xanadu.home","threadId":"9631","inReplyTo":"9e4733910708241238n1899f332j4fafbd6d7ccc48b9@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-08-24T20:23:10Z","receivedAt":"2007-08-24T20:23:10Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 24 Aug 2007, Jon Smirl wrote:\n\n> I'm still trying to debug git-daemon\n> \n> I do find it surprising that git-index-pack can't be happy with in\n> 20MB of RAM and it has to continuously swap it's 30MB of virtual. My\n> disk is chattering itself to death. It stayed that way for 40 minutes.\n> \n> I'm practicing on the kernel tree.\n\nYou hope for miracles, do you?  ;-)\n\nPlease stop hammering that poor little NSLU2 with such a workset, or \nhack some additional 224MB of RAM into it.  There is no magical \nsolution.\n\n\nNicolas\n"},{"id":"51470","messageId":"9e4733910708241327l664d5077k175c13a2c386dbaa@mail.gmail.com","threadId":"9631","inReplyTo":"9e4733910708241238n1899f332j4fafbd6d7ccc48b9@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-24T20:27:34Z","receivedAt":"2007-08-24T20:27:34Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Not sure what I did but I have git-daemon working on the NSLU2 now.\n\nIt is unusable with 32MB physical memory.  I am 2hrs into the clone of\nthe kernel repository and it has only counted 9,500 objects and used\n100min CPU time. There are 540,000 objects in the repository.\n\nDisk is chattering insanely, I'm way IO bound.\n\nprocs -----------memory---------- ---swap-- -----io---- -system-- ----cpu----\n r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa\n 6  2  37960    972    168  11952  160   64  1748    64 2224 2233  5 28  0 67\n 4  2  37960   1012    176  11756  168    0  2424     0 2517 2780 10 29  0 61\n 2  2  37960    944    200  11792  152    0  1456    88 2102 2067  6 21  0 73\n 2  2  37960   1120    180  11620  120    0  1180     0 2106 2122  4 21  0 75\n 2  2  37960   1044    180  11788   76   28  1800    28 2255 2275  7 27  0 66\n 4  3  37960   1144    176  11436   68    0  1896    12 2384 2553  7 23  0 70\n 4  1  37972    992    188  11932   44  188  1148   188 1910 1731  3 18  0 79\n 3  2  37976    804    196  12008  336   16  2104   112 2353 2490 13 22  0 65\n 2  2  37976   1068    164  11720   96    8  2008     8 2502 2731  5 36  0 59\n 2  2  37976   1280    184  11528  140    8  1332    36 2054 1956  7 26  0 67\n 4  2  37976   1028    200  11552  264   16   956    16 1855 1710  4 20  0 76\n 2  2  37976    844    192  11680  144    8  1576     8 2206 2307  5 31  0 64\n 3  1  37984   1304    172  11264   92   28  1444    52 1998 1887  5 23  0 72\n 2  2  38000   1012    168  11680  124   84  1896   192 2385 2486  3 30  0 67\n 5  2  38008    928    164  11916  136   20  1776    20 2256 2308 11 22  0 67\n 2  3  38008   1168    184  11704  144   20  1820    32 2163 2186  5 24  0 71\n 4  4  38016    816    156  11784  248   32  1828    44 2328 2422  2 24  0 74\n 4  1  38020   1476    160  11448  152  104  2080   116 1925 1728  3 24  0 73\n 2  5  38028    828    192  12140  240  140  1768   232 2319 2226  4 29  0 68\n 2  2  38020   1136    172  11880  156   16  1764    72 2081 2020  3 20  0 77\n 2  3  38060   1040    172  12016  188  140  2056   140 2180 2182  6 26  0 68\n\nroot     11241  0.3  0.0    104    24 ?        Ss   06:54   0:07 runsv\ngit-daemon\ngitlog   11242  0.0  0.1    124    40 ?        S    06:54   0:01\nsvlogd -tt /var/log/git-daemon\nroot     11335  0.0  0.4   1620   140 pts/0    S+   06:56   0:00\nstrace git-daemon --verbose --export-all /home/git\nroot     11336  0.0  0.4   1808   144 pts/0    S+   06:56   0:00\ngit-daemon --verbose --export-all /home/git\nroot     11344  0.1  1.0  60240   328 pts/0    S+   06:56   0:02\n/usr/local/bin/git-upload-pack --strict --timeout=0 .\nroot     11349  6.5 50.8 171868 15240 pts/0    D+   06:56   2:09\n/usr/local/bin/git-upload-pack --strict --timeout=0 .\nroot     11350  0.6 14.6  16392  4380 pts/0    S+   06:56   0:12\n/usr/local/bin git-pack-objects --stdout --progress\n--delta-base-offset\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51473","messageId":"9e4733910708241417l44c55306xaa322afda69c6beb@mail.gmail.com","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708241618070.16727@xanadu.home","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-24T21:17:42Z","receivedAt":"2007-08-24T21:17:42Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/24/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 24 Aug 2007, Jon Smirl wrote:\n>\n> > I'm still trying to debug git-daemon\n> >\n> > I do find it surprising that git-index-pack can't be happy with in\n> > 20MB of RAM and it has to continuously swap it's 30MB of virtual. My\n> > disk is chattering itself to death. It stayed that way for 40 minutes.\n> >\n> > I'm practicing on the kernel tree.\n>\n> You hope for miracles, do you?  ;-)\n\nWe're going something wrong in git-daemon. I can clone the tree in\nfive minutes using the http protocol. Using the git protocol would\ntake 24hrs if I let it finish.\n\n\n> Please stop hammering that poor little NSLU2 with such a workset, or\n> hack some additional 224MB of RAM into it.  There is no magical\n> solution.\n>\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51476","messageId":"alpine.LFD.0.999.0708241749040.16727@xanadu.home","threadId":"9631","inReplyTo":"9e4733910708241417l44c55306xaa322afda69c6beb@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-08-24T21:54:02Z","receivedAt":"2007-08-24T21:54:02Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 24 Aug 2007, Jon Smirl wrote:\n\n> On 8/24/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Fri, 24 Aug 2007, Jon Smirl wrote:\n> >\n> > > I'm still trying to debug git-daemon\n> > >\n> > > I do find it surprising that git-index-pack can't be happy with in\n> > > 20MB of RAM and it has to continuously swap it's 30MB of virtual. My\n> > > disk is chattering itself to death. It stayed that way for 40 minutes.\n> > >\n> > > I'm practicing on the kernel tree.\n> >\n> > You hope for miracles, do you?  ;-)\n> \n> We're going something wrong in git-daemon. I can clone the tree in\n> five minutes using the http protocol. Using the git protocol would\n> take 24hrs if I let it finish.\n\nThe http protocol is merely only a dumb file copy with no packing \noptimization what so ever.\n\nThe native protocol performs a whole more to provide clients with only \nthe minimum data needed.\n\nTry running \"git repack -a\" directly on the NSLU2.  You should have the \nsame performance problems as with a clone.\n\n\nNicolas\n"},{"id":"51477","messageId":"9e4733910708241506h6eecc11ge41b1dc313022b4b@mail.gmail.com","threadId":"9631","inReplyTo":"9e4733910708241417l44c55306xaa322afda69c6beb@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-24T22:06:22Z","receivedAt":"2007-08-24T22:06:22Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/24/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> We're going something wrong in git-daemon. I can clone the tree in\n> five minutes using the http protocol. Using the git protocol would\n> take 24hrs if I let it finish.\n\n20Mb/s to kernel.org\ntime git clone git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git\nreal    2m34.629s\n\n20Mb/s to kernel.org\ntime git clone http://www.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git\nreal    3m52.203s\n\nSame kernel from my NSLU2 over http (100Mb/s)\ntime git clone http://jonsmirl.is-a-geek.net/apache2-default/mpc.git\nreal    2m36.227s\n\nUsing git protocol to nslu2 takes 24hrs\n\nOn 8/24/07, Nicolas Pitre <nico@cam.org> wrote:\n> Try running \"git repack -a\" directly on the NSLU2.  You should have the\n> same performance problems as with a clone.\n\nThis is true, it would take over 24hrs to finish.\n\nIs their a reason why initial clone hasn't been special cased? Why\ncan't initial clone just blast over the pack file already sitting on\nthe disk?\n\nI also wonder if a little application of some sorting to in-memory\ndata structures could help with the random IO patterns. I'm getting\nthe same data out of a stupid HTTP server and it doesn't go all IO\nbound on me so a solution has to be possible.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51479","messageId":"fanmmk$f5q$1@sea.gmane.org","threadId":"9631","inReplyTo":"9e4733910708241506h6eecc11ge41b1dc313022b4b@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-08-24T22:39:16Z","receivedAt":"2007-08-24T22:39:16Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jon Smirl wrote:\n> On 8/24/07, Nicolas Pitre <nico@cam.org> wrote:\n\n>> Try running \"git repack -a\" directly on the NSLU2.  You should have the\n>> same performance problems as with a clone.\n> \n> This is true, it would take over 24hrs to finish.\n> \n> Is their a reason why initial clone hasn't been special cased? Why\n> can't initial clone just blast over the pack file already sitting on\n> the disk?\n\nThere was idea to special case clone (just concatenate the packs, the\nreceiving side as someone told there can detect pack boundaries; do not\nforget to pack loose objects, first), instead of using generic fetch --all\nfor clone, bnut no code. Code speaks louder than words (although if someone\nwould provide details of pack boundary detection...)\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"51481","messageId":"7vbqcwcze3.fsf@gitster.siamese.dyndns.org","threadId":"9631","inReplyTo":"fanmmk$f5q$1@sea.gmane.org","subject":"Re: git-daemon on NSLU2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-08-24T22:59:32Z","receivedAt":"2007-08-24T22:59:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> There was idea to special case clone (just concatenate the packs, the\n> receiving side as someone told there can detect pack boundaries; do not\n> forget to pack loose objects, first), instead of using generic fetch --all\n> for clone, bnut no code. Code speaks louder than words (although if someone\n> would provide details of pack boundary detection...)\n\nI have to say that \"although ...\" part of that statement\ndisqualifies this to be called an \"idea\".\n\nReally, I find that you (yes, in this case I am not generalizing\nbut talking specifically about you) tend to overuse the word\n\"idea\" when you talk things that are not yet even at that stage\nyet.\n"},{"id":"51483","messageId":"200708250121.33924.jnareb@gmail.com","threadId":"9631","inReplyTo":"7vbqcwcze3.fsf@gitster.siamese.dyndns.org","subject":"Re: git-daemon on NSLU2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-08-24T23:21:33Z","receivedAt":"2007-08-24T23:21:33Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>> There was idea to special case clone (just concatenate the packs, the\n>> receiving side as someone told there can detect pack boundaries; do not\n>> forget to pack loose objects, first), instead of using generic fetch --all\n>> for clone, bnut no code. Code speaks louder than words (although if someone\n>> would provide details of pack boundary detection...)\n> \n> I have to say that \"although ...\" part of that statement\n> disqualifies this to be called an \"idea\".\n\nErmm... if I remember correctly during discussion (single subthread)\nthere were provided details, or at least idea, of how to separate\nconcatented packs into individual packs. Unfortunately I haven't\nsaved the message, and do not remember enogh of it to search archives...\n\nI should have wrote \"remind\" instead of \"provide\" there...\n\n> Really, I find that you (yes, in this case I am not generalizing\n> but talking specifically about you) tend to overuse the word\n> \"idea\" when you talk things that are not yet even at that stage\n> yet.\n\nI'm not native English speaker... ;-)\n\nSeriously, it's a fault of mine...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"51484","messageId":"alpine.LFD.0.999.0708241616390.25853@woody.linux-foundation.org","threadId":"9631","inReplyTo":"9e4733910708241417l44c55306xaa322afda69c6beb@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-24T23:28:58Z","receivedAt":"2007-08-24T23:28:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 24 Aug 2007, Jon Smirl wrote:\n> \n> We're going something wrong in git-daemon.\n\nNope.\n\nOr rather, it's mostly by design.\n\n> I can clone the tree in five minutes using the http protocol. Using the \n> git protocol would take 24hrs if I let it finish.\n\nThe http side doesn't actually do any global verification, the way \ngit-daemon does. So to it, everything is just temporary buffers, and you \ndon't need any memory at all, really.\n\ngit-daemon will create a packfile. That means that it has to generate the \n*global* object reachability, and will then optimize the object packing \netc etc. That's a minimum of something like 48 bytes per object for just \nthe object chains, and the kernel has a *lot* of objects (over half a \nmillion).\n\nIn addition to the object chains yourself, the native protocol will also \nobviously have to actually *look* at and parse all the tree and commit \nobjects while it does all this, so while it doesn't necessarily keep all \nof those in memory all the time, it will need to access them, and if you \ndon't have enough memory to cache them, that will add its own set of IO.\n\nSo I haven't checked exactly how much memory you really want to have to \nserve big projects, but with some handwavy guesstimate, if you actually \nwant to do a good job I'd guess that you really want to have at least as \nmuch memory as the size of largest project you are serving, and probably \nadd at least 10-20% on top of that.\n\nSo for the kernel, at a guess, you'd probably want to have at least 256MB \nof RAM to do a half-way good job. 512MB is likely nicer and allows you to \nactually cache the stuff over multiple accesses.\n\nBut I haven't actually tested. Maybe it might be bearable at 128M.\n\n\t\t\tLinus\n"},{"id":"51485","messageId":"9e4733910708241646x7b285574t94c3d7eb32bb60c9@mail.gmail.com","threadId":"9631","inReplyTo":"fanmmk$f5q$1@sea.gmane.org","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-24T23:46:50Z","receivedAt":"2007-08-24T23:46:50Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/24/07, Jakub Narebski <jnareb@gmail.com> wrote:\n> There was idea to special case clone (just concatenate the packs, the\n> receiving side as someone told there can detect pack boundaries; do not\n> forget to pack loose objects, first), instead of using generic fetch --all\n> for clone, bnut no code. Code speaks louder than words (although if someone\n> would provide details of pack boundary detection...)\n\nA related concept, initial clone of a repository does the equivalent\nof repack -a on the repo before transmitting it. Why aren't we saving\nthose results by switching the repo onto the new pack file? Then the\nnext clone that comes along won't have to do anything but send the\nfile.\n\nBut this logic can be flipped around, if the remote needs any object\nfrom the pack file, just send them the whole pack file and let the\nremote sort it out. Using this logic you can still minimize the IO\nstatistically.\n\nWhen a remote does a fetch you have to pack all of the loose objects.\nWhen the loose object pile reaches 20MB or so, the fetch can trigger a\nrepack of the oldest half into a pack that is kept by the tree and\nreplaces those older loose objects. For future fetches simply apply\nthe rule of sending the whole pack if any object is needed.\n\nThe repack of the 10MB of older objects can be kicked out to another\nprocess and copied into the tree when it is finished. At that point\nthe loose objects can be deleted. The git db can tolerate a process\ncopying in a new packfile and deleting the old objects while other\nprocesses may be using the database, right?\n\nThis model shouldn't statistically change the amount of data very\nmuch. If you haven't synced your tree in a month a few too many\nobjects may get sent to you. However, it should dramatically reduce\nthe IO load on the server cause by git protocol initial clones.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51486","messageId":"7v1wdscwd4.fsf@gitster.siamese.dyndns.org","threadId":"9631","inReplyTo":"9e4733910708241646x7b285574t94c3d7eb32bb60c9@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-08-25T00:04:55Z","receivedAt":"2007-08-25T00:04:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Jon Smirl\" <jonsmirl@gmail.com> writes:\n\n> On 8/24/07, Jakub Narebski <jnareb@gmail.com> wrote:\n>> There was idea to special case clone (just concatenate the packs, the\n>> receiving side as someone told there can detect pack boundaries; do not\n>> forget to pack loose objects, first), instead of using generic fetch --all\n>> for clone, bnut no code. Code speaks louder than words (although if someone\n>> would provide details of pack boundary detection...)\n>\n> A related concept, initial clone of a repository does the equivalent\n> of repack -a on the repo before transmitting it. Why aren't we saving\n> those results by switching the repo onto the new pack file? Then the\n> next clone that comes along won't have to do anything but send the\n> file.\n\nIf the majority of the access to your repository is the initial\nclone request, then it might be a worthwhile thing to do.  In\nfact didn't we use to have such a \"pre-prepared pack\" support?\n\nBut I do not think \"majority is initial clone\" is the norm.\nEven among the people who does an \"initial clone\" (from the\nend-user perspective), what they do may not be the initial full\nclone your special hack helps (and that was one of the reasons\nwe dropped the pre-prepared pack support --- \"been there, done\nthat\" to some extent).\n\n - If your client \"clone\"s only a single branch by doing:\n\n\t$ git init\n\t$ git remote add origin $remote_url\n        $ git pull origin master\n\n   the set of objects you need to send would be different\n   (slightly smaller) than the normal clone.\n\n - Another example would be a client that uses --reference:\n\n\t$ git clone --reference neigh.git git://yourbox/repo.git\n\n   which would give you a request that is different from the\n   usual initial full clone request.\n"},{"id":"51487","messageId":"alpine.LFD.0.999.0708241951100.16727@xanadu.home","threadId":"9631","inReplyTo":"9e4733910708241506h6eecc11ge41b1dc313022b4b@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-08-25T00:10:46Z","receivedAt":"2007-08-25T00:10:46Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 24 Aug 2007, Jon Smirl wrote:\n\n> On 8/24/07, Nicolas Pitre <nico@cam.org> wrote:\n> > Try running \"git repack -a\" directly on the NSLU2.  You should have the\n> > same performance problems as with a clone.\n> \n> This is true, it would take over 24hrs to finish.\n> \n> Is their a reason why initial clone hasn't been special cased? Why\n> can't initial clone just blast over the pack file already sitting on\n> the disk?\n\nWhat is the gain?  You'll get back to the same performance problem \neventually with some fetch operation, unless you intend to serve clients \nwith the whole pack everytime just like the http protocol does.\n\nAlso you don't want people cloning from you getting stuff that sits in \nyour reflog.  The native protocol makes sure that only the needed \nobjects are sent over and no more.\n\n> I also wonder if a little application of some sorting to in-memory\n> data structures could help with the random IO patterns. I'm getting\n> the same data out of a stupid HTTP server and it doesn't go all IO\n> bound on me so a solution has to be possible.\n\nThe http application is, indeed, stupid.  It performs no reachability \nanalysis, no repacking, no nothing except copying the bits over.\n\nAnd yes I did add some sorting optimizations in this round, so if you \ntry 2.5.3-* you should have them.  But there is a limit to what can be \ndone.\n\nPoint is, if you want serious Git serving, and not only _dumb_ protocols \n(http is one of them) then you need more RAM.  The NSLU2 is cool, but \nmaybe not appropriate for serving the Linux kernel natively with Git.\n\n\nNicolas\n"},{"id":"51494","messageId":"85zm0gdr5j.fsf@lola.goethe.zz","threadId":"9631","inReplyTo":"7v1wdscwd4.fsf@gitster.siamese.dyndns.org","subject":"Re: git-daemon on NSLU2","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-25T07:12:08Z","receivedAt":"2007-08-25T07:12:08Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Jon Smirl\" <jonsmirl@gmail.com> writes:\n>\n>> On 8/24/07, Jakub Narebski <jnareb@gmail.com> wrote:\n>>> There was idea to special case clone (just concatenate the packs, the\n>>> receiving side as someone told there can detect pack boundaries; do not\n>>> forget to pack loose objects, first), instead of using generic fetch --all\n>>> for clone, bnut no code. Code speaks louder than words (although if someone\n>>> would provide details of pack boundary detection...)\n>>\n>> A related concept, initial clone of a repository does the equivalent\n>> of repack -a on the repo before transmitting it. Why aren't we saving\n>> those results by switching the repo onto the new pack file? Then the\n>> next clone that comes along won't have to do anything but send the\n>> file.\n>\n> If the majority of the access to your repository is the initial\n> clone request, then it might be a worthwhile thing to do.  In fact\n> didn't we use to have such a \"pre-prepared pack\" support?\n>\n> But I do not think \"majority is initial clone\" is the norm.\n\nWell, as long as the majority is not affected negatively, catering for\na minority better is a strict improvement.  Most repositories will\nnever get cloned and won't be affected.  But there are some\nrepositories with a non-trivial amount of cloning.\n\n> Even among the people who does an \"initial clone\" (from the\n> end-user perspective), what they do may not be the initial full\n> clone your special hack helps (and that was one of the reasons\n> we dropped the pre-prepared pack support --- \"been there, done\n> that\" to some extent).\n\nIf it doesn't get used, its presence does no harm, of course except\nfrom having to be maintained and tested.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"51514","messageId":"9e4733910708250844n7074cb8coa5844fa6c46b40f0@mail.gmail.com","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708241616390.25853@woody.linux-foundation.org","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-25T15:44:07Z","receivedAt":"2007-08-25T15:44:07Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/24/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > I can clone the tree in five minutes using the http protocol. Using the\n> > git protocol would take 24hrs if I let it finish.\n>\n> The http side doesn't actually do any global verification, the way\n> git-daemon does. So to it, everything is just temporary buffers, and you\n> don't need any memory at all, really.\n>\n> git-daemon will create a packfile. That means that it has to generate the\n> *global* object reachability, and will then optimize the object packing\n> etc etc. That's a minimum of something like 48 bytes per object for just\n> the object chains, and the kernel has a *lot* of objects (over half a\n> million).\n\nA large, repeating work load is created in this process when you take\na 200MB pack, repack it to add a few loose objects and then don't save\nthe results. This model makes the NSLU2 unusable, but I also see it at\nmy shared hosting provider. Initial clones of a repo that take 3min\nfrom kernel.org take 25min on a shared host since the RAM is not\ndedicated.\n\nThere are three categories of fetches:\n1) initial clone, fetch all\n2) fetch recent\n3) I haven't fetched in three months\n\n99% of fetches fall in the first two categories.\n\nA very simple solution is to sendfile() existing packs if they contain\nany objects that the client wants and let the client deal with the\nunwanted objects. Yes this does send extra traffic over the net, but\nthe only group significantly impacted is #2 which is the most\ninfrequent group.\n\nLoose objects are handled as they are currently. To optimize this\nscheme you need to let the loose objects build up at the server and\nthen periodically sweep only the older ones into a pack. Packing the\nentire repo into a single pack would cause recent fetches to retrieve\nthe entire pack.\n\nInitial clone can be optimized further by recognizing that the\nreceiving repository is empty and sending them everything; no need to\ncompute which objects are missing at the server. This method will\nspeed up initial clone since the existing pack can be immediately sent\ninstead of waiting on a pack file to be built. Build the loose object\npack in parallel with sending the existing packs.\n\nI recognize that in the case of cloning a single branch or --reference\ntoo many objects will also be transmitted but I believe the benefits\nof reducing the server load outweigh the overhead of transmitting\nextra objects in this case. You can always remove the extra objects on\nthe client side.\n\nOn 8/24/07, Jakub Narebski <jnareb@gmail.com> wrote:\n> There was idea to special case clone (just concatenate the packs, the\n> receiving side as someone told there can detect pack boundaries; do not\n> forget to pack loose objects, first), instead of using generic fetch --all\n> for clone, bnut no code. Code speaks louder than words (although if someone\n> would provide details of pack boundary detection...)\n\nWrite the file name and length into the socket before sending the\npack. Use sendfile() or it's current incarnation to actually send the\npack. Insert these header lines between packs.\n\n> In addition to the object chains yourself, the native protocol will also\n> obviously have to actually *look* at and parse all the tree and commit\n> objects while it does all this, so while it doesn't necessarily keep all\n> of those in memory all the time, it will need to access them, and if you\n> don't have enough memory to cache them, that will add its own set of IO.\n>\n> So I haven't checked exactly how much memory you really want to have to\n> serve big projects, but with some handwavy guesstimate, if you actually\n> want to do a good job I'd guess that you really want to have at least as\n> much memory as the size of largest project you are serving, and probably\n> add at least 10-20% on top of that.\n>\n> So for the kernel, at a guess, you'd probably want to have at least 256MB\n> of RAM to do a half-way good job. 512MB is likely nicer and allows you to\n> actually cache the stuff over multiple accesses.\n>\n> But I haven't actually tested. Maybe it might be bearable at 128M.\n>\n>                         Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51515","messageId":"fapnd0$rpp$1@sea.gmane.org","threadId":"9631","inReplyTo":"7v1wdscwd4.fsf@gitster.siamese.dyndns.org","subject":"Re: git-daemon on NSLU2","fromName":"Salikh Zakirov","fromEmail":"salikh@gmail.com","sentAt":"2007-08-25T17:02:57Z","receivedAt":"2007-08-25T17:02:57Z","isPatch":false,"sender":{"key":"salikh@gmail.com","avatar":"https://gravatar.com/avatar/952c102bb1dcf721dab8de4f5a11d276756a65d301d021f755e265cc3251efae?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> But I do not think \"majority is initial clone\" is the norm.\n> Even among the people who does an \"initial clone\" (from the\n> end-user perspective), what they do may not be the initial full\n> clone your special hack helps (and that was one of the reasons\n> we dropped the pre-prepared pack support --- \"been there, done\n> that\" to some extent).\n\nFWIW, on my previous job release engineering team used git\nin a special way involving lots of initial clones. \n\nThe project itself was kept under SVN, and several machines\nwere doing continuous builds, starting from scratch.\nUnfortunately, doing from scratch checkouts from SVN was not\nan option because of high SVN checkout overhead, and machines\ndid a git-clone of imported repository instead.\n\nObviously using --reference would have saved even more on initial clone,\nbut the release team consisting of a pregnant woman and an\nintern student had neither time nor inclination to learn\ngit any deeper than were strictly necessary to get the job done.\nApparently, pure git-clone performance was good enough.\n"},{"id":"51552","messageId":"20070826093331.GC30474@coredump.intra.peff.net","threadId":"9631","inReplyTo":"9e4733910708250844n7074cb8coa5844fa6c46b40f0@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-08-26T09:33:31Z","receivedAt":"2007-08-26T09:33:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Aug 25, 2007 at 11:44:07AM -0400, Jon Smirl wrote:\n\n> A very simple solution is to sendfile() existing packs if they contain\n> any objects that the client wants and let the client deal with the\n> unwanted objects. Yes this does send extra traffic over the net, but\n> the only group significantly impacted is #2 which is the most\n> infrequent group.\n>\n> Loose objects are handled as they are currently. To optimize this\n> scheme you need to let the loose objects build up at the server and\n> then periodically sweep only the older ones into a pack. Packing the\n> entire repo into a single pack would cause recent fetches to retrieve\n> the entire pack.\n\nI was about to write \"but then 'fetch recent' clients will have to get\nthe entire repo after the upstream does a 'git-repack -a -d'\" but you\nseem to have figured that out already.\n\nI'm unclear: are you proposing new behavior for git-daemon in general,\nor a special mode for resource-constrained servers? If general behavior,\nare you suggesting that we never use 'git-repack -a' on repos which\nmight be cloned?\n\n-Peff\n"},{"id":"51574","messageId":"9e4733910708260934i1381e73ftb31c7de0d23f6cae@mail.gmail.com","threadId":"9631","inReplyTo":"20070826093331.GC30474@coredump.intra.peff.net","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-26T16:34:29Z","receivedAt":"2007-08-26T16:34:29Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/26/07, Jeff King <peff@peff.net> wrote:\n> On Sat, Aug 25, 2007 at 11:44:07AM -0400, Jon Smirl wrote:\n>\n> > A very simple solution is to sendfile() existing packs if they contain\n> > any objects that the client wants and let the client deal with the\n> > unwanted objects. Yes this does send extra traffic over the net, but\n> > the only group significantly impacted is #2 which is the most\n> > infrequent group.\n> >\n> > Loose objects are handled as they are currently. To optimize this\n> > scheme you need to let the loose objects build up at the server and\n> > then periodically sweep only the older ones into a pack. Packing the\n> > entire repo into a single pack would cause recent fetches to retrieve\n> > the entire pack.\n>\n> I was about to write \"but then 'fetch recent' clients will have to get\n> the entire repo after the upstream does a 'git-repack -a -d'\" but you\n> seem to have figured that out already.\n>\n> I'm unclear: are you proposing new behavior for git-daemon in general,\n> or a special mode for resource-constrained servers? If general behavior,\n> are you suggesting that we never use 'git-repack -a' on repos which\n> might be cloned?\n\nThis would be a new general behavior. There are cases where git-daemon\nis very resource hungry, rearranging things a little can remove this\nneed for everyone.\n\nThere are several ways to address the repack -a problem. But the\nsimplest solution may be the best, send existing packs only on an\ninitial clone. In all other cases continue with the current algorithm.\nWe could work on methods for making the middle case better but it is\nso infrequent it is probably not worth bothering with.\n\nChanging git-daemon only for the initial clone case also means that\npeople don't need to change the way they manage packs.\n\nPosters have been saying, why worry about initial clone since it isn't\ndone that often. I agree that it isn't done that often, but if it is\ndone all on my NSLU2 it will take about 40hrs to complete. We can\neasily see the impact of changing the the initial clone algorithm, the\nhttp clone takes 3min.\n\nBTW, if the NSLU2 needs a repack -a I can do it on another machine and\ncopy it over. Or maybe someone will write a repack that is happy in\n20MB. The NSLU2 is a great home server, it is usually fast enough.\nPower consumption is a tiny 8W, fine to leave on 24/7, My NSLU2 is as\npowerful as the average desktop machine in the early 90's, how quickly\nwe forget.\n\n\n>\n> -Peff\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51575","messageId":"alpine.LFD.0.999.0708260959050.25853@woody.linux-foundation.org","threadId":"9631","inReplyTo":"9e4733910708260934i1381e73ftb31c7de0d23f6cae@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-26T17:15:24Z","receivedAt":"2007-08-26T17:15:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 26 Aug 2007, Jon Smirl wrote:\n> \n> Changing git-daemon only for the initial clone case also means that\n> people don't need to change the way they manage packs.\n\nI do agree that we might want to do some special-case handling for the \ninitial clone (because it *is* kind of special), but it's not necessarily \nas easy as just re-using an existing pack.\n\nAt a minimum, we'd need to have something that knows how to make a single \npack out of several packs and some loose objects. That shouldn't be \n*hard*, but it's certainly nontrivial, especially in the presense of the \nsame objects possibly being available more than once in different packs.\n\n[ The \"duplicate object\" thing does actually happen: even if you use only \n  \"git native\" protocols, you can get duplicate objects because a file was \n  changed back to an earlier version. The incremental packs you get from \n  push/pull'ing between two repositories try to send the minimal \n  incremental changes, but the keyword here is _try_: they will \n  potentially send objects that the receiver already has, if it's not \n  obvious that the receiver has them from the \"commit boundary\" cases ]\n\nMaybe the client side will handle a pack with duplicate objects perfectly \nfine, and it's not an issue. Maybe. It might even be likely (I can't think \nof anything that would obviously break). But at a minimum, it would be \nsomething that needs some code on the sending side, and a lot of \nverification that the end result works ok on the receiving side.\n\nAnd there's actually a deeper problem: the current native protocol \nguarantees that the objects sent over are only those that are reachable. \nThat matters. It matters for subtle security issues (maybe you are \nexporting some repository that was rebased, and has objects that you \ndidn't *intend* to make public!), but it also matters for issues like git \n\"alternates\" files.\n\nIf you only ever look at a single repo, you'll never see the alternates \nissue, but if you're seriously looking at serving git repositories, I \ndon't really see the \"single repo\" case as being at all the most common or \ninteresting case. \n\nAnd if you look at something like kernel.org, the \"alternates\" thing is \n*much* more important than how much memory git-daemon uses! Yes, \nkernel.org would probably be much happier if git-daemon wasn't such a \nmemory pig occasionally, but on the other hand, the win from using \nalternates and being able to share 99% of all objects in all the various \nrelated kernel repositories is actually likely to be a *bigger* memory win \nthan any git-daemon memory usage, because now the disk caching works a \nhell of a lot better!\n\nSo it's not actually clear how the initial clone thing can be optimized on \nthe server side.\n\nIt's easier to optimize on the *client* side: just do the initial clone \nwith rsync/http (and \"git gc\" it on the client afterwards), and then \nchange it to the git native protocol after the clone.\n\nThat may not sound very user-friendly, but let's face it, I think there is \nexactly one person in the whole universe that tries to use an NSLU2 as a \ngit server. So the \"client-side workaround\" is likely to affect a very \nlimited number of clients ;)\n\n\t\tLinus\n"},{"id":"51577","messageId":"9e4733910708261106u3fecde67m8045ddba3aa57650@mail.gmail.com","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708260959050.25853@woody.linux-foundation.org","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-26T18:06:25Z","receivedAt":"2007-08-26T18:06:25Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/26/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> And there's actually a deeper problem: the current native protocol\n> guarantees that the objects sent over are only those that are reachable.\n> That matters. It matters for subtle security issues (maybe you are\n> exporting some repository that was rebased, and has objects that you\n> didn't *intend* to make public!), but it also matters for issues like git\n> \"alternates\" files.\n\nAre these objects visible through the other protocols? It seems\ndangerous to leave something on an open server that you want to keep\nhidden.\n\n> If you only ever look at a single repo, you'll never see the alternates\n> issue, but if you're seriously looking at serving git repositories, I\n> don't really see the \"single repo\" case as being at all the most common or\n> interesting case.\n>\n> And if you look at something like kernel.org, the \"alternates\" thing is\n> *much* more important than how much memory git-daemon uses! Yes,\n> kernel.org would probably be much happier if git-daemon wasn't such a\n> memory pig occasionally, but on the other hand, the win from using\n> alternates and being able to share 99% of all objects in all the various\n> related kernel repositories is actually likely to be a *bigger* memory win\n> than any git-daemon memory usage, because now the disk caching works a\n> hell of a lot better!\n\nDoesn't kernel.org use alternates or something equivalent for serving\nup all those nearly identical kernel trees?\n\nI've been handling the problem locally by using remotes and fetching\nall the repos I'm interested in into a single git db.\n\n>\n> So it's not actually clear how the initial clone thing can be optimized on\n> the server side.\n>\n> It's easier to optimize on the *client* side: just do the initial clone\n> with rsync/http (and \"git gc\" it on the client afterwards), and then\n> change it to the git native protocol after the clone.\n\nEven better, get them to clone from kernel.org and then just fetch in\nthe differences from my server. It's an educational problem.\n\nHow about changing initial clone to refuse to use the git protocol?\n\n>\n> That may not sound very user-friendly, but let's face it, I think there is\n> exactly one person in the whole universe that tries to use an NSLU2 as a\n> git server. So the \"client-side workaround\" is likely to affect a very\n> limited number of clients ;)\n\nI'll send you one and double the size of the user base. I have this\nfancy new 20Mb FIOS connection and I can't come up with anything to\nuse the bandwidth on.\n\nAnyway, I already gave up and moved on to a hosting provider. Repo is\nhere: http://git.digispeaker.com/ There's nothing there yet but a\nclone of the 2.6 tree.\nI don't  think there is a solution for running a git daemon on a shared host.\n\nPetr pointed out to me that an NSLU2 is late 90's equivalent not early\nso my memory if faulty too.\n\n\n>\n>                 Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51578","messageId":"alpine.LFD.0.999.0708261118260.25853@woody.linux-foundation.org","threadId":"9631","inReplyTo":"9e4733910708261106u3fecde67m8045ddba3aa57650@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-26T18:26:07Z","receivedAt":"2007-08-26T18:26:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 26 Aug 2007, Jon Smirl wrote:\n>\n> On 8/26/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > And there's actually a deeper problem: the current native protocol\n> > guarantees that the objects sent over are only those that are reachable.\n> > That matters. It matters for subtle security issues (maybe you are\n> > exporting some repository that was rebased, and has objects that you\n> > didn't *intend* to make public!), but it also matters for issues like git\n> > \"alternates\" files.\n> \n> Are these objects visible through the other protocols? It seems\n> dangerous to leave something on an open server that you want to keep\n> hidden.\n\nThey'd be visible to any stupid walker, yes. But if you're \nsecurity-conscious, you'd simply not *allow* any stupid walkers.\n\nOne of the goals of \"git-daemon\" was to have a simple service that was \n\"obviously secure\". Now, it's debatable just how obvious the daemon is, \nbut it really is pretty simple, and I do think it should be possible to \nalmost statically validate that it only ever reads files, and that it will \nonly ever read files that act like valid *git* data. \n\nSome people may care about that kind of thing. I don't know how many,  but \nit really was one of the design criteria (which is why, for example, git \ndaemon will just silently close the connection if it finds something \nfishy: no fishing expeditions with bad clients trying to figure out what \nfiles exist on a server allowed!).\n\nSo the fact that a web server or rsync will expose everything is kind of \nirrelevant - those are *designed* to expose everything. git-daemon was \ndesigned *not* to do that.\n\n> Doesn't kernel.org use alternates or something equivalent for serving\n> up all those nearly identical kernel trees?\n\nAbsolutely. And that's the point. \"git-daemon\" will serve a nice \nindividualized pack, even though any particular repository doesn't have \none, but is really a combination of \"the base Linus pack + extensions\".\n\n> > So it's not actually clear how the initial clone thing can be optimized on\n> > the server side.\n> >\n> > It's easier to optimize on the *client* side: just do the initial clone\n> > with rsync/http (and \"git gc\" it on the client afterwards), and then\n> > change it to the git native protocol after the clone.\n> \n> Even better, get them to clone from kernel.org and then just fetch in\n> the differences from my server. It's an educational problem.\n\nYes. \n\n> How about changing initial clone to refuse to use the git protocol?\n\nAbsolutely not. It's quite often the best one to use (the ssh protocol \nhas the exact same issues, and is the only secure protocol).\n\nBut on a SNLU2, maybe *you* want to make your server side refuse it? I \nwould be easy enough: if the client doesn't report any existing SHA1's, \nyou just say \"I'm not going to work with you\".\n\n\t\t\tLinus\n"},{"id":"51580","messageId":"9e4733910708261200m5e4c3019g490ffc29b171ef08@mail.gmail.com","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708261118260.25853@woody.linux-foundation.org","subject":"Re: git-daemon on NSLU2","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-26T19:00:18Z","receivedAt":"2007-08-26T19:00:18Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/26/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > Doesn't kernel.org use alternates or something equivalent for serving\n> > up all those nearly identical kernel trees?\n>\n> Absolutely. And that's the point. \"git-daemon\" will serve a nice\n> individualized pack, even though any particular repository doesn't have\n> one, but is really a combination of \"the base Linus pack + extensions\".\n\nA really simple change to the git protocol would be to make the client\nloop on the request. On the first request the server would see that\nthe client has no objects and send the \"base Linus pack\". The client\nwould then loop around and repeat the process which will trigger the\ncurrent pack building process.\n\nDo pack files contain enough information about the heads of the object\nchains for this to work? The client needs to be able to determine it's\nstate after receiving the pack and send the info back in the next\nround.\n\nI'm not buying the security argument. If you want something kept\nhidden get it out of the public db. If I know the sha of the hidden\nobject can't I just add a head for it and git-deamon will happily send\nit and the chain up to it to me?\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"51583","messageId":"alpine.LFD.0.999.0708261318140.25853@woody.linux-foundation.org","threadId":"9631","inReplyTo":"9e4733910708261200m5e4c3019g490ffc29b171ef08@mail.gmail.com","subject":"Re: git-daemon on NSLU2","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-26T20:19:31Z","receivedAt":"2007-08-26T20:19:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 26 Aug 2007, Jon Smirl wrote:\n> \n> A really simple change to the git protocol would be to make the client\n> loop on the request. On the first request the server would see that\n> the client has no objects and send the \"base Linus pack\". The client\n> would then loop around and repeat the process which will trigger the\n> current pack building process.\n\nJon, just give it up. The fact is, the git protocol works the right way \nalready.\n\n> I'm not buying the security argument. If you want something kept hidden \n> get it out of the public db. If I know the sha of the hidden object \n> can't I just add a head for it and git-deamon will happily send it and \n> the chain up to it to me?\n\nThat's a particularly idiotic statement.\n\nIf you know the SHA1, there can *by*definition* not be any hidden objects. \nThe SHA1 depends on the object chain.\n\n\t\tLinus\n"},{"id":"51605","messageId":"7vtzqmng8r.fsf@gitster.siamese.dyndns.org","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708261318140.25853@woody.linux-foundation.org","subject":"Re: git-daemon on NSLU2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-08-26T21:22:12Z","receivedAt":"2007-08-26T21:22:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n>> I'm not buying the security argument. If you want something kept hidden \n>> get it out of the public db. If I know the sha of the hidden object \n>> can't I just add a head for it and git-deamon will happily send it and \n>> the chain up to it to me?\n>\n> That's a particularly idiotic statement.\n>\n> If you know the SHA1, there can *by*definition* not be any hidden objects. \n> The SHA1 depends on the object chain.\n\nI think what we have is even stronger --- upload-pack does not\nallow asking for an arbitrary commit.  The requesting fetch-pack\nside needs to pick from what are offerred, and upload-pack makes\nsure of that.\n"},{"id":"51606","messageId":"20070826222448.GA24142@istic.org","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708260959050.25853@woody.linux-foundation.org","subject":"Re: git-daemon on NSLU2","fromName":"Daniel Hulme","fromEmail":"st@istic.org","sentAt":"2007-08-26T22:24:48Z","receivedAt":"2007-08-26T22:24:48Z","isPatch":false,"sender":{"key":"st@istic.org","avatar":null},"body":"On Sun, Aug 26, 2007 at 10:15:24AM -0700, Linus Torvalds wrote:\n> It's easier to optimize on the *client* side: just do the initial clone \n> with rsync/http (and \"git gc\" it on the client afterwards), and then \n> change it to the git native protocol after the clone.\n\nWhen I was working on Xen two years ago, they did the same thing with\ntheir Mercurial repository. They had a proper repo that handled all the\npush and fetch traffic, and a cron job would periodically pull from that\ninto a second repo. This second one was served by http. People were\nencouraged to download the seed repo and then do a fetch (from the main\none) immediately.\n\nI don't know whether they still do that, but in any case it shows your\nidea is not unprecedented.\n\n-- \nKanga  said to Roo,  \"Drink up  your milk  first, dear, and  talk after-\nwards.\" So Roo, who was drinking his milk, tried to say that he could do\nboth at once... and had to be  patted on the back  and dried for quite a\nlong time afterwards.                     A. A. Milne, 'Winnie-the-Pooh'\n"},{"id":"51617","messageId":"200708270214.28652.jnareb@gmail.com","threadId":"9631","inReplyTo":"20070826093331.GC30474@coredump.intra.peff.net","subject":"Re: git-daemon on NSLU2","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-08-27T00:14:28Z","receivedAt":"2007-08-27T00:14:28Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sun, Aug 26, 2007, Jeff King wrote:\n> On Sat, Aug 25, 2007 at 11:44:07AM -0400, Jon Smirl wrote:\n> \n>> A very simple solution is to sendfile() existing packs if they contain\n>> any objects that the client wants and let the client deal with the\n>> unwanted objects. Yes this does send extra traffic over the net, but\n>> the only group significantly impacted is #2 which is the most\n>> infrequent group.\n>>\n>> Loose objects are handled as they are currently. To optimize this\n>> scheme you need to let the loose objects build up at the server and\n>> then periodically sweep only the older ones into a pack. Packing the\n>> entire repo into a single pack would cause recent fetches to retrieve\n>> the entire pack.\n> \n> I was about to write \"but then 'fetch recent' clients will have to get\n> the entire repo after the upstream does a 'git-repack -a -d'\" but you\n> seem to have figured that out already.\n> \n> I'm unclear: are you proposing new behavior for git-daemon in general,\n> or a special mode for resource-constrained servers? If general behavior,\n> are you suggesting that we never use 'git-repack -a' on repos which\n> might be cloned?\n\nI think that \"reuse existing packs if sensible\" idea (instead of generating\nalways new pack) is a good one, even if at first limited to the clone case.\n\nThere are nevertheless a few complications.\n\n1. When discussing this idea on git mailing list some time ago somebody\nsaid that we don't need to implement \"multi pack\" extension (which was\nat the beginning in the design, to add later, if I understand correctly),\nit is enough to concatenate packs. The receiving side can then detect\nboundaries between packs and split them appropriately. But is a\nconcatenated a proper pack? If not, then we can send concatenation of\npacks only if the client (receiving side) understands it, and can split it;\nit means checking for protocol extension...\n\n2. How to detect that request is for a clone? git-clone is get all remote\nheads and fetch from just received heads. But because fecthing refs and\nfetching objects is separate, we cannot I think use this sequence for\ndetecting that we want a clone. We can use \"no haves\" as heuristic to\ndetect a clone request, but \"no haves\" occurs also for initial fetching of\nsingle branch (i.e. using: git-remote; git-fetch sequence instead of\ngit-clone).\n\n3. The problem with alternates mentioned by Linus is not much a problem,\nas we can simply consider packs from the alternate repository/repositories.\nFor example if we use single alternate, we would send concatenation of\npacks from this repository, and from alternate (and pack of loose objects\nfrom this repository).\n\n\nWe would probably want to have some heuristic (besides configuring\ngit-daemon) to choose between reusing existing packs (and sending them\nconcatenated), and generating a pack for sending. Note that for dumb\ntransports we have the opposite problem and opposite idea: we always\nsend full packs for dumb transports; the idea was to use range downloading\n(available at least for http and ftp protocols) to download only needed\nfragments of packs. Perhaps if some % of pack (number of objects in the\npack or size of pack) is to be send then we reuse the pack, and remove\nobjects in the pack from consideration. No idea of how to implement that,\nthough. Or if number of objects in pack to be send crosses some threshold,\nor generating pack/doing reachability analysis takes to loong, then reuse\nexisting packs.\n\nOr you can wait fro the GitTorrent protocol to be implemented, or implement\nit yourself... ;-)\n\n-- \nJakub Narebski\nPoland\n"},{"id":"51683","messageId":"20070827110318.GB4680@thunk.org","threadId":"9631","inReplyTo":"alpine.LFD.0.999.0708261118260.25853@woody.linux-foundation.org","subject":"Re: git-daemon on NSLU2","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-08-27T11:03:18Z","receivedAt":"2007-08-27T11:03:18Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sun, Aug 26, 2007 at 11:26:07AM -0700, Linus Torvalds wrote:\n> > How about changing initial clone to refuse to use the git protocol?\n> \n> Absolutely not. It's quite often the best one to use (the ssh protocol \n> has the exact same issues, and is the only secure protocol).\n> \n> But on a SNLU2, maybe *you* want to make your server side refuse it? I \n> would be easy enough: if the client doesn't report any existing SHA1's, \n> you just say \"I'm not going to work with you\".\n\nWhat if the server sends a message which current clients interprets as\nan error, and which newer clients could interpret as, \"do a clone from\n<this> URL, and then come back and talk to me\".  Basically an\nautomated redirect to get the \"Linus base pack\" somewhere else, and\nthen to go back to the original server.  It certainly doesn't make\nsense to change anything about the low-level protocol, but maybe a\nhigher level redirect would make sense, just as a user convenience thing.\n\n       \t     \t      \t    \t \t- Ted\n"},{"id":"51728","messageId":"alpine.LFD.0.999.0708270923590.25853@woody.linux-foundation.org","threadId":"9631","inReplyTo":"20070827110318.GB4680@thunk.org","subject":"Re: git-daemon on NSLU2","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-27T16:26:23Z","receivedAt":"2007-08-27T16:26:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 27 Aug 2007, Theodore Tso wrote:\n> \n> What if the server sends a message which current clients interprets as\n> an error, and which newer clients could interpret as, \"do a clone from\n> <this> URL, and then come back and talk to me\".  Basically an\n> automated redirect to get the \"Linus base pack\" somewhere else, and\n> then to go back to the original server.  It certainly doesn't make\n> sense to change anything about the low-level protocol, but maybe a\n> higher level redirect would make sense, just as a user convenience thing.\n\nI agree, a redirect might be a good idea regardless of whether it's \nsomething like \"I'm a poor little NSLU2, please don't do anything but \nincremental updates\", or whether it's something like \"this repository has \nmoved, use address xyz instead\".\n\nAnd it should be pretty easy from a high-level protocol, although it does \nobviously need both server and client support.\n\n\t\t\tLinus\n"}]}