{"thread":{"id":"2669","subject":"Re: Linux 2.6.15-rc2","startedAt":"2005-11-24T12:37:57Z","lastAt":"2005-11-25T08:42:16Z","messageCount":10,"participants":["Ed Tomlinson","Andreas Ericsson","Linus Torvalds","Junio C Hamano","Nick Hengeveld"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"12696","messageId":"200511240737.59153.tomlins@cam.org","threadId":"2669","inReplyTo":"Pine.LNX.4.64.0511191934210.8552@g5.osdl.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Ed Tomlinson","fromEmail":"tomlins@cam.org","sentAt":"2005-11-24T12:37:57Z","receivedAt":"2005-11-24T12:37:57Z","isPatch":false,"sender":{"key":"tomlins@cam.org","avatar":null},"body":"On Saturday 19 November 2005 22:40, Linus Torvalds wrote:\n> There it is (or will soon be - the tar-ball and patches are still \n> uploading, and mirroring can obviously take some time after that).\n\nSomething strange here.   After a cg-update, I had no tag for rc2.   Checking\nshowed no problems so I used cg-clone to get another copy of the repository.\nStill no rc2.\n\ned@grover:/usr/src/2.6$ cg-version\ncogito-0.16rc2 (73874dddeec2d0a8e5cd343eec762d98314def63)\ned@grover:/usr/src/2.6$ git --version\ngit version 0.99.9.GIT\n\ncg-clone http://www.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6\n\nIt looks to be the tag that is missing, gitk show commits after Nov 19.\n\nBoth git and cg were  updated just prior to the cg-update (~Nov 22 8pm EST).\n\nWhat is happening?\n\nTIA\nEd Tomlinson\n"},{"id":"12697","messageId":"4385BAFC.7070906@op5.se","threadId":"2669","inReplyTo":"200511240737.59153.tomlins@cam.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2005-11-24T13:07:08Z","receivedAt":"2005-11-24T13:07:08Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Ed Tomlinson wrote:\n> Something strange here.   After a cg-update, I had no tag for rc2.   Checking\n> showed no problems so I used cg-clone to get another copy of the repository.\n> Still no rc2.\n> \n> ed@grover:/usr/src/2.6$ cg-version\n> cogito-0.16rc2 (73874dddeec2d0a8e5cd343eec762d98314def63)\n> ed@grover:/usr/src/2.6$ git --version\n> git version 0.99.9.GIT\n> \n> cg-clone http://www.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6\n> \n\nThis happened a while ago to someone else too. Apparently the http \ntransport needs serverside help (git-update-server-info or some such \nmust be run on the remote side).\n\nUnless you're restricted by firewalls and other you could try\n\ngit clone \ngit://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6\n\nwhich works flawlessly for me although it takes quite some time to \ntransfer all the data.\n\nLinus, HPA: Are the packs cached on kernel.org? It seems to be at least \na minute before the transfers start.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"12699","messageId":"Pine.LNX.4.64.0511241020050.13959@g5.osdl.org","threadId":"2669","inReplyTo":"200511240737.59153.tomlins@cam.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-24T18:37:15Z","receivedAt":"2005-11-24T18:37:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 24 Nov 2005, Ed Tomlinson wrote:\n> \n> What is happening?\n\nThe http transport isn't very good for git, so git adds various special \nfiles to make it work at all. They need to be specially updated, and I \nhadn't done that.\n\nUsing the native git protocol through git://git.kernel.org/.. gets around \nit, as does using rsync. \n\nI just repacked and updated it now, so how http should work too, although \ninefficiently (because it will get a whole new pack - just one of the \ndisadvantages of the non-native protocols).\n\n\t\tLinus\n"},{"id":"12700","messageId":"Pine.LNX.4.64.0511241037400.13959@g5.osdl.org","threadId":"2669","inReplyTo":"4385BAFC.7070906@op5.se","subject":"Re: Linux 2.6.15-rc2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-24T18:44:14Z","receivedAt":"2005-11-24T18:44:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 24 Nov 2005, Andreas Ericsson wrote:\n> \n> git clone git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git 2.6\n> \n> which works flawlessly for me although it takes quite some time to transfer\n> all the data.\n\nThe initial clone is very expensive for the native git protocol: the \nprotocol is designed to scale well for incremental updates (ie you have a \n_huge_ repository that has changed just a bit, and the protocol should \nwork well for that), and that makes the initial clone quite expensive as \nit marshalls the whole damn repository into this nice packed format.\n\nSo it's often nicer (certainly on the remote server) to use \"rsync\" for \nthe initial clone, and then only after that start using the git protocol.\n\n(This is in no way really fundamental, and the server could cache the \npacks it generates for initial clones, but that isn't implemented yet, and \nprobably won't be for some times).\n\nOf course, especially if you're mostly bandwidth-constrained and the \nserver side is not under a big load, using the native git protocol may \nactually be faster anyway. Because it's always going to generate the \nnicest packing, while rsync:// will just use whatever packing that the \nserver happens to have at that point (but I do repack every few weeks, so \nrsync for the initial clone should never be horribly bad - and since I \njust repacked, it should get that \"perfect\" pack too).\n\n\t\tLinus\n"},{"id":"12702","messageId":"7v4q61suhi.fsf@assigned-by-dhcp.cox.net","threadId":"2669","inReplyTo":"Pine.LNX.4.64.0511241037400.13959@g5.osdl.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-24T19:42:33Z","receivedAt":"2005-11-24T19:42:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> (This is in no way really fundamental, and the server could cache the \n> packs it generates for initial clones, but that isn't implemented yet, and \n> probably won't be for some times).\n\nPerformance perceived by cloners is helped by\n\n    $ mkdir -p .git/pack-cache\n    $ git-rev-list --objects --all | git-pack-objects .git/pack-cache/pack\n\non the server side.  This exact example of preparing by the\nrepository maintainer is optimizing for a wrong case, and I do\nnot think it is worth doing in practice, but this will give you\nthe lower bound when server side cache is implemented to do it\non demand.\n"},{"id":"12703","messageId":"20051124195256.GR3968@reactrix.com","threadId":"2669","inReplyTo":"Pine.LNX.4.64.0511241020050.13959@g5.osdl.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2005-11-24T19:52:56Z","receivedAt":"2005-11-24T19:52:56Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Thu, Nov 24, 2005 at 10:37:15AM -0800, Linus Torvalds wrote:\n\n> I just repacked and updated it now, so how http should work too, although \n> inefficiently (because it will get a whole new pack - just one of the \n> disadvantages of the non-native protocols).\n\nThere's room to improve on that particular inefficiency.  The http\ncommit walker could use Range: headers to fetch loose objects directly\nfrom inside a pack if it didn't make sense to fetch the entire pack.\nFor this to work, pack fetches would need to be deferred until the\nentire tree had been walked, and the commit walker could decide whether\nto fetch the pack or loose objects based on the percentage of packed\nobjects it needed to fetch.  It would also need to fetch all\ntag/commit/tree objects using ranges to be able to fully walk the tree.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"12704","messageId":"Pine.LNX.4.64.0511241154340.13959@g5.osdl.org","threadId":"2669","inReplyTo":"7v4q61suhi.fsf@assigned-by-dhcp.cox.net","subject":"Re: Linux 2.6.15-rc2","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-24T19:57:20Z","receivedAt":"2005-11-24T19:57:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 24 Nov 2005, Junio C Hamano wrote:\n> \n> Performance perceived by cloners is helped by\n> \n>     $ mkdir -p .git/pack-cache\n>     $ git-rev-list --objects --all | git-pack-objects .git/pack-cache/pack\n\nThat really doesn't work very well. I push to that tree often several \ntimes a day, and you'd have to re-do the cache each time.\n\nSo it would be much better if git-pack-objects would just always cache its \noutput in .git/pack-cache - along with some logic to just get rid of old \nones regularly.\n\nSince git-pack-objects has to generate the pack _anyway_, it might as well \nsave it away when it does - so that if you have lots of people doing \nclones or pulling, you'd only need to run it once for a particular set of \nobjects, and you'd not have to do any extra (or unnecessary) maintenance.\n\n\t\tLinus\n"},{"id":"12705","messageId":"7vveyhpxmy.fsf@assigned-by-dhcp.cox.net","threadId":"2669","inReplyTo":"Pine.LNX.4.64.0511241154340.13959@g5.osdl.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-24T21:02:45Z","receivedAt":"2005-11-24T21:02:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Since git-pack-objects has to generate the pack _anyway_, it might as well \n> save it away when it does - so that if you have lots of people doing \n> clones or pulling, you'd only need to run it once for a particular set of \n> objects, and you'd not have to do any extra (or unnecessary) maintenance.\n\nCaching itself is relatively easy (just implement an equivalent\nof tee inside pack-objects ourselves).  More problematic is\npruning.  We could do it from cron based on atime _if_ the\nfilesystem is not mounted noatime but without arranging a\nreasonably way for automated pruning this would become a disk\nhog and extra maintenance burden, which is why I did not\nimplement the dynamic caching part in the initial round.\n\nSince git-daemon would be the primary user of pack-cache/, this\nimplies a repository writable by git-daemon user on public\nmachine (not master), which is an extra thing to note.\n"},{"id":"12718","messageId":"200511242151.00162.tomlins@cam.org","threadId":"2669","inReplyTo":"20051124195256.GR3968@reactrix.com","subject":"Re: Linux 2.6.15-rc2","fromName":"Ed Tomlinson","fromEmail":"tomlins@cam.org","sentAt":"2005-11-25T02:50:59Z","receivedAt":"2005-11-25T02:50:59Z","isPatch":false,"sender":{"key":"tomlins@cam.org","avatar":null},"body":"On Thursday 24 November 2005 14:52, Nick Hengeveld wrote:\n> On Thu, Nov 24, 2005 at 10:37:15AM -0800, Linus Torvalds wrote:\n> \n> > I just repacked and updated it now, so how http should work too, although \n> > inefficiently (because it will get a whole new pack - just one of the \n> > disadvantages of the non-native protocols).\n> \n> There's room to improve on that particular inefficiency.  The http\n> commit walker could use Range: headers to fetch loose objects directly\n> from inside a pack if it didn't make sense to fetch the entire pack.\n> For this to work, pack fetches would need to be deferred until the\n> entire tree had been walked, and the commit walker could decide whether\n> to fetch the pack or loose objects based on the percentage of packed\n> objects it needed to fetch.  It would also need to fetch all\n> tag/commit/tree objects using ranges to be able to fully walk the tree.\n\nAlternately, when creating a new archive the client could ask the server\nwhat protocols are active.  It could then use the best one for the clone and\nupdate the .git/origin files with the optimal one for incremental pulls.\n\nThoughts?\nEd Tomlinson\n"},{"id":"12719","messageId":"4386CE68.1020200@op5.se","threadId":"2669","inReplyTo":"200511242151.00162.tomlins@cam.org","subject":"Re: Linux 2.6.15-rc2","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2005-11-25T08:42:16Z","receivedAt":"2005-11-25T08:42:16Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Ed Tomlinson wrote:\n> On Thursday 24 November 2005 14:52, Nick Hengeveld wrote:\n> \n>>On Thu, Nov 24, 2005 at 10:37:15AM -0800, Linus Torvalds wrote:\n>>\n>>\n>>>I just repacked and updated it now, so how http should work too, although \n>>>inefficiently (because it will get a whole new pack - just one of the \n>>>disadvantages of the non-native protocols).\n>>\n>>There's room to improve on that particular inefficiency.  The http\n>>commit walker could use Range: headers to fetch loose objects directly\n>>from inside a pack if it didn't make sense to fetch the entire pack.\n>>For this to work, pack fetches would need to be deferred until the\n>>entire tree had been walked, and the commit walker could decide whether\n>>to fetch the pack or loose objects based on the percentage of packed\n>>objects it needed to fetch.  It would also need to fetch all\n>>tag/commit/tree objects using ranges to be able to fully walk the tree.\n> \n> \n> Alternately, when creating a new archive the client could ask the server\n> what protocols are active.  It could then use the best one for the clone and\n> update the .git/origin files with the optimal one for incremental pulls.\n> \n\nThis would only work with the git protocol, and since that's the fastest \nprotocol (theoretically that is, Pasky seems to have gotten other \nfigures but I'm not sure I believe those) it should really only ever \nreturn itself which wouldn't make much sense.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"}]}