{"thread":{"id":"26211","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","startedAt":"2011-01-06T02:29:57Z","lastAt":"2011-01-11T01:56:09Z","messageCount":22,"participants":["Zenaan Harkness","Shawn Pearce","Nicolas Pitre","Jeff King","Ilari Liusvaara","Nguyen Thai Ngoc Duy","John Wyzer","Sam Vilain","J.H."],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"159016","messageId":"AANLkTikv+L5Da7A5VM7BAgnue=m0O_-nHmHchJzfGxJa@mail.gmail.com","threadId":"26211","inReplyTo":null,"subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Zenaan Harkness","fromEmail":"zen@freedbms.net","sentAt":"2011-01-06T02:29:57Z","receivedAt":"2011-01-06T02:29:57Z","isPatch":false,"sender":{"key":"zen@freedbms.net","avatar":null},"body":"Bittorrent requires some stability around torrent files.\n\nCan packs be generated deterministically?\nIf not by two separate repos, what about by one particular repo?\n\nFor Linus' linux-2.6.git, that repo is considered 'canonical' by many.\n\nPack-torrents could be ~1MiB, ~10MiB, ~100Mib, ~1GiB, or as configured\nin a particular repo, which repo is the canonical location for\npack-torrents for all who consider that particular repo as canonical.\n\nPerhaps a heuristic/ algorithm: once ten 10MiB (sequentially\ngenerated) pack-torrents are floating around,\nthey could be simply concatenated to create a 100MiB pack-torrent,\nwith a deterministic name and SHA etc,\nso that all those 10MiB pack-torrent files that torrent clients have,\ncan be re-used and locally combined into the 100MiB torrent as needed,\non demand.\n\nSame for 100MiB -> 1GiB pack-torrents.\n\nIndividual extra commits:\nWhile \"small\" number of additional commits go into a repo, clients\nfall back to git-fetch, _after .\n\nIf Linus linus-2.6.git (currently configured \"canonical\" repo) goes\noffline, simply configure a new remote canonical repo.\n\nBranches:\nOther \"branches\" repos of linux-2.6.git could create their own\nconsistent 50MiB (or as configured) pack-torrents which are\ncommits-only-missing-from-linux-2.6 pack-torrents (ie, those missing\nfrom that repo's \"canonical\" upstream).\n\nThis would require clients have a recursive torrent locator (I start\nat linux-net.git, which requires linux-2.6.git, so I go get those\npacks as well as the linux-net.git packs).\n\nPerhaps have a system-wide or user-wide git repo/ torrent config, or\ncheck with user running git-clone linux-net.git \"Do you have an\nexisting git.vger.kernel.org/linux-2.6.git archive?\".\n\nZen\n"},{"id":"159039","messageId":"AANLkTinc12H01Us1mkKieZo75hwjgTCZth_wFvRNscMq@mail.gmail.com","threadId":"26211","inReplyTo":"AANLkTikv+L5Da7A5VM7BAgnue=m0O_-nHmHchJzfGxJa@mail.gmail.com","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-01-06T17:05:00Z","receivedAt":"2011-01-06T17:05:00Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Jan 5, 2011 at 18:29, Zenaan Harkness <zen@freedbms.net> wrote:\n> Bittorrent requires some stability around torrent files.\n>\n> Can packs be generated deterministically?\n\nNo.  We have been trying to avoid doing that, because it ties us into\none particular compression scheme.  We can't tune the algorithm and\nget better compression later, because it would generate a different\npack.  We also rely on the system's libz to generate the compressed\ndata.  A version change to libz may generate a different encoding for\nthe same uncompressed data, simply because they made a tweak to how\nthe compression was performed.  Likewise our own delta compression\ncode can be tweaked to produce a different (but logically identical)\ndelta between the same two objects.\n\nRight now packs aren't deterministic because they use multiple threads\nto generate the deltas, the thread scheduling impacts which base\nobjects deltas are tried against because threads can steal work from\neach other if one finishes before the other one.  Disabling threading\nentirely slows down delta compression considerably on multi-core\nmachines, but does remove this work-stealing, making the pack\ndeterministic... but only for this exact Git binary, with this same\nshared libz.  If the system libz or Git changes, all bets are off.\n\nWe've been down this road before; we don't want to box ourselves into\na tight corner by setting for all time these tunable portions of the\ncompression algorithms.\n\n-- \nShawn.\n"},{"id":"159070","messageId":"alpine.LFD.2.00.1101061552580.22191@xanadu.home","threadId":"26211","inReplyTo":"AANLkTikv+L5Da7A5VM7BAgnue=m0O_-nHmHchJzfGxJa@mail.gmail.com","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-06T21:09:00Z","receivedAt":"2011-01-06T21:09:00Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Jan 2011, Zenaan Harkness wrote:\n\n> Bittorrent requires some stability around torrent files.\n> \n> Can packs be generated deterministically?\n\nThey _could_, but we do _not_ want to do that.\n\nThe only thing which is stable in Git is the canonical representation of \nobjects, and the objects they depend on, expressed by their SHA1 \nsignature.  Any BitTorrent-alike design for Git must be based on that \nproperty and not the packed representation of those objects which is not \nmeant to be stable.\n\nIf you don't want to design anything and simply reuse current BitTorrent \ncodebase then simply create a Git bundle from some release version and \nseed that bundle for a sufficiently long period to be worth it.  Then \nfalling back to git fetch in order to bring the repo up to date with the \nvery latest commits should be small and quick.  When that clone gets too \nbig then it's time to start seeding another more up-to-date bundle.\n\n\nNicolas\n"},{"id":"159090","messageId":"AANLkTikgzqoG2cymNJ0NN03RsTRJi22R9M+0LFJ8U2yB@mail.gmail.com","threadId":"26211","inReplyTo":"alpine.LFD.2.00.1101061552580.22191@xanadu.home","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Zenaan Harkness","fromEmail":"zen@freedbms.net","sentAt":"2011-01-07T02:36:31Z","receivedAt":"2011-01-07T02:36:31Z","isPatch":false,"sender":{"key":"zen@freedbms.net","avatar":null},"body":"On Fri, Jan 7, 2011 at 08:09, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Thu, 6 Jan 2011, Zenaan Harkness wrote:\n>\n>> Bittorrent requires some stability around torrent files.\n>>\n>> Can packs be generated deterministically?\n>\n> They _could_, but we do _not_ want to do that.\n>\n> The only thing which is stable in Git is the canonical representation of\n> objects, and the objects they depend on, expressed by their SHA1\n> signature.  Any BitTorrent-alike design for Git must be based on that\n> property and not the packed representation of those objects which is not\n> meant to be stable.\n>\n> If you don't want to design anything and simply reuse current BitTorrent\n> codebase then simply create a Git bundle from some release version and\n> seed that bundle for a sufficiently long period to be worth it.  Then\n> falling back to git fetch in order to bring the repo up to date with the\n> very latest commits should be small and quick.  When that clone gets too\n> big then it's time to start seeding another more up-to-date bundle.\n\nThanks guys for the explanations.\n\nSo, we don't _want_ to generate packs deterministically.\nBUT, we _can_ reliably unpack a pack (duh).\n\nSo if my configured \"canonical upstream\" decides on a particular\ncompression etc, I (my git client) doesn't care what has been chosen\nby my upstream.\n\nWhat is important for torrent-able packs though is stability over some\ntime period, no matter what the format.\n\nThere's been much talk of caching, invalidating of caches, overlapping\ntorrent-packs etc.\n\nIn every case, for torrents to work, the P2P'd files must have some\nstability over some time period.\n(If this assumption is incorrect, please clarify, not counting\nevery-file-is-a-torrent and every-commit-is-a-torrent.)\n\nSo, torrentable options:\n- torrent per commit\n- torrent per pack\n- torrent per torrent-archive - new file format\n\nTorrent per commit - too small, too many torrents; we need larger\np2p-able sizes in general.\n\nTorrent per pack - packs non-deterministically created, both between\nhosts and even intra-host (libz upgrade, nr_threads change, git pack\nalgorithm optimization).\n\nA new torrent format, if \"close enough\" to current git pack\nperformance (cpu load, threadability, size) is potential for new\nversion of git pack file format - we don't want to store two sets of\npack files on disk, if sensible to not do so; unlikely to happen - I\ncan't conceive that a torrentable format would be anything but worse\nthan pack files and therefore would be rejected from git master.\n\nCan we can relax the perceived requirement to deterministically create\npack files?\nWell, over what time period are pack files stable in a particular git?\nOver what time period do we require stable files for torrenting?\n\nCan we simply configure our local git to keep specified pack files for\nspecified time period?\nAnd use those for torrent-packs?\nPerhaps the torrent file could have a UseBy date?\n\nZen\n"},{"id":"159094","messageId":"alpine.LFD.2.00.1101062221480.22191@xanadu.home","threadId":"26211","inReplyTo":"AANLkTikgzqoG2cymNJ0NN03RsTRJi22R9M+0LFJ8U2yB@mail.gmail.com","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-07T04:33:51Z","receivedAt":"2011-01-07T04:33:51Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Jan 2011, Zenaan Harkness wrote:\n\n> On Fri, Jan 7, 2011 at 08:09, Nicolas Pitre <nico@fluxnic.net> wrote:\n> > On Thu, 6 Jan 2011, Zenaan Harkness wrote:\n> >\n> >> Bittorrent requires some stability around torrent files.\n> >>\n> >> Can packs be generated deterministically?\n> >\n> > They _could_, but we do _not_ want to do that.\n> >\n> > The only thing which is stable in Git is the canonical representation of\n> > objects, and the objects they depend on, expressed by their SHA1\n> > signature.  Any BitTorrent-alike design for Git must be based on that\n> > property and not the packed representation of those objects which is not\n> > meant to be stable.\n> >\n> > If you don't want to design anything and simply reuse current BitTorrent\n> > codebase then simply create a Git bundle from some release version and\n> > seed that bundle for a sufficiently long period to be worth it.  Then\n> > falling back to git fetch in order to bring the repo up to date with the\n> > very latest commits should be small and quick.  When that clone gets too\n> > big then it's time to start seeding another more up-to-date bundle.\n> \n> Thanks guys for the explanations.\n> \n> So, we don't _want_ to generate packs deterministically.\n> BUT, we _can_ reliably unpack a pack (duh).\n\nOf course.\n\n> So if my configured \"canonical upstream\" decides on a particular\n> compression etc, I (my git client) doesn't care what has been chosen\n> by my upstream.\n\nIndeed.  This is like saying: I'm sending you the value 52, but I chose \nto use the representation \"24 + 28\", while someone else might decide to \nencode that value as \"13 * 4\" instead.  You still are able to decode it \nto the same result in both cases.\n\n> What is important for torrent-able packs though is stability over some\n> time period, no matter what the format.\n\nHence my suggestion to simply seed a Git bundle over BitTorrent. Bundles \nare files which are designed to be used by completely ad hoc transports \nand you can fetch from them just like if they were a remote repository.\n\n> There's been much talk of caching, invalidating of caches, overlapping\n> torrent-packs etc.\n\nAnd in my humble opinion this is just all crap.  All those suggestions \nare fragile, create administrative issues, eat up server resources, and \nthey all are suboptimal in the end. No one ever implemented a working \nprototype so far either.\n\nWe don't want caches.  Fundamentally, we do not need any cache.  Caches \nare a pain to administrate on a busy server anyway as they eat disk \nspace and they also represent a much bigger security risk compared to a \nread-only operation.\n\nFurthermore, a cache is good only for the common case that everyone \nwant.  but with Git, you cannot presume that everyone is at the same \nversion locally.  So either you do a custom transfer for each client to \nminimize transfers and caching the result in that case might not benefit \nthat many people, or you make the cached data bigger so to cover more \ncases while making the transfer suboptimal.\n\nFinally, we do have a cache already, and that's the existing packs \nthemselves.  During a clone, the vast majority of the transferred data \nis streamed without further processing straight of those existing packs \nas we try to reuse as much data as possible from those packs so not to \nrecompute/recompress that data all the time.\n\n> In every case, for torrents to work, the P2P'd files must have some\n> stability over some time period.\n> (If this assumption is incorrect, please clarify, not counting\n> every-file-is-a-torrent and every-commit-is-a-torrent.)\n> \n> So, torrentable options:\n> - torrent per commit\n> - torrent per pack\n> - torrent per torrent-archive - new file format\n> \n> Torrent per commit - too small, too many torrents; we need larger\n> p2p-able sizes in general.\n> \n> Torrent per pack - packs non-deterministically created, both between\n> hosts and even intra-host (libz upgrade, nr_threads change, git pack\n> algorithm optimization).\n> \n> A new torrent format, if \"close enough\" to current git pack\n> performance (cpu load, threadability, size) is potential for new\n> version of git pack file format - we don't want to store two sets of\n> pack files on disk, if sensible to not do so; unlikely to happen - I\n> can't conceive that a torrentable format would be anything but worse\n> than pack files and therefore would be rejected from git master.\n> \n> Can we can relax the perceived requirement to deterministically create\n> pack files?\n> Well, over what time period are pack files stable in a particular git?\n> Over what time period do we require stable files for torrenting?\n> \n> Can we simply configure our local git to keep specified pack files for\n> specified time period?\n> And use those for torrent-packs?\n> Perhaps the torrent file could have a UseBy date?\n\nAgain, this is just too much complexity for so little gain.\n\nHere's what I suggest:\n\n\tcd my_project\n\tBUNDLENAME=my_project_$(date \"+%s\").gitbundle\n\tgit bundle create $BUNDLENAME --all\n\tmaketorrent-console your_favorite_tracker $BUNDLENAME\n\nThen start seeding that bundle, and upload $BUNDLENAME.torrent as \nbundle.torrent inside my_project.git on your server.\n\nNow... Git clients could be improved to first check for the availability \nof the file \"bundle.torrent\" on the remote side, either directly in \nmy_project.git, or through some Git protocol extension.  Or even better, \nthe torrent hash could be stored in a Git ref, such as \nrefs/bittorrent/bundle and the client could use that to retrieve the \nbundle.torrent file through some other means.\n\nWhen the bundle.torrent file is retrieved, then just pull the torrent \ncontent (and seed it some more to be nice).  Then simply run \"git clone\" \nusing the original arguments but with the obtained bundle instead of the \noriginal URL.  Then replace the remote URL in .git/config with the \nactual remote URL instead of the bundle file path.  And finally perform \na \"git pull\" to bring the new commits that were added to the remote \nrepository since the bundle was created.  That final pull will be small \nand quick.\n\nAfter a while, that final pull will get bigger as the difference between \nthe bundled version and the current tip in the remote repository will \ngrow.  So every so often, say 3 months, it might be a good idea to \ncreate a new bundle so that the latest commits are included into it in \norder to make that final pull small and quick again.\n\nIsn't that sufficient?\n\n\nNicolas\n"},{"id":"159095","messageId":"20110107052207.GA23128@sigill.intra.peff.net","threadId":"26211","inReplyTo":"alpine.LFD.2.00.1101062221480.22191@xanadu.home","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-07T05:22:07Z","receivedAt":"2011-01-07T05:22:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 06, 2011 at 11:33:51PM -0500, Nicolas Pitre wrote:\n\n> Here's what I suggest:\n> \n> \tcd my_project\n> \tBUNDLENAME=my_project_$(date \"+%s\").gitbundle\n> \tgit bundle create $BUNDLENAME --all\n> \tmaketorrent-console your_favorite_tracker $BUNDLENAME\n> \n> Then start seeding that bundle, and upload $BUNDLENAME.torrent as \n> bundle.torrent inside my_project.git on your server.\n> \n> Now... Git clients could be improved to first check for the availability \n> of the file \"bundle.torrent\" on the remote side, either directly in \n> my_project.git, or through some Git protocol extension.  Or even better, \n> the torrent hash could be stored in a Git ref, such as \n> refs/bittorrent/bundle and the client could use that to retrieve the \n> bundle.torrent file through some other means.\n\nI really like the simplicity of this idea. It could even be generalized\nto handle more traditional mirrors, too. Just slice up the refs/mirrors\nnamespace to provide different methods of fetching some initial set of\nobjects. For example, upload-pack might advertise (in addition to the\nusual refs):\n\n  refs/mirrors/bundle/torrent\n  refs/mirrors/bundle/http\n  refs/mirrors/fetch/git\n  refs/mirrors/fetch/http\n\nand the client can decide its preferred way of getting data: a bundle by\nhttp or by torrent, or connecting directly to some other git repository\nby git protocol or http. It would fetch the appropriate ref, which would\ncontain a blob in some method-specific format. For torrent, it would be\na torrent file. For the others, probably a newline-delimited set of\nURLs. You could also provide a torrent-magnet ref if you didn't even\nwant to distribute the torrent file.\n\nAnd no matter what the method used, at the end you have some set of refs\nand objects, and you can re-try your (now much smaller fetch). And there\nare a few obvious optimizations:\n\n  1. When you get the initial set of refs from the master, remember\n     them. If the mirror actually satisfies everything you were going to\n     fetch, then you don't even have to reconnect for the final fetch.\n\n  2. You can optionally cache the mirror list, and go straight to a\n     mirror for future fetches instead of checking the master. This is\n     only a reasonable thing to do if the mirrors are kept up to date,\n     and provide good incremental access (i.e., only actual git-protocol\n     mirrors, not torrent or http file).\n\n-Peff\n"},{"id":"159096","messageId":"20110107053119.GA23177@sigill.intra.peff.net","threadId":"26211","inReplyTo":"20110107052207.GA23128@sigill.intra.peff.net","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-07T05:31:19Z","receivedAt":"2011-01-07T05:31:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 07, 2011 at 12:22:07AM -0500, Jeff King wrote:\n\n>   refs/mirrors/bundle/torrent\n>   refs/mirrors/bundle/http\n>   refs/mirrors/fetch/git\n>   refs/mirrors/fetch/http\n> \n> and the client can decide its preferred way of getting data: a bundle by\n> http or by torrent, or connecting directly to some other git repository\n> by git protocol or http. It would fetch the appropriate ref, which would\n> contain a blob in some method-specific format. For torrent, it would be\n> a torrent file. For the others, probably a newline-delimited set of\n> URLs. You could also provide a torrent-magnet ref if you didn't even\n> want to distribute the torrent file.\n> \n> And no matter what the method used, at the end you have some set of refs\n> and objects, and you can re-try your (now much smaller fetch).\n\nAnd I think it is probably obvious to you, Nicolas, since these are\nproblems you have been thinking about for some time, but the reason I am\ninterested in this expanded definition of mirroring is for a few\nfeatures people have been asking for:\n\n  1. restartable clone; any bundle format is easily restartable using\n     standard protocols\n\n  2. avoid too-big clones; I remember the gentoo folks wanting to\n     disallow full clones from their actual dev machines and push people\n     off to some more static method of pulling. I think not just because\n     of restartability, but because of the load on the dev machines\n\n  3. people on low-bandwidth servers who fork major projects; if I write\n     three kernel patches and host a git server, I would really like\n     people to only fetch my patches from me and get the rest of it from\n     kernel.org\n\n-Peff\n"},{"id":"159098","messageId":"AANLkTikWKdUE1JoBPmQF-NyacTtoKKD0NZYu_arsqoiW@mail.gmail.com","threadId":"26211","inReplyTo":"20110107053119.GA23177@sigill.intra.peff.net","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Zenaan Harkness","fromEmail":"zen@freedbms.net","sentAt":"2011-01-07T10:04:49Z","receivedAt":"2011-01-07T10:04:49Z","isPatch":false,"sender":{"key":"zen@freedbms.net","avatar":null},"body":"On Fri, Jan 7, 2011 at 16:31, Jeff King <peff@peff.net> wrote:\n> On Fri, Jan 07, 2011 at 12:22:07AM -0500, Jeff King wrote:\n> the reason I am\n> interested in this expanded definition of mirroring is for a few\n> features people have been asking for:\n>\n>  1. restartable clone; any bundle format is easily restartable using\n>     standard protocols\n\nThis is very important to me. I have failed to establish an initial\nrepo for a few larger projects, some apache projects and opentaps most\nrecently. It is getting _really_ frustrating.\n\n\n>  2. avoid too-big clones; I remember the gentoo folks wanting to\n>     disallow full clones from their actual dev machines and push people\n>     off to some more static method of pulling. I think not just because\n>     of restartability, but because of the load on the dev machines\n\nAnd of course the lack of restartability causes an ongoing increase in\nthe load on the machines delivering those large clones.\n\n\n>  3. people on low-bandwidth servers who fork major projects; if I write\n>     three kernel patches and host a git server, I would really like\n>     people to only fetch my patches from me and get the rest of it from\n>     kernel.org\n\nThis is not so much of a problem - can already be handled by cloning\nyour linux-full.git to a private dir, and only publishing your shallow\n\"personal patches only\" clone, or better still, just a tar-ball of\nyour 3 patches, or email them, or etc.\n\n\nSo I agree with the big issues being restartable large clones and\nlowering server loads.\n\nZen\n"},{"id":"159135","messageId":"20110107185218.GA16645@LK-Perkele-VI.localdomain","threadId":"26211","inReplyTo":"20110107053119.GA23177@sigill.intra.peff.net","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Ilari Liusvaara","fromEmail":"ilari.liusvaara@elisanet.fi","sentAt":"2011-01-07T18:52:18Z","receivedAt":"2011-01-07T18:52:18Z","isPatch":false,"sender":{"key":"ilari.liusvaara@elisanet.fi","avatar":null},"body":"On Fri, Jan 07, 2011 at 12:31:19AM -0500, Jeff King wrote:\n> \n>   3. people on low-bandwidth servers who fork major projects; if I write\n>      three kernel patches and host a git server, I would really like\n>      people to only fetch my patches from me and get the rest of it from\n>      kernel.org\n\nOne client-side-only feature that could be useful:\n\nAbility to contact multiple servers in sequence, each time advertising\neverything obtained so far. Then treat the new repo as clone of the last\naddress.\n\nThis would e.g. be very handy if you happen to have local mirror of say, Linux\nkernel and want to fetch some related project without messing with alternates\nor downloading everything again:\n\ngit clone --use-mirror=~/repositories/linux-2.6 git://foo.example/linux-foo\n\nThis would first fetch everything from local source and then update that\nfrom remote, likely being vastly faster.\n\n-Ilari\n"},{"id":"159138","messageId":"20110107191719.GA6175@sigill.intra.peff.net","threadId":"26211","inReplyTo":"20110107185218.GA16645@LK-Perkele-VI.localdomain","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-07T19:17:19Z","receivedAt":"2011-01-07T19:17:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 07, 2011 at 08:52:18PM +0200, Ilari Liusvaara wrote:\n\n> On Fri, Jan 07, 2011 at 12:31:19AM -0500, Jeff King wrote:\n> > \n> >   3. people on low-bandwidth servers who fork major projects; if I write\n> >      three kernel patches and host a git server, I would really like\n> >      people to only fetch my patches from me and get the rest of it from\n> >      kernel.org\n> \n> One client-side-only feature that could be useful:\n> \n> Ability to contact multiple servers in sequence, each time advertising\n> everything obtained so far. Then treat the new repo as clone of the last\n> address.\n> \n> This would e.g. be very handy if you happen to have local mirror of say, Linux\n> kernel and want to fetch some related project without messing with alternates\n> or downloading everything again:\n> \n> git clone --use-mirror=~/repositories/linux-2.6 git://foo.example/linux-foo\n> \n> This would first fetch everything from local source and then update that\n> from remote, likely being vastly faster.\n\nI'm not clear in your example what ~/repositories/linux-2.6 is. Is it a\nrepo? In that case, isn't that basically the same as --reference? Or is\nit a local mirror list?\n\nIf the latter, then yeah, I think it is a good idea. Clients should\ndefinitely be able to ignore, override, or add to mirror lists provided\nby servers. The server can provide hints about useful mirrors, but it is\nup to the client to decide which methods are useful to it and which\nmirrors are closest.\n\nOf course there are some servers who will want to do more than hint\n(e.g., the gentoo case where they really don't want people cloning from\nthe main machine). For those cases, though, I think it is best to\nprovide the hint and to reject clients who don't follow it (e.g., by\nbarfing on somebody who tries to do a full clone). You have to implement\nthat rejection layer anyway for older clients.\n\n-Peff\n"},{"id":"159162","messageId":"20110107214501.GA29959@LK-Perkele-VI.localdomain","threadId":"26211","inReplyTo":"20110107191719.GA6175@sigill.intra.peff.net","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Ilari Liusvaara","fromEmail":"ilari.liusvaara@elisanet.fi","sentAt":"2011-01-07T21:45:01Z","receivedAt":"2011-01-07T21:45:01Z","isPatch":false,"sender":{"key":"ilari.liusvaara@elisanet.fi","avatar":null},"body":"On Fri, Jan 07, 2011 at 02:17:19PM -0500, Jeff King wrote:\n> On Fri, Jan 07, 2011 at 08:52:18PM +0200, Ilari Liusvaara wrote:\n> \n> > \n> > git clone --use-mirror=~/repositories/linux-2.6 git://foo.example/linux-foo\n> > \n> > This would first fetch everything from local source and then update that\n> > from remote, likely being vastly faster.\n> \n> I'm not clear in your example what ~/repositories/linux-2.6 is. Is it a\n> repo? In that case, isn't that basically the same as --reference? Or is\n> it a local mirror list?\n\nYes, it is a repo. No, it isn't the same as --reference. It is list\nof mirrors to try first before connecting to final repository and can\nbe any type of repository URL (local, true smart transport, smart HTTP,\ndumb HTTP, etc...)\n\nIdea is that you have list of mirrors that are faster than the final\nrepository, but not necressarily complete. You want to download most of\nthe stuff from there.\n\n> If the latter, then yeah, I think it is a good idea. Clients should\n> definitely be able to ignore, override, or add to mirror lists provided\n> by servers. The server can provide hints about useful mirrors, but it is\n> up to the client to decide which methods are useful to it and which\n> mirrors are closest.\n\nThis is essentially adding mirrors to mirror list (modulo that mirrors\nare not assumed to be complete).\n\nSecurity:\n\nConfidentiality: The connection to mirror must transverse only trusted\nlinks or be encrypted if material from mirror is sensitive.\n\nIntegerity: The same integerity as the connection to final repo (assuming\nSHA-1 can't be collided) due to fact that git object naming is securely\nunique.\n\n> Of course there are some servers who will want to do more than hint\n> (e.g., the gentoo case where they really don't want people cloning from\n> the main machine). For those cases, though, I think it is best to\n> provide the hint and to reject clients who don't follow it (e.g., by\n> barfing on somebody who tries to do a full clone). You have to implement\n> that rejection layer anyway for older clients.\n\nWith option like this, a client could do:\n\ngit clone --use-mirror=http://git.example.org/base/foo git://git.example.org/foo\n\nTo first grab stuff via HTTP (well-packed dumb HTTP is very light on the\nserver) and then continue via git:// (now much cheaper because client is\nrelatively up to date).\n\n-Ilari\n"},{"id":"159164","messageId":"20110107215631.GA10343@sigill.intra.peff.net","threadId":"26211","inReplyTo":"20110107214501.GA29959@LK-Perkele-VI.localdomain","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-07T21:56:31Z","receivedAt":"2011-01-07T21:56:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jan 07, 2011 at 11:45:01PM +0200, Ilari Liusvaara wrote:\n\n> > I'm not clear in your example what ~/repositories/linux-2.6 is. Is it a\n> > repo? In that case, isn't that basically the same as --reference? Or is\n> > it a local mirror list?\n> \n> Yes, it is a repo. No, it isn't the same as --reference. It is list\n> of mirrors to try first before connecting to final repository and can\n> be any type of repository URL (local, true smart transport, smart HTTP,\n> dumb HTTP, etc...)\n\nOK, I understand what you mean. I was thrown off by your example using a\nlocal repository (in which case you probably would want --reference to\nsave disk space, unless the burden of alternates management is too\nmuch).\n\nSo yeah, I think we are on the same page, except that you were proposing\nto pass the mirror directly, and I was proposing passing a mirror file\nwhich would contain a list of mirrors. Yours is much simpler and would\nprobably be what people want most of the time.\n\n> > If the latter, then yeah, I think it is a good idea. Clients should\n> > definitely be able to ignore, override, or add to mirror lists provided\n> > by servers. The server can provide hints about useful mirrors, but it is\n> > up to the client to decide which methods are useful to it and which\n> > mirrors are closest.\n> \n> This is essentially adding mirrors to mirror list (modulo that mirrors\n> are not assumed to be complete).\n\nI think there should always be an assumption that mirrors are not\nnecessarily complete. That is necessary for bundle-like mirrors to be\nfeasible, since updating the bundle for every commit defeats the\npurpose.\n\nIt would be nice for there to be a way for some mirrors to be marked as\n\"should be considered complete and authoritative\", since we can optimize\nout the final check of the master in that case (as well as for future\nfetches). But that's a future feature. My plan was to leave space in the\nmirror list for arbitrary metadata of that sort.\n\n> > Of course there are some servers who will want to do more than hint\n> > (e.g., the gentoo case where they really don't want people cloning from\n> > the main machine). For those cases, though, I think it is best to\n> > provide the hint and to reject clients who don't follow it (e.g., by\n> > barfing on somebody who tries to do a full clone). You have to implement\n> > that rejection layer anyway for older clients.\n> \n> With option like this, a client could do:\n> \n> git clone --use-mirror=http://git.example.org/base/foo git://git.example.org/foo\n> \n> To first grab stuff via HTTP (well-packed dumb HTTP is very light on the\n> server) and then continue via git:// (now much cheaper because client is\n> relatively up to date).\n\nYes, exactly.\n\n-Peff\n"},{"id":"159169","messageId":"20110107222133.GA2377@LK-Perkele-VI.localdomain","threadId":"26211","inReplyTo":"20110107215631.GA10343@sigill.intra.peff.net","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Ilari Liusvaara","fromEmail":"ilari.liusvaara@elisanet.fi","sentAt":"2011-01-07T22:21:33Z","receivedAt":"2011-01-07T22:21:33Z","isPatch":false,"sender":{"key":"ilari.liusvaara@elisanet.fi","avatar":null},"body":"On Fri, Jan 07, 2011 at 04:56:31PM -0500, Jeff King wrote:\n> On Fri, Jan 07, 2011 at 11:45:01PM +0200, Ilari Liusvaara wrote:\n> \n> \n> I think there should always be an assumption that mirrors are not\n> necessarily complete. That is necessary for bundle-like mirrors to be\n> feasible, since updating the bundle for every commit defeats the\n> purpose.\n\nAlso add protocol that grabs a bundle from HTTP and then opens that\nup? :-)\n\n> It would be nice for there to be a way for some mirrors to be marked as\n> \"should be considered complete and authoritative\", since we can optimize\n> out the final check of the master in that case (as well as for future\n> fetches). But that's a future feature. My plan was to leave space in the\n> mirror list for arbitrary metadata of that sort.\n\nThe first thing one should get/do when connecting to another repository\nis its list of references. One can see from there if what one has got\nis complete or not (with --use-mirror that only allows skipping commit\nnegotiation and fetch, not the whole connection due to the fact that the\nrepositories are contacted in order)...\n\n-Ilari\n"},{"id":"159171","messageId":"20110107222704.GA10583@sigill.intra.peff.net","threadId":"26211","inReplyTo":"20110107222133.GA2377@LK-Perkele-VI.localdomain","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-01-07T22:27:04Z","receivedAt":"2011-01-07T22:27:04Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Jan 08, 2011 at 12:21:33AM +0200, Ilari Liusvaara wrote:\n\n> On Fri, Jan 07, 2011 at 04:56:31PM -0500, Jeff King wrote:\n> > On Fri, Jan 07, 2011 at 11:45:01PM +0200, Ilari Liusvaara wrote:\n> > \n> > \n> > I think there should always be an assumption that mirrors are not\n> > necessarily complete. That is necessary for bundle-like mirrors to be\n> > feasible, since updating the bundle for every commit defeats the\n> > purpose.\n> \n> Also add protocol that grabs a bundle from HTTP and then opens that\n> up? :-)\n\nWell, yes, that still needs to be implemented. But it's all client-side,\nso the server just has to provide the bundle somewhere.\n\n> > It would be nice for there to be a way for some mirrors to be marked as\n> > \"should be considered complete and authoritative\", since we can optimize\n> > out the final check of the master in that case (as well as for future\n> > fetches). But that's a future feature. My plan was to leave space in the\n> > mirror list for arbitrary metadata of that sort.\n> \n> The first thing one should get/do when connecting to another repository\n> is its list of references. One can see from there if what one has got\n> is complete or not (with --use-mirror that only allows skipping commit\n> negotiation and fetch, not the whole connection due to the fact that the\n> repositories are contacted in order)...\n\nYes, but it would be cool to be able to skip even that connect in some\ncases (e.g., mirrors can be useful not just to take load off the master,\nbut also when the master isn't available, either for downtime or because\nthe client is behind a firewall). But the default should definitely be\nto double-check that the master is right, and we can leave more advanced\ncases for later (we just need to be aware of leaving room for them now).\n\nI'm going to start working on a patch series for this, so hopefully\nwe'll see how it's shaping up in a day or two.\n\n-Peff\n"},{"id":"159263","messageId":"AANLkTinVYWit95O9Y0r5BKJiMGJRAOvgPqZ0s8Eez7KJ@mail.gmail.com","threadId":"26211","inReplyTo":"alpine.LFD.2.00.1101062221480.22191@xanadu.home","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-01-10T11:48:22Z","receivedAt":"2011-01-10T11:48:22Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Jan 7, 2011 at 11:33 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> Here's what I suggest:\n>\n>        cd my_project\n>        BUNDLENAME=my_project_$(date \"+%s\").gitbundle\n>        git bundle create $BUNDLENAME --all\n>        maketorrent-console your_favorite_tracker $BUNDLENAME\n>\n> Then start seeding that bundle, and upload $BUNDLENAME.torrent as\n> bundle.torrent inside my_project.git on your server.\n\nI was about to ask if we could put more \"trailer\" sha-1 checksums to\nthe bundle, so we can verify which part is corrupt without\nredownloading the whole thing (this is over http/ftp.. not torrent).\n\nBut I realize it's just easier to split the bundle into multiple\npacks, so we can verify and redownload only corrupt packs. Logically\nit is still a single pack. Splitting help put more sha-1 checksums in\nwithout changing pack format. The packs will be merged back into one\nwith \"index-pack --pack-stream\" patch I sent elsewhere.\n-- \nDuy\n"},{"id":"159268","messageId":"alpine.LFD.2.00.1101100846220.6883@xanadu.home","threadId":"26211","inReplyTo":"AANLkTinVYWit95O9Y0r5BKJiMGJRAOvgPqZ0s8Eez7KJ@mail.gmail.com","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-01-10T13:50:44Z","receivedAt":"2011-01-10T13:50:44Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 10 Jan 2011, Nguyen Thai Ngoc Duy wrote:\n\n> On Fri, Jan 7, 2011 at 11:33 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> > Here's what I suggest:\n> >\n> >        cd my_project\n> >        BUNDLENAME=my_project_$(date \"+%s\").gitbundle\n> >        git bundle create $BUNDLENAME --all\n> >        maketorrent-console your_favorite_tracker $BUNDLENAME\n> >\n> > Then start seeding that bundle, and upload $BUNDLENAME.torrent as\n> > bundle.torrent inside my_project.git on your server.\n> \n> I was about to ask if we could put more \"trailer\" sha-1 checksums to\n> the bundle, so we can verify which part is corrupt without\n> redownloading the whole thing (this is over http/ftp.. not torrent).\n\nAren't HTTP and FTP based on TCP which is meant to be a reliable \ntransport protocol already?  In this case, isn't the final SHA1 embedded \nin the bundle/pack sufficient enough?  Normally, your HTTP/FTP client \nshould get you all data or partial data, but not wrong data.\n\n\nNicolas\n"},{"id":"159270","messageId":"4D2B3643.2070106@gmx.de","threadId":"26211","inReplyTo":"AANLkTinc12H01Us1mkKieZo75hwjgTCZth_wFvRNscMq@mail.gmail.com","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"John Wyzer","fromEmail":"john.wyzer@gmx.de","sentAt":"2011-01-10T16:39:31Z","receivedAt":"2011-01-10T16:39:31Z","isPatch":false,"sender":{"key":"john.wyzer@gmx.de","avatar":null},"body":"On 06/01/11 18:05, Shawn Pearce wrote:\n> On Wed, Jan 5, 2011 at 18:29, Zenaan Harkness<zen@freedbms.net>  wrote:\n>> Bittorrent requires some stability around torrent files.\n>>\n>> Can packs be generated deterministically?\n>\n\nI hope that I don't get something technically wrong (did not read any \ncode, only skimmed the docs) and that this question is not redundant:\n\nWhy not provide an alternative mode for the git:// protocoll that \ninstead of retrieving a big packaged blob breaks this down to the \nsmallest atomic objects from the repository? Those are not changing and \nshould be able to survive partial transfers.\nWhile this might not be as efficient network traffic-wise it would \nprovide a solution for those behind breaking connections.\n"},{"id":"159288","messageId":"4D2B7522.9050400@vilain.net","threadId":"26211","inReplyTo":"20110107185218.GA16645@LK-Perkele-VI.localdomain","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2011-01-10T21:07:46Z","receivedAt":"2011-01-10T21:07:46Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 08/01/11 07:52, Ilari Liusvaara wrote:\n> Ability to contact multiple servers in sequence, each time advertising\n> everything obtained so far. Then treat the new repo as clone of the last\n> address.\n>\n> This would e.g. be very handy if you happen to have local mirror of say, Linux\n> kernel and want to fetch some related project without messing with alternates\n> or downloading everything again:\n>\n> git clone --use-mirror=~/repositories/linux-2.6 git://foo.example/linux-foo\n>\n> This would first fetch everything from local source and then update that\n> from remote, likely being vastly faster.\n\nComing to this discussion a little late, I'll summarise the previous\nresearch.\n\nFirst, the idea of applying the straight BitTorrent protocol to the pack\nfiles was raised, but as Nicolas mentions, this is not useful because\nthe pack files are not deterministic.  The protocol was revisited based\naround the part which is stable, object manifests.  The RFC is at\nhttp://utsl.gen.nz/gittorrent/rfc.html and the prototype code (an\nunsuccessful GSoC project) is at http://repo.or.cz/w/VCS-Git-Torrent.git\n\nAfter some thought, I decided that the BitTorrent protocol itself is all\ncruft and that trying to cut it down to be useful was a waste of time. \nSo, this is where the idea of \"automatic mirroring\" came from.  With\nAutomatic Mirroring, the two main functions of P2P operation - peer\ndiscovery and partial transfer - are broken into discrete features.\n\nI wrote this patch series so far, for \"client-side mirroring\":\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/133626/focus=133628\n\nThe later levels are roughly discussed on this page:\n\nhttp://code.google.com/p/gittorrent/wiki/MirrorSync\n\nThe \"mirror sync\" part is the complicated one, and as others have noted\nno truly successful prototype has yet been built.  Actually the Perl\ngittorrent implementation did manage to perform an incremental clone; it\njust didn't wrap it up nicely.  But I won't go into that too much. \nThere was also another GSoC program to look at caching the object list\ngeneration, the most expensive part of the process in the Perl\nimplementation.  This was a generic mechanism for accelerating object\ngraph traversal and showed promise, however unfortunately was never merged.\n\nThe client-side mirroring patch, in its current form, already supports\nout-of-date mirrors.  It saves refs first into\n'refs/mirrors/hostname/...' and finally contacts the main server to\ncheck what objects it is still missing.  So, if there was a regular\nbittorrent+bundle transport available, it would be a useful way to\nsupport an incremental clone; the client would first clone the (static)\nbittorrent bundle, unpack it with its refs into the 'refs/mirrors/xxx/'\nnamespace, making the subsequent 'git fetch' to get the most recent\nobjects a much more efficient operation.\n\nHope that helps!\n\nCheers,\nSam\n"},{"id":"159290","messageId":"4D2B7D3E.7090400@vilain.net","threadId":"26211","inReplyTo":"4D2B3643.2070106@gmx.de","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2011-01-10T21:42:22Z","receivedAt":"2011-01-10T21:42:22Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 11/01/11 05:39, John Wyzer wrote:\n> Why not provide an alternative mode for the git:// protocoll that\n> instead of retrieving a big packaged blob breaks this down to the\n> smallest atomic objects from the repository? Those are not changing\n> and should be able to survive partial transfers.\n> While this might not be as efficient network traffic-wise it would\n> provide a solution for those behind breaking connections.\n\nTo put this into numbers, for perl.git that might mean transferring 2GB\nof data instead of 70MB of pack.\n\nSam\n"},{"id":"159295","messageId":"AANLkTimh1RRnjXjg-fw_-RQxNW_fLbSYis8n2BvNaCc+@mail.gmail.com","threadId":"26211","inReplyTo":"4D2B3643.2070106@gmx.de","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-01-11T00:03:42Z","receivedAt":"2011-01-11T00:03:42Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Jan 10, 2011 at 11:39 PM, John Wyzer <john.wyzer@gmx.de> wrote:\n> Why not provide an alternative mode for the git:// protocoll that instead of\n> retrieving a big packaged blob breaks this down to the smallest atomic\n> objects from the repository? Those are not changing and should be able to\n> survive partial transfers.\n> While this might not be as efficient network traffic-wise it would provide a\n> solution for those behind breaking connections.\n\nThat's what I'm getting to, except that I'll send deltas as much as I can.\n-- \nDuy\n"},{"id":"159296","messageId":"4D2BAB0A.1060909@kernel.org","threadId":"26211","inReplyTo":"AANLkTimh1RRnjXjg-fw_-RQxNW_fLbSYis8n2BvNaCc+@mail.gmail.com","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"J.H.","fromEmail":"warthog9@kernel.org","sentAt":"2011-01-11T00:57:46Z","receivedAt":"2011-01-11T00:57:46Z","isPatch":false,"sender":{"key":"warthog9@kernel.org","avatar":"https://avatars.githubusercontent.com/u/2334704?v=4"},"body":"On 01/10/2011 04:03 PM, Nguyen Thai Ngoc Duy wrote:\n> On Mon, Jan 10, 2011 at 11:39 PM, John Wyzer <john.wyzer@gmx.de> wrote:\n>> Why not provide an alternative mode for the git:// protocoll that instead of\n>> retrieving a big packaged blob breaks this down to the smallest atomic\n>> objects from the repository? Those are not changing and should be able to\n>> survive partial transfers.\n>> While this might not be as efficient network traffic-wise it would provide a\n>> solution for those behind breaking connections.\n> \n> That's what I'm getting to, except that I'll send deltas as much as I can.\n\nWhile I think we need to come up with a mechanism to allow for resumable\nfetches (I'm thinking slow sporadic links and larger repos like the\nkernel for instance), but breaking the repo up into too small a chunks\nwill very adversely affect the overall transfer and could cause just as\nmuch system thrash on the upstream provider.\n\nI'd be curious to see what the system impact numbers and performance\ndifferences are though, as I do think getting some sort of resumability\nis important, but resumability at the expense of being able to get the\ndata out quickly and efficiently is not going to be a good trade off :-/\n\n- John 'Warthog9' Hawley\n"},{"id":"159297","messageId":"AANLkTimcmxMxgswrQcqex8b8M717LmTctj=3i7jEHOhZ@mail.gmail.com","threadId":"26211","inReplyTo":"4D2BAB0A.1060909@kernel.org","subject":"Re: Resumable clone/Gittorrent (again) - stable packs?","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-01-11T01:56:09Z","receivedAt":"2011-01-11T01:56:09Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jan 11, 2011 at 7:57 AM, J.H. <warthog9@kernel.org> wrote:\n> On 01/10/2011 04:03 PM, Nguyen Thai Ngoc Duy wrote:\n>> On Mon, Jan 10, 2011 at 11:39 PM, John Wyzer <john.wyzer@gmx.de> wrote:\n>>> Why not provide an alternative mode for the git:// protocoll that instead of\n>>> retrieving a big packaged blob breaks this down to the smallest atomic\n>>> objects from the repository? Those are not changing and should be able to\n>>> survive partial transfers.\n>>> While this might not be as efficient network traffic-wise it would provide a\n>>> solution for those behind breaking connections.\n>>\n>> That's what I'm getting to, except that I'll send deltas as much as I can.\n>\n> While I think we need to come up with a mechanism to allow for resumable\n> fetches (I'm thinking slow sporadic links and larger repos like the\n> kernel for instance), but breaking the repo up into too small a chunks\n> will very adversely affect the overall transfer and could cause just as\n> much system thrash on the upstream provider.\n>\n> I'd be curious to see what the system impact numbers and performance\n> differences are though, as I do think getting some sort of resumability\n> is important, but resumability at the expense of being able to get the\n> data out quickly and efficiently is not going to be a good trade off :-/\n\nYeah, I'm interested in those numbers too. Let me get a prototype\nworking, then we'll have numbers to discuss.\n-- \nDuy\n"}]}