{"thread":{"id":"24935","subject":"git pack/unpack over bittorrent - works!","startedAt":"2010-09-01T14:36:16Z","lastAt":"2010-09-07T00:29:25Z","messageCount":88,"participants":["Luke Kenneth Casson Leighton","Nguyen Thai Ngoc Duy","Ævar Arnfjörð Bjarmason","A Large Angry SCM","Jeff King","Nicolas Pitre","Casey Dahlin","Shawn O. Pearce","Brandon Casey","Jakub Narebski","Theodore Tso","Junio C Hamano","Ted Ts'o","Artur Skawina","Kyle Moffett","Tomas Carnecky"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"149559","messageId":"AANLkTik-w6jWgrt_kwAk2uNGhF_=3tMEpTZs3nyF_zGA@mail.gmail.com","threadId":"24935","inReplyTo":null,"subject":"git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-01T14:36:16Z","receivedAt":"2010-09-01T14:36:16Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"http://gitorious.org/python-libbittorrent/pybtlib\n\nhurrah - success!  git fsck shows a \"dangling commit\"!\n\nso, as a proof-of-concept, 400 lines of python code plus a bittorrent\nlibrary shows that it's possible to create a peer-to-peer distributed\nversion of \"git fetch\", by treating the pack objects as \"files to be\nshared\".\n\nas this code has only existed for less than three days, there are a\nlot of loose ends.  such as the pack object files being cached\nin-memory, and it being strictly speaking unnecessary to obtain\nabsolutely every single pack object from the git repository at\nstart-up time, but to optimise that to being \"on-demand\" requires some\nferreting around in the bittorrent library, etc. etc.\n\nif anyone is interested in helping out, or knows of a way to get this\nsponsored and completed, please speak up.  especially the sponsorship\nbit because i am in a severely critical and ridiculous financial\nsituation and am in urgent and immediate need of money.\n\nl.\n"},{"id":"149609","messageId":"AANLkTinu=RoGfq93d+yjHiQwCt0HXx4YtqfvhXyZdO=F@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTik-w6jWgrt_kwAk2uNGhF_=3tMEpTZs3nyF_zGA@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2010-09-01T22:04:03Z","receivedAt":"2010-09-01T22:04:03Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Sep 2, 2010 at 12:36 AM, Luke Kenneth Casson Leighton\n<luke.leighton@gmail.com> wrote:\n> http://gitorious.org/python-libbittorrent/pybtlib\n>\n> hurrah - success!  git fsck shows a \"dangling commit\"!\n>\n> so, as a proof-of-concept, 400 lines of python code plus a bittorrent\n> library shows that it's possible to create a peer-to-peer distributed\n> version of \"git fetch\", by treating the pack objects as \"files to be\n> shared\".\n\nYou should have a look at gittorrent [1] (and finish it too if you are\ninterested). There were discussions whether a pack is stable enough to\nbe shared like this, one of the reason \"commit reel\" was introduced.\n\n[1] http://code.google.com/p/gittorrent/\n-- \nDuy\n"},{"id":"149640","messageId":"AANLkTimpE6rf0azHtrz6BFK5d7YojF+G1YuSA1gusSC=@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTinu=RoGfq93d+yjHiQwCt0HXx4YtqfvhXyZdO=F@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T13:37:30Z","receivedAt":"2010-09-02T13:37:30Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Wed, Sep 1, 2010 at 11:04 PM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n> On Thu, Sep 2, 2010 at 12:36 AM, Luke Kenneth Casson Leighton\n> <luke.leighton@gmail.com> wrote:\n>> http://gitorious.org/python-libbittorrent/pybtlib\n>>\n>> hurrah - success!  git fsck shows a \"dangling commit\"!\n>>\n>> so, as a proof-of-concept, 400 lines of python code plus a bittorrent\n>> library shows that it's possible to create a peer-to-peer distributed\n>> version of \"git fetch\", by treating the pack objects as \"files to be\n>> shared\".\n>\n> You should have a look at gittorrent [1] (and finish it too if you are\n> interested).\n\n it's in perl, and has been shelved afaik.  sam (hi sam, you still\nhere? :) abandoned bittorrent as the underlying mechanism, and i\ndisagree with that decision, hence why i created what i have, to prove\nthat it's viable.\n\n so a) i don't do perl so would need to re-create what's been done\n(which i don't understand, and, because it's incomplete, i can't do a\nperl-to-python translation and \"have something working\") b) i don't\nbelieve in reinventing the wheel ESPECIALLY on something as complex as\npeer-to-peer file distribution.\n\n cameron dale created apt-p2p which is a recreation of a peer-to-peer\nfile distribution mechanism, and, not surprisingly, it's slow and\nproblematic.\n\n> There were discussions whether a pack is stable enough to\n> be shared like this,\n\n it seems to be.  as long as each version of git produces the exact\nsame pack object, off of the command \"git pack-objects --all --stdout\n--thin {ref} < {objref}\"\n\n is that what you're referring to?\n\n because if it isn't, then yes, the sharing of files (named by a\nvirtual filename of packs/{ref}/{objref} of course) which _might_ have\ndifferences, yeah, it becomes a bit of a fuck-up.\n\n> one of the reason \"commit reel\" was introduced.\n\n ah _ha_.  i need to look that up.  will get back to you.\n\n l.\n"},{"id":"149641","messageId":"AANLkTims3F997XrCDN+rRnSPb=LFCU1_CCqTRj6oZyAg@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTimpE6rf0azHtrz6BFK5d7YojF+G1YuSA1gusSC=@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T13:53:54Z","receivedAt":"2010-09-02T13:53:54Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":">> one of the reason \"commit reel\" was introduced.\n>\n>  ah _ha_.  i need to look that up.  will get back to you.\n>\n>  l.\n>\n\nhttp://www.mail-archive.com/gittorrent@lists.utsl.gen.nz/msg00032.html\n\n.... nnnope.  don't understand it.  at all.  there's no context.\n\nhttp://www.mail-archive.com/gittorrent@lists.utsl.gen.nz/msg00003.html -->\n\nhttp://gittorrent.utsl.gen.nz/rfc.html#anchor7\n\n\"Service Not Available\".\n\ngreeat.\n\nhttp://www.mail-archive.com/gittorrent@lists.utsl.gen.nz/msg00023.html\n\nah HA!\n+<t hangText=\"References:\">\n+\n+       The References are the Git refs of the repository being shared.\n+\n+</t>\n+<t hangText=\"Commit reel offset:\">\n+\n+       An offset into the list of all revisions sorted by their natural\n+       topological order, their commit date and their SHA-1.\n+\n+</t>\n+<t hangText=\"Commit reel:\">\n+\n+       A commit reel consists of two commit reel offsets, and all the objects\n+       that are reachable from the revisions between the two offsets, but\n+       not the revisions after the second offset.  In Git terms, a commit\n+       reel is a bundle.\n+\n+</t>\n+<t hangText=\"Block:\">\n+\n+       A block is the actual content of a commit reel, i.e. the objects.\n+       In Git terms, it is a pack.\n+\n+</t>\n\noh - right.... so, translating that: the concern is not that the git\npack-objects might be different (could someone pleaaaase confirm\nthat!) - the concern is that the order of _unpacking_ *has* to be done\nin the specific order in which they were (originally) committed.\n\n if that's all it is, then yes, i thought that that was plainly\nobvious, and had taken it into consideration already, by creating a\nvirtual file which contains the order of the commits.  this is\nachieved merely by making the contents of \"git rev-list\" available.\nvoila, dead-simple: you now have enough information to be able to\napply the pack objects in the right order.\n\nl.\n"},{"id":"149642","messageId":"AANLkTim0QgpGd2aNCF1cUfwwFmntNTFzwNzDkV39Hxgr@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTimpE6rf0azHtrz6BFK5d7YojF+G1YuSA1gusSC=@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-09-02T14:08:46Z","receivedAt":"2010-09-02T14:08:46Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Thu, Sep 2, 2010 at 13:37, Luke Kenneth Casson Leighton\n<luke.leighton@gmail.com> wrote:\n\n>  cameron dale created apt-p2p which is a recreation of a peer-to-peer\n> file distribution mechanism, and, not surprisingly, it's slow and\n> problematic.\n\napt is somewhat of a bad fit for bittorrent. You can get a lot of\nthroughput over torrent, but connecting to the swarm and beginning the\ndownload generally takes much longer than downloading & installing\ndozens of packages using the normal apt transport.\n\nAlso as a matter of implementation they're using a really resource\nhungry Python implementation of BitTorrent instead of something like\nlibtorrent, which is why I stopped running it.\n\nBut presumably most uses for GitTorrent (and what you're doing)\nwouldn't suffer so badly from the latency, you could just leave some\ndaemon on which would download commits in the background.\n\nSo e.g. if someone submitted a series against git.git to the list 30\nminutes ago and he/I were running some git-p2p thingy I could rely on\nthose commits having made it to my repository by now.\n\nJust a thought.\n"},{"id":"149648","messageId":"4C7FC3DC.3060907@gmail.com","threadId":"24935","inReplyTo":"AANLkTimpE6rf0azHtrz6BFK5d7YojF+G1YuSA1gusSC=@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2010-09-02T15:33:48Z","receivedAt":"2010-09-02T15:33:48Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 09/02/2010 09:37 AM, Luke Kenneth Casson Leighton wrote:\n> On Wed, Sep 1, 2010 at 11:04 PM, Nguyen Thai Ngoc Duy<pclouds@gmail.com>  wrote:\n[...]\n>> There were discussions whether a pack is stable enough to\n>> be shared like this,\n>\n>   it seems to be.  as long as each version of git produces the exact\n> same pack object, off of the command \"git pack-objects --all --stdout\n> --thin {ref}<  {objref}\"\n\nThis is not guaranteed.\n"},{"id":"149649","messageId":"AANLkTikBnKQJmgOms2wK1+6fCLtHWiWkhuCVMN7kKLXP@mail.gmail.com","threadId":"24935","inReplyTo":"4C7FC3DC.3060907@gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T15:42:55Z","receivedAt":"2010-09-02T15:42:55Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 4:33 PM, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> On 09/02/2010 09:37 AM, Luke Kenneth Casson Leighton wrote:\n>>\n>> On Wed, Sep 1, 2010 at 11:04 PM, Nguyen Thai Ngoc Duy<pclouds@gmail.com>\n>>  wrote:\n>\n> [...]\n>>>\n>>> There were discussions whether a pack is stable enough to\n>>> be shared like this,\n>>\n>>  it seems to be.  as long as each version of git produces the exact\n>> same pack object, off of the command \"git pack-objects --all --stdout\n>> --thin {ref}<  {objref}\"\n>\n> This is not guaranteed.\n\n ok.  greeeat.\n\n so, some sensible questions:\n\n * what _can_ be guaranteed?\n\n * diffs?\n\n * git-format-patches? (which i am aware can do binary files and also\nrms)?\n\n* individual files in the .git/objects directory?\n\n and, asking perhaps some silly questions:\n\n* why is it not guaranteed?\n\n* under what circumstances is it not guaranteed?  and, crucially, is\nit necessary to care?   i.e. if someone does a shallow git clone, i\ncouldn't give a stuff.\n\n* is it possible to _make_ the repository guaranteed to produce\nidentical pack objects?\n\n* does for example \"git gc\" change the object store in such a way such\nthat one git repo will produce a different pack-object from the same\nref?  if so, can running \"git gc\" prior to producing the pack-objects\ngurantee that the pack-objects will be the same?\n\n* is it a versioning issue?  is it because there are different\nversions (2 and 3)?  if so, that's ok, you just force people to use\nthe same pack-object versions.\n\netc. etc.\n\nl.\n"},{"id":"149651","messageId":"AANLkTi=W3QwWSrNTie-K4QDDrucSVGQa5e3Ldy7m7ihy@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTikBnKQJmgOms2wK1+6fCLtHWiWkhuCVMN7kKLXP@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T15:51:46Z","receivedAt":"2010-09-02T15:51:46Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 4:42 PM, Luke Kenneth Casson Leighton\n<luke.leighton@gmail.com> wrote:\n> On Thu, Sep 2, 2010 at 4:33 PM, A Large Angry SCM <gitzilla@gmail.com> wrote:\n\n> * is it possible to _make_ the repository guaranteed to produce\n> identical pack objects?\n\n i.e. looking at these options, listed from\nDocumentation/technical/protocol.txt:\n\n    00887217a7c7e582c46cec22a130adf4b9d7d950fba0 HEAD\\0multi_ack\nthin-pack side-band side-band-64k ofs-delta shallow no-progress\ninclude-tag\n\n is it possible to use shallow, thin-pack, side-band or side-band-64k\nto guarantee that the pack object will be identical?\n\n another important question:\n\n* if after performing a \"git unpack\" of one pack-object, can it be\nguaranteed that performing a \"git pack-object\" on the *exact* same ref\nand the *exact* same object-ref, will produce the *exact* same\npack-object that was used by \"git unpack\", as long as the exact same\narguments are used?  if not, why not, and if not under _some_\ncircumstances, under what circumstances _can_ the exact same\npack-object be retrieved that was just used?\n\nif there is absolutely absolutely no way to guarantee that the\npack-objects can be the same, under no circumstances or combinations\nof arguments or by forcing only compatible versions to communicate\netc. etc., a rather awful work-around can be applied which is to share\nand permanently cache every single pack-object, rather than use what's\ngone into the repo.\n\nl.\n"},{"id":"149653","messageId":"20100902155810.GB14508@sigill.intra.peff.net","threadId":"24935","inReplyTo":"AANLkTikBnKQJmgOms2wK1+6fCLtHWiWkhuCVMN7kKLXP@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-09-02T15:58:13Z","receivedAt":"2010-09-02T15:58:13Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Sep 02, 2010 at 04:42:55PM +0100, Luke Kenneth Casson Leighton wrote:\n\n> >>  it seems to be.  as long as each version of git produces the exact\n> >> same pack object, off of the command \"git pack-objects --all --stdout\n> >> --thin {ref}<  {objref}\"\n> >\n> > This is not guaranteed.\n> [...]\n> * under what circumstances is it not guaranteed?  and, crucially, is\n> it necessary to care?   i.e. if someone does a shallow git clone, i\n> couldn't give a stuff.\n\npack-objects will reuse previously found deltas. So the deltas you have\nin your existing packs matter. The deltas you have in your existing\npacks depend on many things. At least:\n\n  1. Options you used when packing (e.g., --depth and --window).\n\n  2. Probably exactly _when_ you packed. You could find a good delta\n     from A to B. Later, object C comes into existence, and would\n     provide a better delta base for B. I don't think we will ever try A\n     against C, unless --no-reuse-delta is set.\n\n     You have a different pack than somebody who packed after A, B, and\n     C all existed.\n\n     In practice, this tends not to happen much because the best deltas\n     are usually going backwards in time to a previous version. But it\n     can happen.\n\n-Peff\n"},{"id":"149658","messageId":"alpine.LFD.2.00.1009021233190.19366@xanadu.home","threadId":"24935","inReplyTo":"20100902155810.GB14508@sigill.intra.peff.net","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-02T16:41:45Z","receivedAt":"2010-09-02T16:41:45Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 2 Sep 2010, Jeff King wrote:\n\n> On Thu, Sep 02, 2010 at 04:42:55PM +0100, Luke Kenneth Casson Leighton wrote:\n> \n> > >>  it seems to be.  as long as each version of git produces the exact\n> > >> same pack object, off of the command \"git pack-objects --all --stdout\n> > >> --thin {ref}<  {objref}\"\n> > >\n> > > This is not guaranteed.\n> > [...]\n> > * under what circumstances is it not guaranteed?  and, crucially, is\n> > it necessary to care?   i.e. if someone does a shallow git clone, i\n> > couldn't give a stuff.\n> \n> pack-objects will reuse previously found deltas. So the deltas you have\n> in your existing packs matter. The deltas you have in your existing\n> packs depend on many things. At least:\n> \n>   1. Options you used when packing (e.g., --depth and --window).\n> \n>   2. Probably exactly _when_ you packed. You could find a good delta\n>      from A to B. Later, object C comes into existence, and would\n>      provide a better delta base for B. I don't think we will ever try A\n>      against C, unless --no-reuse-delta is set.\n> \n>      You have a different pack than somebody who packed after A, B, and\n>      C all existed.\n> \n>      In practice, this tends not to happen much because the best deltas\n>      are usually going backwards in time to a previous version. But it\n>      can happen.\n\nI would go as far as stating that this is never guaranteed by design.  \nAnd I will oppose any attempt to introduce such restrictions as this \nwill only prevent future enhancements to packing heuristics.\n\nFor example, right now you already can't rely on having the exact same \npack output even on the same machine using the same arguments and the \nsame inputs simply by using threads.  As soon as you're using more than \none thread (most people do these days) then your pack output becomes non \ndeterministic.\n\n\nNicolas\n"},{"id":"149661","messageId":"4C7FD7A6.9090402@gmail.com","threadId":"24935","inReplyTo":"AANLkTikBnKQJmgOms2wK1+6fCLtHWiWkhuCVMN7kKLXP@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2010-09-02T16:58:14Z","receivedAt":"2010-09-02T16:58:14Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 09/02/2010 11:42 AM, Luke Kenneth Casson Leighton wrote:\n> On Thu, Sep 2, 2010 at 4:33 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>> On 09/02/2010 09:37 AM, Luke Kenneth Casson Leighton wrote:\n>>>\n>>> On Wed, Sep 1, 2010 at 11:04 PM, Nguyen Thai Ngoc Duy<pclouds@gmail.com>\n>>>   wrote:\n>>\n>> [...]\n>>>>\n>>>> There were discussions whether a pack is stable enough to\n>>>> be shared like this,\n>>>\n>>>   it seems to be.  as long as each version of git produces the exact\n>>> same pack object, off of the command \"git pack-objects --all --stdout\n>>> --thin {ref}<    {objref}\"\n>>\n>> This is not guaranteed.\n>\n>   ok.  greeeat.\n>\n>   so, some sensible questions:\n>\n>   * what _can_ be guaranteed?\n>\n>   * diffs?\n\nGiven a pre-image and a post-image, the diff/delta created will recreate \nthe post-image from the pre-image. The bit level representation of the \ndiff/delta is not guaranteed.\n\n>   * git-format-patches? (which i am aware can do binary files and also\n> rms)?\n\nSee above.\n\n> * individual files in the .git/objects directory?\n\nUncompressed: yes. Compressed: no.\n\n>   and, asking perhaps some silly questions:\n>\n> * why is it not guaranteed?\n\nThe pack format was created to move objects from one repository to \nanother. To do that efficiently, it uses many heuristics to decide how \nmuch information is _sufficient_ to do the job but but leaves it to the \nimplementation and user to decide the various trade offs. For instance, \nthere is no canonical the order of the object information in a pack or \nthat the pack must be minimal. This also allows for the multi-threaded \npack implementation.\n\n> * under what circumstances is it not guaranteed?  and, crucially, is\n> it necessary to care?   i.e. if someone does a shallow git clone, i\n> couldn't give a stuff.\n\nPretty much all of the time.\n\n> * is it possible to _make_ the repository guaranteed to produce\n> identical pack objects?\n\nIdentical code on identical systems with identical repositories without \nmulti-threading _might_ work.\n\n> * does for example \"git gc\" change the object store in such a way such\n> that one git repo will produce a different pack-object from the same\n> ref?  if so, can running \"git gc\" prior to producing the pack-objects\n> gurantee that the pack-objects will be the same?\n\nSee above.\n\n> * is it a versioning issue?  is it because there are different\n> versions (2 and 3)?  if so, that's ok, you just force people to use\n> the same pack-object versions.\n\nNo.\n\n> etc. etc.\n>\n> l.\n>\n"},{"id":"149662","messageId":"4C7FD998.9080900@gmail.com","threadId":"24935","inReplyTo":"AANLkTi=W3QwWSrNTie-K4QDDrucSVGQa5e3Ldy7m7ihy@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2010-09-02T17:06:32Z","receivedAt":"2010-09-02T17:06:32Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 09/02/2010 11:51 AM, Luke Kenneth Casson Leighton wrote:\n> On Thu, Sep 2, 2010 at 4:42 PM, Luke Kenneth Casson Leighton\n> <luke.leighton@gmail.com>  wrote:\n>> On Thu, Sep 2, 2010 at 4:33 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>\n>> * is it possible to _make_ the repository guaranteed to produce\n>> identical pack objects?\n>\n>   i.e. looking at these options, listed from\n> Documentation/technical/protocol.txt:\n>\n>      00887217a7c7e582c46cec22a130adf4b9d7d950fba0 HEAD\\0multi_ack\n> thin-pack side-band side-band-64k ofs-delta shallow no-progress\n> include-tag\n>\n>   is it possible to use shallow, thin-pack, side-band or side-band-64k\n> to guarantee that the pack object will be identical?\n\nNo. Looking to use identical packs created on different systems is not \nsomething that git guarantees or, likely, will ever guarantee. If you \nneed that, you need to create and implement the canonical pack-like \ndefinition for your transfer protocol.\n\n>   another important question:\n>\n> * if after performing a \"git unpack\" of one pack-object, can it be\n> guaranteed that performing a \"git pack-object\" on the *exact* same ref\n> and the *exact* same object-ref, will produce the *exact* same\n> pack-object that was used by \"git unpack\", as long as the exact same\n> arguments are used?  if not, why not, and if not under _some_\n> circumstances, under what circumstances _can_ the exact same\n> pack-object be retrieved that was just used?\n\nWrite your own \"packer\" is really the best answer.\n\n> if there is absolutely absolutely no way to guarantee that the\n> pack-objects can be the same, under no circumstances or combinations\n> of arguments or by forcing only compatible versions to communicate\n> etc. etc., a rather awful work-around can be applied which is to share\n> and permanently cache every single pack-object, rather than use what's\n> gone into the repo.\n\nSounds ugly and inefficient.\n"},{"id":"149663","messageId":"4C7FDA32.5050009@gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021233190.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2010-09-02T17:09:06Z","receivedAt":"2010-09-02T17:09:06Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 09/02/2010 12:41 PM, Nicolas Pitre wrote:\n\n[...]\n\n> I would go as far as stating that this is never guaranteed by design.\n> And I will oppose any attempt to introduce such restrictions as this\n> will only prevent future enhancements to packing heuristics.\n>\n> For example, right now you already can't rely on having the exact same\n> pack output even on the same machine using the same arguments and the\n> same inputs simply by using threads.  As soon as you're using more than\n> one thread (most people do these days) then your pack output becomes non\n> deterministic.\n\nFinally, the real pack expert weighs in!\n"},{"id":"149665","messageId":"alpine.LFD.2.00.1009021249510.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTikBnKQJmgOms2wK1+6fCLtHWiWkhuCVMN7kKLXP@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-02T17:21:24Z","receivedAt":"2010-09-02T17:21:24Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 2 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> On Thu, Sep 2, 2010 at 4:33 PM, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> > On 09/02/2010 09:37 AM, Luke Kenneth Casson Leighton wrote:\n> >>\n> >> On Wed, Sep 1, 2010 at 11:04 PM, Nguyen Thai Ngoc Duy<pclouds@gmail.com>\n> >>  wrote:\n> >\n> > [...]\n> >>>\n> >>> There were discussions whether a pack is stable enough to\n> >>> be shared like this,\n> >>\n> >>  it seems to be.  as long as each version of git produces the exact\n> >> same pack object, off of the command \"git pack-objects --all --stdout\n> >> --thin {ref}<  {objref}\"\n> >\n> > This is not guaranteed.\n> \n>  ok.  greeeat.\n> \n>  so, some sensible questions:\n> \n>  * what _can_ be guaranteed?\n\nYou can guarantee that if the SHA1 name of different packs is the same \nthen they contain the same set of objects.  Obviously their packed \nencoding will be different, and even the pack sizes might be quite \ndifferent too.\n\n>  * diffs?\n\nAgain that depends.  Over the evolution of Git, its diff library was \nmodified resulting in slightly different but valid equivalent diff \noutputs.\n\n>  * git-format-patches? (which i am aware can do binary files and also\n> rms)?\n\nSame as above.\n\n> * individual files in the .git/objects directory?\n\nWell, even then you can't guarantee they will be identical from one \nsystem to another.  That may depend on the zlib library version used for \nexample.\n\n>  and, asking perhaps some silly questions:\n> \n> * why is it not guaranteed?\n\nBecause it doesn't need to.\n\n> * under what circumstances is it not guaranteed?  and, crucially, is\n> it necessary to care?   i.e. if someone does a shallow git clone, i\n> couldn't give a stuff.\n\nLike I said, even repeating some repacking on the same machine with same \ninput is likely to produce slightly different packs because of \nthreading.  This is because the work set is divided between threads, and \nsince thread scheduling is not deterministic then some threads might not \nhave the same amount of CPU cycles given to them in relation with the \nother threads.  And when a thread is done with its work set, it will go \nand steal half of the work set from another thread with the most \namount of work \nstill left.  This has the effect of changing the delta pairing outcome \non the workset edges.\n\n> * is it possible to _make_ the repository guaranteed to produce\n> identical pack objects?\n\nSure, but performance will suck.\n\n> * does for example \"git gc\" change the object store in such a way such\n> that one git repo will produce a different pack-object from the same\n> ref?  if so, can running \"git gc\" prior to producing the pack-objects\n> gurantee that the pack-objects will be the same?\n\nNo.  The gc operation will combine multiple small packs into one and try \nto reuse as much data from those existing packs as possible without \nrecomputing it.  So you'll end up reusing whatever delta pairing you \nwere given from your peer the last time you cloned a repo or fetched an \nupdate.  And of course that clone/fetch was the result of a pack \ncombining operation on the sending end which itself tried to reuse as \nmuch of the existing data from different packs without recomputing it \ntoo.  Only the edges between different packs will be delta compressed in \nthose cases, using the particular heuristics that happen to be \nimplemented in the involved Git versions. So you may end up with a \ntotally different pack content containing data segments that originated \nfrom wildly random places on the net.\n\nThe only way to get a bit-for-bit reproducible pack one one specific \nsystem is to use 'git repack' with the -f switch, and limit it to only \none thread.\n\n> * is it a versioning issue?  is it because there are different\n> versions (2 and 3)?  if so, that's ok, you just force people to use\n> the same pack-object versions.\n\nNot at all.  FYI version 3 never was actually deployed so there is \neffectively only version 2 in play.  There are \"features\" such as \nOFS_DELTA that are negotiated when a pack is transferred over the git \nprotocol and if the receiver doesn't advertise them then the sender will \nconvert them on the fly into a compatible form.\n\nBut as the actual pack bitstream goes, it is totally unstable for all \nthe reasons I've stated so far.  Of course, Git being distributed must \nrely on some stable and universal representation of object content, \nhence their SHA1 references.  But their encoding doesn't have to be when \nall peers can cope with all the variations.\n\nI'm sorry as this isn't going to help you much unfortunately.\n\n\n\n\n\n\n\n\n\n\n> \n> etc. etc.\n> \n> l.\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n"},{"id":"149666","messageId":"alpine.LFD.2.00.1009021326290.19366@xanadu.home","threadId":"24935","inReplyTo":"4C7FDA32.5050009@gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-02T17:31:39Z","receivedAt":"2010-09-02T17:31:39Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 2 Sep 2010, A Large Angry SCM wrote:\n\n> On 09/02/2010 12:41 PM, Nicolas Pitre wrote:\n> \n> > For example, right now you already can't rely on having the exact same\n> > pack output even on the same machine using the same arguments and the\n> > same inputs simply by using threads.  As soon as you're using more than\n> > one thread (most people do these days) then your pack output becomes non\n> > deterministic.\n> \n> Finally, the real pack expert weighs in!\n\nBTW I just have a little time to quickly scan through my git mailing \nlist backlog these days, and stumbled on this by luck.  So if people \nwant my opinion on such matters it is safer to CC me directly.\n\n\nNicolas\n"},{"id":"149667","messageId":"AANLkTi=kO9USQYoTLQZyCRrjCHWRtPtd4S5EuFk4-gPv@mail.gmail.com","threadId":"24935","inReplyTo":"20100902155810.GB14508@sigill.intra.peff.net","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T18:07:07Z","receivedAt":"2010-09-02T18:07:07Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 4:58 PM, Jeff King <peff@peff.net> wrote:\n\n> pack-objects will reuse previously found deltas. So the deltas you have\n> in your existing packs matter. The deltas you have in your existing\n> packs depend on many things. At least:\n>\n>  1. Options you used when packing (e.g., --depth and --window).\n>\n>  2. Probably exactly _when_ you packed. You could find a good delta\n>     from A to B. Later, object C comes into existence, and would\n>     provide a better delta base for B. I don't think we will ever try A\n>     against C, unless --no-reuse-delta is set.\n>\n>     You have a different pack than somebody who packed after A, B, and\n>     C all existed.\n>\n>     In practice, this tends not to happen much because the best deltas\n>     are usually going backwards in time to a previous version. But it\n>     can happen.\n\n jeff, thanks for explaining (and to nicolas, i see, since beginning this)\n\n mrhmfffh.  just been reading Documentation/technical/pack-heuristics.txt.\n\n so... options include:\n\n * writing an alternative \"canonical\" pack-object algorithm.  i'm\ninclined to select \"git format-patch\"! :)  but that would be the lazy\nway....\n\n * taking the seeder's pack-objects as the \"canonical\" ones,\nregardless.  cacheing of the results would, sadly, be virtually\nunavoidable, given the situation (multi-threading etc.)\n\n * throw away bittorrent entirely as a transport mechanism.\n\n * force-feed one peer to be \"the\" provider of a particular given\npack.  doesn't matter whom you contact to _obtain_ a pack from, as\nlong as you solely and exclusively get the pack from that particular\npeer.\n\n * slight improvement / variation on the above: if two peers just\ncoincidentally happen to create or have the same pack (as can be shown\nby having the same SHA-1, and/or by having the same data in their\ncache) then ta-daaa, you have a file-sharing network for that\nparticular pack.\n\ni think.... i think i miiight be able, hmmm... i believe it would be\npossible to implement this last option by creating separate .torrents\nfor packs (one each!).  by splitting things down, so that pack objects\nare named as {ref}-{objref}-{SHA-1}.torrent and by providing a \"top\nlevel\" torrent which contains the refs/heads/* and the associate\nrev-list(s)... each set of rev-lists would have the SHA-1 of the\npack-object that happened to be created (and shared) at that\nparticular time, from that particular client: you then genuinely don't\ngive a stuff about who has what, it's all the same, and...\n\nhmmm, i feel a modification to / deviation from the bittorrent\nprotocol coming on :)\n\nl.\n"},{"id":"149668","messageId":"20100902182307.GC9955@fearengine.rdu.redhat.com","threadId":"24935","inReplyTo":"AANLkTi=kO9USQYoTLQZyCRrjCHWRtPtd4S5EuFk4-gPv@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Casey Dahlin","fromEmail":"cdahlin@redhat.com","sentAt":"2010-09-02T18:23:07Z","receivedAt":"2010-09-02T18:23:07Z","isPatch":false,"sender":{"key":"cdahlin@redhat.com","avatar":null},"body":"On Thu, Sep 02, 2010 at 07:07:07PM +0100, Luke Kenneth Casson Leighton wrote:\n> On Thu, Sep 2, 2010 at 4:58 PM, Jeff King <peff@peff.net> wrote:\n> i think.... i think i miiight be able, hmmm... i believe it would be\n> possible to implement this last option by creating separate .torrents\n> for packs (one each!).  by splitting things down, so that pack objects\n> are named as {ref}-{objref}-{SHA-1}.torrent and by providing a \"top\n> level\" torrent which contains the refs/heads/* and the associate\n> rev-list(s)... each set of rev-lists would have the SHA-1 of the\n> pack-object that happened to be created (and shared) at that\n> particular time, from that particular client: you then genuinely don't\n> give a stuff about who has what, it's all the same, and...\n> \n\nYou seem to have some misconceptions about refs too. refs are absolutely\nnot common between two machines. The branch \"master\" on my box might\nmatch the branch \"linux-2.6\" on git.kernel.org. And it might be from two\ndays ago.\n\nA ref is just an annotation in /one particular repository/ that a\nparticular commit is interesting to it for some reason.\n\n--CJD\n"},{"id":"149671","messageId":"AANLkTi=Q7EfeUDB6PuSa88PDtaBZSMMuaMqh8hU25ECb@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021326290.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T19:17:39Z","receivedAt":"2010-09-02T19:17:39Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 6:31 PM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Thu, 2 Sep 2010, A Large Angry SCM wrote:\n>\n>> On 09/02/2010 12:41 PM, Nicolas Pitre wrote:\n>>\n>> > For example, right now you already can't rely on having the exact same\n>> > pack output even on the same machine using the same arguments and the\n>> > same inputs simply by using threads.  As soon as you're using more than\n>> > one thread (most people do these days) then your pack output becomes non\n>> > deterministic.\n>>\n>> Finally, the real pack expert weighs in!\n>\n> BTW I just have a little time to quickly scan through my git mailing\n> list backlog these days, and stumbled on this by luck.  So if people\n> want my opinion on such matters it is safer to CC me directly.\n\n appreciated nicolas.  will keep it short.  ish :)\n\n * based on what you kindly mentioned about \"git repack -f\", would a\n(well-written!) patch to git pack-objects to add a\n\"--single-thread-only\" option be acceptable?\n\n * would you, or anyone else with enough knowledge of how this stuff\nreaallly works, be willing to put some low-priority back-of-mind\nthought into how to create a \"canonical\" pack format - one that can be\nenabled with a command-line-option?  the reason i ask is because if i\neven attempted such a task, i'd die of laughing (probably manically)\nif it was ever accepted.  i'd rather live :)\n\n\n questions (not necessarily for nicolas) - can anyone think of any\ngood reasons _other_ than for multiple file-sharing to have a\n\"canonical\" pack-object?\n\noff the top of my head i can think of one: rsync if the transfer is\ninterrupted.  if the pack-objects are large - and not guaranteed to be\nthe same - then an interrupted rsync transfer would be a bit of a\nwaste of bandwidth.  however if the pack-object could always be made\nthe same, the partial transfer could carry on.   musing a bit\nfurther... mmm... i supooose the same thing applies equally to http\nand ftp.  it's a bit lame, i know: can anyone think of any better\nreasons?\n\nl.\n"},{"id":"149672","messageId":"20100902192910.GJ32601@spearce.org","threadId":"24935","inReplyTo":"AANLkTi=Q7EfeUDB6PuSa88PDtaBZSMMuaMqh8hU25ECb@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2010-09-02T19:29:10Z","receivedAt":"2010-09-02T19:29:10Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Luke Kenneth Casson Leighton <luke.leighton@gmail.com> wrote:\n> \n>  * based on what you kindly mentioned about \"git repack -f\", would a\n> (well-written!) patch to git pack-objects to add a\n> \"--single-thread-only\" option be acceptable?\n\nProbably not.  I can't think of a good reason to limit the number\nof threads that get used.  We already have pack.threads as a\nconfiguration variable to support controlling this for the system,\nbut that's about the only thing that really makes sense.\n \n>  * would you, or anyone else with enough knowledge of how this stuff\n> reaallly works, be willing to put some low-priority back-of-mind\n> thought into how to create a \"canonical\" pack format\n\nWe have.  We've even talked about it on the mailing list.  Multiple\ntimes.  Most times about how to support a p2p Git transport.\nThat whole Gittorrent thing you are ignoring, we put some effort\ninto coming up with a pack-like format that would be more stable,\nat the expense of being larger in total size.\n\n>  questions (not necessarily for nicolas) - can anyone think of any\n> good reasons _other_ than for multiple file-sharing to have a\n> \"canonical\" pack-object?\n\nYes, its called resuming a clone over git://.\n\nRight now if you abort git:// you break the pack stream, and it\ncannot be restarted.  If we had a more consistent encoding we may\nbe able to restart an aborted clone.\n\nBut we can't solve it.  Its a _very_ hard problem.\n\nNico, myself, and a whole lot of other very smart folks who really\nunderstand how Git works today have failed to identify a way to do\nthis that we actually want to write, include in git, and maintain\nlong-term.  Sure, anyone can come up with a specification that says\n\"put this here, that there, break ties this way\".  But we don't\nwant to bind our hands and maintain those rules.\n \n> off the top of my head i can think of one: rsync if the transfer is\n> interrupted.  if the pack-objects are large - and not guaranteed to be\n> the same - then an interrupted rsync transfer would be a bit of a\n> waste of bandwidth.  however if the pack-object could always be made\n> the same, the partial transfer could carry on.   musing a bit\n> further... mmm... i supooose the same thing applies equally to http\n> and ftp.  it's a bit lame, i know: can anyone think of any better\n> reasons?\n\nWe already do with this http:// and ftp:// during fetch or clone.\nWe try to resume with a byte range request, and validate the SHA-1\ntrailer on the end of the pack file after download.  If it doesn't\nmatch, we throw the file away and restart the entire thing.\n\nIn general pack files don't change that often, so there are fairly\ngood odds that resuming an aborted clone only a few hours after\nit aborted would succeed by simply resuming the file download.\nBut every week or two (or even nightly!) its common for packs to\nbe completely rewritten (when the repository owner does `git gc`),\nso we really cannot rely on packs being stable long-term.\n\n-- \nShawn.\n"},{"id":"149690","messageId":"AANLkTinFPxsY6frVnga8u15aovQarfWreBYJfri6ywoK@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021249510.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T19:41:57Z","receivedAt":"2010-09-02T19:41:57Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"nicolas, thanks for responding: you'll see this some time in the\nfuture when you catch up, it's not a high priority, nothing new, just\nthinking out loud, for benefit of archives.\n\nOn Thu, Sep 2, 2010 at 6:21 PM, Nicolas Pitre <nico@fluxnic.net> wrote:\n>>  * what _can_ be guaranteed?\n>\n> You can guarantee that if the SHA1 name of different packs is the same\n> then they contain the same set of objects.  Obviously their packed\n> encoding will be different, and even the pack sizes might be quite\n> different too.\n\n ack.  ok, so the idea of creating lots and lots of 2nd level\n.torrents by name {ref}-{commitref}-{SHA-1}.torrent is about the only\nway to get around that.\n\n@begin lots of no, no and hell no...\n> [...]\n@end\n\n dang.  diffs, versions, threads and zlibs as well, all conspiring against me :)\n\n>> * is it possible to _make_ the repository guaranteed to produce\n>> identical pack objects?\n>\n> Sure, but performance will suck.\n\n that's fiiine :)  as i've learned on the pyjamas project, it's rare\nthat you have speed and interoperability at the same time...\n\n> The only way to get a bit-for-bit reproducible pack one one specific\n> system is to use 'git repack' with the -f switch, and limit it to only\n> one thread.\n\n whew - a way out, at last.  you had me going, for a minute :)\n\n>> * is it a versioning issue?  is it because there are different\n>> versions (2 and 3)?  if so, that's ok, you just force people to use\n>> the same pack-object versions.\n>\n> Not at all.  FYI [....]\n\n appreciated.\n\n\n> I'm sorry as this isn't going to help you much unfortunately.\n\n neeh, i'm flexible.  it looks like i'm going to need to deviate from\nbittorrent, after all, start adding new commands over which the git\nrev-list gets transferred, rather than as a VFS layer.  the reason is\nthat bittorrent depends on the files and the data in the files being\nall the same, so that a hash can be taken of the whole lot and the\nend-result verified.\n\n if the pack-objects are going to vary, then the VFS layer idea is\nblown completely out the water, except for the absolute basic\nmeta-info such as \"refs/heads/*\".  so i might as well just use\n\"actual\" bittorrent to transfer packs via\n{ref}-{commitref}-{SHA-1}.torrent.\n\nho hum, drawing board we come...\n\n\nl.\n"},{"id":"149701","messageId":"AANLkTi=D7t+cNSdmZORwBEUU6wy-4QXHpw6YcXe61ADa@mail.gmail.com","threadId":"24935","inReplyTo":"20100902192910.GJ32601@spearce.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T19:51:32Z","receivedAt":"2010-09-02T19:51:32Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 8:29 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Luke Kenneth Casson Leighton <luke.leighton@gmail.com> wrote:\n>>\n>>  * based on what you kindly mentioned about \"git repack -f\", would a\n>> (well-written!) patch to git pack-objects to add a\n>> \"--single-thread-only\" option be acceptable?\n>\n> Probably not.  I can't think of a good reason to limit the number\n> of threads that get used.  We already have pack.threads as a\n> configuration variable to support controlling this for the system,\n> but that's about the only thing that really makes sense.\n>\n>>  * would you, or anyone else with enough knowledge of how this stuff\n>> reaallly works, be willing to put some low-priority back-of-mind\n>> thought into how to create a \"canonical\" pack format\n>\n> We have.  We've even talked about it on the mailing list.  Multiple\n> times.  Most times about how to support a p2p Git transport.\n> That whole Gittorrent thing you are ignoring,\n\n i'm not ignoring it - it was abandoned and sam created mirrorsync\ninstead!  and i can't ignore something when all the damn information\non it has been withdrawn from the internet!  i _have_ been looking,\nand just can't darn well find anything.   fortunately, i'm reasonably\nbright, catch on fast, and listen well.  ok.  _sometimes_ i listen\nwell :)\n\n> we put some effort\n> into coming up with a pack-like format that would be more stable,\n> at the expense of being larger in total size.\n\n ahhh goood.\n\n> Nico, myself, and a whole lot of other very smart folks who really\n> understand how Git works today have failed to identify a way to do\n> this that we actually want to write, include in git, and maintain\n> long-term.\n\n bugger.  *sigh* ok.  so, scratch that question, nico (the\ncanonical-pack question but not the --single-thread one)\n\n so, this, and...\n\n> In general pack files don't change that often, so there are fairly\n\n ... this, all tend to point towards the idea of sharing packs by\n{ref}-{commitref}-{SHA1}.torrent as being a reasonabe and \"good\nenough\" idea.  on the basis that anyone who happens to be doing\ngit-sharing _right now_ is likely to end up sharing the exact same\npack object generated by the same one (original) seed.\n\ni'd better start looking at bittornado in more detail...\n\n l.\n\np.s. thank you to everyone who's responding, i dunno about you but\nthis is fascinating.\n"},{"id":"149702","messageId":"4C800062.7020707@gmail.com","threadId":"24935","inReplyTo":"AANLkTinFPxsY6frVnga8u15aovQarfWreBYJfri6ywoK@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2010-09-02T19:52:02Z","receivedAt":"2010-09-02T19:52:02Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 09/02/2010 03:41 PM, Luke Kenneth Casson Leighton wrote:\n\n[...]\n\n>   neeh, i'm flexible.  it looks like i'm going to need to deviate from\n> bittorrent, after all, start adding new commands over which the git\n> rev-list gets transferred, rather than as a VFS layer.  the reason is\n> that bittorrent depends on the files and the data in the files being\n> all the same, so that a hash can be taken of the whole lot and the\n> end-result verified.\n>\n>   if the pack-objects are going to vary, then the VFS layer idea is\n> blown completely out the water, except for the absolute basic\n> meta-info such as \"refs/heads/*\".  so i might as well just use\n> \"actual\" bittorrent to transfer packs via\n> {ref}-{commitref}-{SHA-1}.torrent.\n>\n> ho hum, drawing board we come...\n\nYou can treat the git object store (both loose and packed objects) as a \nVFS and use the refs of interest (the I-needs and the I-haves) to create \na view...\n"},{"id":"149704","messageId":"AANLkTimDi=KYZ7Bs4C+WEGoP8y-yzjynddWpkxohWoix@mail.gmail.com","threadId":"24935","inReplyTo":"20100902192910.GJ32601@spearce.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T20:06:22Z","receivedAt":"2010-09-02T20:06:22Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 8:29 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Luke Kenneth Casson Leighton <luke.leighton@gmail.com> wrote:\n>>\n>>  * based on what you kindly mentioned about \"git repack -f\", would a\n>> (well-written!) patch to git pack-objects to add a\n>> \"--single-thread-only\" option be acceptable?\n>\n> Probably not.  I can't think of a good reason to limit the number\n> of threads that get used.\n\n i can - so that git pack-objects, after \"git repack -f\", returns a\ncanonical pack! :)\n\n> We already have pack.threads as a\n> configuration variable to support controlling this for the system,\n> but that's about the only thing that really makes sense.\n\n ookaaay, so that would work, but would force that particular repo to\nbe _entirely_ single-threaded, which is non-optimal when all that's\nneeded is \"git pack-objects\" to be single-threaded.  sure, i could\nwrite a hack with some shell-script which takes the current option /\nvalue for \"pack.threads\", stores it, changes it to 1, runs \"git\npack-objects\" and then changes it back again, but... eeuw.  (not to\nmention race-conditions....)\n\n l.\n"},{"id":"149709","messageId":"FyEIt68YHr11lsX_CGcHmYfITTgX-iSs9tVNIBMG7FQ_WhGc4ttvXw@cipher.nrlssc.navy.mil","threadId":"24935","inReplyTo":"20100902192910.GJ32601@spearce.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Brandon Casey","fromEmail":"brandon.casey.ctr@nrlssc.navy.mil","sentAt":"2010-09-02T20:28:28Z","receivedAt":"2010-09-02T20:28:28Z","isPatch":false,"sender":{"key":"brandon.casey.ctr@nrlssc.navy.mil","avatar":null},"body":"On 09/02/2010 02:29 PM, Shawn O. Pearce wrote:\n> Luke Kenneth Casson Leighton <luke.leighton@gmail.com> wrote:\n>>\n>>  * based on what you kindly mentioned about \"git repack -f\", would a\n>> (well-written!) patch to git pack-objects to add a\n>> \"--single-thread-only\" option be acceptable?\n> \n> Probably not.  I can't think of a good reason to limit the number\n> of threads that get used.  We already have pack.threads as a\n> configuration variable to support controlling this for the system,\n> but that's about the only thing that really makes sense.\n\nI think pack-objects already has a --threads option allowing to\nspecify the number of threads to use.\n\n   --threads=1\n\nshould do it.\n\n-Brandon\n"},{"id":"149715","messageId":"m3y6bjnadu.fsf@localhost.localdomain","threadId":"24935","inReplyTo":"20100902192910.GJ32601@spearce.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-09-02T20:45:43Z","receivedAt":"2010-09-02T20:45:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> Luke Kenneth Casson Leighton <luke.leighton@gmail.com> wrote:\n\n> >  * would you, or anyone else with enough knowledge of how this stuff\n> > reaallly works, be willing to put some low-priority back-of-mind\n> > thought into how to create a \"canonical\" pack format\n> \n> We have.  We've even talked about it on the mailing list.  Multiple\n> times.  Most times about how to support a p2p Git transport.\n> That whole GitTorrent thing you are ignoring, we put some effort\n> into coming up with a pack-like format that would be more stable,\n> at the expense of being larger in total size.\n\nIf I remember it correctly the main idea of GitTorrent was instead of\ndividing file into pieces of data like in BitTorrent (pieces being\ndownloaded in parallel from different peers) it divides set of objects\ninto \"reels\" (which are special case of bundles, IIRC).\n\n> >  questions (not necessarily for nicolas) - can anyone think of any\n> > good reasons _other_ than for multiple file-sharing to have a\n> > \"canonical\" pack-object?\n> \n> Yes, its called resuming a clone over git://.\n> \n> Right now if you abort git:// you break the pack stream, and it\n> cannot be restarted.  If we had a more consistent encoding we may\n> be able to restart an aborted clone.\n> \n> But we can't solve it.  Its a _very_ hard problem.\n> \n> Nico, myself, and a whole lot of other very smart folks who really\n> understand how Git works today have failed to identify a way to do\n> this that we actually want to write, include in git, and maintain\n> long-term.  Sure, anyone can come up with a specification that says\n> \"put this here, that there, break ties this way\".  But we don't\n> want to bind our hands and maintain those rules.\n\nIf I remember the discussion stalled (i.e. no working implementation),\nand one of the latest proposals was to have some way of recovering\nobjects from partially downloaded file, and a way to request packfile\nwithout objects that got already downloaded.\n\nIIRC.\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"149716","messageId":"AANLkTin-gwN+QXjjwr7UPMi=QwSFdDjRBqmx9jgunhJM@mail.gmail.com","threadId":"24935","inReplyTo":"FyEIt68YHr11lsX_CGcHmYfITTgX-iSs9tVNIBMG7FQ_WhGc4ttvXw@cipher.nrlssc.navy.mil","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T20:48:44Z","receivedAt":"2010-09-02T20:48:44Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 9:28 PM, Brandon Casey\n<brandon.casey.ctr@nrlssc.navy.mil> wrote:\n> I think pack-objects already has a --threads option allowing to\n> specify the number of threads to use.\n>\n>   --threads=1\n>\n> should do it.\n\n git pack-objects --help\n ....\n  ....\n       --threads=<n>\n           Specifies the number of threads to spawn when searching for best\n           delta matches.\n\n wha-hey!  thank youuu brandon.\n\n okaay... so maaaybeee there's a workaround (not involving patches to git. whew)\n\nl.\n"},{"id":"149720","messageId":"AANLkTikSHXivniUk-1KU30Ws23ebnbDhOmjKmpmVH-Y9@mail.gmail.com","threadId":"24935","inReplyTo":"m3y6bjnadu.fsf@localhost.localdomain","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T21:10:55Z","receivedAt":"2010-09-02T21:10:55Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Thu, Sep 2, 2010 at 9:45 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n\n> If I remember the discussion stalled (i.e. no working implementation),\n> and one of the latest proposals was to have some way of recovering\n> objects from partially downloaded file, and a way to request packfile\n> without objects that got already downloaded.\n\n oo.  ouch.  i can understand why things stalled, then.  you're\neffectively adding an extra layer in, and even if you could add a\nunique naming scheme on those objects (if one doesn't already exist?),\nthose object might (or might not!) come up the second time round (for\nreasons mentioned already - threads resulting in different deltas\nbeing picked etc.) ... and if they weren't picked for the re-generated\npack, you'd have to _delete_ them from the receiving end so as to\navoid polluting the recipient's object store haaarrgh *spit*, *cough*.\n\n ok.\n\n what _might_ work however iiiiIiis... to split the pack-object into\ntwo parts.  or, to add an \"extra part\", to be more precise:\n\na) complete list of all objects.  _just_ the list of objects.\nb) existing pack-object format/structure.\n\nin this way, the sender having done all the hard work already of\ndetermining what objects are to go into a pack-object, transfers that\n*first*.  _theeen_ you begin transferring the pack-object.  theeeen,\nif the pack-object transfer is ever interrupted, you simply send back\nthat list of objects, and ask \"uhh, you know that list of objects we\nwere talking about?  well, here it is *splat* - are you able to\nrecreate the pack-object from that, for me, and if so please gimme\nagain\"\n\nand, 10^N-1 times out of 10^N, for reasons that shawn kindly\nexplained, i bet you the answer would be \"yes\".\n\n... um... in fact... um... i believe i'm merely talking about the .idx\nindex file, aren't i?  because... um... the index file contains the\nlist of object refs in the pack, yes?\n\nsooo.... taking a wild guess, here: if you were to parse the .idx file\nand extract the list of object-refs, and then pass that to \"git\npack-objects --window=0 --delta=0\", would you end up with the exact\nsame pack file, because you'd forced git pack-objects to only return\nthat specific list of object-refs?\n\nl.\n"},{"id":"149723","messageId":"AANLkTimTe8_P7uHA6Ytm3+3Ha7rygWHzrSFOUq7fdX-L@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTikSHXivniUk-1KU30Ws23ebnbDhOmjKmpmVH-Y9@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-02T21:19:48Z","receivedAt":"2010-09-02T21:19:48Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"sorry, shouldn't have hit send so quick.\n\nOn Thu, Sep 2, 2010 at 10:10 PM, Luke Kenneth Casson Leighton\n<luke.leighton@gmail.com> wrote:\n\n> were talking about?  well, here it is *splat* - are you able to\n> recreate the pack-object from that, for me, and if so please gimme\n> again\"\n\n> sooo.... taking a wild guess, here: if you were to parse the .idx file\n> and extract the list of object-refs, and then pass that to \"git\n> pack-objects --window=0 --delta=0\", would you end up with the exact\n> same pack file, because you'd forced git pack-objects to only return\n> that specific list of object-refs?\n\n becauuuuse... if soooo... then that's a solution!  you get the .idx\nfile first, you make sure that that's cached (and it's small, isn't\nit, so that would be ok), then you transfer it around the network, and\nyou ask each repo if they can re-create the pack-object from the\ncontents of the .idx, if they can, blam, they're a seed/sharer for\nthat pack-object; if they can't, tough titty, they've probably had git\ngc run, or had some more commits added, or whatever, and can't\nparticipate, wow big deal, no great loss.\n\ndoes that fly? :)  if it doesn't fly (with --window=0 --delta=0) would\nit be easy to add an option to say \"oi, gimme a pack with nothing but\nthese refs, in exactly this order, no quibbling, no questions, just\ngimme\"?  and you could call it \"the same as the original\" because,\nduh, it would contain exactly the same objects as the original.\n\nam i even vaaaguely along the right lines? *scratches head*.\n\nl.\n"},{"id":"149735","messageId":"alpine.LFD.2.00.1009021624170.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTinFPxsY6frVnga8u15aovQarfWreBYJfri6ywoK@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-02T23:09:45Z","receivedAt":"2010-09-02T23:09:45Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 2 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> nicolas, thanks for responding: you'll see this some time in the\n> future when you catch up, it's not a high priority, nothing new, just\n> thinking out loud, for benefit of archives.\n\nWell, I might as well pay more attention to *you* now.  :-)\n\n> >> * is it possible to _make_ the repository guaranteed to produce\n> >> identical pack objects?\n> >\n> > Sure, but performance will suck.\n> \n>  that's fiiine :)  as i've learned on the pyjamas project, it's rare\n> that you have speed and interoperability at the same time...\n\nWell, did you hear about this thing called Git?  It appears that those \nGit developers are performance freaks.  :-)  Yet, Git is interoperable \nacross almost all versions ever released because we made sure that only \nfundamental things are defined and relied upon.  And that excludes \nactual delta pairing and pack object ordering.  That's why a pack file \nmay have many different byte sequences and yet still represent the same \ncanonical data.\n\n>  if the pack-objects are going to vary, then the VFS layer idea is\n> blown completely out the water, except for the absolute basic\n> meta-info such as \"refs/heads/*\".  so i might as well just use\n> \"actual\" bittorrent to transfer packs via\n> {ref}-{commitref}-{SHA-1}.torrent.\n\nFor the archive benefit, here's what I think on the whole idea.\n\nThe BitTorrent model is simply unappropriate for Git.  It doesn't fit to \nthe Git model at all as BitTorrent works on stable and static data, and \nrequires a lot of people wanting that same data.\n\nWhen you perform a fetch, Git does actually negociate with the server to \nfigure out what's missing locally, and the server does produce a custom \npack for you that is optimized so that only what's needed for you to be \nup to date is transferred.\n\nEven if you try to cache a set of packs to suit the BitTorrent static \ndata model, you'll need so many packs to cover all the possible gaps \nbetween a server and a random number of clients each with a random \nrepository state.  Of course it is possible to have bigger packs \ncovering larger gaps, but then you lose the biggest advantage that the \nsmart Git protocol has.  And with smaller, more fine grained packs, \nyou'll end up with so many of them that finding a live torrent for the \nactual one you need is going to be difficult.\n\n> ho hum, drawing board we come...\n\nYep.  Instead of transferring packs, a BitTorrent-alike transfer should \nbe based on the transfer of _objects_.  Therefore you can make the \ncorrespondance between file chunks in BitTorrent with objects in a Git \naware system.  So, when contacting a peer, you could negociate what is \nthe set of objects that the peer has that you don't, and vice versa.  \nObjects in Git are stable and immutable, and they all have a unique SHA1 \nsignature.  And to optimize the negociation, the pack index content can \nbe used, first by exchanging the content of the first level \nfan-out table and ignoring those entries that are equal.  This for each \npeer.\n\nThen, each peer make requests to connected peers for objects that those \npeers have but that isn't available locally, just like chunks in \nBitTorrent.\n\nBut here's the twist to make this scale well.  Since the object sender \nknows what objects the receiver already has, it therefore can choose the \nobject encoding.  Meaning that the sender can simply *reuse* a delta \nencoding for an object it is requested to send if the requestor already \nhas the base object for this delta.\n\nSo in most cases, the object to send will be small, especially if it is \na delta object.  That should fit the \nchunk model.  But if an object is bigger than a certain treshold, then \nits transfer could be chunked across multiple peers just like classic \nBitTorrent.  In this case, the chunking would need to be done on \nthe non delta uncompressed object data as this is the only thing that is \nuniversally stable (doesn't mean that the _transfer_ of those chunks \ncan't be compressed).\n\nNow this design has many open questions, such as finding out what is the \nlatest set of refs amongst all the peers, whether or not what we have \nlocally are ancestors of the remote refs, etc.\n\nAnd of course, while this will make for a speedy object transfer, the \nresulting mess on the receiver's end will have to be validated \nand repacked in the end.  So overall this might not end up being faster \noverall for the fetcher.\n\n\nNicolas\n"},{"id":"149742","messageId":"alpine.LFD.2.00.1009021931340.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTikSHXivniUk-1KU30Ws23ebnbDhOmjKmpmVH-Y9@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-03T00:29:26Z","receivedAt":"2010-09-03T00:29:26Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 2 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> On Thu, Sep 2, 2010 at 9:45 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n> \n> > If I remember the discussion stalled (i.e. no working implementation),\n> > and one of the latest proposals was to have some way of recovering\n> > objects from partially downloaded file, and a way to request packfile\n> > without objects that got already downloaded.\n> \n>  oo.  ouch.  i can understand why things stalled, then.  you're\n> effectively adding an extra layer in, and even if you could add a\n> unique naming scheme on those objects (if one doesn't already exist?),\n> those object might (or might not!) come up the second time round (for\n> reasons mentioned already - threads resulting in different deltas\n> being picked etc.) ... and if they weren't picked for the re-generated\n> pack, you'd have to _delete_ them from the receiving end so as to\n> avoid polluting the recipient's object store haaarrgh *spit*, *cough*.\n\nWell, actually there is no need to delete anything.  Git can cope with \nduplicated objects just fine.  A subsequent gc will get rid of the \nduplicates automatically.\n\n>  what _might_ work however iiiiIiis... to split the pack-object into\n> two parts.  or, to add an \"extra part\", to be more precise:\n> \n> a) complete list of all objects.  _just_ the list of objects.\n> b) existing pack-object format/structure.\n> \n> in this way, the sender having done all the hard work already of\n> determining what objects are to go into a pack-object, transfers that\n> *first*.  _theeen_ you begin transferring the pack-object.  theeeen,\n> if the pack-object transfer is ever interrupted, you simply send back\n> that list of objects, and ask \"uhh, you know that list of objects we\n> were talking about?  well, here it is *splat* - are you able to\n> recreate the pack-object from that, for me, and if so please gimme\n> again\"\n\nWell, it isn't that simple.\n\nFirst, a resumable clone is useful only when there is a big transfer in \nplay.  Otherwise it isn't worth the trouble.\n\nSo, if the clone is big, then this list of objects can be in the \nmillions.  For example my linux kernel repo with a couple branches \ncurrently has:\n\n$ git rev-list --all --objects | wc -l\n2808136\n\nSo, 2808136 objects, with 20-byte SHA1 for each of them, and you have a \n54 MB object list to transfer already.  This is a significant overhead \nthat we prefer to avoid, given the actual pack transfer which is:\n\n$ git pack-objects --all --stdout --progress < /dev/null | wc -c\nCounting objects: 2808136, done.\nCompressing objects: 100% (384219/384219), done.\n645201934\nTotal 2808136 (delta 2422420), reused 2788225 (delta 2402700)\n\nThe output from wc is 645201934 = 615 MB for this repository.  Hence the \nlist of object alone is quite significant.\n\nAnd even then, what if the transfer crashes during that object list \ntransfer?  On flaky connections this might happen within 54 MB.\n\n> and, 10^N-1 times out of 10^N, for reasons that shawn kindly\n> explained, i bet you the answer would be \"yes\".\n\nFor the list of objects, sure.  But that isn't a big deal.  It is easy \nenough to tell the remote about the commits we already have and ask for \nthe rest.  With a commit SHA1, the remote can figure out all the objects \nwe have. But all is in that determination of the latest commit we have.  \nIf we get a partial pack, it is possible to somehow salvage as many \nobjects from it, and determine what top commit(s) that correspond to.  \nIt is possible to set your local repo just as if you had requested a \nshallow clone and then the resume would simply be a deepening of that \nshallow clone.\n\nBut usually the very first commit in a pack is huge as it typically \nisn't delta compressed (a delta chain has to start somewhere).  And this \nfirst commit will roughly represent the same size as a tarball for that \ncommit.  And if you don't get at least that first commit then you are \nscrewed.  Or if you don't get a complete second commit when deepening a \nclone you are still screwed.\n\nAnother issue is what to do with objects that are themselves huge.\n\nYet another issue: what to do with all those objects I've got in my \npartial pack, but that I can't connect to any commit yet.  We don't want \nthem transferred again but it isn't easy to tell the remote about them.\n\nYou could tell the remote: \"I have this pack for this commit from this \ncommit but I got only this amount of bytes from it, please resume \ntransfer here.\"  But as mentioned before the pack stream is not \ndeterministic, and we really don't want to make it single-threaded on a \nserver.  Furthermore this is a lot of work for the server as even if the \npack stream is deterministic, then the server still has to recreate the \nfirst part of the pack just to throw it away until the desired offset is \nreached.  And caching pack results also has all sorts of implications \nwe've prefered to avoid on a server for security reasons (better keep \nserving operations read-only).\n\n> ... um... in fact... um... i believe i'm merely talking about the .idx\n> index file, aren't i?  because... um... the index file contains the\n> list of object refs in the pack, yes?\n\nIn one pack, yes.  You might have multiple packs.  And that doesn't mean \nthat all the objects from a pack are all relevant to the actual branches \nyou are willing to export.\n\n> sooo.... taking a wild guess, here: if you were to parse the .idx file\n> and extract the list of object-refs, and then pass that to \"git\n> pack-objects --window=0 --delta=0\", would you end up with the exact\n> same pack file, because you'd forced git pack-objects to only return\n> that specific list of object-refs?\n\nIf you do this i.e. turn off delta compression, then the 615 MB \nrepository above will turn itself into a multi-gigabyte pack!\n\n\nNicolas\n"},{"id":"149743","messageId":"alpine.LFD.2.00.1009022033520.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTimDi=KYZ7Bs4C+WEGoP8y-yzjynddWpkxohWoix@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-03T00:36:31Z","receivedAt":"2010-09-03T00:36:31Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 2 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> On Thu, Sep 2, 2010 at 8:29 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> > Luke Kenneth Casson Leighton <luke.leighton@gmail.com> wrote:\n> >>\n> >>  * based on what you kindly mentioned about \"git repack -f\", would a\n> >> (well-written!) patch to git pack-objects to add a\n> >> \"--single-thread-only\" option be acceptable?\n> >\n> > Probably not.  I can't think of a good reason to limit the number\n> > of threads that get used.\n> \n>  i can - so that git pack-objects, after \"git repack -f\", returns a\n> canonical pack! :)\n\nBut did you try it?  The -f means \"don't reuse any existing pack data \nand recompute every delta from scratch to find the best matches\".  This \nis a very very costly operation that most people happily live without.\n\n\nNicolas\n"},{"id":"149746","messageId":"AANLkTik1hfe3jVWy236611d7hdP=yt+d3vCBiGvDa26H@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021931340.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2010-09-03T02:48:21Z","receivedAt":"2010-09-03T02:48:21Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Sep 3, 2010 at 10:29 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> But usually the very first commit in a pack is huge as it typically\n> isn't delta compressed (a delta chain has to start somewhere).  And this\n> first commit will roughly represent the same size as a tarball for that\n> commit.  And if you don't get at least that first commit then you are\n> screwed.  Or if you don't get a complete second commit when deepening a\n> clone you are still screwed.\n\nElijah's recent work on \"rev-list --objects -- pathspec\" [1] may help\nsplit a commit into many parts that can be sent separately.\n\n[1] http://mid.gmane.org/1282803711-10253-1-git-send-email-newren@gmail.com\n\n> Another issue is what to do with objects that are themselves huge.\n\nFor big blobs, it's probably best sending them separately so they can\nbe resumed.\n-- \nDuy\n"},{"id":"149752","messageId":"AANLkTi=r21khJLMvwH7h2M_Yqv8QFAD-L2RX8NrbEjoc@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021931340.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-03T10:23:41Z","receivedAt":"2010-09-03T10:23:41Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"very quick reply am off out\n\nOn Fri, Sep 3, 2010 at 1:29 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n>> sooo.... taking a wild guess, here: if you were to parse the .idx file\n>> and extract the list of object-refs, and then pass that to \"git\n>> pack-objects --window=0 --delta=0\", would you end up with the exact\n>> same pack file, because you'd forced git pack-objects to only return\n>> that specific list of object-refs?\n>\n> If you do this i.e. turn off delta compression, then the 615 MB\n> repository above will turn itself into a multi-gigabyte pack!\n\n sorry nicolas, i believe you may have misunderstood.  first obtain\n.idx which will return a delta-compressed list of objects, yes?  then\nuse --window=0 --delta=0 with exact same list, surely you will end up\nwith the exact same list which you gave to the *previous* command,\nyes?\n\n$ git pack-objects\n>> .pack and .idx\n$ mv .pack .pack2\n$ extract_pack_objects_from_idx.sh .idx > foo\n$ git pack-objects --window=0 --delta=0 < foo\n$ diff .pack .pack2\n\nso no, of _course_ not just ask --window=0 --delta=0 for the very\nfirst run: if that's what i'd said, that would indeed be dumb.\n\nl.\n"},{"id":"149753","messageId":"AANLkTimFwSZY5b=w32oiX05u--=tFw0-A8kXTtghMpSg@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009022033520.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-03T10:34:21Z","receivedAt":"2010-09-03T10:34:21Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Fri, Sep 3, 2010 at 1:36 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n\n>>  i can - so that git pack-objects, after \"git repack -f\", returns a\n>> canonical pack! :)\n>\n> But did you try it?  The -f means \"don't reuse any existing pack data\n> and recompute every delta from scratch to find the best matches\".\n\n yehh, i tried it - on a small 2mb repo.  whoopsie.  ok.  still one or\ntwo ideas left.\n"},{"id":"149754","messageId":"B757A854-C7BF-4CBF-9132-91D205344606@mit.edu","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021624170.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2010-09-03T10:37:12Z","receivedAt":"2010-09-03T10:37:12Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"\nOn Sep 2, 2010, at 7:09 PM, Nicolas Pitre wrote:\n\n> The BitTorrent model is simply unappropriate for Git.  It doesn't fit to \n> the Git model at all as BitTorrent works on stable and static data, and \n> requires a lot of people wanting that same data.\n\nI wonder if this would work.   Assume for the moment that once every N days (where N might be one month), the \"Gittorrent master\" generates a \"canonical pack construction file\" that contains information about the objects, their delta pairing, the version of diff and zlib required, etc., and the SHA1 of the expected resulting encoding of this pack file.   This would contain the canonical pack for the entire repository, and for the Linux kernel, it might be released on www.kernel.org, and contain all of the objects from the beginning of time to the latest commit in Linus Torvalds' repository.\n\nAlthough everybody's git repository will have different packs that they use and store natively, by following the \"canonical pack construction file\" they will be able to build up a canonical encoding for the pack that can then be used to more quickly allow newbie developers who are pulling the full clone of Linus's tree for the first time to download using a peer2peer Bittorrent style download.   So people who are willing to participate as part of the peer2peer network can download the instructions for how to make the canonical pack once a month, and use it to create the canonical pack.  If the \"Gittorrent master\" has spent a lot of time to carefully compute the most efficient set of delta pairings, they will get the slight benefit of a more efficient pack which they could use instead of th\n eir local one without having to use large values of --window and --depth to \"git repack\". \n\nThis allows the peer2peer download to be used where it most matters --- for the bulk download for people who are cloning from the \"canonical repository\" for the first time.  After that, they will no doubt find it far more efficient to download incremental uploads using the git protocol. \n\n-- Ted\n"},{"id":"149755","messageId":"AANLkTimoz6Ux18mtmm4oVr21p6RmKURgcWWK8JkWvmu5@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009021931340.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-03T10:54:12Z","receivedAt":"2010-09-03T10:54:12Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"have a bit more time.\n\nOn Fri, Sep 3, 2010 at 1:29 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n>> pack, you'd have to _delete_ them from the receiving end so as to\n>> avoid polluting the recipient's object store haaarrgh *spit*, *cough*.\n>\n> Well, actually there is no need to delete anything.  Git can cope with\n> duplicated objects just fine.  A subsequent gc will get rid of the\n> duplicates automatically.\n\n excellent.  good to hear.\n\n>>  what _might_ work however iiiiIiis... to split the pack-object into\n>> two parts.  or, to add an \"extra part\", to be more precise:\n>>\n>> a) complete list of all objects.  _just_ the list of objects.\n>> b) existing pack-object format/structure.\n>>\n>> in this way, the sender having done all the hard work already of\n>> determining what objects are to go into a pack-object, transfers that\n>> *first*.  _theeen_ you begin transferring the pack-object.  theeeen,\n>> if the pack-object transfer is ever interrupted, you simply send back\n>> that list of objects, and ask \"uhh, you know that list of objects we\n>> were talking about?  well, here it is *splat* - are you able to\n>> recreate the pack-object from that, for me, and if so please gimme\n>> again\"\n>\n> Well, it isn't that simple.\n>\n> First, a resumable clone is useful only when there is a big transfer in\n> play.  Otherwise it isn't worth the trouble.\n>\n> So, if the clone is big, then this list of objects can be in the\n> millions.  For example my linux kernel repo with a couple branches\n> currently has:\n>\n> $ git rev-list --all --objects | wc -l\n> 2808136\n>\n> So, 2808136 objects, with 20-byte SHA1 for each of them, and you have a\n> 54 MB object list to transfer already.\n\n ok:\n\n a) that's fine.  first time, you have to do that, you have to do that.\n\n b) i have some ideas in mind, to say things like \"i already have the\nfollowing objects up to here, please give me a list of everything\nsince then\".\n\n > And even then, what if the transfer crashes during that object list\n> transfer?  On flaky connections this might happen within 54 MB.\n\n that's fine: i envisage the object list being cached at the remote\nend (by the first seed), and also being a \"shared file\", such that\nthere may even be complete copies of that \"file\" out there already,\nsuch that resumption is a non-issue.\n\n>> and, 10^N-1 times out of 10^N, for reasons that shawn kindly\n>> explained, i bet you the answer would be \"yes\".\n>\n> For the list of objects, sure.  But that isn't a big deal.  It is easy\n> enough to tell the remote about the commits we already have and ask for\n> the rest.  With a commit SHA1, the remote can figure out all the objects\n> we have. But all is in that determination of the latest commit we have.\n> If we get a partial pack, it is possible to somehow salvage as many\n> objects from it, and determine what top commit(s) that correspond to.\n> It is possible to set your local repo just as if you had requested a\n> shallow clone and then the resume would simply be a deepening of that\n> shallow clone.\n\n i'll need to re-read this when i have more time.  apologies.\n\n> Another issue is what to do with objects that are themselves huge.\n\n that's fine, too: in fact, that's the perfect scenario where a\nfile-sharing protocol excels.\n\n> Yet another issue: what to do with all those objects I've got in my\n> partial pack, but that I can't connect to any commit yet.  We don't want\n> them transferred again but it isn't easy to tell the remote about them.\n>\n> You could tell the remote: \"I have this pack for this commit from this\n> commit but I got only this amount of bytes from it, please resume\n> transfer here.\"\n\n ok i have a couple of ideas/thoughts\n\na) one of which was to send the commit index list back to the remote\nend, but that would be baaaaad as it could be 54mb as you say, so it\nwould be necessary to say \"here is the SHA1 of the index file you gave\nme earlier, do you still have it, if so please can we resume\n\nb) as long as _somebody_ has a complete copy, distributed throughout\nthe file-sharing network, of that complete pack, \"resume transfer\nhere\" isn't .... the concept is moot.  the bittorrent protocol covers\nthat concept of \"resume\" very very easily.\n\n> But as mentioned before the pack stream is not\n> deterministic,\n\n one a one-off basis, it is; and even then, i believe that it could be\nmade to \"not matter\".  so you ask a server for a pack object and get a\ndifferent SHA-1?  so what, you just make that part of the\nfile-sharing-network unique key: {ref}-{objref}-{SHA-1} instead of\njust {ref}-{objref}.  if the connection's lost, wow big deal, you just\nask again and you end up with a different SHA-1.  you're back to a\nsituation which is actually no different from and no less efficient\nthan the present http transfer system.\n\n ... but i'd rather avoid this scenario, if possible.\n\n> and we really don't want to make it single-threaded on a\n> server.  Furthermore this is a lot of work for the server as even if the\n> pack stream is deterministic, then the server still has to recreate the\n> first part of the pack just to throw it away until the desired offset is\n> reached.  And caching pack results also has all sorts of implications\n> we've prefered to avoid on a server for security reasons (better keep\n> serving operations read-only).\n\n i've already thrown out the idea of cacheing the pack objects\nthemselves, but am still exploring the concept of cacheing the .idx\nfile, even for short periods of time.\n\n so the server does a lot of work creating that .idx file, but it\ncontains the complete list of all objects, which you _could_ just\nobtain again by just asking explicitly for each and every single one\nof those objects, no more, no less, no deltas, no windows - the list,\nthe whole list and nothing but the list.\n\n>> ... um... in fact... um... i believe i'm merely talking about the .idx\n>> index file, aren't i?  because... um... the index file contains the\n>> list of object refs in the pack, yes?\n>\n> In one pack, yes.  You might have multiple packs.  And that doesn't mean\n> that all the objects from a pack are all relevant to the actual branches\n> you are willing to export.\n\n yes that's fine.  multiple packs are considered to be independent\nfiles of the \"VFS layer\" in the file-sharing network.  that's taken\ncare of.  what i need to know is: can you recreate a pack object given\nthe list of objects in its .idx file?\n\n>> sooo.... taking a wild guess, here: if you were to parse the .idx file\n>> and extract the list of object-refs, and then pass that to \"git\n>> pack-objects --window=0 --delta=0\", would you end up with the exact\n>> same pack file, because you'd forced git pack-objects to only return\n>> that specific list of object-refs?\n>\n> If you do this i.e. turn off delta compression, then the 615 MB\n> repository above will turn itself into a multi-gigabyte pack!\n\n ok this was covered in my previous post, hope it's clearer.  perhaps\n\"git pack-objects --window=0 --delta=0 <\n{list-of-objects-extracted-from-the-idx-file}\" isn't the way to\nachieve what i envisage - if not, does anyone have any ideas on how\nextracting the exact list of objects as previously given by a .idx\nfile can be achieved?\n\nl.\n"},{"id":"149756","messageId":"AANLkTi=ZKVFYH8GnhBmTSJbqsP9_c6ZGzWSyJMf2BSXM@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTik1hfe3jVWy236611d7hdP=yt+d3vCBiGvDa26H@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-03T10:55:57Z","receivedAt":"2010-09-03T10:55:57Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Fri, Sep 3, 2010 at 3:48 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n> On Fri, Sep 3, 2010 at 10:29 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n>> But usually the very first commit in a pack is huge as it typically\n>> isn't delta compressed (a delta chain has to start somewhere).  And this\n>> first commit will roughly represent the same size as a tarball for that\n>> commit.  And if you don't get at least that first commit then you are\n>> screwed.  Or if you don't get a complete second commit when deepening a\n>> clone you are still screwed.\n>\n> Elijah's recent work on \"rev-list --objects -- pathspec\" [1] may help\n> split a commit into many parts that can be sent separately.\n\n thank you nguyen, will take a look later in the day at that.   much\nappreciated.\n"},{"id":"149757","messageId":"AANLkTim4tkrshTOw8-hg8J8Tq6vb2xCzJfN4XrUyE7XD@mail.gmail.com","threadId":"24935","inReplyTo":"B757A854-C7BF-4CBF-9132-91D205344606@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-03T11:04:52Z","receivedAt":"2010-09-03T11:04:52Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Fri, Sep 3, 2010 at 11:37 AM, Theodore Tso <tytso@mit.edu> wrote:\n\n> So people who are willing to participate as part of the peer2peer\n> network can download the instructions for how to make the\n> canonical pack once a month, and use it to create the canonical pack.\n\n yes.\n\nalso it's already becoming clear that people mayy need to run a\n\"front\" copy of the peertopeer git repository, to which they locally\nperform \"git push\", for various reasons including not wishing to\nexpose \"random experimentation commits\" out onto the wider internet\nuntil they're actually ready to do so.  i realise that merging can get\nround this (flattening many patches into one big commit) but there is\nalso the issue of having to run a daemon on the peertopeer repo - you\nmiiight not want to interfere with that.\n\nso, yes, running a \"special\" git repo is probably a good idea.\n\n thanks theo.\n\nl.\n"},{"id":"149764","messageId":"7vsk1qzrn0.fsf@alter.siamese.dyndns.org","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009022033520.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-09-03T17:03:15Z","receivedAt":"2010-09-03T17:03:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n>>  i can - so that git pack-objects, after \"git repack -f\", returns a\n>> canonical pack! :)\n>\n> But did you try it?  The -f means \"don't reuse any existing pack data \n> and recompute every delta from scratch to find the best matches\".  This \n> is a very very costly operation that most people happily live without.\n\nAlso \"the best matches\" will change, hopefully in a better way, with newer\nvintage of git.  Optimization like c83f032 (apply delta depth bias to\nalready deltified objects, 2007-07-12) that is based on heuristics derived\nfrom empirical statistics can and should be allowed to happen.\n"},{"id":"149765","messageId":"7voccezr7m.fsf@alter.siamese.dyndns.org","threadId":"24935","inReplyTo":"B757A854-C7BF-4CBF-9132-91D205344606@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-09-03T17:12:29Z","receivedAt":"2010-09-03T17:12:29Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Theodore Tso <tytso@MIT.EDU> writes:\n\n> ...  So people who are willing\n> to participate as part of the peer2peer network can download the\n> instructions for how to make the canonical pack once a month, and use it\n> to create the canonical pack.  If the \"Gittorrent master\" has spent a\n> lot of time to carefully compute the most efficient set of delta\n> pairings, they will get the slight benefit of a more efficient pack\n> which they could use instead of th eir local one without having to use\n> large values of --window and --depth to \"git repack\".\n\nHmm, is the idea essentially to tell people \"Here is a snapshot of Linus\nrepository as of a few weeks ago, carefully repacked.  Instead of running\n\"git clone\" yourself, please bootstrap your repository by copying it over\nbittorrent and then \"git pull\" to update it\"?\n"},{"id":"149774","messageId":"20100903183120.GA4887@thunk.org","threadId":"24935","inReplyTo":"7voccezr7m.fsf@alter.siamese.dyndns.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2010-09-03T18:31:20Z","receivedAt":"2010-09-03T18:31:20Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Sep 03, 2010 at 10:12:29AM -0700, Junio C Hamano wrote:\n> Theodore Tso <tytso@MIT.EDU> writes:\n> \n> > ...  So people who are willing\n> > to participate as part of the peer2peer network can download the\n> > instructions for how to make the canonical pack once a month, and use it\n> > to create the canonical pack.  If the \"Gittorrent master\" has spent a\n> > lot of time to carefully compute the most efficient set of delta\n> > pairings, they will get the slight benefit of a more efficient pack\n> > which they could use instead of th eir local one without having to use\n> > large values of --window and --depth to \"git repack\".\n> \n> Hmm, is the idea essentially to tell people \"Here is a snapshot of Linus\n> repository as of a few weeks ago, carefully repacked.  Instead of running\n> \"git clone\" yourself, please bootstrap your repository by copying it over\n> bittorrent and then \"git pull\" to update it\"?\n\nEssentially, yes.  I just don't think bittorrent makes sense for\nanything else, because the git protocol is so much more efficient for\ntiny incremental updates...\n\nSo the only other part of my idea is that we could construct a special\nset of instructions that would allow them to recreate the carefully\nrepacked snapshot of Linus's repository without having to download it\nfrom a central seed site.  Instead, they could download a small set of\ninstructions, and use that in combination with the objects already in\ntheir repository, to create a bit-identical version of the carefully\nrepacked Linus repository.  It's basically rip-off of jigdo, but\napplied to git repositories instead of Debian .iso files.\n\n\t       \t\t    \t       - Ted\n"},{"id":"149782","messageId":"alpine.LFD.2.00.1009031522590.19366@xanadu.home","threadId":"24935","inReplyTo":"20100903183120.GA4887@thunk.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-03T19:41:26Z","receivedAt":"2010-09-03T19:41:26Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 3 Sep 2010, Ted Ts'o wrote:\n\n> On Fri, Sep 03, 2010 at 10:12:29AM -0700, Junio C Hamano wrote:\n> > Theodore Tso <tytso@MIT.EDU> writes:\n> > \n> > > ...  So people who are willing\n> > > to participate as part of the peer2peer network can download the\n> > > instructions for how to make the canonical pack once a month, and use it\n> > > to create the canonical pack.  If the \"Gittorrent master\" has spent a\n> > > lot of time to carefully compute the most efficient set of delta\n> > > pairings, they will get the slight benefit of a more efficient pack\n> > > which they could use instead of th eir local one without having to use\n> > > large values of --window and --depth to \"git repack\".\n> > \n> > Hmm, is the idea essentially to tell people \"Here is a snapshot of Linus\n> > repository as of a few weeks ago, carefully repacked.  Instead of running\n> > \"git clone\" yourself, please bootstrap your repository by copying it over\n> > bittorrent and then \"git pull\" to update it\"?\n> \n> Essentially, yes.  I just don't think bittorrent makes sense for\n> anything else, because the git protocol is so much more efficient for\n> tiny incremental updates...\n> \n> So the only other part of my idea is that we could construct a special\n> set of instructions that would allow them to recreate the carefully\n> repacked snapshot of Linus's repository without having to download it\n> from a central seed site.  Instead, they could download a small set of\n> instructions, and use that in combination with the objects already in\n> their repository, to create a bit-identical version of the carefully\n> repacked Linus repository.  It's basically rip-off of jigdo, but\n> applied to git repositories instead of Debian .iso files.\n\nSmall?  Well...\n\nLet's see what such instructions for how to make the canonical pack \nmight look like:\n\nFirst you need the full ordered list of objects.  That's a 20-byte SHA1\nper object.  The current Linux repo has 1704556 objects, therefore this\nlist is 33MB already.\n\nThen you need to identify which of those objects are deltas, and against\nwhich object.  Assuming we can index in the list of objects, that means,\nsay, one bit to identify a delta, and 31 bits for indexing the base. In\nmy case this is currently 1393087 deltas, meaning 5.3 MB of additional\ninformation.\n\nBut then, the deltas themselves can have variations in their encoding.\nAnd we did change the heuristics for the actual delta encoding in the\npast too (while remaining backward compatible), but for a canonical pack\ncreation we'd need to describe that in order to make things totally\nreproducible.\n\nSo there are 2 choices here: Either we specify the Git version to make \nsure identical delta code is used, but that will put big pressure on \nthat code to remain stable and not improve anymore as any behavior \nchange will create a compatibility issue forcing people to upgrade their \nGit version all at the same time.  That's not something I want to see \nthe world rely upon.\n\nThe other choice is to actually provide the delta output as part of the \ninstruction for the canonical pack creation.\n\nIn my case, the delta output represents:\n\n$ git verifi-pack -v .git/objects/pack/*.pack | \\\n  awk --posix  '/^[0-9a-f]{40}/ && $6 { tot += 1; size += $4 } \\\n                END { print tot, size }'\n1393087 155022247\n\nWe therefore have 148 MB of purely delta data here.\n\nSo that makes for a grand total of 33 MB + 148 MB = 181 MB of data just\nto be able to unambiguously reproduce a pack with a full guarantee of\nperfect reproducibility.\n\nBut even with the presumption of stable delta code, the recipee would \nstill take 38 MB that everyone would have to download every month which \nis far more than what a monthly incremental update of a kernel repo \nrequires.  Of course you could create a delta between consecutive \nrecipees, but that is becoming rather awkward.\n\nI still think that if someone really want to apply the P2P principle à \nla BitTorrent to Git, then it should be based on the distributed \nexchange of _objects_ as I outlined in a previous email, and not file \nchunks like BitTorrent does.  The canonical Git _objects_ are fully \ndefined, while their actual encoding may change.\n\n\nNicolas\n"},{"id":"149786","messageId":"AANLkTi=sC3NMNzPRQM5RKwnZQyRq-gq6+7wdiT5LGDrc@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009031522590.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-03T21:11:22Z","receivedAt":"2010-09-03T21:11:22Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Fri, Sep 3, 2010 at 8:41 PM, Nicolas Pitre <nico@fluxnic.net> wrote:\n\n> I still think that if someone really want to apply the P2P principle à\n> la BitTorrent to Git, then it should be based on the distributed\n> exchange of _objects_ as I outlined in a previous email, and not file\n> chunks like BitTorrent does.  The canonical Git _objects_ are fully\n> defined, while their actual encoding may change.\n\nok - missed it.  let's go back... ah _ha_ - with this:\n\n\"Yep.  Instead of transferring packs, a BitTorrent-alike transfer should\nbe based on the transfer of _objects_.  Therefore you can make the\ncorrespondance between file chunks in BitTorrent with objects in a Git\naware system.  So, when contacting a peer, you could negociate what is\nthe set of objects that the peer has that you don't, and vice versa.\nObjects in Git are stable and immutable, and they all have a unique SHA1\nsignature.  And to optimize the negociation, the pack index content can\nbe used, first by exchanging the content of the first level\nfan-out table and ignoring those entries that are equal.  This for each\npeer.\"\n\nok, so, great!  it does actually seem that, despite us using different\nterminologies, we're thinking along the same sort of lines.  i'm\nmarginally hampered by being unfamiliar with git, for which i\napologise.\n\nso, when i mentioned extracting the objects from the index file of\n\"git pack-object\", i was debating whether to then use that to\nre-create the pack object (in some nebulous way) - that's sort-of the\nsame thing.  i was also debating whether to mention the idea of using\ngit pack-object to extract one and only one object ( there is likely a\nmore efficient way of doing that ).   but, yes: i was thinking of\nmaking the vfs-layer expose individual objects, i just hadn't\nmentioned it yet (and missed your earlier reply, nicolas, for which i\napologise).\n\nbtw the idea of parsing the fan-out table would not have occurred to\nme in a miiiilllion years :)\n\n i'll take a look at that.  but whilst i'm doing that, the main\nquestion i really need to know is: how do you get one single explicit\nobject out of git?\n\ntia,\n\nl.\n"},{"id":"149804","messageId":"AANLkTinoyehduhdHSEm5yGTLvU6C-ViE885yLd63iQU0@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTi=sC3NMNzPRQM5RKwnZQyRq-gq6+7wdiT5LGDrc@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2010-09-04T00:24:24Z","receivedAt":"2010-09-04T00:24:24Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Sep 4, 2010 at 7:11 AM, Luke Kenneth Casson Leighton\n<luke.leighton@gmail.com> wrote:\n>  i'll take a look at that.  but whilst i'm doing that, the main\n> question i really need to know is: how do you get one single explicit\n> object out of git?\n\ngit cat-file <type> <sha-1>\n\nHowever if you are going to send objects, one by one, it is extremely\ninefficient. I think Nico has pointed that out. Individual object\nsending should only be done for large blobs.\n-- \nDuy\n"},{"id":"149808","messageId":"AANLkTimCca8PJQSQks4fvktVL7mE8gtFXpJKvCz9A+Wh@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTinoyehduhdHSEm5yGTLvU6C-ViE885yLd63iQU0@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2010-09-04T00:57:24Z","receivedAt":"2010-09-04T00:57:24Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Sep 4, 2010 at 10:24 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n> On Sat, Sep 4, 2010 at 7:11 AM, Luke Kenneth Casson Leighton\n> <luke.leighton@gmail.com> wrote:\n>>  i'll take a look at that.  but whilst i'm doing that, the main\n>> question i really need to know is: how do you get one single explicit\n>> object out of git?\n>\n> git cat-file <type> <sha-1>\n>\n> However if you are going to send objects, one by one, it is extremely\n> inefficient. I think Nico has pointed that out. Individual object\n> sending should only be done for large blobs.\n\nOn second thought, I don't know, maybe it could work if this sort of\nindividual object sending is based on bup. Bup splits files into small\npieces. Common parts of files are likely shared, reducing the need to\ndelta. But then you would deal with a huge number of small blobs.\n-- \nDuy\n"},{"id":"149810","messageId":"4C81A67B.2060400@gmail.com","threadId":"24935","inReplyTo":"AANLkTinoyehduhdHSEm5yGTLvU6C-ViE885yLd63iQU0@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2010-09-04T01:52:59Z","receivedAt":"2010-09-04T01:52:59Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Hmm, taking a few steps back, what is the expected usage of git-p2p?\nNote it's a bit of a trick question; what i'm really asking is what _else_,\nother than pulling/tracking Linus' kernel tree will/can be done with it?\n\nBecause once you accept that all peers are equal, but some peers are more\nequal than others, deriving a canonical representation of the object store\nbecomes relatively simple. Then, it's just a question of fetching the missing\nbits, whether using a dumb (rsync-like) transport, or a git-aware protocol.\n(I've no idea why you'd want to base a transfer protocol on the unstable packs,\nbuilding it on top of objects seems to be the only sane choice)\n\nI'm mostly git-ignorant and i'm assuming the following two things -- if someone\nmore familiar w/ git internals could confirm/deny, that would be great:\n\n1) \"git pull git:...\" would (or could be made to) work w/ a client that asks for\n   \"A..E\", but also tells the server to omit \"B,C and D\" from the wire traffic.    \n\n2) Git doesn't use chained deltas. IOW given commits \"A --d1-> B --d2-> C\",\n   \"C\" can be represented as a delta against \"A\" or \"B\", but _not_ against \"d1\". \n   (Think of the case where \"C\" reverts /part of/ \"B\")\n\n\nThen there are security implications... Which pretty much mandate having \"special\"\npeers anyway, at least for transferring heads (branches/tags etc). Which means\nthe second paragraph above applies. And as the \"special peer\" in practice can be\njust a signed tag/commit, like \"v2.6.35\", it's not such a big limitation like it\nmay seem at first...\n\nartur\n"},{"id":"149811","messageId":"04755B03-EE1D-48FA-8894-33AA8E2661C0@mit.edu","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009031522590.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2010-09-04T01:57:10Z","receivedAt":"2010-09-04T01:57:10Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"\nOn Sep 3, 2010, at 3:41 PM, Nicolas Pitre wrote:\n\n> \n> Let's see what such instructions for how to make the canonical pack \n> might look like:\n\nBut we don't need to replicate any particular pack.  We just need to provide instructions that can be replicated everywhere to provide *a* canonical pack.\n\n> \n> First you need the full ordered list of objects.  That's a 20-byte SHA1\n> per object.  The current Linux repo has 1704556 objects, therefore this\n> list is 33MB already.\n\nAssume the people creating this \"gitdo\" pack (i.e., much like jigdo) have a superset of Linus's objects.  So if we have all of the branches in Linus's repository, we can construct all of the necessary objects going back in time to constitute his repository.   If Linus has only one branch in his repo, we only need a single 20-byte SHA1 branch identifier.   For git, presumbly we would need three (one for next, maint, and master).\n\nWhat abort the order of the objects in the pack?  Well, ordering doesn't matter, right?   So let's assume the pack is sorted by hash id.   Is there any downside to that?  I can't think of any, but you're the pack expert...\n\nIf we do that, we would thus only need to send 20 bytes instead of 33MB.  \n\n> Then you need to identify which of those objects are deltas, and against\n> which object.  Assuming we can index in the list of objects, that means,\n> say, one bit to identify a delta, and 31 bits for indexing the base. In\n> my case this is currently 1393087 deltas, meaning 5.3 MB of additional\n> information.\n\nOK, this we'll need which means 5.3MB.\n\n\n> \n> But then, the deltas themselves can have variations in their encoding.\n> And we did change the heuristics for the actual delta encoding in the\n> past too (while remaining backward compatible), but for a canonical pack\n> creation we'd need to describe that in order to make things totally\n> reproducible.\n> \n> So there are 2 choices here: Either we specify the Git version to make \n> sure identical delta code is used, but that will put big pressure on \n> that code to remain stable and not improve anymore as any behavior \n> change will create a compatibility issue forcing people to upgrade their \n> Git version all at the same time.  That's not something I want to see \n> the world rely upon.\n\nI don't think the choice is that stark.  It does mean that in addition to whatever pack encoding format is used by git natively, the code would also need to preserve one version of the delta hueristics for \"Canonical pack version 1\". After this version is declared, it's true that you might come up with a stunning new innovation that saves some disk space.  How much is that likely to be?  3%?  5%?   Worst case, it means that (1) the bittorent-distributed packs might not be as efficient, and (2) the code would be made more complex because we would either need to (a) keep multiple versions of the code, or (b) the code might need to have some conditionals:\n\n\tif (canonical pack v1)\n\t\tdo_this_code;\n\telse\n\t\tdo_this_more_clever_code;\n\nIs that really that horrible?  And certainly we should be able to set things up so that it won't be a brake on innovation...\n\n> \n> The other choice is to actually provide the delta output as part of the \n> instruction for the canonical pack creation.\n> \n> So that makes for a grand total of 33 MB + 148 MB = 181 MB of data just\n> to be able to unambiguously reproduce a pack with a full guarantee of\n> perfect reproducibility.\n\nSo if we use the methods I've suggested, we would only need to send 5.3MB instead of 33MB or 181MB....\n\n> \n> But even with the presumption of stable delta code, the recipee would \n> still take 38 MB that everyone would have to download every month which \n> is far more than what a monthly incremental update of a kernel repo \n> requires.  Of course you could create a delta between consecutive \n> recipees, but that is becoming rather awkward.\n\nThe \"recipee\" would only need to download this if they are willing to participate as being one of the \"seeders\" in the BitTorrent network.   People who are willing to do this are presumably willing to transmit many more megabytes of data than 5MB or 33MB or 181MB.  Given the draconian policies of various ISP such as Comcast, it's not clear to me how many people will be willing to be seeders.   But if they are, I don't think downloading 5.3MB of instructions to generate a 600MB canonical pack to be distributed to hundreds or thousands of strangers will stop them.  :-)\n\n> I still think that if someone really want to apply the P2P principle à \n> la BitTorrent to Git, then it should be based on the distributed \n> exchange of _objects_ as I outlined in a previous email, and not file \n> chunks like BitTorrent does.  The canonical Git _objects_ are fully \n> defined, while their actual encoding may change.\n\nThe advantages of sending a canonical pack is that it's relatively less code to write, since we can reuse the standard BitTorrent clients and servers to transmit the git repository.  The downsides are that it's mainly useful for downloading the entire repository, but I think that's the most useful place for peer2peer anyway.\n\nThe advantage of a distributed exchange of _objects_ is you can use it to update a random repository --- but normally it's so efficient to download an incremental set of objects from github or kernel.org, so I'm not sure what would be the point.   It would also require exchanging a lot more metadata, since presumably the client would first have to receive all of the object id's (which would be 33MB), and then use that to decide how to distribute asking for those objects from his/her peers.   Which means sending object lists to different servers.   If these peers do not yet have a complete set of objects, they'll have to nack some of the object requests.   Furthermore with only a partial set of objects downloaded, how will the client do the delta compression?   Which means the client will need to store all of the objects in a decompressed form, and only when it has received all of the objects  will it be able to compress them.   So it's going to be fairly inefficient in terms of disk space.  I suppose the server could send the delta information (another 5.3MB) and then the client could use that to prioritize its object request lists.   But still, this is quite different form the git protocol, and a lot will have to be written from scratch.\n\n-- Ted\n"},{"id":"149815","messageId":"alpine.LFD.2.00.1009032304560.19366@xanadu.home","threadId":"24935","inReplyTo":"4C81A67B.2060400@gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-04T04:39:56Z","receivedAt":"2010-09-04T04:39:56Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 4 Sep 2010, Artur Skawina wrote:\n\n> Hmm, taking a few steps back, what is the expected usage of git-p2p?\n> Note it's a bit of a trick question; what i'm really asking is what _else_,\n> other than pulling/tracking Linus' kernel tree will/can be done with it?\n\nDunno.\n\n> Because once you accept that all peers are equal, but some peers are more\n> equal than others, deriving a canonical representation of the object store\n> becomes relatively simple.\n\nThat depends what you consider a canonical representation.  I don't \nthink the actual object store should ever be \"canonicalized\".\n\n> Then, it's just a question of fetching the missing\n> bits, whether using a dumb (rsync-like) transport, or a git-aware protocol.\n\nBut Git does that already.\n\n> (I've no idea why you'd want to base a transfer protocol on the unstable packs,\n> building it on top of objects seems to be the only sane choice)\n\nThere seems to be quite some confusion around objects and packs.\n\nThe Git \"database\" is _only_ a big pile of objects that is content \naddressable i.e. each object has a name which is derived from its \ncontent.  This is the 40 hexadecimal string.\n\nThere are only 4 types of objects. Roughly they are:\n\n1) A \"blob\" object contains plain data, usually used for file content.\n\n2) A \"tree\" object contains a list of entries made of a file or \n   directory name, and the object name that corresponds to it.  For \n   files, the referenced objects are \"blobs\". For directories, the \n   referenced objects are some other \"trees\".  This is how the file and \n   directory hierarchy are represented.\n\n3) A \"commit\" object contains a reference to the top tree object \n   corresponding to the root directory of the project, a reference to \n   the previous \"commit\" object, and a text message to describe this \n   commit.  If this commit represents a merge, then there will be more \n   than one reference to previous commits.  This is how the commit \n   history is represented.\n\n4) And finally a \"tag\" object contains a reference to any other object \n   and a text message.  Most of the time, only commit objects are \n   referenced that way.  This is used to identify some particular \n   commits.\n\nAnd finally, there are a few files, one for each \"branch\", used to \ncontain a reference to the latest commit object for each of those \nbranches.\n\nThat's it!  Here you have the *whole* architecture of Git!\n\nNow... one way to store those objects on disk is to simply deflate them \nwith zlib and put the result in a file, one file per object.  The first \n2 chars from the object name are used to create ssubdirectories under \n.git/objects/ and the remaining 38 chars are used for the actual file \nname within those subdirectories.  This is the \"loose\" object format or \nencoding.\n\nAnother way to store those objects is to cram them together in one or \nmultiple (or many) pack files.  The advantage with the pack file is that \nwe can encode any object as a delta against any other object in the same \npack file.  This is the \"packed\" object format or encoding.\n\n> I'm mostly git-ignorant and i'm assuming the following two things -- if someone\n> more familiar w/ git internals could confirm/deny, that would be great:\n> \n> 1) \"git pull git:...\" would (or could be made to) work w/ a client that asks for\n>    \"A..E\", but also tells the server to omit \"B,C and D\" from the wire traffic.    \n\nWhat Git does when transferring data on the wire is actually to create a \nspecial pack file that contains _only_ those objects that the sender has \nbut that the receiver doesn't, and stream that over the net.  So if the \nclient tells the server that it already has commit A, then the server \nwill create a pack that contains only those objects that were created \nafter commit A, and omit all the objects that can be reached through \ncommit A that are also used by later commits (think unchanged files).  \nIf you also have commits B, C and D, then the server will also exclude \nall the objects that are reachable through those commits from that \nspecial pack.\n\nOn the receiving end, Git simply writes the received pack into a file \nalong with the other existing packs, and compute a pack index for it.\n\n> 2) Git doesn't use chained deltas. IOW given commits \"A --d1-> B --d2-> C\",\n>    \"C\" can be represented as a delta against \"A\" or \"B\", but _not_ against \"d1\". \n>    (Think of the case where \"C\" reverts /part of/ \"B\")\n\nGit does use chained deltas indeed.  But deltas are used only at the \nobject level within a pack file.  Any blob object can be represented as \na delta against any other blob in the pack, regardless of the commit(s) \nthose blob objects belong to.  Same thing for tree objects.  So you can \nhave deltas going in total random directions if you look them from a \ncommit perspective.  So \"C\" can have some of its objects being deltas \nagainst objects from \"B\", or \"A\", or any other commit for that matter, \nor even objects belonging to the same commit \"C\". And some other objects \nfrom \"B\" can delta against objects from \"C\" too. There is simply no \nrestrictions at all on the actual delta direction.  The only rule is \nthat an object may only delta against another object of the same type.\n\nOf course we don't try to delta each object against all the other \navailable objects as that would be a O(n^2) operation (imagine with n = \n1.7 million objects).  So we use many heuristics to make this delta \npacking efficient without taking an infinite amount of time.\n\nFor example, if we have objects X and Y that need to be packed together \nand sent to a client over the net, and we find that Y is already a delta \nagainst X in one pack that exists locally, then we simply and literally \ncopy the delta representation of Y from that local pack file and send it \nout without recomputing that delta.\n\n> Then there are security implications... Which pretty much mandate having \"special\"\n> peers anyway, at least for transferring heads (branches/tags etc). Which means\n> the second paragraph above applies.\n\nWell... Actually, all you need is only one trusted peer to provide those \nheads i.e. the top commit SHA1 name for each branches you need.  From \nthat one SHA1 name per branch, you can validate the entire repository as \nevery object reference throughout is based on the content of the object \nit refers to.  For example, to validate the authenticity of everything \nfrom a random copy of the Linux kernel repository, I need only 20 bytes \nfrom a trusted source.  No need to have this information distributed \namongst multiple peers.\n\nAnd even if the delta encoding is different from the one used in Linus' \nrepository, or even if the packing is done differently (different number \nof packs, etc.) then the final SHA1 will always be the same.  This is \nbecause the actual content from all referenced objects is the same \nregardless of their effective encoding or format.\n\n\nNicolas\n"},{"id":"149817","messageId":"AANLkTikVf=X8cLP9s6W9VGOt0EHE4J5MYsBpgKYhrAri@mail.gmail.com","threadId":"24935","inReplyTo":"04755B03-EE1D-48FA-8894-33AA8E2661C0@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Kyle Moffett","fromEmail":"kyle@moffetthome.net","sentAt":"2010-09-04T05:23:54Z","receivedAt":"2010-09-04T05:23:54Z","isPatch":false,"sender":{"key":"kyle@moffetthome.net","avatar":null},"body":"Ted,\n\nI think your \"canonical pack\" idea has value, but I'd be inclined to\ntry to optimize more for the \"common case\" of developing on a fast\nlocal network with many local checkouts, where you occasionally\npush/fetch external sources via a slow link.\n\nSpecifically, let's look at the very reasonable scenario of a\ndeveloper working over a slow DSL or dialup connection.  He's probably\ngot many copies of various GIT repositories cloned all over the place\n(hey, disk is cheap!), but right now he just wants a fresh clean copy\nof somebody else's new tree with whatever its 3 feature branches are.\nFurthermore, he's probably even got 80% of the commit objects from\nthat tree archived in his last clone from linux-next.\n\nIn theory he could very carefully arrange his repositories with\njudicious use of alternate object directories.  From personal\nexperience, though, such arrangements are *VERY* prone to accidentally\npurging wanted objects; unless you *never* ever delete a branch in the\n\"reference\" repository.\n\nSo I think the real problem to solve would be:  Given a collection of\nlocal computers each with many local repositories, what is the best\nway to optimize a clone of a \"new\" remote repository (over a slow\nlink) by copying most of the data from other local repositories\naccessible via a fast link?\n\nThe goal would be to design a P2P protocol capable of rapidly and\nefficiently building distributed searchable indexes of ordered commits\nthat identify which peer(s) contain that each commit.\n\nWhen you attempt to perform a \"git fetch --peer\" from a repository, it\nwould quickly connect to a few of the metadata index nodes in the P2P\nnetwork and use them to negotiate \"have\"s with the upstream server.\nThe client would then sequentially perform the local \"fetch\"\noperations necessary to obtain all the objects it used to minimize the\ncommit range with the server.  Once all of those \"fetch\" operations\ncompleted, it could proceed to fetch objects from the server normally.\n\nSome amount of design and benchmarking would need to be done in order\nto figure out the most efficient indexing algorithm for finding a\nminimal set of \"have\"s of potentially thousands of refs, many with\nindependent root commits.  For example if the index was grouped\naccording to \"root commit\" (of which there may be more than one), you\n*should* be able to quickly ask the server about a small list of root\ncommits and then only continue asking about commits whose roots are\nall known to the server.\n\nThe actual P2P software would probably involve 2 different daemon\nprocesses.  The first would communicate with each other and with the\nrepositories, maintaining the ref and commit indexes.  These daemons\nwould advertise themselves with Avahi, or alternatively in an\nenterprise environment they would be managed by your sysadmins and be\nautomatically discovered using DNS-SD.  Clients looking to perform a\nP2P fetch would first ask these.\n\nThe second daemon would be a modified git-daemon that connects to the\nadvertised \"index\" daemons and advertises its own refs and commit\nlists, as well as its IP address and port.\n\nMy apologies if there are any blatant typos or thinkos, it's a bit\nlater here than I would normally be writing about technical topics.\n\nCheers,\nKyle Moffett\n"},{"id":"149818","messageId":"alpine.LFD.2.00.1009040040030.19366@xanadu.home","threadId":"24935","inReplyTo":"04755B03-EE1D-48FA-8894-33AA8E2661C0@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-04T05:40:35Z","receivedAt":"2010-09-04T05:40:35Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 3 Sep 2010, Theodore Tso wrote:\n\n> \n> On Sep 3, 2010, at 3:41 PM, Nicolas Pitre wrote:\n> \n> > \n> > Let's see what such instructions for how to make the canonical pack \n> > might look like:\n> \n> But we don't need to replicate any particular pack.  We just need to \n> provide instructions that can be replicated everywhere to provide *a* \n> canonical pack.\n\nBut that canonical pack could be any particular pack.\n\n> > First you need the full ordered list of objects.  That's a 20-byte SHA1\n> > per object.  The current Linux repo has 1704556 objects, therefore this\n> > list is 33MB already.\n> \n> Assume the people creating this \"gitdo\" pack (i.e., much like jigdo) \n> have a superset of Linus's objects.  So if we have all of the branches \n> in Linus's repository, we can construct all of the necessary objects \n> going back in time to constitute his repository.  If Linus has only \n> one branch in his repo, we only need a single 20-byte SHA1 branch \n> identifier.  For git, presumbly we would need three (one for next, \n> maint, and master).\n\nSure, but that's not sufficient.  All this 20-byte SHA1 gives you is a \nset of objects.  That says nothing about their encoding.\n\n> What about the order of the objects in the pack?  Well, ordering \n> doesn't matter, right?  So let's assume the pack is sorted by hash id.  \n> Is there any downside to that?  I can't think of any, but you're the \n> pack expert...\n\nOrdering does matter a big deal.  Since object IDs are the SHA1 of their \ncontent, those IDs are totally random.  So if you store objects \naccording to their sorted IDs, then the placement of objects belonging \nto, say, the top commit will be totally random.  And since you are the \nfilesystem expert, I don't have to tell you what performance impacts \nthis random access of small segments of data scattered throughout a \n400MB file will have on a checkout operation.\n\n> If we do that, we would thus only need to send 20 bytes instead of 33MB.  \n> \n> > Then you need to identify which of those objects are deltas, and against\n> > which object.  Assuming we can index in the list of objects, that means,\n> > say, one bit to identify a delta, and 31 bits for indexing the base. In\n> > my case this is currently 1393087 deltas, meaning 5.3 MB of additional\n> > information.\n> \n> OK, this we'll need which means 5.3MB.\n> \n> > \n> > But then, the deltas themselves can have variations in their encoding.\n> > And we did change the heuristics for the actual delta encoding in the\n> > past too (while remaining backward compatible), but for a canonical pack\n> > creation we'd need to describe that in order to make things totally\n> > reproducible.\n> > \n> > So there are 2 choices here: Either we specify the Git version to make \n> > sure identical delta code is used, but that will put big pressure on \n> > that code to remain stable and not improve anymore as any behavior \n> > change will create a compatibility issue forcing people to upgrade their \n> > Git version all at the same time.  That's not something I want to see \n> > the world rely upon.\n> \n> I don't think the choice is that stark.  It does mean that in addition \n> to whatever pack encoding format is used by git natively, the code \n> would also need to preserve one version of the delta hueristics for \n> \"Canonical pack version 1\". After this version is declared, it's true \n> that you might come up with a stunning new innovation that saves some \n> disk space.  How much is that likely to be?  3%?  5%?  Worst case, it \n> means that (1) the bittorent-distributed packs might not be as \n> efficient, and (2) the code would be made more complex because we \n> would either need to (a) keep multiple versions of the code, or (b) \n> the code might need to have some conditionals:\n> \n> \tif (canonical pack v1)\n> \t\tdo_this_code;\n> \telse\n> \t\tdo_this_more_clever_code;\n> \n> Is that really that horrible?  And certainly we should be able to set things up so that it won't be a brake on innovation...\n\nWell, this would still be a non negligible maintenance cost.  And for \nwhat purpose already? What is the real advantage?\n\n> The advantages of sending a canonical pack is that it's relatively \n> less code to write, since we can reuse the standard BitTorrent clients \n> and servers to transmit the git repository.  The downsides are that \n> it's mainly useful for downloading the entire repository, but I think \n> that's the most useful place for peer2peer anyway.\n\nSure.  But I don't think it is worth making Git less flexible just for \nthe purpose of ensuring that people could independently create identical \npacks.  I'd advocate for \"no code to write at all\" instead, and simply \nhave one person create and seed the reference pack.\n\nAnd if you are willing to participate in the seeding of such a torrent, \nthen you better not be bandwidth limited, meaning that you certainly can \nafford to download that reference pack in the first place.\n\nAnd that reference pack doesn't have to change that often either.  If \nyou update it only on every major kernel releases, then you'll need to \nfetch it about once every 3 months.  Incremental updates from those \npoints should be relatively small.\n\nYet... it should be possible in practice to produce identical packs, \ngiven that the Git version is specified, the zlib version is specified, \nthe number of threads for the repack is equal to 1, the -f flag is used \nmeaning a full repack is performed, the delta depth and window size is \nspecified, and the head branches are specified.  Given that torrents are \nalso identified by a hash of their content, it should be pretty easy to \nsee if the attempt to reproduce the reference pack worked, and start \nseeding right away if it did.\n\nBut again, I don't think it is worth freezing the pack format into a \ncanonical encoding for this purpose.\n\n\nNicolas\n"},{"id":"149819","messageId":"4C81DC34.2090800@gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009032304560.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2010-09-04T05:42:12Z","receivedAt":"2010-09-04T05:42:12Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":">> 2) Git doesn't use chained deltas. IOW given commits \"A --d1-> B --d2-> C\",\n>>    \"C\" can be represented as a delta against \"A\" or \"B\", but _not_ against \"d1\". \n>>    (Think of the case where \"C\" reverts /part of/ \"B\")\n> \n> Git does use chained deltas indeed.  But deltas are used only at the \n> object level within a pack file.  Any blob object can be represented as \n> a delta against any other blob in the pack, regardless of the commit(s) \n> those blob objects belong to.  Same thing for tree objects.  So you can \n> have deltas going in total random directions if you look them from a \n> commit perspective.  So \"C\" can have some of its objects being deltas \n> against objects from \"B\", or \"A\", or any other commit for that matter, \n> or even objects belonging to the same commit \"C\". And some other objects \n> from \"B\" can delta against objects from \"C\" too. There is simply no \n> restrictions at all on the actual delta direction.  The only rule is \n> that an object may only delta against another object of the same type.\n> \n> Of course we don't try to delta each object against all the other \n> available objects as that would be a O(n^2) operation (imagine with n = \n> 1.7 million objects).  So we use many heuristics to make this delta \n> packing efficient without taking an infinite amount of time.\n> \n> For example, if we have objects X and Y that need to be packed together \n> and sent to a client over the net, and we find that Y is already a delta \n> against X in one pack that exists locally, then we simply and literally \n> copy the delta representation of Y from that local pack file and send it \n> out without recomputing that delta.\n\nWhat i meant by 'chained deltas' is a representation that takes delta#1 and\napplies delta#2 to the first delta, and applies the result to the source of\ndelta#1. Which could be a more compact representation of eg. a partial revert.\n\nIOW, if I have commits A..Y, ask (via git pull) for commits X and Z, then I'm\nguaranteed to receive them either raw, or as a delta vs commits A..X, right?\nWhat I'm really asking is, if a (modified) git-upload-pack skips transferring\ncommit X, and just sends me commit Z (possibly as delta vs 'X'), _and_ I \nobtain commit 'X\" in some other way, I will be able to reconstruct 'Z', correct?\n\nTIA, \n\nartur\n"},{"id":"149820","messageId":"alpine.LFD.2.00.1009040153280.19366@xanadu.home","threadId":"24935","inReplyTo":"4C81DC34.2090800@gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-04T06:13:23Z","receivedAt":"2010-09-04T06:13:23Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 4 Sep 2010, Artur Skawina wrote:\n\n> >> 2) Git doesn't use chained deltas. IOW given commits \"A --d1-> B --d2-> C\",\n> >>    \"C\" can be represented as a delta against \"A\" or \"B\", but _not_ against \"d1\". \n> >>    (Think of the case where \"C\" reverts /part of/ \"B\")\n> > \n> > Git does use chained deltas indeed.  But deltas are used only at the \n> > object level within a pack file.  Any blob object can be represented as \n> > a delta against any other blob in the pack, regardless of the commit(s) \n> > those blob objects belong to.  Same thing for tree objects.  So you can \n> > have deltas going in total random directions if you look them from a \n> > commit perspective.  So \"C\" can have some of its objects being deltas \n> > against objects from \"B\", or \"A\", or any other commit for that matter, \n> > or even objects belonging to the same commit \"C\". And some other objects \n> > from \"B\" can delta against objects from \"C\" too. There is simply no \n> > restrictions at all on the actual delta direction.  The only rule is \n> > that an object may only delta against another object of the same type.\n> > \n> > Of course we don't try to delta each object against all the other \n> > available objects as that would be a O(n^2) operation (imagine with n = \n> > 1.7 million objects).  So we use many heuristics to make this delta \n> > packing efficient without taking an infinite amount of time.\n> > \n> > For example, if we have objects X and Y that need to be packed together \n> > and sent to a client over the net, and we find that Y is already a delta \n> > against X in one pack that exists locally, then we simply and literally \n> > copy the delta representation of Y from that local pack file and send it \n> > out without recomputing that delta.\n> \n> What i meant by 'chained deltas' is a representation that takes delta#1 and\n> applies delta#2 to the first delta, and applies the result to the source of\n> delta#1. Which could be a more compact representation of eg. a partial revert.\n\nYou can have deltas on top of deltas.  And they can be in any direction \ni.e. object #1 can be a delta that only takes half of a bigger object \n#2, or object #2 can be a delta copying a smaller object #1 and adding \nmore data to it.  In the end both representation will take more or less \nthe same amount of space.\n\n> IOW, if I have commits A..Y, ask (via git pull) for commits X and Z, then I'm\n> guaranteed to receive them either raw, or as a delta vs commits A..X, right?\n\nIn such case you will receive only those new objects that commits X and \nZ introduced, and those new objects may indeed be encoded either as \nwhole objects, or as deltas against either objects that you already \nhave, or even as deltas against objects that are part of the transfer.\n\n> What I'm really asking is, if a (modified) git-upload-pack skips transferring\n> commit X, and just sends me commit Z (possibly as delta vs 'X'), _and_ I \n> obtain commit 'X\" in some other way, I will be able to reconstruct 'Z', correct?\n\nYes.  Although it is 'git pack-objects' that decides what objects to \nsend, not 'git-upload-pack'.\n\n\nNicolas\n"},{"id":"149829","messageId":"5D61124F-21D7-4A98-A1B8-6A013C8CB199@mit.edu","threadId":"24935","inReplyTo":"AANLkTikVf=X8cLP9s6W9VGOt0EHE4J5MYsBpgKYhrAri@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2010-09-04T11:46:15Z","receivedAt":"2010-09-04T11:46:15Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"\nOn Sep 4, 2010, at 1:23 AM, Kyle Moffett wrote:\n\n> \n> Specifically, let's look at the very reasonable scenario of a\n> developer working over a slow DSL or dialup connection.  He's probably\n> got many copies of various GIT repositories cloned all over the place\n> (hey, disk is cheap!), but right now he just wants a fresh clean copy\n> of somebody else's new tree with whatever its 3 feature branches are.\n> Furthermore, he's probably even got 80% of the commit objects from\n> that tree archived in his last clone from linux-next.\n> \n> In theory he could very carefully arrange his repositories with\n> judicious use of alternate object directories.  From personal\n> experience, though, such arrangements are *VERY* prone to accidentally\n> purging wanted objects; unless you *never* ever delete a branch in the\n> \"reference\" repository.\n\nThis is a reasonable problem, but there's a much easier solution that's worked for me, and it's very simple.  Simply use Linus's repository as the reference repository.  IOW, whenever I start working on a new computer, I'll do keep a bare repository containing only Linus's repo.  I might pull that down fresh from kernel.org, or if I'm on a local computer, it's usually much easier just to scp -r over a local copy of /usr/projects/linux/bare to my local machine.    That's something like 98% of all of the needed commits right there.  And since Linus never rewinds any branches on his repo.  I'm done.   It works for me.  :-)\n\nFor a repo like git which does have a 'pu' branch which rewinds, life is a bit more tricky, I'll admit.   OTOH, the fact that there is a 'pu' branch means that there are few people who keep externally published trees, since it's much easier to let Junio collect people's various patchsets into the 'pu' branch, at which point it's all nicely integrated.  In some ways, the 'pu' branch is really the equivalent of the linux-next tree.  :-)\n\nOK, suppose we have a project which is large enough to have large numbers of downstream trees, but which also has one or more branches that happen to be rewinding.  I can't think of any project which has those characteristics, but OK, it's worth exploring.  Now what?   Well, if I'm doing this all on one machine (the use case is my laptop), and I have a *very* slow external connection (the use case I can think of it is when I am on a cruise ship, where not only the bandwidth slow, but there is also a very long latency since it's a satellite link), what I'd do is find all my trees that might be related to the tree that I want to clone, and hard link them into my target repository.   I'd then create a temp .git/refs/hack directory, and create branch pointers to the tips of all of the branches\n  from the donating repositories.  I'd then just do a straightforward git pull, which would transfer over just the objects I need, then delete the .git/refs/hack directory, and then do a git gc.\n\nOk, so how do we generalize this if we have a large number of local machines?   The first insight is that we want to treat local machines very differently from remote machines, since we obviously want to do as many of the object transfers from the local machines.   The second is to note the reason why I had to hackishly create the .git/refs/hack directory.   That's because git fetch already optimizes things based on branch heads to minimize the need to send vast object sets back and forth to figure out what objects are required.\n\nSo I suspect we might be able to do something where the peer2peer downloader first contacts all of its local \"buddies\", and gets their branch heads, and then, using the standard git protocol, contacts the remote git server (i.e., github, git.kernel.org, etc.), and sends it all of its available branch heads.  It will then be able to assemble what it needs from remote server and from its local buddies.   In effect, this is basically an \"octopus fetch\", and it seems like most of the machinery can be reused from the git protocol, with only some minor modifications.\n\nI agree this is a very different use case than the \"initial download\" case.   At least for the Linux kernel, I have a solution that works pretty much well enough.    I suspect for the git tree, I could just as easily solve it by just doing a brute force copy of Junio's tree and then pulling down the extra objects using \"git fetch\".   The reality is most projects have a canonical tree, and 99% of the time people's subtrees contain very few objects that aren't already in the canonical tree.  So while we could create some complicated design to try to optimize picking a few commits that weren't in Linus's tree, but were in linux-next, when downloading someone's subsystem development tree, how many commits do we expect will be unique in someone's subsystem tree?   A hundred, at most?   Even wit\n h a slow link, simply doing a git clone --bare of a local copy of Linus's tree, and then doing a straightforward \"git fetch\" is probably going to good enough, even if you are stuck on a cruise ship with a laptop that you smuggled aboard without your partner/spouse knowing so you can get that hacking fix.  :-)\n\n-- Ted\n"},{"id":"149831","messageId":"AANLkTi=7jUSCNiPf+HfEQuxaf16Jt06--bFE7=Of9wp=@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009040153280.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T11:58:31Z","receivedAt":"2010-09-04T11:58:31Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"ok, there's one other question i need some info on (thank you to\nnguyen for answering about git cat-file): is there a way to make\ngit-pack-objects _just_ give the index file, rather than go to all the\ntrouble of creating the pack file as well?\n\nthe reason i ask is because i would like to get the index file, parse\nit for objects (as nicolas recommends) then present the contents of\n\"git cat-file\" as files within subdirectories, via this\npseudo-VFS-layer over bittorrent.\n\ntia,\n\nl.\n"},{"id":"149832","messageId":"5B5470E5-57E6-48D2-981B-CE77FA43546F@mit.edu","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009040040030.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2010-09-04T12:00:40Z","receivedAt":"2010-09-04T12:00:40Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"\nOn Sep 4, 2010, at 1:40 AM, Nicolas Pitre wrote:\n>> What about the order of the objects in the pack?  Well, ordering \n>> doesn't matter, right?  So let's assume the pack is sorted by hash id.  \n>> Is there any downside to that?  I can't think of any, but you're the \n>> pack expert...\n> \n> Ordering does matter a big deal.  Since object IDs are the SHA1 of their \n> content, those IDs are totally random.  So if you store objects \n> according to their sorted IDs, then the placement of objects belonging \n> to, say, the top commit will be totally random.  And since you are the \n> filesystem expert, I don't have to tell you what performance impacts \n> this random access of small segments of data scattered throughout a \n> 400MB file will have on a checkout operation.\n\nDoes git repack optimize the order so that certain things (like checkouts for example) are really fast?  I admit I hadn't noticed.  Usually until the core packs are in my page cache, it has always seemed to me that things are pretty slow.   And of course, the way objects and grouped together and ordered for \"gitk\" or \"git log\" to be fast won't be the safe as a checkout operation...\n\n> Sure.  But I don't think it is worth making Git less flexible just for \n> the purpose of ensuring that people could independently create identical \n> packs.  I'd advocate for \"no code to write at all\" instead, and simply \n> have one person create and seed the reference pack.\n\nI don't think it's a matter of making Git \"less flexible\", it's just simply a code maintenance headache of needing to be able to support encoding both a canonical format as well as the latest bleeding-edge, most efficient encoding format.   And how often are you changing/improving the encoding process, anyway?  It didn't seem to me like that part fo the code was constantly being tweaked/improved. \n\nStill, you're right, it might not be worth it.  To be honest, I was more interested about the fact that this might also be used to give people hints about how to better repack their local repositories so that they didn't have to run git repack with large --window and --depth arguments.  But that would only provide very small improvements in storage space in most cases, so it's probably not even worth it for that.\n\nQuite frankly, I'm a little dubious about how critical peer2peer really is, for pretty much any use case.  Most of the time, I can grab the base \"reference\" tree and drop it on my laptop before I go off the grid and have to rely on EDGE or some other slow networking technology.  And if the use case is some small, but illegal-in-some-jurisdiction code, such as ebook DRM liberation scripts (the kind which today are typically distributed via pastebin's :-), my guess is that zipping up a git repository and dropping it on a standard bittorrent server run by the Swedish Pirate party is going to be much more effective.   :-)\n\n-- Ted\n"},{"id":"149835","messageId":"AANLkTik8S708qOBaZDPdPheqinYKhBP71w=u=9BFhyjA@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009040040030.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T12:33:10Z","receivedAt":"2010-09-04T12:33:10Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 6:40 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n\n> But again, I don't think it is worth freezing the pack format into a\n> canonical encoding for this purpose.\n\n ... and, even if different seeds produce different pack index files,\nor even if the same seed produces a different pack index file, it\ndoesn't matter because the objects will still be there, can still be\nshared, can still be used to reconstruct the commit.\n\n so even if you end up asking for and sharing different objects\nbecause you're using a different index file, it really really doesn't\nmatter: all roads lead to rome.\n\nyou might end up not _having_ some particular objects (due to git gc)\nbut that's ok: just re-request the index file again (and go through\nthe process of evaluating / comparing the index fan-objects again\n*sigh* you can't have everything).\n\nl.\n"},{"id":"149836","messageId":"AANLkTi=F4RLbCPwoUGTAUFzqBuPFuM4qiAdpkrrGmntn@mail.gmail.com","threadId":"24935","inReplyTo":"5B5470E5-57E6-48D2-981B-CE77FA43546F@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T12:44:30Z","receivedAt":"2010-09-04T12:44:30Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 1:00 PM, Theodore Tso <tytso@mit.edu> wrote:\n\n> Quite frankly, I'm a little dubious about how critical peer2peer really is, for pretty much any use case.  Most of the time, I can grab the base \"reference\" tree and drop it on my laptop before I go off the grid and have to rely on EDGE or some other slow networking technology.\n\n yes... but if you meet up with other people, where you have a fast\nLAN segment or set up your own private wifi mesh, are you able to a)\nsync up the mailing lists using git b) sync up the project's bug-list\nusing git c) sync up the wiki (if there is one) using git d)\nseamlessly continue to appear to be talking (with the people you're\nmeeting) on the \"mailing lists\" as if you actually had a decent\nconnection to the server e) report, comment on, change the status of\nand share bugs between all of the people you're meeting as if you had\na decent connection to the bugtracker server...\n\nyou see how it's not just about \"The Source Code\"?  sure, yes, you or\nanyone else _on their own_ can do code development, isolated from\neveryone else and the internet...\n\nexample: i went to europython, and met the moinmoin developers (nice\npeople).  i wanted to help with the sprint after hours: it turned out\nthat we'd got the day wrong, so we then went \"oh well, let's find\nsomewhere to do a bit of hacking\", and _immediately_ the discussion\nturned into \"how we can all of us find internet connectivity\".  i said\nthat i had a 3G USB but i didn't have ndiswrapper installed for the\nbcm4328 on my laptop, so i couldn't offer a mesh; someone else said\nthat their laptop's WIFI didn't even _do_ master-mode, and we then had\na nice discussion about various little network routers that ran\nopenwrt and lamented the fact that none of us had brought one along.\n\nneedless to say, we didn't do any hacking :)\n\nl.\n"},{"id":"149837","messageId":"AANLkTik9awEd40s3r-O8t9DwZBh34Z0ozsxMm1QNjNoT@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTi=7jUSCNiPf+HfEQuxaf16Jt06--bFE7=Of9wp=@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T13:14:36Z","receivedAt":"2010-09-04T13:14:36Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 12:58 PM, Luke Kenneth Casson Leighton\n<luke.leighton@gmail.com> wrote:\n> ok, there's one other question i need some info on (thank you to\n> nguyen for answering about git cat-file): is there a way to make\n> git-pack-objects _just_ give the index file, rather than go to all the\n> trouble of creating the pack file as well?\n>\n> the reason i ask is because i would like to get the index file, parse\n> it for objects (as nicolas recommends) then present the contents of\n> \"git cat-file\" as files within subdirectories, via this\n> pseudo-VFS-layer over bittorrent.\n\npack-objects.c - write_pack_file():\n...\n        if (pack_to_stdout) {\n            sha1close(f, sha1, CSUM_CLOSE);\n        } else if (nr_written == nr_remaining) {\n            sha1close(f, sha1, CSUM_FSYNC);\n        } else {\n            int fd = sha1close(f, sha1, 0);\n            fixup_pack_header_footer(fd, sha1, pack_tmp_name,\n                         nr_written, sha1, offset);\n            close(fd);\n        }\n\n        if (!pack_to_stdout) {\n            struct stat st;\n            const char *idx_tmp_name;\n            char tmpname[PATH_MAX];\n\n            idx_tmp_name = write_idx_file(NULL, written_list,\n                              nr_written, sha1);\n\n....\n\nahh, whoops.  any advances on that?\n\ngrep write_idx_file *.c */*.c\n\n* git-index-pack requires a pack file in order to re-create the index:\ni don't want that\n* git-pack-objects appears to have no way of telling it \"just gimme\nindex file please\"\n* fast-import.c appears not to be what's needed either.\n\nso - any other methods for just getting the index file (exclusively?)\nany other commands i've missed?  if not, are there any other ways of\ngetting a pack's index of objects without err... getting the index\nfile?  (i believe the answer to be no, but i'm just making sure) and\non that basis i believe it is safe to ask: any objections to a patch\nwhich adds \"--index-only\" to builtin/pack-objects.c?\n\nl.\n"},{"id":"149838","messageId":"4C824CAD.9070509@gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009040153280.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2010-09-04T13:42:05Z","receivedAt":"2010-09-04T13:42:05Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"On 09/04/10 08:13, Nicolas Pitre wrote:\n> On Sat, 4 Sep 2010, Artur Skawina wrote:\n>> What I'm really asking is, if a (modified) git-upload-pack skips transferring\n>> commit X, and just sends me commit Z (possibly as delta vs 'X'), _and_ I \n>> obtain commit 'X\" in some other way, I will be able to reconstruct 'Z', correct?\n> \n> Yes.  Although it is 'git pack-objects' that decides what objects to \n> send, not 'git-upload-pack'.\n\nThank you very much for the detailed answers.\n\nAFAIU both previously mentioned assumptions hold, so here's an example of\ngit-p2p-v3 use, simplified and with most boring stuff (p2p,ref and error\nhandling omitted.\n(the first version made a canonical, shared, virtual representation of the\nobject store, the second added more git-awareness to the transport, and then\nI started wondering if all of that is actually necessary; hence...).\n\nLet's say I'm a git repo tracking Linus' tree, right now the newest commit that\ni have is \"v2.6.33\" (but it could be anything, including \"\" for a fresh, empty\nclone) and I want to become up to date.\n\n1) I fetch a list of IPs of well known seeds, eg from kernel.org.\n\n2) I send an UDP packet to some of them, containing the repo \n   (\"git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux-2.6.git\"),\n   the ref that I'm interested in (\"master\") and the hash of the last commit\n   that I have (\"60b341b7\").\n   This is enough to start participating in the cloud by serving \"..\"60b341b7\",\n   but we'll skip the server part in this example.\n\n3) I receive answers to some of the above queries, containing the status of\n   these peers wrt to the given repo and ref, ie the same data I sent above.\n   Plus a list of random other live peers known to be tracking this ref, which\n   I'll use to repeat step #2 and #3 until I have a list of enough peers to\n   continue.\n\n4) Now i know of 47 peers that already have the tag or commit \"v2.6.37\" (either\n   I already knew that I wanted this one, or determined it during #3 and/or #1;\n   ref handling omitted from this example for brevity).\n\n   So i connect to one of the peers, and basically ask for the equivalent of\n   \"git fetch peer01 v2.6.37\". \n   But that would pull all new objects from that one peer, and that isn't what\n   i want. So i need to make it not only send me a thin pack, but also to omit\n   some of the objects. As at this point i don't actually know anything about\n   the objects in between \"v2.6.33\" and \"v2.6.37\" I can not split the request\n   into smaller ones.\n\n   So I'll cheat -- I'll take the number of available peers (\"47\") and the\n   number of this peer (\"0\"), send these two integers over and ask the other\n   side to skip transferring me any object whose \n   (HASH%available_peers)!=this_peer .\n\n5) for (int this_peer=1; this_peer<available_peers; this_peer++)\n     Repeat#4(this_peer);\n   /* in parallel until i saturate the link */\n\n6) Now i have 47 different packs, which probably do not make any sense\n   individually, because they contain deltas vs nonexisting objects, but\n   as a whole can be used to reconstruct the full tree.\n   6a) Except of course if there are circular dependencies, which can\n       occur eg. if peer#1 decided to encode object A as delta(B) and \n       peer#2 did B=delta(A), but this will be rare, and I'll just need\n       to refetch either A or B to break the cycle, this time with real\n       I-HAVES, hence this is guaranteed to succeed.\n\nWhat am I missing?\n\nartur\n"},{"id":"149840","messageId":"AANLkTim1XMY6Qe+h9LpqfoBzFE+B5AobcOpHx1rDfXwZ@mail.gmail.com","threadId":"24935","inReplyTo":"AANLkTikVf=X8cLP9s6W9VGOt0EHE4J5MYsBpgKYhrAri@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T14:06:29Z","receivedAt":"2010-09-04T14:06:29Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 6:23 AM, Kyle Moffett <kyle@moffetthome.net> wrote:\n> So I think the real problem to solve would be:  Given a collection of\n> local computers each with many local repositories, what is the best\n> way to optimize a clone of a \"new\" remote repository (over a slow\n> link) by copying most of the data from other local repositories\n> accessible via a fast link?\n\n the most immediate algorithm that occurs to me that would be ideal\nwould be - rsync!  whilst i had the privilege of being able to listen\n10 years ago to tridge describe rsync in detail, so am aware that it\nis parallelisable (which was the whole point of his thesis), a) i'm\nnot sure how to _distribute_ it b) i wouldn't have a clue where to\nstart!\n\nso, i believe that a much simpler algorithm is to follow nicolas' advice, and:\n\n* split up a pack-index file by its fanout (1st byte of SHAs in the idx)\n* create SHA1s of the list of object-refs within an individual fanout\n* compare the per-fanout SHA1s remote and local\n* if same, deduce \"oh look, we have that per-fanout list already\"\n* grab the per-fanout object-ref list using standard p2p filesharing\n\nin this way you'd end up breaking down e.g. 50mb of pack-index (for\ne.g. linux-2.6.git) into rouughly 200k chunks, and you'd exchange\nrouughly 50k of network traffic to find out that you'd got some of\nthose fanout object-ref-lists already.  which is nice.\n\n(see Documentation/technical/pack-format.txt, \"Pack Idx File\" for\ndescription of fanouts, but according to gitdb/pack.py it's just the\n1st byte of the SHA1s it points to)\n\n> The goal would be to design a P2P protocol capable of rapidly and\n> efficiently building distributed searchable indexes of ordered commits\n> that identify which peer(s) contain that each commit.\n\n yyyyup.\n\n> When you attempt to perform a \"git fetch --peer\" from a repository, it\n> would quickly connect to a few of the metadata index nodes in the P2P\n> network and use them to negotiate \"have\"s with the upstream server.\n> The client would then sequentially perform the local \"fetch\"\n> operations necessary to obtain all the objects it used to minimize the\n> commit range with the server.  Once all of those \"fetch\" operations\n> completed, it could proceed to fetch objects from the server normally.\n\n why stop at just fetching objects only from the server?  why not have\nthe objects distributed as well?  after all, if one peer has just gone\nto all the trouble of getting an object, surely it can share it, too?\n\n or am i misunderstanding what you're describing?\n\n> Some amount of design and benchmarking would need to be done in order\n> to figure out the most efficient indexing algorithm for finding a\n> minimal set of \"have\"s of potentially thousands of refs, many with\n> independent root commits.  For example if the index was grouped\n> according to \"root commit\" (of which there may be more than one), you\n> *should* be able to quickly ask the server about a small list of root\n> commits and then only continue asking about commits whose roots are\n> all known to the server.\n\n intuitively i follow what you're saying.\n\n> The actual P2P software would probably involve 2 different daemon\n> processes.  The first would communicate with each other and with the\n> repositories, maintaining the ref and commit indexes.  These daemons\n> would advertise themselves with Avahi,\n\n NO.\n\n ok.  more to the point: you want to waste time forcing people to\ninstall a pile of shite called d-bus, just so that people can use git,\ngo ahead.\n\ncan we gloss quickly over the mention of avahi as that *sweet-voice*\nmost delightful be-all and solve-all solution *normal-voice* and move\non?\n\n> or alternatively in an\n> enterprise environment they would be managed by your sysadmins and be\n> automatically discovered using DNS-SD.\n\n and what about on the public hostile internet?  no - i feel that the\nbasis should be something that's proven already, that's had at least\nten years to show for itself.  deviating from that basis: fine - at\nleast there's not a massive amount of change required for\nre-implementors to rewrite a compatible version (in c, or *shudder*\njava)\n\n> The second daemon would be a modified git-daemon that connects to the\n> advertised \"index\" daemons and advertises its own refs and commit\n> lists, as well as its IP address and port.\n\n yes, there are definitely two distinct purposes.  i'm not sure it's\nnecessary - or a good idea - to split the two out into separate\ndaemons, for the reason that you may, just like bittorrent can use a\nsingle port to share multiple torrents comprising multiple files, wish\nto use a single daemon to serve multiple git repositories.\n\n if you start from the basis of splitting things out then you have a\nbit of a headache on your hands wrt publishing multiple git repos.\n\nl.\n"},{"id":"149843","messageId":"AANLkTi==yv2CkgKEPJbTLf0P2XMtLmny1t6Zqhwh8wbV@mail.gmail.com","threadId":"24935","inReplyTo":"5B5470E5-57E6-48D2-981B-CE77FA43546F@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T14:50:29Z","receivedAt":"2010-09-04T14:50:29Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 1:00 PM, Theodore Tso <tytso@mit.edu> wrote:\n\n> such as ebook DRM liberation scripts (the kind which today\n> are typically distributed via pastebin's :-), my guess is that\n> zipping up a git repository and dropping it on a standard\n> bittorrent server run by the Swedish Pirate party is going to\n> be much more effective.   :-)\n\n:)  the legality or illegality isn't interesting - or is a... red\nherring, being one of the unfortunate anarchistic-word-associations\nwith the concept of \"file sharing\".  the robustness and convenience\naspects - to developers not users - is where it gets reaaally\ninteresting.\n\n i do not know of a single free software development tool - not a\nsingle one - which is peer-to-peer distributed.  just... none.  what\ndoes that say??  and we have people bitching about how great but\nnon-free skype is.  there seems to be a complete lack of understanding\nof the benefits of peer-to-peer infrastructure in the free software\ncommunity as a whole, and a complete lack of interest in the benefits,\ntoo - perhaps for reasons no more complex than the tools don't exist\nso it's catch-22, and the fact that the word \"distributed\" is\n_already_ associated with the likes of SMTP, DNS and \"git\" so\neverybody thinks \"we're okay, jack, go play with your nice dreams of\np2p networking, we're gonna write _real_ code now\".\n\n... mmmm :)\n\nl.\n"},{"id":"149856","messageId":"4C82809A.3020101@gmail.com","threadId":"24935","inReplyTo":"20100904155638.GA17606@pcpool00.mathematik.uni-freiburg.de","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2010-09-04T17:23:38Z","receivedAt":"2010-09-04T17:23:38Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"On 09/04/10 17:56, Bernhard R. Link wrote:\n> While your approach looks like it could work with one commonly looked\n> for branch/tag, it might be worthwhile to look at some more complicated\n> cases and see if the protocol can be extended to also support those.\n> \n> Assume you are looking to update abranch B built on top of some\n> say arm-specific branch A based on top of the main-line kernel L.\n> \n> There you have 6 types of peers:\n> \n> 1) those that have the current branch B\n> 2) those that have an older state of B\n> 3) those that have the current branch A\n> 4) those that have an older state of A\n> 5) those that have the current branch L\n> 6) those that have an older state of L.\n> \n> Assuming you want a quite obscure branch, type 1 and 2 peers will not\n> be that many, so would be nice if there was some way to also get stuff\n> from the others.\n> \n> Peers of type 6 do not interest you, as there will be enough of type 5.\n> \n> But already peers of type 3 might not be much and even less of type 1\n> so types 2 and 4 get interesting.\n> \n> If you get first a tree of all commit-ids you still miss (you only\n> need that information once and every peer of type 1 can give it to you)\n> and have somehow a way to look at the heads of each peer, it should be\n> streight forward to split the needed objects using your way (not specifying\n> the head you have but the one you hope to get from others), they should be\n> able to send you a partial pack in the way you describe.\n\nI doubt git-p2p would work well w/ any \"quite obscure branch\";\nobviously, the p2p approach works best for popular content...\n\nBut there's the case of _new_ content, that has not already propagated\nthrugh the swarm. And this was the very reason I chose to have 'ref' in\nthe protocol.\n\nLet's say Linus releases a new kernel, I want to fetch it, and find 200\npeers of which only three have already updated.\nHere another field in the initial UDP protocol comes in, that i omitted\nin the original description in order to keep things simple (it's just\nan optimization).\nThe UDP response from the peer, in addition to the latest commit_id that\nit has, also contains an integer \"commits_ahead\" that says how many\ncommits ahead of the _my_ id (that I've sent in the request) this peer is.\nSo in the above situation I immediately find out that eg I'm missing 2000\ncommits. It would be extremely stupid to try to update using just the\nthree seeds, of course. \nBut by looking at all the commit_aheads of all peers I can see\n(make an educated guess, really) that 150 of them already have 1500 of\nthe new commits. So what i do is pick one of the IDs given by one of the\npeers, such that a sufficient number of other sources are likely to\nalready have it (ie they all have a lower 'commit_ahead'). And start\nfetching that commit from the 150 sources, instead of my real target commit.\nOnce I'm done with this, I'll repeat the process again; by now hopefully\nmore peers have already updated. If the target commit is still rare I can\nagain look for an intermediate commit, that has a sufficient numbers of\nseeds. \nNo extra traffic, and it also discourages abuse, as leeching from just the\nfew seeds will likely result in overall slower download.\n\nThis works best if the ref isn't rewound, obviously, but i think that would\nbe the common case.\n\nDoes this address at least part of your concerns above? The 'obscure branch'\ncase I'm not sure is easily solvable; I don't know if having a lot of rarely\nused content widely distributed in the swarm would be a great idea...\n\nOne other interesting case is reusing objects needed by (large) merges, that\ncould already be available in the cloud. I'm not sure it happens often enough\nto care though...\n\nartur\n"},{"id":"149858","messageId":"20100904181405.GB4887@thunk.org","threadId":"24935","inReplyTo":"AANLkTi==yv2CkgKEPJbTLf0P2XMtLmny1t6Zqhwh8wbV@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2010-09-04T18:14:05Z","receivedAt":"2010-09-04T18:14:05Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sat, Sep 04, 2010 at 03:50:29PM +0100, Luke Kenneth Casson Leighton wrote:\n> \n> :)  the legality or illegality isn't interesting - or is a... red\n> herring, being one of the unfortunate anarchistic-word-associations\n> with the concept of \"file sharing\".  the robustness and convenience\n> aspects - to developers not users - is where it gets reaaally\n> interesting.\n\nI ask the question because I think being clear about what goals might\nbe are critically important.  If in fact the goals is to evade\ndetection be spreading out responsibility for code which is illegal in\nsome jurisdictions (even if they are commonly used and approved of by\npeople who aren't spending millions of dollars purchasing\ncongresscritters), there are many additional requirements that are\nimposed on such a system.\n\nIf the goal is speeding up git downloads, then we need to be careful\nabout exactly what problem we are trying to solve.\n\n>  i do not know of a single free software development tool - not a\n> single one - which is peer-to-peer distributed.  just... none.  what\n> does that say??  and we have people bitching about how great but\n> non-free skype is.  there seems to be a complete lack of understanding\n> of the benefits of peer-to-peer infrastructure in the free software\n> community as a whole, and a complete lack of interest in the benefits,\n> too...\n\nMaybe it's because the benefits don't exist for many people?  At least\nwhere I live, my local ISP (Comcast, which is very common internet\nprovider in the States) deliberately degrades the transfer of\npeer2peer downloads.  As a result, it doesn't make sense for me to use\nbittorrent to download the latest Ubuntu or Fedora iso image.  It's in\nfact much faster for me to download it from an ftp site or a web site.\n\nAnd git is *extremely* efficient about its network usage, since it\nsends compressed deltas --- especially if you already have a base\nresponsitory estlablished.  For example, I took a git repository which\nI haven't touched since August 4th --- exactly one month ago --- and\ndid a \"git fetch\" to bring it up to date by downloading from\ngit.kernel.org.  How much network traffic was required, after being\none month behind?  2.8MB of bytes received, 133k of bytes transmitted.\n\nThat's not a lot.  And it's well within the capabilities of even a\nreally busy server to handle.  Remember, peer2peer only helps if the\naggregate network bandwidth of the peers is greater than (a) your\ndownload pipe, or (b) a central server's upload pipe.  And if we're\nonly transmitting 2.8MB, and a git.kernel.org has an aggregate\nconnection of over a gigabit per second to the internet --- it's not\nlikely that peer2peer would in fact result in a faster download.  Nor\nis it likely that that git updates are likely to be something which\nthe kernel.org folks would even notice as a sizeable percentage of\ntheir usable network bandwidth.  First of all, ISO image files are\nmuch bigger, and secondly, there are many more users downloading ISO\nfiles than there are developers downloading git updates, and certainly\nrelatively few developers downloading full git repositories (since\neverybody genreally tries really hard to only do this once).\n\n> - perhaps for reasons no more complex than the tools don't exist\n> so it's catch-22, and the fact that the word \"distributed\" is\n> _already_ associated with the likes of SMTP, DNS and \"git\" so\n> everybody thinks \"we're okay, jack, go play with your nice dreams of\n> p2p networking, we're gonna write _real_ code now\".\n\nWell, perhaps it's because what we have right now works pretty well.   :-)\n\nWhich brings me back to my original question --- what problem exactly\nare you trying to solve?  What's the scenario?\n\nIf the answer is \"peer2peer is cool technology, and we want to play\",\nthat's fine.  Put it would confirm the hypothesis that in this case,\npeer2peer is a solution looking for a problem...\n\n\t\t\t\t\t- Ted\n"},{"id":"149859","messageId":"4C829413.1010302@gmail.com","threadId":"24935","inReplyTo":"20100904155638.GA17606@pcpool00.mathematik.uni-freiburg.de","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2010-09-04T18:46:43Z","receivedAt":"2010-09-04T18:46:43Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"On 09/04/10 17:56, Bernhard R. Link wrote:\n> Assume you are looking to update abranch B built on top of some\n> say arm-specific branch A based on top of the main-line kernel L.\n> \n> There you have 6 types of peers:\n> \n> 1) those that have the current branch B\n> 2) those that have an older state of B\n> 3) those that have the current branch A\n> 4) those that have an older state of A\n> 5) those that have the current branch L\n> 6) those that have an older state of L.\n> \n> Assuming you want a quite obscure branch, type 1 and 2 peers will not\n> be that many, so would be nice if there was some way to also get stuff\n> from the others.\n> \n> Peers of type 6 do not interest you, as there will be enough of type 5.\n> \n> But already peers of type 3 might not be much and even less of type 1\n> so types 2 and 4 get interesting.\n> \n> If you get first a tree of all commit-ids you still miss (you only\n> need that information once and every peer of type 1 can give it to you)\n> and have somehow a way to look at the heads of each peer, it should be\n> streight forward to split the needed objects using your way (not specifying\n> the head you have but the one you hope to get from others), they should be\n> able to send you a partial pack in the way you describe.\n\nActually, since the seeds of this quite obscure branch know that this branch\nis rare and are most likely also sharing other, much more popular, refs, they\ncan just add another note to the initial reply, saying \n\"this_ref_is_descendant_of($popular_branch, $common_base)\". Then the client\ncan fetch old..$common_base and $common_base..$rare_head independently (even\nin parallel).\n\nartur\n"},{"id":"149866","messageId":"AANLkTikAfSrfKRaK3ozXV_eT6Rd-VRbXQUQLk3SY8QnJ@mail.gmail.com","threadId":"24935","inReplyTo":"20100904181405.GB4887@thunk.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T20:00:56Z","receivedAt":"2010-09-04T20:00:56Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 7:14 PM, Ted Ts'o <tytso@mit.edu> wrote:\n> At least\n> where I live, my local ISP (Comcast, which is very common internet\n> provider in the States) deliberately degrades the transfer of\n> peer2peer downloads.\n\n if microsoft can add ncacn_http to MSRPC for the exact same sorts of\nreasons, and even skype likewise provides a user-config option to\nspecify \"port 80\" or port \"3128\", then it's perfectly possible to do\nlikewise. ncacn_http actually has HTTP 1.1 headers on it, and, once\nyou've negotiated enough to look like HTTP, the raw socket is hander\nover to MSRPC for it to play with.\n\n> Which brings me back to my original question --- what problem exactly\n> are you trying to solve?  What's the scenario?\n\ni described those in prior messages.  to summarise: they're basically\nreduction of dependence on centralised infrastructure, and to allow\ndevelopers to carry on doing code-sprints using bugtrackers, wikis and\nanything else that can be \"git-able\" as its back-end, _even_ in the\ncases where there is little or absolutely no bandwidth... and _still_\nsync up globally once any one of the developers gets back online.\n\n so i'm _not_ just thinking of the \"code, code, code\" scenario, and\ni'm not just thinking in terms of the single developer \"i code,\ntherefore i am\" scenario.  i'm thinking of scenarios which increase\nthe productivity of collaborative development even in the face of\nunreliable or non-existent connectivity [carrier pigeons...]\n\nl.\n"},{"id":"149869","messageId":"m3tym5mfce.fsf@localhost.localdomain","threadId":"24935","inReplyTo":"20100904181405.GB4887@thunk.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-09-04T20:20:42Z","receivedAt":"2010-09-04T20:20:42Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Ted Ts'o <tytso@mit.edu> writes:\n> On Sat, Sep 04, 2010 at 03:50:29PM +0100, Luke Kenneth Casson Leighton wrote:\n> > \n> > :)  the legality or illegality isn't interesting - or is a... red\n> > herring, being one of the unfortunate anarchistic-word-associations\n> > with the concept of \"file sharing\".  the robustness and convenience\n> > aspects - to developers not users - is where it gets reaaally\n> > interesting.\n> \n> I ask the question because I think being clear about what goals might\n> be are critically important.  If in fact the goals is to evade\n> detection be spreading out responsibility for code which is illegal in\n> some jurisdictions (even if they are commonly used and approved of by\n> people who aren't spending millions of dollars purchasing\n> congresscritters), there are many additional requirements that are\n> imposed on such a system.\n> \n> If the goal is speeding up git downloads, then we need to be careful\n> about exactly what problem we are trying to solve.\n> \n> >  i do not know of a single free software development tool - not a\n> > single one - which is peer-to-peer distributed.  just... none.  what\n> > does that say??  and we have people bitching about how great but\n> > non-free skype is.  there seems to be a complete lack of understanding\n> > of the benefits of peer-to-peer infrastructure in the free software\n> > community as a whole, and a complete lack of interest in the benefits,\n> > too...\n\nLuke, you don't have to be peer-to-peer to be decentralized and\ndistributed.  People from what I understand bitch most about\ncentralized (and closed) services.\n \n> Maybe it's because the benefits don't exist for many people?  At least\n> where I live, my local ISP (Comcast, which is very common internet\n> provider in the States) deliberately degrades the transfer of\n> peer2peer downloads.  As a result, it doesn't make sense for me to use\n> bittorrent to download the latest Ubuntu or Fedora iso image.  It's in\n> fact much faster for me to download it from an ftp site or a web site.\n> \n> And git is *extremely* efficient about its network usage, since it\n> sends compressed deltas --- especially if you already have a base\n> responsitory established.  For example, I took a git repository which\n> I haven't touched since August 4th --- exactly one month ago --- and\n> did a \"git fetch\" to bring it up to date by downloading from\n> git.kernel.org.  How much network traffic was required, after being\n> one month behind?  2.8MB of bytes received, 133k of bytes transmitted.\n\nI think the major problem git-p2p wants to solve is if base repository\nis *not* established, i.e. the initial fetch / full clone operation.\n \nNote that with --reference argument to git clone, if you have similar\nrelated repository, you don't have to do a full fetch cloning a fork\nof repository you already have (e.g. you have Linus repo, and want to\nfetch linux-next).\n\n> That's not a lot.  And it's well within the capabilities of even a\n> really busy server to handle.  Remember, peer2peer only helps if the\n> aggregate network bandwidth of the peers is greater than (a) your\n> download pipe, or (b) a central server's upload pipe.  And if we're\n> only transmitting 2.8MB, and a git.kernel.org has an aggregate\n> connection of over a gigabit per second to the internet --- it's not\n> likely that peer2peer would in fact result in a faster download.  Nor\n> is it likely that that git updates are likely to be something which\n> the kernel.org folks would even notice as a sizeable percentage of\n> their usable network bandwidth.  First of all, ISO image files are\n> much bigger, and secondly, there are many more users downloading ISO\n> files than there are developers downloading git updates, and certainly\n> relatively few developers downloading full git repositories (since\n> everybody genreally tries really hard to only do this once).\n\nWell, full initial clone of Linux kernel repository (or any other\nlarge project with long history) is quite large.  Also, not all\nprojects have big upload pipe.\n\nAdditional problem is that clone is currrently non-resumable (at all),\nso if you have flaky web connection it might be hard to do initial\nclone.\n\n\nOne way of solving this problem that (as I have heard some projects\nuse) is to prepare \"initial\" bundle; this bundle can be downloaded via\nHTTP or FTP resumably, or be shared via ordinary P2P like BitTorrent.\n\nThe initial pack could be 'kept' (nt subject to repacking); with some\ncode it could serve as canonical starting packfile for cloning, I\nthink.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"149871","messageId":"AANLkTimhCi2vWWnHGwT5ToRtFbjkxTgVYVvYLR3UCb2S@mail.gmail.com","threadId":"24935","inReplyTo":"m3tym5mfce.fsf@localhost.localdomain","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T20:47:19Z","receivedAt":"2010-09-04T20:47:19Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 9:20 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n\n> Luke, you don't have to be peer-to-peer to be decentralized and\n> distributed.  People from what I understand bitch most about\n> centralized (and closed) services.\n\ni've covered this in the FAQ i wrote:\n\nFAQ:\n\nQ: is git a \"distributed source control system\"?\nA: yeees, but the \"distribution\" part has to be done the hard way,\n   by setting up servers, forcing developers and users to configure\n   git to use those [single-point-of-failure] servers.  so it's\n   \"more correct\" to say that git is a \"distributable\" source control\n   system.\n\n\nso if you believe that git is \"distributed\" just because people can\nset up a server.... mmm :)  i'd say that your administrative and\ntechnical skills are way above the average persons' capabilities.\nproper peer-to-peer networking infrastructure takes care of things\nlike firewall-busting, by using UPnP automatically, as part of the\ninfrastructure.\n\ncome on, people - _think_.  we're so used to being able to run our own\ninfrastructure and workaround problems or server down-time, but most\npeople still use \"winzip\", if they have any kind of revision control\n_at all_.\n\nl.\n"},{"id":"149872","messageId":"201009042316.38655.jnareb@gmail.com","threadId":"24935","inReplyTo":"AANLkTimhCi2vWWnHGwT5ToRtFbjkxTgVYVvYLR3UCb2S@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-09-04T21:16:30Z","receivedAt":"2010-09-04T21:16:30Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, Sep 4, 2010, Luke Kenneth Casson Leighton wrote:\n> On Sat, Sep 4, 2010 at 9:20 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n> \n> > Luke, you don't have to be peer-to-peer to be decentralized and\n> > distributed.  People from what I understand bitch most about\n> > centralized (and closed) services.\n> \n> i've covered this in the FAQ i wrote:\n> \n> FAQ:\n> \n> Q: is git a \"distributed source control system\"?\n> A: yeees, but the \"distribution\" part has to be done the hard way,\n>    by setting up servers, forcing developers and users to configure\n>    git to use those [single-point-of-failure] servers.  so it's\n>    \"more correct\" to say that git is a \"distributable\" source control\n>    system.\n \n\"Distributed\" is not equivalent to \"peer to peer\".\n \n> so if you believe that git is \"distributed\" just because people can\n> set up a server.... mmm :)  i'd say that your administrative and\n> technical skills are way above the average persons' capabilities.\n\nGit is distributed at least in the sense that it is opposite\nto centralized version control systems (like Subversion)\nwhere you have and can have only single server with full repository.\n\nSetting up server (git, smart HTTP, ssh) is not that hard.\n\n> proper peer-to-peer networking infrastructure takes care of things\n> like firewall-busting, by using UPnP automatically, as part of the\n> infrastructure.\n\nWith \"smart\" HTTP transport support there is no need for any \nfirewall-busting.\n \n> come on, people - _think_.  we're so used to being able to run our own\n> infrastructure and workaround problems or server down-time, but most\n> people still use \"winzip\", if they have any kind of revision control\n> _at all_.\n\nIn what way this paragraph refers to and is relewant with respect to\ndiscussion in this subthread?\n\n::plonk::\n-- \nJakub Narebski\nPoland\n"},{"id":"149873","messageId":"AANLkTi=-+XMPOqaeiE_z4s_tCKHhK-rpyCgZgNmOjmZ1@mail.gmail.com","threadId":"24935","inReplyTo":"201009042316.38655.jnareb@gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-04T21:24:37Z","receivedAt":"2010-09-04T21:24:37Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 10:16 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n> On Sat, Sep 4, 2010, Luke Kenneth Casson Leighton wrote:\n>> On Sat, Sep 4, 2010 at 9:20 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n>>\n>> > Luke, you don't have to be peer-to-peer to be decentralized and\n>> > distributed.  People from what I understand bitch most about\n>> > centralized (and closed) services.\n>>\n>> i've covered this in the FAQ i wrote:\n>>\n>> FAQ:\n>>\n>> Q: is git a \"distributed source control system\"?\n>> A: yeees, but the \"distribution\" part has to be done the hard way,\n>>    by setting up servers, forcing developers and users to configure\n>>    git to use those [single-point-of-failure] servers.  so it's\n>>    \"more correct\" to say that git is a \"distributable\" source control\n>>    system.\n>\n> \"Distributed\" is not equivalent to \"peer to peer\".\n\n correct.  exactly.\n\n> Setting up server (git, smart HTTP, ssh) is not that hard.\n\n for you and me, and for the majority of people developing git, this is correct.\n\n>> proper peer-to-peer networking infrastructure takes care of things\n>> like firewall-busting, by using UPnP automatically, as part of the\n>> infrastructure.\n>\n> With \"smart\" HTTP transport support there is no need for any\n> firewall-busting.\n\n the assumption is that the users are capable of deploying a server\n(at all), and are happy to reconfigure to use that server.\n\n ok.\n\n this is getting off-topic and is distracting, both for me and for\npeople wishing to read about git.  with respect, jakob, i'm going to\ndo something which i don't normally do, and that's begin to be\nselective about what i reply to.  apologies, but i'm on very short\ntimescales to get this code working, for financial reasons.\n\nl.\n"},{"id":"149894","messageId":"20100904224139.GD4887@thunk.org","threadId":"24935","inReplyTo":"AANLkTikAfSrfKRaK3ozXV_eT6Rd-VRbXQUQLk3SY8QnJ@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2010-09-04T22:41:39Z","receivedAt":"2010-09-04T22:41:39Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sat, Sep 04, 2010 at 09:00:56PM +0100, Luke Kenneth Casson Leighton wrote:\n> > Which brings me back to my original question --- what problem exactly\n> > are you trying to solve?  What's the scenario?\n> \n> i described those in prior messages.  to summarise: they're basically\n> reduction of dependence on centralised infrastructure, and to allow\n> developers to carry on doing code-sprints using bugtrackers, wikis and\n> anything else that can be \"git-able\" as its back-end, _even_ in the\n> cases where there is little or absolutely no bandwidth... and _still_\n> sync up globally once any one of the developers gets back online.\n\nSo at all of the code sprints I've been at, the developers all have\nlocally very good bandwidth between each other.  And if they don't\nhave wifi, what *will* they have?  In the example you gave, you never\nwere able to bring up a local area network, because you had one or two\nlamers who couldn't even do wifi in adhoc mode.  Hell, even if you had\nto hook up someone's laptop using an RS-232 line and PPP, that would\nbe plenty of bandwidth for git.  So you weren't specific enough in\nyour scenario.  How could it happen?  And is it really all that realistic?\n\nEven if it did, it wouldn't be hard to just set up a git server on one\nof the laptop.  What makes peer2peer so critically important in this\nuse case?  (And no, carrier pigeons are not particularly realistic for\na code sprint....)\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"149896","messageId":"20100904224726.GE4887@thunk.org","threadId":"24935","inReplyTo":"AANLkTimhCi2vWWnHGwT5ToRtFbjkxTgVYVvYLR3UCb2S@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2010-09-04T22:47:26Z","receivedAt":"2010-09-04T22:47:26Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sat, Sep 04, 2010 at 09:47:19PM +0100, Luke Kenneth Casson Leighton wrote:\n> \n> Q: is git a \"distributed source control system\"?\n> A: yeees, but the \"distribution\" part has to be done the hard way,\n>    by setting up servers, forcing developers and users to configure\n>    git to use those [single-point-of-failure] servers.  so it's\n>    \"more correct\" to say that git is a \"distributable\" source control\n>    system.\n\nIs typing the command \"git daemon\" on the command line really that\nhard?  For bonus points it could register with Avahi (i.e., the\nZeroconf protocol), which would make it easier to set up adhoc sharing\narrangements.  But if that's your goal, I'd suggest making \"git\ndaemon\" simpler to set up --- since it's pretty trivial to set up as\nit is.  I'd then suggest adding to \"git gui\" an option to allow a user\nto browse local git servers who have advertised themselves using the\nZeroconf protocol and do a git fetch from the local server.\n\nBut if you want to work on peer2peer because it's \"cool\", feel free.\nBut I suspect telling people that all if you have to do is run \"git\ndaemon\" on one machine, and using \"git gui\" to browse through the\navailable git servers on the other machines, it really isn't going to\nget any easier....\n\n\t\t\t\t\t\t- Ted\n"},{"id":"149918","messageId":"alpine.LFD.2.00.1009041107180.19366@xanadu.home","threadId":"24935","inReplyTo":"5B5470E5-57E6-48D2-981B-CE77FA43546F@mit.edu","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-05T01:18:16Z","receivedAt":"2010-09-05T01:18:16Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 4 Sep 2010, Theodore Tso wrote:\n\n> \n> On Sep 4, 2010, at 1:40 AM, Nicolas Pitre wrote:\n> >> What about the order of the objects in the pack?  Well, ordering \n> >> doesn't matter, right?  So let's assume the pack is sorted by hash id.  \n> >> Is there any downside to that?  I can't think of any, but you're the \n> >> pack expert...\n> > \n> > Ordering does matter a big deal.  Since object IDs are the SHA1 of their \n> > content, those IDs are totally random.  So if you store objects \n> > according to their sorted IDs, then the placement of objects belonging \n> > to, say, the top commit will be totally random.  And since you are the \n> > filesystem expert, I don't have to tell you what performance impacts \n> > this random access of small segments of data scattered throughout a \n> > 400MB file will have on a checkout operation.\n> \n> Does git repack optimize the order so that certain things (like \n> checkouts for example) are really fast?  I admit I hadn't noticed.  \n\nIt does indeed.  The object ordering in the pack is so that checking out \nthe most recent commit will perform an almost perfect linear read from \nthe beginning of the pack file.  Then the further back you go in the \ncommit history the more random the access in the pack will be.  But \nthat's OK as the most accessed commits are the recent ones, so we \noptimize object placement for that case.\n\n> Usually until the core packs are in my page cache, it has always \n> seemed to me that things are pretty slow.\n\nCold cache numbers could be much much worse.  In theory, checking out \nthe latest commit after a repack should be similar to extracting the \nequivalent source tree from a tarball.\n\n> And of course, the way objects and grouped together and ordered for \n> \"gitk\" or \"git log\" to be fast won't be the safe as a checkout \n> operation...\n\nIndeed. But that case is optimized too, as all the commit objects are \nput together at the front of the pack so the walking of the commit \nhistory has good access locality too.\n\n> > Sure.  But I don't think it is worth making Git less flexible just for \n> > the purpose of ensuring that people could independently create identical \n> > packs.  I'd advocate for \"no code to write at all\" instead, and simply \n> > have one person create and seed the reference pack.\n> \n> I don't think it's a matter of making Git \"less flexible\", it's just \n> simply a code maintenance headache of needing to be able to support \n> encoding both a canonical format as well as the latest bleeding-edge, \n> most efficient encoding format.\n\nIndeed.  But what I'm saying is that one would have to put some efforts \ninto that canonical format by removing the current flexibility that Git \nhas in producing a pack.  For example, right now Git is relying on that \nflexibility when using threads. The algorithm for load balancing ends up \nmaking slight differences between different repack invocations even when \nthey're started with the same input.  With multiple threads, the output \nfrom pack-objects is therefore not deterministic.\n\n> And how often are you changing/improving the encoding process, anyway?  \n> It didn't seem to me like that part of the code was constantly being \n> tweaked/improved.\n\nIt used to change quite often. It is true that these days things have \nslowed down in that area, but I prefer to keep options open unless there \nis a really convincing reason not to.\n\n> Still, you're right, it might not be worth it.  To be honest, I was \n> more interested about the fact that this might also be used to give \n> people hints about how to better repack their local repositories so \n> that they didn't have to run git repack with large --window and \n> --depth arguments.  But that would only provide very small \n> improvements in storage space in most cases, so it's probably not even \n> worth it for that.\n> \n> Quite frankly, I'm a little dubious about how critical peer2peer \n> really is, for pretty much any use case.\n\nI agree.  So far it has been an interesting topic for discussion, but in \npractice I doubt the actual benefits will justify the required efforts \nand/or constraints on the protocol. Otherwise we would have a working \nimplementation in use already.  People tried in the past, and so far \nnone of those attempts passed the reality test nor kept people motivated \nenough to work on them further.\n\n> Most of the time, I can grab \n> the base \"reference\" tree and drop it on my laptop before I go off the \n> grid and have to rely on EDGE or some other slow networking \n> technology.  And if the use case is some small, but \n> illegal-in-some-jurisdiction code, such as ebook DRM liberation \n> scripts (the kind which today are typically distributed via pastebin's \n> :-), my guess is that zipping up a git repository and dropping it on a \n> standard bittorrent server run by the Swedish Pirate party is going to \n> be much more effective.  :-)\n\nWho cares about the development history for such scripts anyway?\n\n\nNicolas\n"},{"id":"149919","messageId":"alpine.LFD.2.00.1009042119570.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTim1XMY6Qe+h9LpqfoBzFE+B5AobcOpHx1rDfXwZ@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-05T01:32:04Z","receivedAt":"2010-09-05T01:32:04Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 4 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> so, i believe that a much simpler algorithm is to follow nicolas' advice, and:\n> \n> * split up a pack-index file by its fanout (1st byte of SHAs in the idx)\n> * create SHA1s of the list of object-refs within an individual fanout\n> * compare the per-fanout SHA1s remote and local\n> * if same, deduce \"oh look, we have that per-fanout list already\"\n> * grab the per-fanout object-ref list using standard p2p filesharing\n> \n> in this way you'd end up breaking down e.g. 50mb of pack-index (for\n> e.g. linux-2.6.git) into rouughly 200k chunks, and you'd exchange\n> rouughly 50k of network traffic to find out that you'd got some of\n> those fanout object-ref-lists already.  which is nice.\n\nScrap that idea -- this won't work.  The problem is that, by nature, \nSHA1 is totally random.  So if you have, say, 256 objects to transfer \n(and 256 objects is not that much) then, statistically, the probability \nthat the SHA1s for those objects end up uniformly distributed across all \nthe 256 fanouts is quite high.  the algorithm I mentioned completely \nbreaks down in that case.\n\n\nNicolas\n"},{"id":"149920","messageId":"4C82F5CA.2080502@dbservice.com","threadId":"24935","inReplyTo":"20100904224726.GE4887@thunk.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2010-09-05T01:43:38Z","receivedAt":"2010-09-05T01:43:38Z","isPatch":false,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"On 9/5/10 12:47 AM, Ted Ts'o wrote:\n> hard?  For bonus points it could register with Avahi (i.e., the\n> Zeroconf protocol), which would make it easier to set up adhoc sharing\n> arrangements.  But if that's your goal, I'd suggest making \"git\n\nhttp://github.com/toolmantim/bananajour, and there's another project\nwith similar goals, but I forgot its name and can't find it right now.\n\ntom\n"},{"id":"149923","messageId":"alpine.LFD.2.00.1009042132500.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTik9awEd40s3r-O8t9DwZBh34Z0ozsxMm1QNjNoT@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-05T02:16:02Z","receivedAt":"2010-09-05T02:16:02Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 4 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> * git-index-pack requires a pack file in order to re-create the index:\n> i don't want that\n> * git-pack-objects appears to have no way of telling it \"just gimme\n> index file please\"\n> * fast-import.c appears not to be what's needed either.\n> \n> so - any other methods for just getting the index file (exclusively?)\n> any other commands i've missed?  if not, are there any other ways of\n> getting a pack's index of objects without err... getting the index\n> file?  (i believe the answer to be no, but i'm just making sure) and\n> on that basis i believe it is safe to ask: any objections to a patch\n> which adds \"--index-only\" to builtin/pack-objects.c?\n\nNo patch is needed.\n\nFirst, what you want is an index of objects you are willing to share, \nand not the index of whatever pack file you might have on your disk, \nespecially if you have multiple packs which is typical.\n\nTry this instead:\n\n    git rev-list --objects HEAD | cut -c -40 | sort\n\nThat will give you a sorted list of all objects reachable from the \ncurrent branch.  With the Linux repo, you may replace \"HEAD\" with \n\"v2.6.34..v2.6.35\" if you wish, and that would give you the list of the \nnew objects that were introduced between v2.6.34 and v2.6.35.  This will \nprovide you with 84642 objects instead of the 1.7 million objects that \nthe Linux repo contains (easier when testing stuff).\n\nThat sorted list of objects is more or less what the pack index file \ncontains, plus an offset in the pack for each entry.  It is used to \nquickly find the offset for a given object in the corresponding pack \nfile, and the fanout is only a way to cut 3 iterations in the binary \nsearch.\n\nBut anyway, what you want is really to select the precise set of objects \nyou wish to share, and not blindly using the pack index file.  If you \nhave a public branch and a private branch in your repository, then \nobjects from both branches may end up in the same pack and you probably \ndon't want to publish those objects from the private branch. The only \nreliable way to generate a list of object is to use the output from 'git \nrev-list'.  Those objects may come from one or multiple packs, or be \nloose in the object subdirectories, or even borrowed from another \nrepository through the alternates mechanism.  But rev-list will dig \nthose object SHA1s for you and only those you asked for.\n\nYou should look at the Git documentation for plumbing commands.  The \nplumbing is actually a toolset that allows you to manipulate and extract \ninformation from a Git repository.  This is really handy for prototyping \nnew functionalities. Initially, the Git user interface was all \nimplemented in shell scripts on top of that plumbing.\n\nBack to that rev-list output... OK, you want the equivalent of a fanout \ntable.  You may do something like this then:\n\n    git rev-list --objects v2.6.34..v2.6.35 | cut -c -2 | sort | uniq -c\n\nAnd so on.\n\n\nNicolas\n"},{"id":"149973","messageId":"AANLkTimZA=VpGjcZEjoRVJUZcwnYoPQF5bNHyM2J8byE@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009042119570.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-05T17:16:26Z","receivedAt":"2010-09-05T17:16:26Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sun, Sep 5, 2010 at 2:32 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Sat, 4 Sep 2010, Luke Kenneth Casson Leighton wrote:\n>\n>> so, i believe that a much simpler algorithm is to follow nicolas' advice, and:\n>>\n>> * split up a pack-index file by its fanout (1st byte of SHAs in the idx)\n>> * create SHA1s of the list of object-refs within an individual fanout\n>> * compare the per-fanout SHA1s remote and local\n>> * if same, deduce \"oh look, we have that per-fanout list already\"\n>> * grab the per-fanout object-ref list using standard p2p filesharing\n>>\n>> in this way you'd end up breaking down e.g. 50mb of pack-index (for\n>> e.g. linux-2.6.git) into rouughly 200k chunks, and you'd exchange\n>> rouughly 50k of network traffic to find out that you'd got some of\n>> those fanout object-ref-lists already.  which is nice.\n>\n> Scrap that idea -- this won't work.  The problem is that, by nature,\n> SHA1 is totally random.  So if you have, say, 256 objects to transfer\n> (and 256 objects is not that much) then, statistically, the probability\n> that the SHA1s for those objects end up uniformly distributed across all\n> the 256 fanouts is quite high.  the algorithm I mentioned completely\n> breaks down in that case.\n\n mmm... that's no so baad.  requesting a table/pseudo-file with 1\nfanout or 256 fanouts is still only one extra round-trip.  if i split\nit into pseudo-subdirectories _then_ yes you have 256 requests.  that\ncan be avoided with a bit of work.  so, no biggie :)\n\nl.\n"},{"id":"149974","messageId":"AANLkTikgiO21M7a7Ovz5nB2kW60dE7wQe5gc-4O+wbER@mail.gmail.com","threadId":"24935","inReplyTo":"20100904224139.GD4887@thunk.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-05T17:22:59Z","receivedAt":"2010-09-05T17:22:59Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sat, Sep 4, 2010 at 11:41 PM, Ted Ts'o <tytso@mit.edu> wrote:\n> On Sat, Sep 04, 2010 at 09:00:56PM +0100, Luke Kenneth Casson Leighton wrote:\n>> > Which brings me back to my original question --- what problem exactly\n>> > are you trying to solve?  What's the scenario?\n>>\n>> i described those in prior messages.  to summarise: they're basically\n>> reduction of dependence on centralised infrastructure, and to allow\n>> developers to carry on doing code-sprints using bugtrackers, wikis and\n>> anything else that can be \"git-able\" as its back-end, _even_ in the\n>> cases where there is little or absolutely no bandwidth... and _still_\n>> sync up globally once any one of the developers gets back online.\n>\n> So at all of the code sprints I've been at, the developers all have\n> locally very good bandwidth between each other.  And if they don't\n\n ted - with respect, much as i'd like to debate the merits or\notherwise of the purpose of this work, i'd far rather actually focus\non actually doing it. can i leave it to you and the other people here\non the list to debate both sides - actually three sides because there\nis a case for helping casey to get \"git hive\" going, as well?\n\n i look forward to seeing lots more ideas and use-cases beyond those\nwhich i can envisage, and am grateful to the people who have been\nprivately contacting me to express gratitude at potentially having\nsomething which makes software development and free software\ninvolvement easier, in circumstance such as difficult or expensive\ninternet connectivity.\n\n l.\n"},{"id":"149975","messageId":"AANLkTim8XLB5SjV3JtWT-ARN_XuofKDjYRSYT8kPxEvq@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009041107180.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-05T17:25:42Z","receivedAt":"2010-09-05T17:25:42Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sun, Sep 5, 2010 at 2:18 AM, Nicolas Pitre <nico@fluxnic.net> wrote\n> I agree.  So far it has been an interesting topic for discussion, but in\n> practice I doubt the actual benefits will justify the required efforts\n> and/or constraints on the protocol. Otherwise we would have a working\n> implementation in use already.  People tried in the past, and so far\n> none of those attempts passed the reality test nor kept people motivated\n> enough to work on them further.\n\n then i'm all the more grateful that you continue to drop technical\nhints in my direction.  thank you for not judging.\n\n l.\n"},{"id":"149984","messageId":"AANLkTi=YLx6MqbWd_N0geXbuXLdqAUOneGoym75dfthL@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009042132500.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-05T18:05:57Z","receivedAt":"2010-09-05T18:05:57Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Sun, Sep 5, 2010 at 3:16 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Sat, 4 Sep 2010, Luke Kenneth Casson Leighton wrote:\n>\n>> * git-index-pack requires a pack file in order to re-create the index:\n>> i don't want that\n>> * git-pack-objects appears to have no way of telling it \"just gimme\n>> index file please\"\n>> * fast-import.c appears not to be what's needed either.\n>>\n>> so - any other methods for just getting the index file (exclusively?)\n>> any other commands i've missed?  if not, are there any other ways of\n>> getting a pack's index of objects without err... getting the index\n>> file?  (i believe the answer to be no, but i'm just making sure) and\n>> on that basis i believe it is safe to ask: any objections to a patch\n>> which adds \"--index-only\" to builtin/pack-objects.c?\n>\n> No patch is needed.\n>\n> First, what you want is an index of objects you are willing to share,\n> and not the index of whatever pack file you might have on your disk,\n> especially if you have multiple packs which is typical.\n\n blast.  so *sigh* ignoring the benefits that can be obtained by the\ndelta-compression thing, somewhat; ignoring the fact that perhaps less\ntraffic miight be transferred by happening to borrow objects from\nanother branch (which is the situation that, i believe, happens with\n\"git pull\" over http:// or git://); ignoring the fact that i actually\nimplemented using the .idx file yesterday ... :)\n\n ... there is a bit of a disadvantage to using pack index files that\nit goes all the way down (if i am reading things correctly) and cannot\nbe told \"give me just the objects related to a particular commit\"....\n\n\n> Try this instead:\n>\n>    git rev-list --objects HEAD | cut -c -40 | sort\n>\n> That will give you a sorted list of all objects reachable from the\n> current branch.  With the Linux repo, you may replace \"HEAD\" with\n> \"v2.6.34..v2.6.35\" if you wish, and that would give you the list of the\n> new objects that were introduced between v2.6.34 and v2.6.35.\n\n ... unlike this, which is in fact much more along the lines of what i\nwas looking for (minus the loveliness of the delta compression oh\nwell)\n\n> This will\n> provide you with 84642 objects instead of the 1.7 million objects that\n> the Linux repo contains (easier when testing stuff).\n\n hurrah! :)  [but, then if you actually want to go back and get alll\ncommits, that's ... well, we'll not worry about that too much, given\nthe benefits of being able to get smaller chunks.]\n\n> That sorted list of objects is more or less what the pack index file\n> contains, plus an offset in the pack for each entry.  It is used to\n> quickly find the offset for a given object in the corresponding pack\n> file, and the fanout is only a way to cut 3 iterations in the binary\n> search.\n>\n> But anyway, what you want is really to select the precise set of objects\n> you wish to share, and not blindly using the pack index file.  If you\n> have a public branch and a private branch in your repository, then\n> objects from both branches may end up in the same pack\n\n slightly confused: are you of the belief that i intend to ignore\nrefs/branches/* starting points?\n\n> and you probably\n> don't want to publish those objects from the private branch.\n\n ahh, i wondered where i'd seen the bit about \"confusing\" two\nbranches, i thought it was in another message.  so many flying back &\nforth :)  from what i can gather, this is exactly what happens with\ngit fetch from http:// or git:// so what's the big deal about that?\nwhy stop gitp2p from benefitting from the extra compression that could\nresult from \"borrowing\" bits of another branch's objects, neh?\n\n or .. have i misunderstood?\n\n> The only\n> reliable way to generate a list of object is to use the output from 'git\n> rev-list'.  Those objects may come from one or multiple packs, or be\n> loose in the object subdirectories, or even borrowed from another\n> repository through the alternates mechanism.  But rev-list will dig\n> those object SHA1s for you and only those you asked for.\n\n excellent.  that's proobably what i need right now.\n\n> You should look at the Git documentation for plumbing commands.  The\n> plumbing is actually a toolset that allows you to manipulate and extract\n> information from a Git repository.  This is really handy for prototyping\n> new functionalities. Initially, the Git user interface was all\n> implemented in shell scripts on top of that plumbing.\n\n i'm using gitdb (ok don't need that any more, if i don't walk the\npack-index file *sigh*) and python-git - am quite happy with the speed\nat which i can knock stuff together, using it.  the only tricky wobbly\nmoment i had was not being able to pass in a file-handle to stdin (git\npack-objects) and i got round that with \"input = os.tmpfile();\ninput.write(objref+\"\\n\"); input.seek(0)\".\n\n> Back to that rev-list output... OK, you want the equivalent of a fanout\n> table.  You may do something like this then:\n>\n>    git rev-list --objects v2.6.34..v2.6.35 | cut -c -2 | sort | uniq -c\n\n  ack.  got it.\n\n thanks nicolas.\n\nl.\n"},{"id":"150012","messageId":"alpine.LFD.2.00.1009051820100.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTi=YLx6MqbWd_N0geXbuXLdqAUOneGoym75dfthL@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-05T23:52:23Z","receivedAt":"2010-09-05T23:52:23Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 5 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> On Sun, Sep 5, 2010 at 3:16 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> > On Sat, 4 Sep 2010, Luke Kenneth Casson Leighton wrote:\n> >\n> >> * git-index-pack requires a pack file in order to re-create the index:\n> >> i don't want that\n> >> * git-pack-objects appears to have no way of telling it \"just gimme\n> >> index file please\"\n> >> * fast-import.c appears not to be what's needed either.\n> >>\n> >> so - any other methods for just getting the index file (exclusively?)\n> >> any other commands i've missed?  if not, are there any other ways of\n> >> getting a pack's index of objects without err... getting the index\n> >> file?  (i believe the answer to be no, but i'm just making sure) and\n> >> on that basis i believe it is safe to ask: any objections to a patch\n> >> which adds \"--index-only\" to builtin/pack-objects.c?\n> >\n> > No patch is needed.\n> >\n> > First, what you want is an index of objects you are willing to share,\n> > and not the index of whatever pack file you might have on your disk,\n> > especially if you have multiple packs which is typical.\n> \n>  blast.  so *sigh* ignoring the benefits that can be obtained by the\n> delta-compression thing, somewhat; ignoring the fact that perhaps less\n> traffic miight be transferred by happening to borrow objects from\n> another branch (which is the situation that, i believe, happens with\n> \"git pull\" over http:// or git://); ignoring the fact that i actually\n> implemented using the .idx file yesterday ... :)\n\nPlease, let's get it slow.\n\nThere are 2 concepts you really need to master in order to come up with \na solution.  And those concepts are completely independent from \neach other, but at the moment you are blending them up together and \nthat's not good.\n\nThe first one is all about object enumeration.  And object enumeration \nis all about 'git rev-list'.  This is important when offering objects to \nthe outside world that you actually do offer _all_ the needed objects, \nbut _only_ the needed objects.  If some objects are missing you get a \nbroken repository.  But more objects can also be a security problem as \nthose extra objects may contain confidential data that you never \nintended to publish.\n\nAnd object enumeration has absolutely nothing to do with packs, nor .idx \nfiles for that matter.  As I said, the objects you want might be split \nacross multiple packs, and also in loose form, and also in some \nalternate location that is shared amongst many repositories on the same \nfilesystem.  But a single pack may also contain more than what you want \nto offer, and it is extremely important that you do _not_ offer those \nobjects that are not reachable from the branch you want to publish.\n\nFollowing me so far?\n\nThe second concept is all about object _representation_ or _encoding_.  \nThat's where the deltas come into play.  So the idea is to grab the list \nof objects you want to publish, and then look into existing packs to see \nif you could find them in delta form.  So, for each object, if you do \nfind them in delta form, and the objec the delta is made against is 1) \nalso part of the list of objects you want to send, or 2) is already \navailable at the remote end, then you may simply reuse that delta data \nas is from the pack.  Finding if a particular pack has the wanted object \nis easy: you just need to look it up in the .idx file.  Then, in the \ncorresponding pack file you parse the object header to find out if it is \na delta, and what its base object is.\n\n>  ... there is a bit of a disadvantage to using pack index files that\n> it goes all the way down (if i am reading things correctly) and cannot\n> be told \"give me just the objects related to a particular commit\"....\n\nExact.  The .idx file gives you a list of objects that exists in the \ncorresponding pack.  That list of object might belong to a totally \nrandom number of random commits.  You may also have a random number of \npacks across which some or all objects are distributed.  Because, of \ncourse, not all the objects you need are always packed.\n\nSo... I hope you understand now that there is no relation between \ncommits and .idx files.  The only exception is when you do create a \ncustom pack with 'git pack-objects'.\n\n> > Try this instead:\n> >\n> >    git rev-list --objects HEAD | cut -c -40 | sort\n> >\n> > That will give you a sorted list of all objects reachable from the\n> > current branch.  With the Linux repo, you may replace \"HEAD\" with\n> > \"v2.6.34..v2.6.35\" if you wish, and that would give you the list of the\n> > new objects that were introduced between v2.6.34 and v2.6.35.\n> \n>  ... unlike this, which is in fact much more along the lines of what i\n> was looking for (minus the loveliness of the delta compression oh\n> well)\n\nAgain, delta compression is a _separate_ issue.\n\n> > This will\n> > provide you with 84642 objects instead of the 1.7 million objects that\n> > the Linux repo contains (easier when testing stuff).\n> \n>  hurrah! :)  [but, then if you actually want to go back and get alll\n> commits, that's ... well, we'll not worry about that too much, given\n> the benefits of being able to get smaller chunks.]\n\nIf you want all commits then you just need --all instead of HEAD.\n\n> > That sorted list of objects is more or less what the pack index file\n> > contains, plus an offset in the pack for each entry.  It is used to\n> > quickly find the offset for a given object in the corresponding pack\n> > file, and the fanout is only a way to cut 3 iterations in the binary\n> > search.\n> >\n> > But anyway, what you want is really to select the precise set of objects\n> > you wish to share, and not blindly using the pack index file.  If you\n> > have a public branch and a private branch in your repository, then\n> > objects from both branches may end up in the same pack\n> \n>  slightly confused: are you of the belief that i intend to ignore\n> refs/branches/* starting points?\n\nI don't know what your exact understanding of Git is, and although I \nknow one or two things about the Git storage model, I get confused \nmyself by some of your comments, such as this one above.\n\n> > and you probably\n> > don't want to publish those objects from the private branch.\n> \n>  ahh, i wondered where i'd seen the bit about \"confusing\" two\n> branches, i thought it was in another message.  so many flying back &\n> forth :)  from what i can gather, this is exactly what happens with\n> git fetch from http:// or git:// so what's the big deal about that?\n> why stop gitp2p from benefitting from the extra compression that could\n> result from \"borrowing\" bits of another branch's objects, neh?\n\nNo.  git:// will _never_ ever transfer any object that is not part of \nthe published branch(es).  If an object that does get transmitted is \nactually a delta against an object that is only part of a branch that is \nnot published, then the delta will be expanded and redone against \nanother suitable object before transmission.\n\n\nNicolas\n"},{"id":"150014","messageId":"alpine.LFD.2.00.1009051952390.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTim8XLB5SjV3JtWT-ARN_XuofKDjYRSYT8kPxEvq@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-06T00:05:26Z","receivedAt":"2010-09-06T00:05:26Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 5 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> On Sun, Sep 5, 2010 at 2:18 AM, Nicolas Pitre <nico@fluxnic.net> wrote\n> > I agree.  So far it has been an interesting topic for discussion, but in\n> > practice I doubt the actual benefits will justify the required efforts\n> > and/or constraints on the protocol. Otherwise we would have a working\n> > implementation in use already.  People tried in the past, and so far\n> > none of those attempts passed the reality test nor kept people motivated\n> > enough to work on them further.\n> \n>  then i'm all the more grateful that you continue to drop technical\n> hints in my direction.  thank you for not judging.\n\nWell, either you'll come to the same conclusion as the other people \nbefore you (myself included), or you'll surprise us all with some clever \nsolution.  But that part is up to you.  In either cases, I think it is a \ngood thing if I can help you grasp the technical limitations and issues \nfaster so you don't waste your time on false assumptions.\n\n\nNicolas\n"},{"id":"150080","messageId":"AANLkTi=CEOj40Sj+zegvX+ry8-y6p7UwsyqdtoHB1d-T@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009051820100.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-06T13:23:48Z","receivedAt":"2010-09-06T13:23:48Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Mon, Sep 6, 2010 at 12:52 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n\n>> another branch (which is the situation that, i believe, happens with\n>> \"git pull\" over http:// or git://); ignoring the fact that i actually\n>> implemented using the .idx file yesterday ... :)\n>\n> Please, let's get it slow.\n\n ack :)\n\n> There are 2 concepts you really need to master in order to come up with\n> a solution.  And those concepts are completely independent from\n> each other, but at the moment you are blending them up together and\n> that's not good.\n\n i kinda get it - but i realise that's not good enough: i need to be\nable to _say_ i get it, in a way that satisfies you.\n\n> The first one is all about object enumeration.  And object enumeration\n> is all about 'git rev-list'.  This is important when offering objects to\n> the outside world that you actually do offer _all_ the needed objects,\n> but _only_ the needed objects.  If some objects are missing you get a\n> broken repository.  But more objects can also be a security problem as\n> those extra objects may contain confidential data that you never\n> intended to publish.\n\n ack.\n\n> And object enumeration has absolutely nothing to do with packs, nor .idx\n> files for that matter.\n\n mmm packs not being to do with object enumeration i get.  i\nunderstand that .idx files contain \"lists of objects\" which isn't the\nsame thing (and also happen to contain pointers/offsets to the objects\nof its associated .pack)\n\n at some point i'd really like to know what the object list is (not\nthe objects themselves) that comes out of \"git pack-objects --thin\"\nbut my curiosity can wait.\n\n> As I said, the objects you want might be split\n> across multiple packs, and also in loose form, and also in some\n> alternate location that is shared amongst many repositories on the same\n> filesystem.\n\n ok - this tells me (and it's confirmed, below) that you're describing\nthe situation based on what can be found in .git - _not_ what comes\nout of \"git pack-objects\".  i wouldn't _dream_ of digging around in a\n.git/ location looking for packs or idx files, but because i have\nmentioned them _without_ prefixing every mention with \"the\ncustom-generated .idx and/or .pack as generated by git pack-objects\",\nyou may have got the wrong impression, for which i apologise.\n\n>  But a single pack may also contain more than what you want\n> to offer, and it is extremely important that you do _not_ offer those\n> objects that are not reachable from the branch you want to publish.\n>\n> Following me so far?\n\n yep :)\n\n> The second concept is all about object _representation_ or _encoding_.\n> That's where the deltas come into play.  So the idea is to grab the list\n> of objects you want to publish, and then look into existing packs to see\n> if you could find them in delta form.  So, for each object, if you do\n> find them in delta form, and the objec the delta is made against is 1)\n> also part of the list of objects you want to send, or 2) is already\n> available at the remote end, then you may simply reuse that delta data\n> as is from the pack.  Finding if a particular pack has the wanted object\n> is easy: you just need to look it up in the .idx file.  Then, in the\n> corresponding pack file you parse the object header to find out if it is\n> a delta, and what its base object is.\n\n ok.  all of this makes sense - but it's enough for me to be able to\nask questions, rather than \"do\", if you know what i mean.\n\n>>  ... there is a bit of a disadvantage to using pack index files that\n>> it goes all the way down (if i am reading things correctly) and cannot\n>> be told \"give me just the objects related to a particular commit\"....\n>\n> Exact.  The .idx file gives you a list of objects that exists in the\n> corresponding pack.  That list of object might belong to a totally\n> random number of random commits.  You may also have a random number of\n> packs across which some or all objects are distributed.  Because, of\n> course, not all the objects you need are always packed.\n>\n> So... I hope you understand now that there is no relation between\n> commits and .idx files.  The only exception is when you do create a\n> custom pack with 'git pack-objects'.\n\n yes.  ahh... that's what i've been doing: using \"git pack-objects\n--thin\".  and the reason for that is because i've seen it used in the\nhttp implementation of \"git fetch\".\n\n so, my questions up until now regarding .pack and .idx have all been\ntargetted at that, and based on that context, _not_ the packs+idx\nfiles that are in .git/\n\n>> > Try this instead:\n>> >\n>> >    git rev-list --objects HEAD | cut -c -40 | sort\n>> >\n>> > That will give you a sorted list of all objects reachable from the\n>> > current branch.  With the Linux repo, you may replace \"HEAD\" with\n>> > \"v2.6.34..v2.6.35\" if you wish, and that would give you the list of the\n>> > new objects that were introduced between v2.6.34 and v2.6.35.\n>>\n>>  ... unlike this, which is in fact much more along the lines of what i\n>> was looking for (minus the loveliness of the delta compression oh\n>> well)\n>\n> Again, delta compression is a _separate_ issue.\n>\n>> > This will\n>> > provide you with 84642 objects instead of the 1.7 million objects that\n>> > the Linux repo contains (easier when testing stuff).\n>>\n>>  hurrah! :)  [but, then if you actually want to go back and get alll\n>> commits, that's ... well, we'll not worry about that too much, given\n>> the benefits of being able to get smaller chunks.]\n>\n> If you want all commits then you just need --all instead of HEAD.\n\n no, i want commits separated and individual and \"compoundable\".  the plan is:\n\n* to get the ref associated with refs/heads/master\n* to get the list of all commits associated with that master ref\n* to work out how far local deviates from remote along that list of commits\n* to get the objects which will make up the missing commits (if they\naren't already in the local store)\n* to apply those commits in the correct order\n\nin other words, the plan is to follow what git http fetch and/org git\ngit:// fetch does as much as possible (ok, perhaps not).\n\nthe reason for getting the objects individually (blobs etc.) should be\nclear: prior commits _could_ have resulted in that exact object having\nbeen obtained already.\n\nso far i have implemented:\n\n* get the master ref using git for-each-ref\n* get the list of all commits using git rev-list\n* enumerate the list of objects associated with an individual commit by:\n    i) creating a CUSTOM pack+idx using git pack-objects {ref}\n    ii) *parsing* the idx file using gitdb's FileIndex to get the list\nof objects\n    iii) transferring that list to the local machine\n* requesting *individual* objects from the enumerated list out of the idx file\n   by using a CUSTOM \"git pack-objects --thin {ref} < {ref}\" command\n\nthat's as far as i've got, before you mentioned that it would be\nbetter to use \"git rev-list --objects commit1..commit2\" and to use\n\"git cat-file\" to obtain the actual object [what's not clear in this\nplan is how to store that cat'ed file at the local end, hence the\ncontinued use of git pack-objects --thin {ref} < {ref}]\n\nthe prior implementation was to treat the custom pack-object as if it\nwas \"the atomic leaf-node operation\" instead of individual objects\n(blobs, trees).\n\n>> > That sorted list of objects is more or less what the pack index file\n>> > contains, plus an offset in the pack for each entry.  It is used to\n>> > quickly find the offset for a given object in the corresponding pack\n>> > file, and the fanout is only a way to cut 3 iterations in the binary\n>> > search.\n>> >\n>> > But anyway, what you want is really to select the precise set of objects\n>> > you wish to share, and not blindly using the pack index file.  If you\n>> > have a public branch and a private branch in your repository, then\n>> > objects from both branches may end up in the same pack\n>>\n>>  slightly confused: are you of the belief that i intend to ignore\n>> refs/branches/* starting points?\n>\n> I don't know what your exact understanding of Git is, and although I\n> know one or two things about the Git storage model, I get confused\n> myself by some of your comments, such as this one above.\n\n soorree.  i believe the source of the confusion is that you believed\nthat i intend to \"blindly use a pack index file\" as in \"blindly go\nrummaging around in .git/ at the remote end\" when i have absolutely no\nintention of doing so.\n\n what i _have_ been doing however is custom-generating pack-objects\nand associated pack-indexes (just like git http fetch) _including_\nusing the --thin option because that's what git http fetch does.\n\n i believe that this results in the concerns that you raised (about\nhaving access to unauthorised data) being dealt with.\n\n>> > don't want to publish those objects from the private branch.\n>>\n>>  ahh, i wondered where i'd seen the bit about \"confusing\" two\n>> branches, i thought it was in another message.  so many flying back &\n>> forth :)  from what i can gather, this is exactly what happens with\n>> git fetch from http:// or git:// so what's the big deal about that?\n>> why stop gitp2p from benefitting from the extra compression that could\n>> result from \"borrowing\" bits of another branch's objects, neh?\n>\n> No.  git:// will _never_ ever transfer any object that is not part of\n> the published branch(es).\n\n ... because it uses, from what i can gather, git pack-objects --thin\n\n> If an object that does get transmitted is\n> actually a delta against an object that is only part of a branch that is\n> not published, then the delta will be expanded and redone against\n> another suitable object before transmission.\n\n and that's handled by git pack-objects --thin (am i right?)\n\n ok.\n\n so.  we have a hierarchical plan: get the commit list, get a\nper-commit object-list, get the objects (if needed), store the\nobjects.\n\n problem: despite looking through virtually every single builtin/*.c\nfile which uses write_sha1_file (which i believe i have correctly\nidentified, from examining git unpack-objects, as being the function\nwhich stores actual objects, including their type), i do not see a git\ncommand (yet) which performs the reverse operation of \"git cat-file\".\n\nbuiltin/apply.c - that's for patches\nbuiltin/checkout.c - that's for the merge result.\nbuiltin/notes.c - creating a note\nbuiltin/tag.c - creating a tag\nbuiltin/mktree.c - creating a tree object but *only* from a text listing\n\nok - maybe this is one of the ones that i need, but only if i use \"git\ncat-file -p\" to pretty-print the output of tree objects but i don't\nthink that's a good idea.\n\nwhat else...\n\ncache_tree.c - nope.\ncommit.c - nope.\nread-cache.c - beh? nope.  has blank args \"\", 0\n\nso... um... unless i actually manually create a pack object (perhaps\nusing python-gitdb to construct it) out of the data obtained by \"git\ncat-file\" i don't see how this would work.\n\nl.\n"},{"id":"150094","messageId":"alpine.LFD.2.00.1009061025210.19366@xanadu.home","threadId":"24935","inReplyTo":"AANLkTi=CEOj40Sj+zegvX+ry8-y6p7UwsyqdtoHB1d-T@mail.gmail.com","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-06T16:51:47Z","receivedAt":"2010-09-06T16:51:47Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Sep 2010, Luke Kenneth Casson Leighton wrote:\n\n> On Mon, Sep 6, 2010 at 12:52 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> \n> > And object enumeration has absolutely nothing to do with packs, nor .idx\n> > files for that matter.\n> \n>  mmm packs not being to do with object enumeration i get.  i\n> understand that .idx files contain \"lists of objects\" which isn't the\n> same thing (and also happen to contain pointers/offsets to the objects\n> of its associated .pack)\n> \n>  at some point i'd really like to know what the object list is (not\n> the objects themselves) that comes out of \"git pack-objects --thin\"\n\nYou need to feed 'git pack-objects' a list of objects in the first \nplace for it to pack anything.  So you must have that list even before \npack-objects can produce any output.  And that list is usually generated \nby 'git rev-list'.  So, typically, you'd do:\n\n\tgit rev-list --objects <commit_range> | git pack-objects foo\n\nBut these days the ability to enumerate objects was integrated into \npack-objects directly, so you can do:\n\n\techo \"<commit_range>\" | git pack-objects --revs foo\n\nBut you should get the idea.\n\n> > So... I hope you understand now that there is no relation between\n> > commits and .idx files.  The only exception is when you do create a\n> > custom pack with 'git pack-objects'.\n> \n>  yes.  ahh... that's what i've been doing: using \"git pack-objects\n> --thin\".  and the reason for that is because i've seen it used in the\n> http implementation of \"git fetch\".\n\nWell, the HTTP implementation is a rather tricky example as there are \nactually two implementations: one that is dumb and only slurps packs and \nloose objects out of a remote .git/ directory, and another that is smart \nenough to carry the smarter Git protocol across HTTP requests.  \n\nWhen using the \"smart\" Git protocol, the client tells the server what it \nalready has, and then the server uses pack-objects to produce a pack \nwith only those objects that the client doesn't have, and stream that \npack directly without even storing it on disk.  The server doesn't even \nproduce a .idx file in that case.  It is up to the client to store the \npack on disk, and feed it through 'git index-pack' to construct a .idx \nfile for it locally.\n\n>  so, my questions up until now regarding .pack and .idx have all been\n> targetted at that, and based on that context, _not_ the packs+idx\n> files that are in .git/\n\nTell me if the above clears them up.\n\n> > If you want all commits then you just need --all instead of HEAD.\n> \n>  no, i want commits separated and individual and \"compoundable\".  the plan is:\n> \n> * to get the ref associated with refs/heads/master\n\nYou can do:\n\n\tgit rev-parse refs/heads/master\n\n> * to get the list of all commits associated with that master ref\n\nJust use (without the --objects argument):\n\n\tgit rev-list refs/heads/master\n\n> * to work out how far local deviates from remote along that list of commits\n\nThat's an operation that only the peer with the most recent commits can \ndo, unless you transfer that huge list of commits from above across the \nnetwork.  So, on a server (i.e. the peer sending objects) you'd do:\n\n\tgit rev-list <refs_that_I_publish> --not <refs_that_the_remote_has>\n\n> * to get the objects which will make up the missing commits (if they\n> aren't already in the local store)\n\nAgain, that's a task for the peer with objects to offer.  It just has to \nuse the above rev-list invocation and add the --objects argument to it \n(or feed the equivalent ref specifications to pack-objects directly as \nshown previously).\n\n> * to apply those commits in the correct order\n\nWhy would you care about this?  There is nothing to \"apply\" as all you \nhave to do is simply transfer objects.\n\n> in other words, the plan is to follow what git http fetch and/org git\n> git:// fetch does as much as possible (ok, perhaps not).\n\nWell... I don't think it would be easy to do the same in a P2P context.  \nThose fetch operations are totally stream oriented between 2 peers, and \nnot many to many.\n\n> the reason for getting the objects individually (blobs etc.) should be\n> clear: prior commits _could_ have resulted in that exact object having\n> been obtained already.\n\nSure.  But objects known to exist on the remote side won't be listed by \nrev-list.\n\n> so far i have implemented:\n> \n> * get the master ref using git for-each-ref\n> * get the list of all commits using git rev-list\n\nSo far so good.\n\n> * enumerate the list of objects associated with an individual commit by:\n>     i) creating a CUSTOM pack+idx using git pack-objects {ref}\n>     ii) *parsing* the idx file using gitdb's FileIndex to get the list\n> of objects\n\nThat's where you're going so much out of your way to give you trouble.  \nA simple rev-list would give you that list:\n\n\tgit rev-list --objects <this_commit> --not <this_commit''s_parents>\n\nThat's it.\n\n>     iii) transferring that list to the local machine\n> * requesting *individual* objects from the enumerated list out of the idx file\n>    by using a CUSTOM \"git pack-objects --thin {ref} < {ref}\" command\n\nThat's where you'll have to get your hands real dirty and write actual \ncode to serve individual objects but not through cat-file.  In a P2P \nsetup you'd want to transfer as little amount of data as possible, \nmeaning that you'd want to serve deltas as much as possible.  It's then \na matter of finding if the requested object exists already in delta \nform, if so then whether or not its base object is something that the \nother end has in which case you send that as is, otherwise figuring out \nif that would be worth creating a delta against another object known to \nexist at the other end.\n\nOn the receiving end, you'd simply have to store those objects in \nthe .git/objects/ directories as loose objects after expanding the \ndeltas, or even stuff everything \ninto a pack and run 'git index-pack' on it when the transfer is \ncomplete (and run 'git repack' to optimize the pack eventually).\n\nOnce the transfer is complete, you do a  reachability and validity check \non the refs you are supposed to have received the objects for, and if \neverything is OK then you update the refs and you're done.\n\n> that's as far as i've got, before you mentioned that it would be\n> better to use \"git rev-list --objects commit1..commit2\" and to use\n> \"git cat-file\" to obtain the actual object [what's not clear in this\n> plan is how to store that cat'ed file at the local end, hence the\n> continued use of git pack-objects --thin {ref} < {ref}]\n> \n> the prior implementation was to treat the custom pack-object as if it\n> was \"the atomic leaf-node operation\" instead of individual objects\n> (blobs, trees).\n\nWell, OK.  But suppose that you have only 2 new commits with a big \namount of objects for each.  Typically the very first commit of a \nproject corresponds to the import of that project into Git, and it is \nequivalent to the whole work tree.  Don't you want to spread the request \nfor those objects across as many peers as possible?\n\n>  what i _have_ been doing however is custom-generating pack-objects\n> and associated pack-indexes (just like git http fetch) _including_\n> using the --thin option because that's what git http fetch does.\n\nWell, let's get back to that HTTP fetch which has a double personality.  \nThe \"smart\" HTTP fetch doesn't involve any pack index at all.  It ends \nup streaming pack-objects stdout's output over the net and the pack \nindex is recreated on the other end.\n\nHowever, what the _dumb_ HTTP fetch does (and that is the same idea for \nthe FTP fetch, or the rsync fetch) is to dig into the remote's .git \ndirectory and grab those .idx files, look into them to see if the \ncorresponding .pack file actually contain the wanted objects, so to only \ndownloads the needed packs afterwards.  And those dumb protocols are \nwhat they are: dumb.  They usually end up transferring way more data \nthan actually necessary.\n\n> > If an object that does get transmitted is\n> > actually a delta against an object that is only part of a branch that is\n> > not published, then the delta will be expanded and redone against\n> > another suitable object before transmission.\n> \n>  and that's handled by git pack-objects --thin (am i right?)\n\nRight.  But the --thin flag here is unrelated to this.\n\nWhat --thin does is to tell pack-objects that it can produce deltas \nagainst objects that will _not_ be included in the produced pack.  That \nis OK only if the consumer of that pack is 1) aware of that fact and\n2) is going to \"fix\" the pack by appending those objects to the pack \nfrom a local copy.  For a pack to be \"valid\" in your .git directory, it \nhas to be self contained with regards to deltas.  It is not allowed to \nhave deltas across different packs as this makes the issue of delta \nloops extremely difficult to deal with in the context of incremental \nrepacks.\n\n>  so.  we have a hierarchical plan: get the commit list, get a\n> per-commit object-list, get the objects (if needed), store the\n> objects.\n> \n>  problem: despite looking through virtually every single builtin/*.c\n> file which uses write_sha1_file (which i believe i have correctly\n> identified, from examining git unpack-objects, as being the function\n> which stores actual objects, including their type), i do not see a git\n> command (yet) which performs the reverse operation of \"git cat-file\".\n\nIt is 'git hash-object'.\n\n\nNicolas\n"},{"id":"150148","messageId":"AANLkTimxyWOd3MUnbXZS0ZdqEXb8oRCUwHDNtavbCpgJ@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009061025210.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-06T22:33:57Z","receivedAt":"2010-09-06T22:33:57Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"nicolas, thank you very brief reply (busy for 2 days)\n\nOn Mon, Sep 6, 2010 at 5:51 PM, Nicolas Pitre <nico@fluxnic.net> wrote:\n\n>> * to work out how far local deviates from remote along that list of commits\n>\n> That's an operation that only the peer with the most recent commits can\n> do, unless you transfer that huge list of commits from above across the\n> network.\n\n i Have A Plan for dealing with that.\n\n> So, on a server (i.e. the peer sending objects) you'd do:\n>\n>        git rev-list <refs_that_I_publish> --not <refs_that_the_remote_has>\n\n sadly that involves telling the sender what the recipient has.\n\n>>  problem: despite looking through virtually every single builtin/*.c\n>> file which uses write_sha1_file (which i believe i have correctly\n>> identified, from examining git unpack-objects, as being the function\n>> which stores actual objects, including their type), i do not see a git\n>> command (yet) which performs the reverse operation of \"git cat-file\".\n>\n> It is 'git hash-object'.\n\n ah _haa_ - thank you!  ok, so i have to create a pack-object-like\nformat, putting the object type at the beginning of the format, then\nput the contents of \"git cat-file\" after it.\n\n so - apologies, will be dealing with some work-related stuff for a\nday or so.  thank you for everything so far nicolas.\n\nl.\n"},{"id":"150151","messageId":"7v8w3etpjr.fsf@alter.siamese.dyndns.org","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009061025210.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-09-06T23:34:00Z","receivedAt":"2010-09-06T23:34:00Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@fluxnic.net> writes:\n\n>> * enumerate the list of objects associated with an individual commit by:\n>>     i) creating a CUSTOM pack+idx using git pack-objects {ref}\n>>     ii) *parsing* the idx file using gitdb's FileIndex to get the list\n>> of objects\n>\n> That's where you're going so much out of your way to give you trouble.  \n> A simple rev-list would give you that list:\n>\n> \tgit rev-list --objects <this_commit> --not <this_commit''s_parents>\n>\n> That's it.\n\nI didn't want to get into this discussion, but where in the above picture\ndoes the usual \"want/ack\" exchange fit?\n\nThe biggest trouble before object transfer actually happens is that the\nsending end needs to find a set of commits that are known to exist at the\nreceiving end, but it needs to do that starting from a state where the tip\ncommits the receiving end has are not known by it.  That is why the\nreceiver must go back in his history and keep asking the sender \"I have\nthis, this, this, this...; now have you heard enough?\"  until the sender\nsees a commit that it knows about.  After that happens, it can list them\non its \"git rev-list --objects <my tips> --not <he has these>\" command\nline in order to enumerate objects it needs to send.\n\nIf the receiver is purely following the sender, not doing any work on its\nown, the tips the receiver has may always be known by the sender, but that\nis not an interesting case at all.\n"},{"id":"150158","messageId":"alpine.LFD.2.00.1009061942150.19366@xanadu.home","threadId":"24935","inReplyTo":"7v8w3etpjr.fsf@alter.siamese.dyndns.org","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-09-06T23:57:43Z","receivedAt":"2010-09-06T23:57:43Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 6 Sep 2010, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@fluxnic.net> writes:\n> \n> >> * enumerate the list of objects associated with an individual commit by:\n> >>     i) creating a CUSTOM pack+idx using git pack-objects {ref}\n> >>     ii) *parsing* the idx file using gitdb's FileIndex to get the list\n> >> of objects\n> >\n> > That's where you're going so much out of your way to give you trouble.  \n> > A simple rev-list would give you that list:\n> >\n> > \tgit rev-list --objects <this_commit> --not <this_commit''s_parents>\n> >\n> > That's it.\n> \n> I didn't want to get into this discussion, but where in the above picture\n> does the usual \"want/ack\" exchange fit?\n\nBefore object enumeration obviously.  But I think that Luke has enough \nto play with already by only assuming the easy case for now.  If Git P2P \nis to be viable, it has to prove itself at least with the easy case \nfirst.\n\n\nNicolas\n"},{"id":"150162","messageId":"AANLkTimq50_suDMu67PSDE0LGJDCLd6TMhJOdSGot+dK@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009061942150.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-07T00:17:24Z","receivedAt":"2010-09-07T00:17:24Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Tue, Sep 7, 2010 at 12:57 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n> On Mon, 6 Sep 2010, Junio C Hamano wrote:\n>\n>> Nicolas Pitre <nico@fluxnic.net> writes:\n>>\n>> >> * enumerate the list of objects associated with an individual commit by:\n>> >>     i) creating a CUSTOM pack+idx using git pack-objects {ref}\n>> >>     ii) *parsing* the idx file using gitdb's FileIndex to get the list\n>> >> of objects\n>> >\n>> > That's where you're going so much out of your way to give you trouble.\n>> > A simple rev-list would give you that list:\n>> >\n>> >     git rev-list --objects <this_commit> --not <this_commit''s_parents>\n>> >\n>> > That's it.\n>>\n>> I didn't want to get into this discussion, but where in the above picture\n>> does the usual \"want/ack\" exchange fit?\n>\n> Before object enumeration obviously.  But I think that Luke has enough\n> to play with already by only assuming the easy case for now.  If Git P2P\n> is to be viable, it has to prove itself at least with the easy case\n> first.\n\n :)   yes.  worry about that later.  optimisation.  time to think.\nidea earlier (from 2 hours ago) unworkable, thought of another one,\nsplit commit list into multi-level \"virtual hierarchical\nsubdirectories\" of say 256 entries each.  can therefore easily trip\ndown each \"subdirectory\" which will quickly get you to the right place\nwhere the commits are different, with only a few roundtrips.  sort-of\nbinary search but 256-way search.  binary search not optimal here\nbecause of multiple network round-trips.  sorry very obtuse will write\nup better.\n\nl.\n"},{"id":"150163","messageId":"AANLkTimy9vYDACKBZ7JuBosukLdZx6-bRwPqVLEJBRcj@mail.gmail.com","threadId":"24935","inReplyTo":"alpine.LFD.2.00.1009061942150.19366@xanadu.home","subject":"Re: git pack/unpack over bittorrent - works!","fromName":"Luke Kenneth Casson Leighton","fromEmail":"luke.leighton@gmail.com","sentAt":"2010-09-07T00:29:25Z","receivedAt":"2010-09-07T00:29:25Z","isPatch":false,"sender":{"key":"luke.leighton@gmail.com","avatar":null},"body":"On Tue, Sep 7, 2010 at 12:57 AM, Nicolas Pitre <nico@fluxnic.net> wrote:\n>  But I think that Luke has enough\n> to play with already by only assuming the easy case for now.\n\n um, yes.\n\n i only just noticed that bittorrent's multi-file mode has blocking\nthat doesn't line up with the bloody file beginnings and ends:\n\n| block 0 256k | block 1 256k | block N 11bytes |\n|file1 | file2          | file3 | file4               |\n\nthis is why cameron dale threw his hands up in horror at bittorrent\nwhen he created debtorrent, but i didn't understand why, fully, as i\nthought that he was creating one .torrent per .deb _anyway_ so if he\nhad it would have been moot.\n\nso i have to do some tests to see if multiple individual torrents can\nbe added using the BitTornado API to the same server, and redesign the\ndamn code so that it adds new \"files\" to be downloaded based on the\nprevious one completing.\n\nso, the first quotes file quotes to be requested will be the\n\"rev-list\", and the \"finished\" callback function will result in the\nnext layer of \"files\" as .torrents to be requested, and then finally\nthe files representing the actual \"objects\" - generated as git\ncat-files on the remote end - get added and requested.\n\nso - all event-driven.  eek!  fuun.  about as understandable as nmbd\nin samba (*) :)\n\nl.\n\n(*) i can say that because i did its 1st major rewrite ha ha\n"}]}