{"thread":{"id":"1062","subject":"\"git-send-pack\"","startedAt":"2005-06-30T17:54:48Z","lastAt":"2005-07-07T03:31:18Z","messageCount":86,"participants":["Linus Torvalds","A Large Angry SCM","Jan Harkes","Mike Taht","Daniel Barkalow","H. Peter Anvin","Junio C Hamano","Dan Holmsand","Matthias Urlichs","Eric W. Biederman","Petr Baudis","Tony Luck","Kevin Smith"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"5466","messageId":"Pine.LNX.4.58.0506301025510.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":null,"subject":"\"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T17:54:48Z","receivedAt":"2005-06-30T17:54:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nOk,\n I'm happy to say that the first cut of my new packed-object-sending thing \nseems to work. I have successfully sent updates both locally and over ssh, \nand it seems to work fine, although it has some limitations.\n\nThe syntax is very simple indeed:\n\n\tgit-send-pack destination\n\nwill go to the destination (which can be either a local directory or a\nremote ssh one, with the remote destination format currently being _only_\nthe \"machine:path\" format), and it will go through all the refs in the \nremote destination, compare them with the local ones, and create a pack \nthat updates from one to the other.\n\nIf the pack/unpack sequence is successful, it then updates the refs at the \nother end, and is done.\n\nMy quick tests were very successful, in the sense that it even performed\nreally well. But I only tested some small updates.\n\nAnyway, what are the limitations? Here's a few obvious ones:\n\n - the code actually contains support for limiting the refs to be updated\n   on the remote end, but I don't actually pass the arguments to the \n   remote git-receive-pack binary yet, so this is currently not \n   functional. Call me lazy.\n\n - the thing currently refuses to create new refs. Again, this is mainly \n   just me being lazy: it should be easy to add support for creating a new \n   branch, it just requires some care to make sure that we take the old \n   branches into account when generating the pack-file so that we don't \n   send too many objects over. \n\n - I really hate how \"ssh\" apparently cannot be told to have alternate \n   paths. For example, on master.kernel.org, I don't control the setup, so \n   I can't install my own git binaries anywhere except in my ~/bin\n   directory, but I also cannot get ssh to accept that that is a valid \n   path. This one really bums me out, and I think it's an ssh deficiency. \n\n   You apparently have to compile in the paths at compile-time into sshd, \n   and PermitUserEnvironment is disabled by default (not that it even \n   seems to work for the PATH environment, but that may have been my \n   testing that didn't re-start sshd).\n\n   That just sucks.\n\n - It doesn't update the working directory at the other end. This is fine \n   for what it's intended for (pushing to a central \"raw\" git archives), \n   so this could be considered a feature, but it's worth pointing out. \n   Only a \"pull\" will update your working directory, and this pack sending \n   really is meant to be used in a kind of \"push to central archive\" way.\n\n - this is also (at least once we've tested it a lot more and added the\n   code to allow it to create new refs on the remote side) meant to be a\n   good way to mirror things out, since clearly rsync isn't scaling. \n\n   However, I don't know what the rules for acceptable mirroring \n   approaches are, and it's entirely possible (nay, probable) that an ssh\n   connection from the \"master\" ain't it. It would be good to know what \n   (of any) would be acceptable solutions..\n\nAnyway, please do give it a test. I think I'll use this to sync up to\nkernel.org, except I _really_ would want to solve that ssh issue some \nother way than hardcoding the /home/torvalds/bin/ path in my local \ncopies.. If somebody knows a good solution, pls holler.\n\n\t\tLinus\n"},{"id":"5468","messageId":"42C438CA.3040507@gmail.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301025510.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2005-06-30T18:24:10Z","receivedAt":"2005-06-30T18:24:10Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Have you tried something like the following?\n\nssh torvalds@master.kernel.org \\\n\t'/bin/sh -c \"export PATH=/tmp/foo:$PATH ; env\"'\n\nLinus Torvalds wrote:\n> \n...\n >\n> Anyway, please do give it a test. I think I'll use this to sync up to\n> kernel.org, except I _really_ would want to solve that ssh issue some \n> other way than hardcoding the /home/torvalds/bin/ path in my local \n> copies.. If somebody knows a good solution, pls holler.\n"},{"id":"5469","messageId":"42C439AB.30002@gmail.com","threadId":"1062","inReplyTo":"42C438CA.3040507@gmail.com","subject":"Re: \"git-send-pack\"","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2005-06-30T18:27:55Z","receivedAt":"2005-06-30T18:27:55Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Damn! That should have been:\n\nssh torvalds@master.kernel.org \\\n\t'/bin/sh -c \"export PATH=~/tmp/foo:$PATH ; env\"'\n\nA Large Angry SCM wrote:\n> Have you tried something like the following?\n> \n> ssh torvalds@master.kernel.org \\\n>     '/bin/sh -c \"export PATH=/tmp/foo:$PATH ; env\"'\n> \n> Linus Torvalds wrote:\n>>\n> ...\n>  >\n>> Anyway, please do give it a test. I think I'll use this to sync up to\n>> kernel.org, except I _really_ would want to solve that ssh issue some \n>> other way than hardcoding the /home/torvalds/bin/ path in my local \n>> copies.. If somebody knows a good solution, pls holler.\n> \n"},{"id":"5470","messageId":"20050630184517.GB28841@delft.aura.cs.cmu.edu","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301025510.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Jan Harkes","fromEmail":"jaharkes@cs.cmu.edu","sentAt":"2005-06-30T18:45:17Z","receivedAt":"2005-06-30T18:45:17Z","isPatch":false,"sender":{"key":"jaharkes@cs.cmu.edu","avatar":"https://gravatar.com/avatar/cf95aecd150ca8ef33d6edc337ac4bb9e13aa4246fc3679257d578c7fddc1633?d=mp&s=160"},"body":"On Thu, Jun 30, 2005 at 10:54:48AM -0700, Linus Torvalds wrote:\n> Anyway, please do give it a test. I think I'll use this to sync up to\n> kernel.org, except I _really_ would want to solve that ssh issue some \n> other way than hardcoding the /home/torvalds/bin/ path in my local \n> copies.. If somebody knows a good solution, pls holler.\n\nI've got a couple of 'export FOO=bar' lines in ~/.bashrc on the\n\"remote-side\" and it looks like they are set correctly when\nI do something like \"ssh remote.host env\".\n\nJan\n"},{"id":"5473","messageId":"42C4419C.7000309@timesys.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301025510.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-06-30T19:01:48Z","receivedAt":"2005-06-30T19:01:48Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"\n>    However, I don't know what the rules for acceptable mirroring \n>    approaches are, and it's entirely possible (nay, probable) that an ssh\n>    connection from the \"master\" ain't it. It would be good to know what \n>    (of any) would be acceptable solutions..\n\nFlute, perhaps\n\nhttp://www.atm.tut.fi/mad/\n\nor fcast\n\nhttp://www.inrialpes.fr/planete/people/roca/mcl/mcl.html\n\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"5472","messageId":"Pine.LNX.4.58.0506301159120.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C438CA.3040507@gmail.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T19:04:51Z","receivedAt":"2005-06-30T19:04:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, A Large Angry SCM wrote:\n>\n> Have you tried something like the following?\n> \n> ssh torvalds@master.kernel.org \\\n> \t'/bin/sh -c \"export PATH=/tmp/foo:$PATH ; env\"'\n\nThe point is that the user does not call \"ssh\" itself, but git-send-pack \ndoes it automatically.\n\nAnd that means that git-send-pack will always do the same thing, for any\nhost it is given. If one host needs a special PATH, that's an effing pain.\n\nHowever, Kees Cook points out that it's driver error: I set up my PATH in\n.bash_profile, and if I just do it in .bashrc instead it all works.\n\nDanke,\n\n\t\tLinus\n"},{"id":"5475","messageId":"Pine.LNX.4.58.0506301233270.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C4419C.7000309@timesys.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T19:42:28Z","receivedAt":"2005-06-30T19:42:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, Mike Taht wrote:\n> \n> >    However, I don't know what the rules for acceptable mirroring \n> >    approaches are, and it's entirely possible (nay, probable) that an ssh\n> >    connection from the \"master\" ain't it. It would be good to know what \n> >    (of any) would be acceptable solutions..\n> \n> Flute, perhaps\n> \n> http://www.atm.tut.fi/mad/\n\nWell, I was hoping for something that has git knowledge, since there are \nissues like updating objects in the right order. \n\nSo \"git-send-pack\" is nice in many ways: it allows you to update any \nnumber of branches (in particular, it allows you to update just a _subset_ \nof the branches, which is nice if you have a shared central repository, \nand some people have write permissions to some branches but not to \nothers), but it also allows for efficient unpacking on the receiver side \nin a way no \"general-purpose\" mirror program can really match.\n\nHowever, that requires the receiver to run a git-aware unpacker (in this\ncase git-receive-pack). I'm hoping that would be acceptable, I'm just\nwondering what kind of safety concerns I'd need to make sure of in order\nto make people comfortable running a special receiver program.\n\nSo the current approach is very flexible: if the pusher has ssh access, he\ncan do it. Safe, secure, and no new security issues. And since the only\nprograms the receiver has to be able to run is two git programs\n(git-receive-pack will run git-unpack-objects), maybe it would be ok to\neven have \"git-receive-pack\" as the shell for the receiver side, so that\nyou don't actually give the mirrorer any shell access at all. But it's\nstill \"push-based\" in the sense that it's kernel.org that is doing the\npushing, and that may simply not be acceptable.\n\n\t\t\tLinus\n"},{"id":"5476","messageId":"Pine.LNX.4.58.0506301242470.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301025510.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T19:44:53Z","receivedAt":"2005-06-30T19:44:53Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, Linus Torvalds wrote:\n> \n> Anyway, please do give it a test. I think I'll use this to sync up to\n> kernel.org\n\nIn fact, the most recent push was gone with a\n\n\tgit-send-pack master.kernel.org:/pub/scm/linux/kernel/git/torvalds/git.git\n\nso if the new commit (\"Do ref matching on the sender side rather than on \nreceiver\") shows up after the mirrors have caught up, then this thing is \nofficially in production use..\n\n\t\tLinus\n"},{"id":"5477","messageId":"Pine.LNX.4.21.0506301403300.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301025510.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-30T19:49:18Z","receivedAt":"2005-06-30T19:49:18Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 30 Jun 2005, Linus Torvalds wrote:\n\n> Anyway, what are the limitations? Here's a few obvious ones:\n> \n>  - I really hate how \"ssh\" apparently cannot be told to have alternate \n>    paths. For example, on master.kernel.org, I don't control the setup, so \n>    I can't install my own git binaries anywhere except in my ~/bin\n>    directory, but I also cannot get ssh to accept that that is a valid \n>    path. This one really bums me out, and I think it's an ssh deficiency. \n> \n>    You apparently have to compile in the paths at compile-time into sshd, \n>    and PermitUserEnvironment is disabled by default (not that it even \n>    seems to work for the PATH environment, but that may have been my \n>    testing that didn't re-start sshd).\n> \n>    That just sucks.\n\nThe easiest thing might be to have a centrally-installed wrapper script\nthat could run programs installed in your home directory. E.g., if\n\"git\" had a \"source ~/.git-env\" at the beginning, and your ~/.git-env\nfixed your PATH, then \"git receive-pack ARGS\" should work, for a generic\ncentrally installed git and special stuff in your home directory.\n\n>  - It doesn't update the working directory at the other end. This is fine \n>    for what it's intended for (pushing to a central \"raw\" git archives), \n>    so this could be considered a feature, but it's worth pointing out. \n>    Only a \"pull\" will update your working directory, and this pack sending \n>    really is meant to be used in a kind of \"push to central archive\" way.\n\nI thought only \"resolve\" (as part of \"fetch\") updated your working\ndirectory, so this is completely consistant.\n\n>  - this is also (at least once we've tested it a lot more and added the\n>    code to allow it to create new refs on the remote side) meant to be a\n>    good way to mirror things out, since clearly rsync isn't scaling. \n> \n>    However, I don't know what the rules for acceptable mirroring \n>    approaches are, and it's entirely possible (nay, probable) that an ssh\n>    connection from the \"master\" ain't it. It would be good to know what \n>    (of any) would be acceptable solutions..\n\nThe right solution probably involves getting each pack file you push to\nthe mirrors as well as to the master. They'll probably update no less\nfrequently than you push, and they should go through a series of states\nwhich matches the master, so it's not necessary to have anything smart on\nmaster sending them, and they only have to unpack the files they get (and\nupdate the refs afterward). That should make the cross-system trust\nrequirements relatively minimal; the mirror can fetch things from master,\nand neither side has to allow the other to specify a command line.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5478","messageId":"Pine.LNX.4.58.0506301302410.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"Pine.LNX.4.21.0506301403300.30848-100000@iabervon.org","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T20:12:08Z","receivedAt":"2005-06-30T20:12:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, Daniel Barkalow wrote:\n> \n> The right solution probably involves getting each pack file you push to\n> the mirrors as well as to the master. They'll probably update no less\n> frequently than you push, and they should go through a series of states\n> which matches the master, so it's not necessary to have anything smart on\n> master sending them, and they only have to unpack the files they get (and\n> update the refs afterward).\n\nHmm, yes. That would work, together with just fetching the heads.\n\nIt won't _really_ solve the problem, since the pushed pack objects will\ngrow at a proportional rate to the current objects - it's just a constant\nfactor (admittedly a potentially fairly _big_ constant factor)  \nimprovement both in size and in number of files.\n\nSo the mirroring ends up getting slowly slower and slower as the number of \npack files go up. In contrast, a git-aware thing can be basically \nconstant-time, and mirroring expense ends up being relative to the size of \nthe change rather than the size of the repository.\n\nBut mirroring just pack-files might solve the problem for the forseeable \nfuture, so..\n\n\"git-receive-pack\" would need to take a flag to tell it to instead of\nunpacking just check the object instead (ie call \"git-unpack-object\" with\nthe \"-n\" flag - it will check that everything looks ok, including the\nembedded protecting SHA1 hash), and write it out to the filesystem (as it\ncomes in) and then rename it to the right place.\n\n\t\t\tLinus\n"},{"id":"5479","messageId":"42C454B2.6090307@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301302410.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-06-30T20:23:14Z","receivedAt":"2005-06-30T20:23:14Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> It won't _really_ solve the problem, since the pushed pack objects will\n> grow at a proportional rate to the current objects - it's just a constant\n> factor (admittedly a potentially fairly _big_ constant factor)  \n> improvement both in size and in number of files.\n> \n\nIf I've understood this correctly, it's not a constant factor \nimprovement in the number of files (in the size, yes); it's changing it \nfrom O(t*c) to O(t) where t is number of trees and c is number of \nchangesets.  That's key.\n\nThe problem we're having (on kernel.org) right now is that there isn't a \nhierarchial time stamp in Unix, so we have to compare on a file-by-file \nlevel.  rsync is quite good at discovering an invariant beginning of a \nfile, but when it comes to a mass of files it has to compare the stamps \non each and every one, each time.  It will only descend into a single \nfile, however, if that file has had its timestamp changed.\n\nFor the purposes of rsync, storing the objects in a single append-only \nfile would be a very efficient method, since the rsync algorithm will \nquickly discover an invariant head and only transmit the tail.  It's not \nideal, and having something git-aware would be better, but I think it's \nreally would be nice to have something which also plays well with rsync. \n  There is a *lot* of infrastructure in rsync which is actually hard to \nreplicate with another tool (including the server architecture); in many \nways it would be easier to convince the rsync developers to create a \nplugin architecture and re-use all that code rather than developing an \nequivalent tool from scratch.\n\n\t-hpa\n"},{"id":"5483","messageId":"7vll4r1sxz.fsf@assigned-by-dhcp.cox.net","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301242470.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-06-30T20:38:16Z","receivedAt":"2005-06-30T20:38:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> In fact, the most recent push was gone with a\n\nLT> \tgit-send-pack master.kernel.org:/pub/scm/linux/kernel/git/torvalds/git.git\n\nCongrats for a job well done.\n\nNow is there anything for us poor mortals who would want to have\na \"pull\" support?  Logging in via ssh and run send-pack on the\nother end is workable but not so pretty ;-).\n"},{"id":"5485","messageId":"Pine.LNX.4.21.0506301611000.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301302410.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-30T20:49:12Z","receivedAt":"2005-06-30T20:49:12Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 30 Jun 2005, Linus Torvalds wrote:\n\n> On Thu, 30 Jun 2005, Daniel Barkalow wrote:\n> > \n> > The right solution probably involves getting each pack file you push to\n> > the mirrors as well as to the master. They'll probably update no less\n> > frequently than you push, and they should go through a series of states\n> > which matches the master, so it's not necessary to have anything smart on\n> > master sending them, and they only have to unpack the files they get (and\n> > update the refs afterward).\n> \n> Hmm, yes. That would work, together with just fetching the heads.\n> \n> It won't _really_ solve the problem, since the pushed pack objects will\n> grow at a proportional rate to the current objects - it's just a constant\n> factor (admittedly a potentially fairly _big_ constant factor)  \n> improvement both in size and in number of files.\n>\n> So the mirroring ends up getting slowly slower and slower as the number of \n> pack files go up. In contrast, a git-aware thing can be basically \n> constant-time, and mirroring expense ends up being relative to the size of \n> the change rather than the size of the repository.\n> \n> But mirroring just pack-files might solve the problem for the forseeable \n> future, so..\n\nWhenever it gets slow, you could replace all the old packs with a single\nnew pack containing all the old objects; and master could repack whenever\nit has a lot of pack files. That's pretty close to O(n) in change size.\n\nAlternatively, having a reverse-ordered list of pack files would mean that\nmirrors could just go through that list until they found one they already\nhad, and stop there, which would really be O(n).\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5484","messageId":"Pine.LNX.4.58.0506301344070.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C454B2.6090307@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T20:52:11Z","receivedAt":"2005-06-30T20:52:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n> If I've understood this correctly, it's not a constant factor \n> improvement in the number of files (in the size, yes); it's changing it \n> from O(t*c) to O(t) where t is number of trees and c is number of \n> changesets.  That's key.\n\nNo, it _is_ a constant factor even in number of files, if you just keep \nthe pack objects around without re-packing them.\n\nBasically, you'd get one new pack-file every time I push. That's better\nthan getting <n> \"raw object\" files (where <n> can be anything from just a\ncouple to several thousand, depending on whether I had pulled things), but\nit's still just a constant factor on both number of files and size of\nfiles.\n\nNow, you could re-pack the objects every once in a while: it would force a\nwhole new \"epoch\", of course and then the mirrorers would have to fetch\nthe whole repacked file, but that might be fine. Especially if you stop\nre-packing after you've hit a certain size (say, a couple of megs), and\nthen start on the next pack.\n\n> For the purposes of rsync, storing the objects in a single append-only \n> file would be a very efficient method, since the rsync algorithm will \n> quickly discover an invariant head and only transmit the tail.\n\nActually, it won't be \"quick\" - it will have to read the whole file and do \nit's hash window thing.\n\nYou _could_ append the pack-files into one single \"superpack\" file (since\nyou can figure out where the pack boundaries are), but it would be\nextremely big after a while, and rsync would spend all its time doing over\nthe hash window. You'd definitely be better off with re-packing.\n\n\t\tLinus\n"},{"id":"5488","messageId":"Pine.LNX.4.21.0506301651250.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"7vll4r1sxz.fsf@assigned-by-dhcp.cox.net","subject":"Re: \"git-send-pack\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-30T21:05:00Z","receivedAt":"2005-06-30T21:05:00Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 30 Jun 2005, Junio C Hamano wrote:\n\n> >>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n> \n> LT> In fact, the most recent push was gone with a\n> \n> LT> \tgit-send-pack master.kernel.org:/pub/scm/linux/kernel/git/torvalds/git.git\n> \n> Congrats for a job well done.\n> \n> Now is there anything for us poor mortals who would want to have\n> a \"pull\" support?  Logging in via ssh and run send-pack on the\n> other end is workable but not so pretty ;-).\n\nI suspect that I'll be able to merge send-pack/receive-pack with\nssh-push/ssh-pull this evening, and then it'll have the feature of not\ncaring too much which side your command line is on.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5487","messageId":"Pine.LNX.4.58.0506301357280.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"7vll4r1sxz.fsf@assigned-by-dhcp.cox.net","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T21:08:08Z","receivedAt":"2005-06-30T21:08:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, Junio C Hamano wrote:\n> \n> Now is there anything for us poor mortals who would want to have\n> a \"pull\" support?  Logging in via ssh and run send-pack on the\n> other end is workable but not so pretty ;-).\n\nI'm thinking about it. You can't actually do send-pack from the other end,\nsince send-pack needs to know what the base is, and the base you have may\nnot even exist in the remote.\n\nSo a \"git-pull-pack\" will follow the objects on the other side until it \nhits one we have, and _then_ it can send a nice pack. It's not hard per \nse, and some of the problems are actually simpler than git-send-pack, but \nit needs more communication (and in order to be efficient you want to not \nping-pong a \"do-you-have-it\" query every time around).\n\nI also want to make sure that the biggest burden is on the pull side, not \nthe push side. I have a plan, though.\n\n\t\tLinus\n"},{"id":"5489","messageId":"42C45FB2.8030206@gmail.com","threadId":"1062","inReplyTo":"7vll4r1sxz.fsf@assigned-by-dhcp.cox.net","subject":"Re: \"git-send-pack\"","fromName":"Dan Holmsand","fromEmail":"holmsand@gmail.com","sentAt":"2005-06-30T21:10:10Z","receivedAt":"2005-06-30T21:10:10Z","isPatch":false,"sender":{"key":"holmsand@gmail.com","avatar":"https://gravatar.com/avatar/5c722084bafd85e754a02efad01fe69107eb6f393253c49232c5c9f7faa974df?d=mp&s=160"},"body":"Junio C Hamano wrote:\n>>>>>>\"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n> \n> \n> LT> In fact, the most recent push was gone with a\n> \n> LT> \tgit-send-pack master.kernel.org:/pub/scm/linux/kernel/git/torvalds/git.git\n> \n> Congrats for a job well done.\n\nAgree totally. And the whole pack thing is really cool. Git is sooo much\nfaster when running from pack-files only on my poor laptop.\n\n> Now is there anything for us poor mortals who would want to have\n> a \"pull\" support?  Logging in via ssh and run send-pack on the\n> other end is workable but not so pretty ;-).\n\nAgreed again :-)\n\nEven cooler would be pack-pulls via http. That would be a bit hard on \nthe servers with the current git-pack-objects, but it ought to be \npossible to create something similar that doesn't re-delta anything, but \ninstead just spits out what's in an existing pack-file, and (perhaps) \ndeltifies objects from the file system.\n\nIf people then re-pack their repositories occasionally, this should be \nplenty fast, the number of files for rsync to deal with could be kept \ndown, as could download times for mortal users.\n\n/dan\n"},{"id":"5491","messageId":"42C462CD.9010909@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301344070.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-06-30T21:23:25Z","receivedAt":"2005-06-30T21:23:25Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n>>For the purposes of rsync, storing the objects in a single append-only \n>>file would be a very efficient method, since the rsync algorithm will \n>>quickly discover an invariant head and only transmit the tail.\n> \n> Actually, it won't be \"quick\" - it will have to read the whole file and do \n> it's hash window thing.\n> \n\nIt does that, but it only have to do that when the actual file has \nchanged.  That's acceptable, at least for the repository sizes we're \nlikely to deal with within the medium term.\n\n\t-hpa\n"},{"id":"5490","messageId":"42C4639F.4040205@zytor.com","threadId":"1062","inReplyTo":"42C462CD.9010909@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-06-30T21:26:55Z","receivedAt":"2005-06-30T21:26:55Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"H. Peter Anvin wrote:\n> Linus Torvalds wrote:\n> \n>>\n>>> For the purposes of rsync, storing the objects in a single \n>>> append-only file would be a very efficient method, since the rsync \n>>> algorithm will quickly discover an invariant head and only transmit \n>>> the tail.\n>>\n>>\n>> Actually, it won't be \"quick\" - it will have to read the whole file \n>> and do it's hash window thing.\n>>\n> \n> It does that, but it only have to do that when the actual file has \n> changed.  That's acceptable, at least for the repository sizes we're \n> likely to deal with within the medium term.\n> \n\nI guess I should clarify a bit here.  I'm concerned with two aspects: \nthe \"keeping mirrors in sync\" problem, where asking people to use a tool \nother than rsync is a really tough sell, and the developer usage \nscenario, in which case something git-aware is obviously the better thing.\n\n\t-hpa\n"},{"id":"5492","messageId":"Pine.LNX.4.58.0506301412470.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"Pine.LNX.4.21.0506301651250.30848-100000@iabervon.org","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T21:29:18Z","receivedAt":"2005-06-30T21:29:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, Daniel Barkalow wrote:\n> \n> I suspect that I'll be able to merge send-pack/receive-pack with\n> ssh-push/ssh-pull this evening, and then it'll have the feature of not\n> caring too much which side your command line is on.\n\nThe simple thing to do is to just get one commit at a time, see if you \nhave it already, parse if it not, and go on to the parents.\n\nThat would fit the current git-pull thing, and may be good enough, but it \nhas the downside that it can need a _lot_ of back-and-forth fecthing of \ncommit objects from the other side until you find the one you want. That's \ngoing to be _very_ slow over a high-latency connection.\n\nSo what I'd suggest is:\n\n - puller starts by just asking \"what's your SHA1 for the ref I want\"\n\n   The puller wants to know this, because a common case may be that it \n   already has it, in which case it doesn't need to do anything. But more \n   importantly, the puller will need to know this anyway if it gets an \n   object-pack, so that the puller can update it's FETCH_HEAD.\n\n - if puller doesn't have it, then the _puller_ does:\n\n\t\"git-rev-list my-current-refs\"\n\n   to generate an in-date-order list of commits it has, and it starts \n   feeding the result in chunks of 100 entries or something to the other\n   end.\n\n - now, the server sees this stream of SHA1's that the client wants, and \n   it can very cheaply just test \"do I have this SHA1\". Now, if the client \n   hasn't made any changes at all, then the first one will be a hit, and \n   we already have sufficient knowledge to tell what the difference \n   between the client and the server is.\n\n   But more importantly, even if the client _has_ made changes, the client \n   likely has more available CPU than the server has, _and_ the client \n   likely has a shorter list of changes than the server has, so it's\n   really the client that should do this. We should burden the server as \n   lightly as possible for this to scale.\n\n - At some point the server sees the first SHA1 it recognizes, and at that \n   point the server will have to start working. It will just send back an \n   \"ok, got it\" message (telling the client to not bother continuing to \n   send it any more commit ID's), and then does\n\n\tgit-rev-list --objects ref-client-wants ^first-common-sha1 |\n\t\tgit-pack-objects --stdout\n\n - the client just unpacks the objects, and if successful, it puts the new \n   top ref it got into FETCH_HEAD. It's now done.\n\nAnd I do _not_ think that it makes a lot of sense to try to be symmetric.  \nFor one thing, while a \"git-send-pack\" should update all the refs\nin-place, a \"git-pull-pack\" should _not_ update the ref, it should just\nset FETCH_HEAD instead and the puller can decide what he wants to do with\nthat ref (possibly merge it, but possibly just make it be a new local\nbranch \"remote-branch\").\n\nSo I think sending and receiving are fundamentally non-symmetric.\n\n\t\tLinus\n"},{"id":"5493","messageId":"Pine.LNX.4.58.0506301432500.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C462CD.9010909@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T21:42:26Z","receivedAt":"2005-06-30T21:42:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n> It does that, but it only have to do that when the actual file has \n> changed.  That's acceptable, at least for the repository sizes we're \n> likely to deal with within the medium term.\n\nWell, realize that \"incremental packs\" deltify a lot worse than a \"big\npack\", since pack-files don't do deltas to objects outside the pack-file.\n\nSo we'd get _some_ compression, but not as much as possible. The current\nkernel compresses down to a single 63 MB pack-file (that's with the 2.6.11\ntree too, not just the HEAD history), but without deltas it weights in at\nabout 177 MB.\n\nSo a \"sum of incremental packs\" should be somewhere in between those two\nvalues, even today. For a single kernel archive.\n\nSo repository sizes aren't exactly trivial. I don't know how expensive\nthat rsync hash thing is, but one thing you lose is the ability to\nhardlink objects, so if you have a few kernel repositories at some point\nit doesn't fit in the cache any more, and then the rsync will have to read\nthat much pack object stuff from disk in addition to doing the hash. Ugh.\n\n\t\tLinus\n"},{"id":"5494","messageId":"42C46A3C.1070104@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301412470.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-06-30T21:55:08Z","receivedAt":"2005-06-30T21:55:08Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"It seems to me that git always defines a DAG of objects, such that if \nyou have a list of terminals (defined as objects not referenced by other \nobjects), you can, given access to the same objects, figure out all \nintervening objects.\n\nThe tricky bit becomes finding the DAG both sides have in common with as \nlittle traffic as possible.\n\nFor producing minimum network traffic, I think something like this would \nwork:\n\na) The sender sends a list of its terminals to the receiver.\n\nb) The receiver sends a list of nodes it needs, plus a list of all its \nown meta-terminals, obtained by pruning its own DAG according to the \nterminals list of the sender.\n\nc) This may have to be performed iteratively?  I need to sit down and \nwork out the exact algorithm for all cases, including branch trees and \nmulti-rooted DAGs.\n\nd) Once the sender knows the subset of its own DAG available to the \nreceiver, it can transmit either all objects that it has the sender does \nnot, or all objects on the path to one or more specific objects (e.g. HEAD.)\n\n\t-hpa\n"},{"id":"5495","messageId":"42C46B86.8070006@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301432500.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-06-30T22:00:38Z","receivedAt":"2005-06-30T22:00:38Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n>>It does that, but it only have to do that when the actual file has \n>>changed.  That's acceptable, at least for the repository sizes we're \n>>likely to deal with within the medium term.\n> \n> \n> Well, realize that \"incremental packs\" deltify a lot worse than a \"big\n> pack\", since pack-files don't do deltas to objects outside the pack-file.\n> \n> So we'd get _some_ compression, but not as much as possible. The current\n> kernel compresses down to a single 63 MB pack-file (that's with the 2.6.11\n> tree too, not just the HEAD history), but without deltas it weights in at\n> about 177 MB.\n> \n> So a \"sum of incremental packs\" should be somewhere in between those two\n> values, even today. For a single kernel archive.\n> \n> So repository sizes aren't exactly trivial. I don't know how expensive\n> that rsync hash thing is, but one thing you lose is the ability to\n> hardlink objects, so if you have a few kernel repositories at some point\n> it doesn't fit in the cache any more, and then the rsync will have to read\n> that much pack object stuff from disk in addition to doing the hash. Ugh.\n> \n\nThe bulk of the cost in doing the hashing comes from having to read the \nfile.\n\nWell, if you grow a single pack file with appending, then you can have \ndelta references to earlier objects within the same pack file.\n\nAt least at this point, we'd handle a few very large files a lot better \nthan an enormous swarm of smaller ones.\n\nIn the end, it might be that the right thing to do for git on kernel.org \nis to have a single, unified object store which isn't accessible by \nanything other than git-specific protocols.  There would have to be some \nway of dealing with, for example, conflicting tags that apply to \ndifferent repositories, though.\n\n\t-hpa\n"},{"id":"5497","messageId":"Pine.LNX.4.21.0506301731220.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301412470.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-30T22:25:43Z","receivedAt":"2005-06-30T22:25:43Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 30 Jun 2005, Linus Torvalds wrote:\n\n> On Thu, 30 Jun 2005, Daniel Barkalow wrote:\n> > \n> > I suspect that I'll be able to merge send-pack/receive-pack with\n> > ssh-push/ssh-pull this evening, and then it'll have the feature of not\n> > caring too much which side your command line is on.\n> \n> The simple thing to do is to just get one commit at a time, see if you \n> have it already, parse if it not, and go on to the parents.\n> \n> That would fit the current git-pull thing, and may be good enough, but it \n> has the downside that it can need a _lot_ of back-and-forth fecthing of \n> commit objects from the other side until you find the one you want. That's \n> going to be _very_ slow over a high-latency connection.\n> \n> So what I'd suggest is:\n> \n> 1- puller starts by just asking \"what's your SHA1 for the ref I want\"\n> \n>    The puller wants to know this, because a common case may be that it \n>    already has it, in which case it doesn't need to do anything. But more \n>    importantly, the puller will need to know this anyway if it gets an \n>    object-pack, so that the puller can update it's FETCH_HEAD.\n\nAlready have this, for the non-pack case.\n\n>  - At some point the server sees the first SHA1 it recognizes, and at that \n>    point the server will have to start working. It will just send back an \n>    \"ok, got it\" message (telling the client to not bother continuing to \n>    send it any more commit ID's), and then does\n> \n> \tgit-rev-list --objects ref-client-wants ^first-common-sha1 |\n> \t\tgit-pack-objects --stdout\n\nRight.\n\n>  - the client just unpacks the objects, and if successful, it puts the new \n>    top ref it got into FETCH_HEAD. It's now done.\n\nOr wherever it's been told to, yes.\n\n> And I do _not_ think that it makes a lot of sense to try to be symmetric.  \n> For one thing, while a \"git-send-pack\" should update all the refs\n> in-place, a \"git-pull-pack\" should _not_ update the ref, it should just\n> set FETCH_HEAD instead and the puller can decide what he wants to do with\n> that ref (possibly merge it, but possibly just make it be a new local\n> branch \"remote-branch\").\n\nMy expectation is that the puller will have a ref \"remote-branch\", and\nwill therefore: (1) want to update it, and (2) know the last commit pulled\nfrom it. In this situation, we can skip figuring out the start (the two\npoints I didn't quote), because we saved it from before.\n\nAt least, this is how I've always done it; I've got a \"linus\" branch that\nfollows the public repo, and I commit changes to a different branch. I\nsuppose one could skip hanging onto this info, but it seems like an\nobviously useful thing to keep, if for no other reason than that I want to\ndiff against it. This is essentially promoting FETCH_HEAD to a refs/heads/\nthing, and having separate ones when you pull from separate sources.\n\nI suppose things are different if you do a lot of one-shot pulls, rather\nthan tracking branches that you pull from; I'll need to think about this\ncase (assuming that's actually what you do).\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5496","messageId":"Pine.LNX.4.58.0506301514240.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C46A3C.1070104@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T22:26:23Z","receivedAt":"2005-06-30T22:26:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n> For producing minimum network traffic, I think something like this would \n> work:\n\nIn the \"minimum traffic\", the thing to look at is number of packets, and \npenalize further for anything that requires a synchronous reply.\n\nThat's why I'd suggest just letting the client stream out the list of\nobjects it has - it may appear wasteful to stream out even a thousand\nSHA1's, but hey, that's just 20kB worth of data, and especially if there\nis no synchronous stuff, that's just 15 ethernet packets.\n\nFor the server side, looking up a thousand SHA's is pretty easy (it's\n_really_ cheap if the server ends up using a few big packed objects: you\ndon't even have to look at the pack data itself, it can look at just the\nindex and say \"yup, I've got it\")\n\nSo I'd go for simple brute force over anything that needs to discuss\nthings and have a back-and-forth between server/client. And making the\nclient do the heavy lifting is the right thing to do (the server will have\nto create the pack, which can be expensive, but you can tune the delta \nwindow for how much CPU the server has)\n\n\t\tLinus\n"},{"id":"5498","messageId":"42C482ED.1010306@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301514240.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-06-30T23:40:29Z","receivedAt":"2005-06-30T23:40:29Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n>>For producing minimum network traffic, I think something like this would \n>>work:\n> \n> In the \"minimum traffic\", the thing to look at is number of packets, and \n> penalize further for anything that requires a synchronous reply.\n> \n> That's why I'd suggest just letting the client stream out the list of\n> objects it has - it may appear wasteful to stream out even a thousand\n> SHA1's, but hey, that's just 20kB worth of data, and especially if there\n> is no synchronous stuff, that's just 15 ethernet packets.\n> \n\nIn your linux-2.6 tree, there are currently 54,204 objects, and that is \nafter less than one full 2.6.x kernel release cycle.  That's a megabyte \nof SHA1s.\n\nIn /pub/scm on kernel.org, there are currently 1,815,573 objects or hard \nlinks to objects, which would take a 36.3 MB list to produce.\n\nAlthough this is better than what rsync does, which is it encodes this \nlist into ASCII with pathnames and all and it ends up being closer to \n200 MB, it isn't fundamentally different.\n\n\t-hpa\n"},{"id":"5499","messageId":"Pine.LNX.4.58.0506301655310.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"Pine.LNX.4.21.0506301731220.30848-100000@iabervon.org","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-30T23:56:29Z","receivedAt":"2005-06-30T23:56:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, Daniel Barkalow wrote:\n> \n> My expectation is that the puller will have a ref \"remote-branch\", and\n> will therefore: (1) want to update it, and (2) know the last commit pulled\n> from it. In this situation, we can skip figuring out the start (the two\n> points I didn't quote), because we saved it from before.\n\nThis is _never_ how I do things, so I think that's a bad expectation. I \nhave other peoples trees \"just show up\", since they are actually based on \nmine..\n\n\t\tLinus\n"},{"id":"5500","messageId":"Pine.LNX.4.58.0506301656570.14331@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C482ED.1010306@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-01T00:02:57Z","receivedAt":"2005-07-01T00:02:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n> In your linux-2.6 tree, there are currently 54,204 objects, and that is \n> after less than one full 2.6.x kernel release cycle.  That's a megabyte \n> of SHA1s.\n\nBut that's _all_ objects. There are \"only\" 4040 commit objects (which are\nalways the starting point for a search). \n\nSo streaming out the commit objects a few hundred at a time is actually \na very simple strategy. \n\nAlso, note that the server is usually _more_ ahead than the client is, and \nthe server is the one that potentially has lots of commits that the \nclient doesn't have. Not the other way around. So if the client makes a \nlist of it's top commits, it almost certainly won't have to make a very \nlong list until the server can tell it \"ok, stop, I've seen it\".\n\nYeah, maybe we want to limit the \"burst\" to 70 sha1's, since that will fit \nin a regular-sized ethernet packet, but whatever - you'd burst out your \ncommits \"latest first\", so you'd never even get to the current 4040 unless \nyou've literally done the kind of work we've done in the git tree for the \nlast 3 months _and_you've_not_pulled_from_that_server_in_the_whole_time_.\n\n\t\tLinus\n"},{"id":"5504","messageId":"42C49B35.3050204@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301656570.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-01T01:24:05Z","receivedAt":"2005-07-01T01:24:05Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 30 Jun 2005, H. Peter Anvin wrote:\n> \n>>In your linux-2.6 tree, there are currently 54,204 objects, and that is \n>>after less than one full 2.6.x kernel release cycle.  That's a megabyte \n>>of SHA1s.\n> \n> \n> But that's _all_ objects. There are \"only\" 4040 commit objects (which are\n> always the starting point for a search). \n> \n\nWell, there are objects that reference commit objects (e.g. tag \nobjects), not the other way around, but your point is well taken.\n\n> So streaming out the commit objects a few hundred at a time is actually \n> a very simple strategy. \n> \n> Also, note that the server is usually _more_ ahead than the client is, and \n> the server is the one that potentially has lots of commits that the \n> client doesn't have. Not the other way around. So if the client makes a \n> list of it's top commits, it almost certainly won't have to make a very \n> long list until the server can tell it \"ok, stop, I've seen it\".\n\nWell, what I proposed was pretty much that except to have the client \n(receiver) start first.\n\nI prefer calling it sender and receiver, because in the case of upload \nand download you have different sides being the \"server\".\n\n> Yeah, maybe we want to limit the \"burst\" to 70 sha1's, since that will fit \n> in a regular-sized ethernet packet, but whatever - you'd burst out your \n> commits \"latest first\", so you'd never even get to the current 4040 unless \n> you've literally done the kind of work we've done in the git tree for the \n> last 3 months _and_you've_not_pulled_from_that_server_in_the_whole_time_.\n\nWell, in the common case (sender has a superset of receiver), what I \nproposed would converge on the first iteration.  I'm not even convinced \nthat the algorithm *ever* needs to iterate.\n\n\t-hpa\n"},{"id":"5506","messageId":"Pine.LNX.4.21.0507010033080.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301655310.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-07-01T05:01:04Z","receivedAt":"2005-07-01T05:01:04Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 30 Jun 2005, Linus Torvalds wrote:\n\n> On Thu, 30 Jun 2005, Daniel Barkalow wrote:\n> > \n> > My expectation is that the puller will have a ref \"remote-branch\", and\n> > will therefore: (1) want to update it, and (2) know the last commit pulled\n> > from it. In this situation, we can skip figuring out the start (the two\n> > points I didn't quote), because we saved it from before.\n> \n> This is _never_ how I do things, so I think that's a bad expectation. I \n> have other peoples trees \"just show up\", since they are actually based on \n> mine..\n\nOkay, so my next task will be to support this case.\n\nWhat I'm doing now is:\n\n - if the source is using an old version, fall back on individual objects\n\n - send one (or more) ids to exclude\n\n - find out if the server recognized any of the ids\n\n - if not, fall back on transferring individual objects (or we could try\n   another batch)\n\n - request a pack for the given hash, excluding whatever we've said to\n   exclude\n\nI've implemented this for the case of updating a head, and got it to\ntransfer a pack of 11 objects. It took 31s (including connecting) to\ntransfer the entire history of git (3973 objects) over a DSL-DSL link with\na 39ms ping time. I sent the same thing with the old method previously,\nand it took ages (wasn't timing it, though).\n\nIt should be possible to notice that we're not updating a ref, send all\nthe refs you have instead, see if the source recognized any, try again\nwith the next 70 commits, check, and repeat. Does this match what you were\nsuggesting?\n\nI can send you the messy version tomorrow if you want to hack on it or\ntest it, and I'll have a clean patch series over the weekend.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5530","messageId":"pan.2005.07.01.09.50.25.353825@smurf.noris.de","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301233270.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-07-01T09:50:28Z","receivedAt":"2005-07-01T09:50:28Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, Linus Torvalds wrote:\n\n> maybe it would be ok to\n> even have \"git-receive-pack\" as the shell for the receiver side, so that\n> you don't actually give the mirrorer any shell access at all.\n\nYou can probably just set the remote command (in ~/.ssh/authorized_keys)\nto git-receive-pack. That also works around any $PATH issues.\n\nOnce this is stable, master.kernel.org should be updated with the\nlatest git.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nPeople are never so ready to believe you as when you say things in dispraise\nof yourself; and you are never so much annoyed as when they take you at your\nword.\n\t\t\t\t\t-- Somerset Maugham\n"},{"id":"5531","messageId":"pan.2005.07.01.10.31.53.906759@smurf.noris.de","threadId":"1062","inReplyTo":"42C46B86.8070006@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-07-01T10:31:53Z","receivedAt":"2005-07-01T10:31:53Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, H. Peter Anvin wrote:\n\n> In the end, it might be that the right thing to do for git on kernel.org \n> is to have a single, unified object store which isn't accessible by \n> anything other than git-specific protocols.\n\nMakes sense.\n\n>  There would have to be some \n> way of dealing with, for example, conflicting tags that apply to \n> different repositories, though.\n>\nIt seems that user-specific subdirectories in refs/heads (and, presumably,\n../tags) mostly work already.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nDon't lock the barn after it is stolen.\n"},{"id":"5533","messageId":"m13bqyk4uh.fsf_-_@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"42C46B86.8070006@zytor.com","subject":"Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-01T13:56:06Z","receivedAt":"2005-07-01T13:56:06Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> In the end, it might be that the right thing to do for git on kernel.org is to\n> have a single, unified object store which isn't accessible by anything other\n> than git-specific protocols.  There would have to be some way of dealing with,\n> for example, conflicting tags that apply to different repositories, though.\n\nAs far as I can tell public distributed tags are not that hard and if\nyou are going to be synching them it is probably worth working on.\n\nThe basic idea is that instead of having one global tag of\n'linux-2.6.13-rc1' you have a global tag of\n'torvalds@osdl.org/linux-2.6.13-rc1'.\n\nThe important part is that the tag namespace is made hierarchical\nwith at least 2 levels.  Where the top level is a globally\nunique tag owner id and the bottom level is the actual tag.  This\nprevents collisions when merging trees because two peoples\ntags are never in the same namespace, as least when\npeople are not actively hostile :)\n\nStill being a complete git dummy I think the trivial mapping is\nto put tags in:\n.git/refs/tags/user@domain/tag\nand then have a symlink at:\n.git/TAGS \nthat points to your default directory of tags.\n\nEric\n"},{"id":"5534","messageId":"20050701144327.GF4821@delft.aura.cs.cmu.edu","threadId":"1062","inReplyTo":"pan.2005.07.01.10.31.53.906759@smurf.noris.de","subject":"Re: \"git-send-pack\"","fromName":"Jan Harkes","fromEmail":"jaharkes@cs.cmu.edu","sentAt":"2005-07-01T14:43:27Z","receivedAt":"2005-07-01T14:43:27Z","isPatch":false,"sender":{"key":"jaharkes@cs.cmu.edu","avatar":"https://gravatar.com/avatar/cf95aecd150ca8ef33d6edc337ac4bb9e13aa4246fc3679257d578c7fddc1633?d=mp&s=160"},"body":"On Fri, Jul 01, 2005 at 12:31:53PM +0200, Matthias Urlichs wrote:\n> > In the end, it might be that the right thing to do for git on kernel.org \n> > is to have a single, unified object store which isn't accessible by \n> > anything other than git-specific protocols.\n> \n> Makes sense.\n> \n> >  There would have to be some \n> > way of dealing with, for example, conflicting tags that apply to \n> > different repositories, though.\n>\n> It seems that user-specific subdirectories in refs/heads (and, presumably,\n> ../tags) mostly work already.\n\nThey work pretty well, the core git commands have no problem with them\nand I just sent off some patches for gitweb and gitk.\n\nAll git/objects directories can be merged into a common repository. The\nrefs/heads and refs/tags be copied to user specific subdirectories.\n\nThen a pull like,\n    git pull http://www.kernel.org/.../torvalds/linux-2.6.git\n\nWould become,\n    git pull http://www.kernel.org/.../linux-2.6.git torvalds/linux-2.6/master\n\nIt would make rsync more expensive for people who are interested in only\na branch or two, but there is only one repository which should be easier\non the mirrors. The http, ssh, and some future 'pack' transfer methods\nwon't see a difference since they only pull the specific commits they\nneed to catch up with a branch.\n\nJan\n"},{"id":"5540","messageId":"42C5714A.1020203@zytor.com","threadId":"1062","inReplyTo":"m13bqyk4uh.fsf_-_@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-01T16:37:30Z","receivedAt":"2005-07-01T16:37:30Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Eric W. Biederman wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n> \n> \n>>In the end, it might be that the right thing to do for git on kernel.org is to\n>>have a single, unified object store which isn't accessible by anything other\n>>than git-specific protocols.  There would have to be some way of dealing with,\n>>for example, conflicting tags that apply to different repositories, though.\n> \n> \n> As far as I can tell public distributed tags are not that hard and if\n> you are going to be synching them it is probably worth working on.\n> \n> The basic idea is that instead of having one global tag of\n> 'linux-2.6.13-rc1' you have a global tag of\n> 'torvalds@osdl.org/linux-2.6.13-rc1'.\n> \n> The important part is that the tag namespace is made hierarchical\n> with at least 2 levels.  Where the top level is a globally\n> unique tag owner id and the bottom level is the actual tag.  This\n> prevents collisions when merging trees because two peoples\n> tags are never in the same namespace, as least when\n> people are not actively hostile :)\n> \n> Still being a complete git dummy I think the trivial mapping is\n> to put tags in:\n> .git/refs/tags/user@domain/tag\n> and then have a symlink at:\n> .git/TAGS \n> that points to your default directory of tags.\n> \n\nUnless you have an authentication mechanism and *enforce* it (you can do \nthat with GPG signatures if *and only if* your disambiguation includes \nyour GPG signature fingerprint) you still have a problem with someone \nintroducing fake tags as a DoS attack.\n\n\t-hpa\n"},{"id":"5541","messageId":"20050701180944.GA14375@pasky.ji.cz","threadId":"1062","inReplyTo":"m13bqyk4uh.fsf_-_@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-07-01T18:09:44Z","receivedAt":"2005-07-01T18:09:44Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Jul 01, 2005 at 03:56:06PM CEST, I got a letter\nwhere \"Eric W. Biederman\" <ebiederm@xmission.com> told me that...\n> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n> \n> > In the end, it might be that the right thing to do for git on kernel.org is to\n> > have a single, unified object store which isn't accessible by anything other\n> > than git-specific protocols.  There would have to be some way of dealing with,\n> > for example, conflicting tags that apply to different repositories, though.\n> \n> As far as I can tell public distributed tags are not that hard and if\n> you are going to be synching them it is probably worth working on.\n> \n> The basic idea is that instead of having one global tag of\n> 'linux-2.6.13-rc1' you have a global tag of\n> 'torvalds@osdl.org/linux-2.6.13-rc1'.\n> \n> The important part is that the tag namespace is made hierarchical\n> with at least 2 levels.  Where the top level is a globally\n> unique tag owner id and the bottom level is the actual tag.  This\n> prevents collisions when merging trees because two peoples\n> tags are never in the same namespace, as least when\n> people are not actively hostile :)\n\nI don't know, I don't consider this very appealing myself. I'd rather\nprefer the private tags to be per-repository rather than per-user, since\nthose ugly \"merged-here\", \"broken\" etc. tags aren't very useful on\nlarger scope than of a repository. OTOH, what tags would be per-user,\nnot per-repository and not global?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"5542","messageId":"42C58D83.9060107@zytor.com","threadId":"1062","inReplyTo":"20050701180944.GA14375@pasky.ji.cz","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-01T18:37:55Z","receivedAt":"2005-07-01T18:37:55Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Petr Baudis wrote:\n> Dear diary, on Fri, Jul 01, 2005 at 03:56:06PM CEST, I got a letter\n> where \"Eric W. Biederman\" <ebiederm@xmission.com> told me that...\n> \n>>\"H. Peter Anvin\" <hpa@zytor.com> writes:\n>>\n>>\n>>>In the end, it might be that the right thing to do for git on kernel.org is to\n>>>have a single, unified object store which isn't accessible by anything other\n>>>than git-specific protocols.  There would have to be some way of dealing with,\n>>>for example, conflicting tags that apply to different repositories, though.\n>>\n>>As far as I can tell public distributed tags are not that hard and if\n>>you are going to be synching them it is probably worth working on.\n>>\n>>The basic idea is that instead of having one global tag of\n>>'linux-2.6.13-rc1' you have a global tag of\n>>'torvalds@osdl.org/linux-2.6.13-rc1'.\n>>\n>>The important part is that the tag namespace is made hierarchical\n>>with at least 2 levels.  Where the top level is a globally\n>>unique tag owner id and the bottom level is the actual tag.  This\n>>prevents collisions when merging trees because two peoples\n>>tags are never in the same namespace, as least when\n>>people are not actively hostile :)\n> \n> \n> I don't know, I don't consider this very appealing myself. I'd rather\n> prefer the private tags to be per-repository rather than per-user, since\n> those ugly \"merged-here\", \"broken\" etc. tags aren't very useful on\n> larger scope than of a repository. OTOH, what tags would be per-user,\n> not per-repository and not global?\n> \n\nHe's talking about global tags, just using a \"globally unique\" \nnamespace.  Which of course only works right if only genuinely can't \ncreate tags outside your assigned namespace.\n\n\t-hpa\n"},{"id":"5543","messageId":"pan.2005.07.01.21.20.11.596336@smurf.noris.de","threadId":"1062","inReplyTo":"42C58D83.9060107@zytor.com","subject":"Re: Tags","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-07-01T21:20:16Z","receivedAt":"2005-07-01T21:20:16Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, H. Peter Anvin wrote:\n\n> Which of course only works right if only genuinely can't \n> create tags outside your assigned namespace.\n\nI'd rather say that you can't *push* the tags to the central server if\ntheir namspace is wrong, but nothing would prevent you from *creating*\narbitrary tags in your own repository.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nHabit is habit, and not to be flung out of the window by any man, but coaxed\ndown-stairs a step at a time.\n\t\t-- Mark Twain\n"},{"id":"5544","messageId":"20050701214230.GA22003@pasky.ji.cz","threadId":"1062","inReplyTo":"42C58D83.9060107@zytor.com","subject":"Re: Tags","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-07-01T21:42:30Z","receivedAt":"2005-07-01T21:42:30Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Jul 01, 2005 at 08:37:55PM CEST, I got a letter\nwhere \"H. Peter Anvin\" <hpa@zytor.com> told me that...\n> Petr Baudis wrote:\n> >Dear diary, on Fri, Jul 01, 2005 at 03:56:06PM CEST, I got a letter\n> >where \"Eric W. Biederman\" <ebiederm@xmission.com> told me that...\n> >\n> >>\"H. Peter Anvin\" <hpa@zytor.com> writes:\n> >>\n> >>\n> >>>In the end, it might be that the right thing to do for git on kernel.org \n> >>>is to\n> >>>have a single, unified object store which isn't accessible by anything \n> >>>other\n> >>>than git-specific protocols.  There would have to be some way of dealing \n> >>>with,\n> >>>for example, conflicting tags that apply to different repositories, \n> >>>though.\n> >>\n> >>As far as I can tell public distributed tags are not that hard and if\n> >>you are going to be synching them it is probably worth working on.\n> >>\n> >>The basic idea is that instead of having one global tag of\n> >>'linux-2.6.13-rc1' you have a global tag of\n> >>'torvalds@osdl.org/linux-2.6.13-rc1'.\n> >>\n> >>The important part is that the tag namespace is made hierarchical\n> >>with at least 2 levels.  Where the top level is a globally\n> >>unique tag owner id and the bottom level is the actual tag.  This\n> >>prevents collisions when merging trees because two peoples\n> >>tags are never in the same namespace, as least when\n> >>people are not actively hostile :)\n> >\n> >\n> >I don't know, I don't consider this very appealing myself. I'd rather\n> >prefer the private tags to be per-repository rather than per-user, since\n> >those ugly \"merged-here\", \"broken\" etc. tags aren't very useful on\n> >larger scope than of a repository. OTOH, what tags would be per-user,\n> >not per-repository and not global?\n> >\n> \n> He's talking about global tags, just using a \"globally unique\" \n> namespace.  Which of course only works right if only genuinely can't \n> create tags outside your assigned namespace.\n\nI doubt that's really useful either. Rather artificial mechanisms for\nprotection of the namespace would have to be deployed, and again, what\nwould it be good for anyway? If you are tagging linux-2.m.n, you are\nprobably whoever you should be - David, Alan, Marcelo, Linus, or whoever\nelse, while if you are tagging linux-2.m.n-cki, you are likely Con\nKolivas. I don't believe there is any (or much) potential for \"natural\"\nconflicts and if you are malicious, you will just fake the namespace;\nbut frequently what's interesting about the tags is not the author at\nall - I would consider it confusing to have to suddenly dive to another\nnamespace when Linus hands maintenance of linux-2.m to someone else.\n\nThe only significant value I can therefore see in the namespaces is\nprevention of user mistakes, but I think the successful strategy here\nwould be just \"upstream will notice\", and make sure the upstream will be\nnoticed properly (perhaps even interactively) about any new tags it\ngets.\n\nOk, I admit that it boils down to me being lazy and that \"it'd be more\ntyping!\"... ;-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"5545","messageId":"42C5BB33.5010304@zytor.com","threadId":"1062","inReplyTo":"20050701214230.GA22003@pasky.ji.cz","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-01T21:52:51Z","receivedAt":"2005-07-01T21:52:51Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Petr Baudis wrote:\n> \n> I doubt that's really useful either. Rather artificial mechanisms for\n> protection of the namespace would have to be deployed, and again, what\n> would it be good for anyway? If you are tagging linux-2.m.n, you are\n> probably whoever you should be - David, Alan, Marcelo, Linus, or whoever\n> else, while if you are tagging linux-2.m.n-cki, you are likely Con\n> Kolivas. I don't believe there is any (or much) potential for \"natural\"\n> conflicts and if you are malicious, you will just fake the namespace;\n> but frequently what's interesting about the tags is not the author at\n> all - I would consider it confusing to have to suddenly dive to another\n> namespace when Linus hands maintenance of linux-2.m to someone else.\n> \n> The only significant value I can therefore see in the namespaces is\n> prevention of user mistakes, but I think the successful strategy here\n> would be just \"upstream will notice\", and make sure the upstream will be\n> noticed properly (perhaps even interactively) about any new tags it\n> gets.\n> \n> Ok, I admit that it boils down to me being lazy and that \"it'd be more\n> typing!\"... ;-)\n> \n\nYou're missing the whole point of the discussion.  Right now the only \nthing that makes a global object store impossible is the potential for a \ntag conflict, either intentional or accidental.\n\n\t-hpa\n"},{"id":"5546","messageId":"Pine.LNX.4.21.0507011818540.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"42C5BB33.5010304@zytor.com","subject":"Re: Tags","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-07-01T22:27:03Z","receivedAt":"2005-07-01T22:27:03Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 1 Jul 2005, H. Peter Anvin wrote:\n\n> You're missing the whole point of the discussion.  Right now the only \n> thing that makes a global object store impossible is the potential for a \n> tag conflict, either intentional or accidental.\n\nIs there some issue remaining with having a global *object* store,\nsymlinked from multiple repositories, each with its own tags and\nsuch? (I'd think that, in the refs, there would be more contention over\nthe heads than the tags, in any case; refs/heads/master is kind of\npopular)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5547","messageId":"m1u0jef8z9.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"42C5714A.1020203@zytor.com","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-01T22:38:02Z","receivedAt":"2005-07-01T22:38:02Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> Eric W. Biederman wrote:\n>> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n>>\n> Unless you have an authentication mechanism and *enforce* it (you can do that\n> with GPG signatures if *and only if* your disambiguation includes your GPG\n> signature fingerprint) you still have a problem with someone introducing fake\n> tags as a DoS attack.\n\nThere is a question of how bad is this.   For releases you certainly\nneed some kind of signature that people can verify and we\nalready have that but I think we can keep spoofing tags\ndown to the same level as spoofing patches.\n\nBasically all this takes is to make your global namespace\nthe committer email address and you have the rule that\nyou can only tag your own commits.  Then when you merge\ntags you never automatically add tags to your own tag namespace.\n\nI think that is enough to make global tags usable in practice.\n\nAnd for those who are typing challenged if all you ever\nlook at are your own tags the you should never need to\nspecify a fully qualified tag name as git should be able\nto find the committer email address through other means.\n\nEric\n"},{"id":"5548","messageId":"42C5C75F.4040100@zytor.com","threadId":"1062","inReplyTo":"m1u0jef8z9.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-01T22:44:47Z","receivedAt":"2005-07-01T22:44:47Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Eric W. Biederman wrote:\n> \n> There is a question of how bad is this.   For releases you certainly\n> need some kind of signature that people can verify and we\n> already have that but I think we can keep spoofing tags\n> down to the same level as spoofing patches.\n> \n> Basically all this takes is to make your global namespace\n> the committer email address and you have the rule that\n> you can only tag your own commits.  Then when you merge\n> tags you never automatically add tags to your own tag namespace.\n> \n\nDoesn't work.  You can trivially generate a key with someone else's \naddress.  It would require a full PKI.\n\n\t-hpa\n"},{"id":"5549","messageId":"20050701225949.GA28011@pasky.ji.cz","threadId":"1062","inReplyTo":"42C5BB33.5010304@zytor.com","subject":"Re: Tags","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-07-01T22:59:49Z","receivedAt":"2005-07-01T22:59:49Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Jul 01, 2005 at 11:52:51PM CEST, I got a letter\nwhere \"H. Peter Anvin\" <hpa@zytor.com> told me that...\n> You're missing the whole point of the discussion.  Right now the only \n> thing that makes a global object store impossible is the potential for a \n> tag conflict, either intentional or accidental.\n\nOk, I was arguing about something a bit different here, sorry.\n\nThe point of refs/tags/ should be to just indicate tags which we have in\nthe current head (remember that this structure comes from the times\nbefore Dave, when the repository:\"master branch\" mapping was 1:1), since\nthat are usually the only objects you have in _your_ repository.  What's\nthe point of having tag linux-1.0.4-ac128 when you don't have the\nlinux-1.0.4-ac branch whatsoever?  The distinction of \"public\" vs\n\"private\" tags here is really only that the \"public\" tags should be\npropagated to your head when you merge the remote head.  This way, each\nhead will have its own set of tags, and it will be only tags which\nactually reference objects relevant to the head.\n\nNow that we can have many branches in a repository, each with its own\nset of tags, we should probably extend the tags hierarchy to\nrefs/tags/<head>/<tagname>. And see, you can actually have that in the\nglobal object store, as long as the head names are unique. But heads\ndon't propagate in any way so that's a purely administrative issue on\nthe global store side.\n\nBTW, I don't think many (most?) heads named \"master\" are big issue.\nThat's how the head is called locally, and noone says that's how the\nhead should be known at the other side too. It's fine to have a head\ncalled \"master\" in your repository and when pushing to the global object\nstore call it \"pasky/linux-l33t\" over there. (If you are using Cogito,\nyou can add that branch using a URL\nproto://global/obj/store#pasky/linux-l33t.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"5550","messageId":"m1ll4qf7mg.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"42C5C75F.4040100@zytor.com","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-01T23:07:19Z","receivedAt":"2005-07-01T23:07:19Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> Eric W. Biederman wrote:\n>> There is a question of how bad is this.   For releases you certainly\n>> need some kind of signature that people can verify and we\n>> already have that but I think we can keep spoofing tags\n>> down to the same level as spoofing patches.\n>> Basically all this takes is to make your global namespace\n>> the committer email address and you have the rule that\n>> you can only tag your own commits.  Then when you merge\n>> tags you never automatically add tags to your own tag namespace.\n>>\n>\n> Doesn't work.  You can trivially generate a key with someone else's address.  It\n> would require a full PKI.\n\nI'm not saying it's provable correct.  I'm simply saying it is as\ncorrect as the rest of the git repository.\n\nIf I really care what developer xyz tagged I will pull from them,\nor a mirror I trust.  And since developer xyz doesn't pull his\nown global tags from other repositories that should be sufficient.\n\nPlus if you pull from a spoofed tag somewhere further along\nwhen you merge your code the merge will fail because what\nyou thought was a common ancestor isn't.  And you will\nalso likely get an error when you have the same tag\ncoming from 2 different sources with different values.\n\nSo all I am really arguing is that using the committer\nemail address is simply sufficient to prevent non-malicious\nconflicts between developers, and it makes it enough\nthat to get a malicious conflict isn't completely trivial.\nSo I think it is good enough.\n\nBut for releases and things lots of people must trust yes you want\na full PKI infrastructure but I don't see a reason any of that\nshould be inherently tied to tags.\n\nEric\n"},{"id":"5551","messageId":"Pine.LNX.4.21.0507011907440.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"m1ll4qf7mg.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-07-01T23:22:04Z","receivedAt":"2005-07-01T23:22:04Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 1 Jul 2005, Eric W. Biederman wrote:\n\n> Plus if you pull from a spoofed tag somewhere further along\n> when you merge your code the merge will fail because what\n> you thought was a common ancestor isn't.  And you will\n> also likely get an error when you have the same tag\n> coming from 2 different sources with different values.\n\nActually, I think it would be beneficial to support multiple tags with the\nsame name in any case: if people are going to use local private tags like\n\"broken\", either we need to support having refs/tags/broken being a list\nof hashes, or any particular user can only have one broken version.\n\nI don't see any major problems with having refs/ files contain potentially\nmultiple hashes (limited by what makes sense to be multiple; i.e., heads/*\nshould have only one value), and this lets the users check the content of\nthe tag objects to figure out what they care about, and either specify\nthings in more detail or discard things they don't like (or, when\nappropriate, use all values). The main issue I see is that rsync wouldn't\nmerge them usefully.\n\n(And it would be useful to have a structure to support keeping a simple\npiece of information about a set of objects.)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5552","messageId":"42C5D553.80905@timesys.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0506301656570.14331@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-07-01T23:44:19Z","receivedAt":"2005-07-01T23:44:19Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"Linus Torvalds wrote:\n\n> Also, note that the server is usually _more_ ahead than the client is, and \n> the server is the one that potentially has lots of commits that the \n> client doesn't have. Not the other way around. So if the client makes a \n> list of it's top commits, it almost certainly won't have to make a very \n> long list until the server can tell it \"ok, stop, I've seen it\".\n> \n> Yeah, maybe we want to limit the \"burst\" to 70 sha1's, since that will fit \n> in a regular-sized ethernet packet, but whatever - you'd burst out your \n> commits \"latest first\", so you'd never even get to the current 4040 unless \n> you've literally done the kind of work we've done in the git tree for the \n> last 3 months _and_you've_not_pulled_from_that_server_in_the_whole_time_.\n\nYou are getting closer and closer to where something like bitTorrent or \na multicast protocol makes sense. The problem isn't just the number of \noutstanding commit objects but the number of machines and developers \nthat want to grab those commits at the same time.\n\n\nMike Taht\nPostCards From The Bleeding Edge\nhttp://the-edge.blogspot.com \"Tempel 1 worth 2.2 million trillion bux\"\n"},{"id":"5553","messageId":"42C5DA77.4030107@zytor.com","threadId":"1062","inReplyTo":"m1ll4qf7mg.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T00:06:15Z","receivedAt":"2005-07-02T00:06:15Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Eric W. Biederman wrote:\n> \n> If I really care what developer xyz tagged I will pull from them,\n> or a mirror I trust.  And since developer xyz doesn't pull his\n> own global tags from other repositories that should be sufficient.\n> \n\nYou're missing something totally and utterly fundamental here: I'm \ntalking about creating an infrastructure (think sourceforge) where there \nis only one git repository for the whole system, period, full stop, end \nof story.\n\n\t-hpa\n"},{"id":"5554","messageId":"42C5DABB.1020505@zytor.com","threadId":"1062","inReplyTo":"42C5D553.80905@timesys.com","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T00:07:23Z","receivedAt":"2005-07-02T00:07:23Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Mike Taht wrote:\n> \n> You are getting closer and closer to where something like bitTorrent or \n> a multicast protocol makes sense. The problem isn't just the number of \n> outstanding commit objects but the number of machines and developers \n> that want to grab those commits at the same time.\n> \n\nNot really.\n\n\t-hpa\n"},{"id":"5555","messageId":"Pine.LNX.4.58.0507011831060.2977@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C5D553.80905@timesys.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T01:56:45Z","receivedAt":"2005-07-02T01:56:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Jul 2005, Mike Taht wrote:\n> \n> You are getting closer and closer to where something like bitTorrent or \n> a multicast protocol makes sense. The problem isn't just the number of \n> outstanding commit objects but the number of machines and developers \n> that want to grab those commits at the same time.\n\nI don't think so. First off, I don't think the decision is kernel- \nspecific, in the sense that I at least use git for sparse and git itself \ntoo, so the solution should make sense for small projects as well.\n\nAlso, even for the kernel, the total dataset right now (after three months\nor whatever) is a 60MB pack. It's not like we're sending DVD's or even\nCD's worth of data around - we're sending the equivalent of 20MB per\n_month_. That's really not a lot of data. You could easily keep up with a \nslow modem.\n\nAlso, the number of people involved isn't _that_ big. We're talking a few\nthousand people who actively would update their trees for a big project,\nand many smaller projects have anything from a couple to maybe a hundred. \nA few mirrors, and you don't have any problem.\n\nSo I think that the problem is actually not that big, and we just need to\nfind an acceptable format. Quite frankly, it might be perfectly acceptable\nfor kernel.org to run a simple packing script once a week which packs\neverything into one single file, and even if that means that the mirrors\nwill have to re-get everything once a week, that actually sounds \nacceptable.\n\nIt's obviously a _stupid_ way to handle the rsync problem, so there's \nbound to be some cleaner solution, but the point is that we can probably \nmake mirroring acceptable even with a really really stupid approach. I'd \nbe a bit ashamed of just how ugly it is, but it would likely _work_ fine.\nYou'd create 52 pack-files in a year, but each pack-file is likely just\nten megabytes each. \n\nOh, each pack-file should also be associated with the list of \"refs\" that\nwere used to generate that pack-file, so make that 104 files per project\nyear (but the list of \"refs\" would usually be something small, like\n\n\trefs/heads/master       4a89a04f1ee21a7c1f4413f1ad7dcfac50ff9b63\n\trefs/tags/v2.6.11       5dc01c595e6c6ec9ccda4f6f69c131c0dd945f8c\n\trefs/tags/v2.6.11-tree  5dc01c595e6c6ec9ccda4f6f69c131c0dd945f8c\n\trefs/tags/v2.6.12       26791a8bcf0e6d33f43aef7682bdb555236d56de\n\trefs/tags/v2.6.12-rc2   9e734775f7c22d2f89943ad6c745571f1930105f\n\trefs/tags/v2.6.12-rc3   0397236d43e48e821cce5bbe6a80a1a56bb7cc3a\n\trefs/tags/v2.6.12-rc4   ebb5573ea8beaf000d4833735f3e53acb9af844c\n\trefs/tags/v2.6.12-rc5   06f6d9e2f140466eeb41e494e14167f90210f89d\n\trefs/tags/v2.6.12-rc6   701d7ecec3e0c6b4ab9bb824fd2b34be4da63b7e\n\trefs/tags/v2.6.13-rc1   733ad933f62e82ebc92fed988c7f0795e64dea62\n\nwhich was trivially generated from my current tree with\n\n\tfor i in refs/*/*; do echo -ne $i\"\\t\"; cat $i; done\n\nso now you can use the refs associated with the previous pack-file as the \nlist of refs you're _not_ interested in, and the current list of refs as \nthe list you _are_ interested in, and generate the new pack-file.\n\nGenerating the pack-file would literally be something like\n\n\tobj=$(git-rev-parse $(cut -f2 new-list) --not $(cut -f2 old-list))\n\tgit-rev-list $obj | git-pack-objects --stdin > new-pack\n\nso a few one-liners like this, run from a cron-job once a week, should\njust do it.\n\n\t\tLinus\n"},{"id":"5556","messageId":"42C61351.10306@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507011831060.2977@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T04:08:49Z","receivedAt":"2005-07-02T04:08:49Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Also, the number of people involved isn't _that_ big. We're talking a few\n> thousand people who actively would update their trees for a big project,\n> and many smaller projects have anything from a couple to maybe a hundred. \n> A few mirrors, and you don't have any problem.\n> \n> So I think that the problem is actually not that big, and we just need to\n> find an acceptable format. Quite frankly, it might be perfectly acceptable\n> for kernel.org to run a simple packing script once a week which packs\n> everything into one single file, and even if that means that the mirrors\n> will have to re-get everything once a week, that actually sounds \n> acceptable.\n> \n> It's obviously a _stupid_ way to handle the rsync problem, so there's \n> bound to be some cleaner solution, but the point is that we can probably \n> make mirroring acceptable even with a really really stupid approach. I'd \n> be a bit ashamed of just how ugly it is, but it would likely _work_ fine.\n> You'd create 52 pack-files in a year, but each pack-file is likely just\n> ten megabytes each. \n> \n\nAny reason not to simply append objects to an existing packfile?  It \nreally seems like an easy solutions, and should have relatively good I/O \npatterns to boot simply because it naturally creates a topological sort \nof the objects.\n\n\t-hpa\n"},{"id":"5557","messageId":"Pine.LNX.4.58.0507012119360.3019@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C61351.10306@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T04:22:30Z","receivedAt":"2005-07-02T04:22:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Jul 2005, H. Peter Anvin wrote:\n> \n> Any reason not to simply append objects to an existing packfile?\n\nWhat happens when somebody screws up in the middle?\n\nThe one thing I care about more than anything else is consistency. We are \ncareful about writing objects in the right order, and we can re-create the \nstate from the originator etc. But if we start appending stuff and \nsomething goes wrong in the middle, I'm just not going to touch it. A \n\"truncate and hope for the best\" algorithm? \n\nBesides, the result is not a valid git archive any more. \n\n\t\tLinus\n"},{"id":"5558","messageId":"42C61818.30109@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507012119360.3019@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T04:29:12Z","receivedAt":"2005-07-02T04:29:12Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 1 Jul 2005, H. Peter Anvin wrote:\n> \n>>Any reason not to simply append objects to an existing packfile?\n> \n> \n> What happens when somebody screws up in the middle?\n> \n> The one thing I care about more than anything else is consistency. We are \n> careful about writing objects in the right order, and we can re-create the \n> state from the originator etc. But if we start appending stuff and \n> something goes wrong in the middle, I'm just not going to touch it. A \n> \"truncate and hope for the best\" algorithm? \n> \n> Besides, the result is not a valid git archive any more. \n> \n\nIt's a log.  It's a standard technique to append entries to a log.  The \nrequirements for this to always be consistent is that a) it's possible \nto know when the entry/entries at the end are inconsistent and b) it's \nalways possible to roll back the log to a consistent state.\n\nThis is normally done with commit records (write data - fdatasync - \nwrite commit record - fdatasync), but in the case of git, the commit \nrecord isn't required because each git record is self-validating.  This \nis an incredibly powerful property.\n\nIf the log is written in topological sort order, then even a truncated \nlog file is a valid (subset) git object store.\n\n\t-hpa\n"},{"id":"5563","messageId":"m1hdfdg0aa.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"42C5DA77.4030107@zytor.com","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-02T07:00:29Z","receivedAt":"2005-07-02T07:00:29Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> Eric W. Biederman wrote:\n>> If I really care what developer xyz tagged I will pull from them,\n>> or a mirror I trust.  And since developer xyz doesn't pull his\n>> own global tags from other repositories that should be sufficient.\n>>\n>\n> You're missing something totally and utterly fundamental here: I'm talking about\n> creating an infrastructure (think sourceforge) where there is only one git\n> repository for the whole system, period, full stop, end of story.\n\nCould be I'm certainly not up to speed on git yet.\n\nHowever all you have to do for your single system git repository is\nto filter tags at creation time.  So for a person to upload something\nyou need a git aware tool and you need authentication so you are certain\nit is the right person creating the tag.  \n\nSince it is a shared repository you probably want rules like you can\nonly create tags that belong to yourself or are owned by people \nwho do not have accounts on the system.\n\nLikewise in a system like sourceforge it is desirable to check all\nof the committer information in commits as well, so you have a reasonable\naudit trail, and it make sense to check little things like the file under\na sha1 key actually matches the sha1 key.\n\nDownstream mirrors can happily rsync just fine.  So long as they\nverify the upstream source.\n\nTags that you mirror are of course suspect but they will always be.\nThe primary tags created by people with accounts should be reliable\nthough.\n\nSo in essence I see nothing with my proposal that is any worse than\nany other part of git.\n\nThat being said, it sounds like there is a slightly more git \nknowledgeable/native version suggested having to do with multiple\nheads.\n\nEric\n"},{"id":"5566","messageId":"pan.2005.07.02.16.00.49.454455@smurf.noris.de","threadId":"1062","inReplyTo":"42C5C75F.4040100@zytor.com","subject":"Re: Tags","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-07-02T16:00:50Z","receivedAt":"2005-07-02T16:00:50Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, H. Peter Anvin wrote:\n\n> Doesn't work.  You can trivially generate a key with someone else's \n> address.  It would require a full PKI.\n\nSo you use the GPG key's fingerprint as the directory name, and add\na few strategically named symlinks for convenience. *Shrug*\n\nBesides, what's wrong with requiring full PKI? Everybody who has\na kernel.org account should be in the strongly connected set...\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nWhat I want is all of the power and none of the responsibility.\n"},{"id":"5567","messageId":"Pine.LNX.4.58.0507021009580.3019@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C61818.30109@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T17:16:46Z","receivedAt":"2005-07-02T17:16:46Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 1 Jul 2005, H. Peter Anvin wrote:\n> \n> It's a log.\n\n..but that's not what we're looking for. I'm not looking for kernel.org to\nbe my distributed backup tape.\n\nFor it to be useful, it must do more than just log all activity and mirror\nit out via rsync. It must also be usable for people pulling on it. Which\nmeans that it has to be a valid git archive or at least easily\nincrementally unpackable, so that people can actually use the end result.\n\nA log of packs that are just incremented is certainly unpackable: you\nteach git-unpack-objects to just unpack several packs after each other.  \nBut since it's not seekable, you'd have to unpack a 100MB compressed\narchive just to get the last tip of it that you don't have unpacked yet.\n\nAlso, it means that it's impossible to efficiently do a git-specific \nthing. I want people to be able to do what we used to be able to do with \nBK: just do a\n\n\tgit pull master.kernel.org:xxxx\n\nand get something useful. And that means _not_ having to pull a 100MB blob \nto get the last objects at the end.\n\nAnd don't tell me \"rsync can efficiently get just the end\". That's true \nfor _mirrors_, but it's not true for users that don't have every single \narchive on kernel.org. I don't have (and I don't want to have) a copy of \nevery single persons log that ever might want to push to me.\n\nSo no, a log simply isn't useful. It _has_ to be a valid git archive to be \nuseful. Thousands of objects satisfy that. Or a \"few packs + few objects\". \nNot a log.\n\n\t\tLinus\n"},{"id":"5568","messageId":"42C6D0E2.8000909@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507021009580.3019@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T17:37:38Z","receivedAt":"2005-07-02T17:37:38Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> ..but that's not what we're looking for. I'm not looking for kernel.org to\n> be my distributed backup tape.\n> \n> For it to be useful, it must do more than just log all activity and mirror\n> it out via rsync. It must also be usable for people pulling on it. Which\n> means that it has to be a valid git archive or at least easily\n> incrementally unpackable, so that people can actually use the end result.\n> \n> A log of packs that are just incremented is certainly unpackable: you\n> teach git-unpack-objects to just unpack several packs after each other.  \n> But since it's not seekable, you'd have to unpack a 100MB compressed\n> archive just to get the last tip of it that you don't have unpacked yet.\n> \n\nAgreed, you also need an index file.  The index file can be recreated \nfrom the log file in case of corruption, but is what you'd use to seek \ndirectly to an object.\n\n\t-hpa\n"},{"id":"5569","messageId":"12c511ca05070210441c0d3a33@mail.gmail.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507021009580.3019@ppc970.osdl.org","subject":"Re: \"git-send-pack\"","fromName":"Tony Luck","fromEmail":"tony.luck@gmail.com","sentAt":"2005-07-02T17:44:47Z","receivedAt":"2005-07-02T17:44:47Z","isPatch":false,"sender":{"key":"tony.luck@gmail.com","avatar":null},"body":"Here's another approach.\n\nTeach the variants of git-pull to look for a file that names an\nalternate repository\nthat should be used to get any object that is referenced in the repository, but\ndoesn't exist in it.\n\nAt least part of the problem for kernel.org is that there around 50 repositories\nthat are tracking the 2.6 kernel.  All of them have 50,000 objects that are\nduplicates of each other ... and a few hundred 'unique' objects that belong\nto just one repo, or are minimally shared.\n\nIf there was a way to specify an alternate repo, then a large GIT server like\nkernel.org could set up a \"git-history\"[1] repo which each of the hosted repos\ncould point to.  Then a cron job could look for duplicates, and move them\noff to the history area.\n\n-Tony\n\n[1] Different projects, like git and sparse, might never have any common\nfiles with the Linux kernel ... but they can all share the same history.\n"},{"id":"5570","messageId":"42C6D318.8050108@zytor.com","threadId":"1062","inReplyTo":"m1hdfdg0aa.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T17:47:04Z","receivedAt":"2005-07-02T17:47:04Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Eric W. Biederman wrote:\n> \n> However all you have to do for your single system git repository is\n> to filter tags at creation time.  So for a person to upload something\n> you need a git aware tool and you need authentication so you are certain\n> it is the right person creating the tag.  \n> \n\nThat's complicated; it pretty much works out to having to have a PKI and \na system of registered IDs, or some such.  That's painful.\n\n\t-hpa\n"},{"id":"5571","messageId":"42C6D36D.4060006@zytor.com","threadId":"1062","inReplyTo":"12c511ca05070210441c0d3a33@mail.gmail.com","subject":"Re: \"git-send-pack\"","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T17:48:29Z","receivedAt":"2005-07-02T17:48:29Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Tony Luck wrote:\n> \n> At least part of the problem for kernel.org is that there around 50 repositories\n> that are tracking the 2.6 kernel.  All of them have 50,000 objects that are\n> duplicates of each other ... and a few hundred 'unique' objects that belong\n> to just one repo, or are minimally shared.\n> \n> If there was a way to specify an alternate repo, then a large GIT server like\n> kernel.org could set up a \"git-history\"[1] repo which each of the hosted repos\n> could point to.  Then a cron job could look for duplicates, and move them\n> off to the history area.\n> \n\nThis is why I've been talking about a global object repository -- \nincluding the problems associated with them.  git as it currently stands \npermit a single global object store, *except* for the issue of duplicate \ntags.\n\n\t-hpa\n"},{"id":"5572","messageId":"m1k6k9drfk.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"42C6D318.8050108@zytor.com","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-02T17:54:39Z","receivedAt":"2005-07-02T17:54:39Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> Eric W. Biederman wrote:\n>> However all you have to do for your single system git repository is\n>> to filter tags at creation time.  So for a person to upload something\n>> you need a git aware tool and you need authentication so you are certain\n>> it is the right person creating the tag.\n>\n> That's complicated; it pretty much works out to having to have a PKI and a\n> system of registered IDs, or some such.  That's painful.\n\n?? Isn't that what ssh is?\n\nTo some extent a lot depends on how active you expect people to\ntry and forge things.  If there is an expectation of honesty\nyou are fine.  \n\nIf you want to build one mondo repository with thousands of developers\nhaving write access you need to be more careful.  But as far as I know\nnone of that is specific to tags.\n\nEric\n"},{"id":"5573","messageId":"42C6D5AD.9070304@zytor.com","threadId":"1062","inReplyTo":"m1k6k9drfk.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T17:58:05Z","receivedAt":"2005-07-02T17:58:05Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Eric W. Biederman wrote:\n> \n> ?? Isn't that what ssh is?\n> \n> To some extent a lot depends on how active you expect people to\n> try and forge things.  If there is an expectation of honesty\n> you are fine.  \n> \n\nI can't afford to have that.\n\n> If you want to build one mondo repository with thousands of developers\n> having write access you need to be more careful.  But as far as I know\n> none of that is specific to tags.\n\nWell, you're wrong.  Tags is the only part of git which cannot be \nprotected by git's own self-validation system.\n\n\t-hpa\n"},{"id":"5574","messageId":"42C6D8F3.7050204@gmail.com","threadId":"1062","inReplyTo":"42C6D36D.4060006@zytor.com","subject":"Re: \"git-send-pack\"","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2005-07-02T18:12:03Z","receivedAt":"2005-07-02T18:12:03Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"H. Peter Anvin wrote:\n> Tony Luck wrote:\n>>\n...\n> \n> This is why I've been talking about a global object repository -- \n> including the problems associated with them.  git as it currently stands \n> permit a single global object store, *except* for the issue of duplicate \n> tags.\n\nSo why not store just the git objects in the global repository and keep\nall the things that reference an object (HEAD, branches/*, refs/*/*,\netc.) in a per project and/or contributor area like it is currently?\n"},{"id":"5575","messageId":"m1fyuxdpq4.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"42C6D5AD.9070304@zytor.com","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-02T18:31:31Z","receivedAt":"2005-07-02T18:31:31Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> Eric W. Biederman wrote:\n>> ?? Isn't that what ssh is?\n>> To some extent a lot depends on how active you expect people to\n>> try and forge things.  If there is an expectation of honesty\n>> you are fine.\n>\n> I can't afford to have that.\n\nSo you are now your requirements are more stringent then sourceforge?\nSourcefore limited things by reducing the scope of commits per\nproject.  But once you had commit access to a project you could do\njust about anything.\n\n>> If you want to build one mondo repository with thousands of developers\n>> having write access you need to be more careful.  But as far as I know\n>> none of that is specific to tags.\n>\n> Well, you're wrong.  Tags is the only part of git which cannot be protected by\n> git's own self-validation system.\n\nWhich is why I suggested having tags in sync with the committer\ninformation, that way you are as valid as the commit record\nin git.  Although I suspect the multiple head solution is\nprobably better, and simply limiting the people who can commit\nto an individual head will achieve what is necessary.  One user\nper head?\n\nOne thing arch has shown is that you can sucessfully move\nauthentication/permission checking to the underlying environment\nif you structure things carefully.\n\nI guess the problem is really we want to structure things so that\na user who has downloaded the code can verify they have the\nrelease/tag is what they are looking for.  You can detect\na spoofed file in objects by simply verifying the sha1 of the file.\n\nFor a file that you can't internally verify that way the traditional\nway to handle that is to create a file with a gpg signature.  So\nis there anything wrong with adding .git/refs/tags/tag-name.sign\nthat is a traditional signature file?   That will at least give\nyou an end to end consistency check.  (Hmm.  Why didn't I suggest\nthis before?)\n\nIf you don't want to mirror and propagate data you need to do\nconsistency checks earlier in the process, and I have probably had\nsome poor suggestions on how to implement those.  But if everything\nis setup so we can verify things once we have the code downloaded,\nwhere you perform the checks is simply a matter of optimization.\n\nEric\n"},{"id":"5576","messageId":"Pine.LNX.4.58.0507021141350.4716@ppc970.osdl.org","threadId":"1062","inReplyTo":"42C6D5AD.9070304@zytor.com","subject":"Re: Tags","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T18:45:52Z","receivedAt":"2005-07-02T18:45:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Jul 2005, H. Peter Anvin wrote:\n> \n> Well, you're wrong.  Tags is the only part of git which cannot be \n> protected by git's own self-validation system.\n\nWell, you _can_ use the tag objects. That's what I do. The namespace isn't\nthe tag name you use (\"v2.6.12\"), it's the name of the tag itself (in this\ncase \"26791a8bcf0e6d33f43aef7682bdb555236d56de\"), and then it does\nactually distribute fine. The symbolic name is encoded within the tag, but \nisn't guaranteed to be unique in any way.\n\nSo no, it doesn't protect the tag _name_ per se. Anybody can create a tag\ncalled \"v2.6.12\", and I don't think there's any way to handle clashes\nsanely. But you can find the tag objects in a pack, and you could index \nthem separately. Then you'd need to let the users decide which ones they \ntrust or want to use.\n\n\t\tLinus\n"},{"id":"5577","messageId":"pan.2005.07.02.19.55.13.345854@smurf.noris.de","threadId":"1062","inReplyTo":"m1fyuxdpq4.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-07-02T19:55:16Z","receivedAt":"2005-07-02T19:55:16Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, Eric W. Biederman wrote:\n\n> So\n> is there anything wrong with adding .git/refs/tags/tag-name.sign\n> that is a traditional signature file?\n\nThe signature is already appended to the tag file itself (or can be).\nSee \"git-tag-script\".\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nDemocracy is that form of government where everybody gets what the majority\ndeserves.\n\t\t\t\t\t-- James Dale Davidson\n"},{"id":"5578","messageId":"20050702203805.GB19206@delft.aura.cs.cmu.edu","threadId":"1062","inReplyTo":"42C5DA77.4030107@zytor.com","subject":"Re: Tags","fromName":"Jan Harkes","fromEmail":"jaharkes@cs.cmu.edu","sentAt":"2005-07-02T20:38:06Z","receivedAt":"2005-07-02T20:38:06Z","isPatch":false,"sender":{"key":"jaharkes@cs.cmu.edu","avatar":"https://gravatar.com/avatar/cf95aecd150ca8ef33d6edc337ac4bb9e13aa4246fc3679257d578c7fddc1633?d=mp&s=160"},"body":"On Fri, Jul 01, 2005 at 05:06:15PM -0700, H. Peter Anvin wrote:\n> Eric W. Biederman wrote:\n> >\n> >If I really care what developer xyz tagged I will pull from them,\n> >or a mirror I trust.  And since developer xyz doesn't pull his\n> >own global tags from other repositories that should be sufficient.\n> >\n> \n> You're missing something totally and utterly fundamental here: I'm \n> talking about creating an infrastructure (think sourceforge) where there \n> is only one git repository for the whole system, period, full stop, end \n> of story.\n\nI'm not entirely sure what you are envisoning, but it is definitely\ndoable in a secure way.\n\n- Assume that each developer will one or more private trees with one or\n  more branches on kernel.org, lets say all these private repositories\n  are stored under /scm/git/<user>/\n\n- Now you create a single 'global repository' which is going to be the\n  publicly visible one that will be mirrored out,\n\n- Then you run the following script (untested)\n  #!/bin/sh\n  GIT_DIR=$global_repo\n  for user in `(cd /scm/git ; ls)`; do\n    for tree in `find /scm/git/$user -name *.git` ; do\n\tfor ref in `find $tree/refs -type f`  ; do\n\t    type=`echo $ref | sed 'sX^.*/refs/\\([^/]*\\)/.*$X\\1X'`\n\t    name=`echo $ref | sed 'sX^.*/refs/[^/]*/\\(.*\\)$X\\1X'`\n\t    git fetch /scm/git/$tree $branch \n\t    mkdir -p $GIT_DIR/refs/$type/$user/$name\n\t    cat $GIT_DIR/FETCH_HEAD > $GIT_DIR/refs/$type/$user/$name\n\tdone\n    done\n  done\n\n- You can repack the global repository whenever you want.\n- Finally, once a user knows that all his changes are available from the\n  global repository, he can remove any objects from his tree and use\n  GIT_ALTERNATE_OBJECT_DIRECTORIES=$global_repo/objects\n  (maybe there should be a flag for git prune to removes local objects\n  that are already available in the alternate object directories)\n\nJan\n"},{"id":"5579","messageId":"42C7043C.9080904@zytor.com","threadId":"1062","inReplyTo":"m1fyuxdpq4.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T21:16:44Z","receivedAt":"2005-07-02T21:16:44Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Eric W. Biederman wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n> \n> \n>>Eric W. Biederman wrote:\n>>\n>>>?? Isn't that what ssh is?\n>>>To some extent a lot depends on how active you expect people to\n>>>try and forge things.  If there is an expectation of honesty\n>>>you are fine.\n>>\n>>I can't afford to have that.\n> \n> So you are now your requirements are more stringent then sourceforge?\n> Sourcefore limited things by reducing the scope of commits per\n> project.  But once you had commit access to a project you could do\n> just about anything.\n> \n\nThey're not using a single global object storage.\n\n\t-hpa\n"},{"id":"5580","messageId":"Pine.LNX.4.58.0507021432370.8247@g5.osdl.org","threadId":"1062","inReplyTo":"42C7043C.9080904@zytor.com","subject":"Re: Tags","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T21:39:23Z","receivedAt":"2005-07-02T21:39:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Jul 2005, H. Peter Anvin wrote:\n> \n> They're not using a single global object storage.\n\nNote that the fact that you use a common object store does not mean that \neverything should be common.\n\nI still contend that tags and branches and things like that should be \npersonal. A \"gitforge\" thing should _not_ try to unify tags. Instead, give \npeople their own private area for keeping their own private references \n(you can limit it to just a few kilobytes per person, so you might as well \njust consider it to be part of their \"user information\" thing along with \nwhatever other preferences they have).\n\nThen, they call all share the objects, and there's never any confusion\nabout tags - everybody has their own tags, and you add a few simple\noperations like \"copy user xxx's tag to my tag-space, and start a new \nbranch from that\".\n\nThere're really no downsides. The only thing you need to have is some nice\ntag-browser (and some simple permission model where developers can say\n\"others can read my tag\" or \"this tag is visible only to me\" - the object \nstore may be shared, but if nobody can see your pointers into the object \nstore, you effectively have a totally private branch - which might be \nwhat some people want).\n\nThere's really never any reason to make tags global. Even in the case of\nthe kernel, people don't want to see a tag like \"v2.6.12\". They want to\nsee what _I_ tagged v2.6.12, so implicit in that whole thing is very much\nthat they want to see _my_ tags. Again, it's a _browsing_ issue, not a\n\"tags should be global\" issue. They should be visible and easily \nfetchable.\n\n\t\tLinus\n"},{"id":"5581","messageId":"42C70A5B.9070606@zytor.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507021432370.8247@g5.osdl.org","subject":"Re: Tags","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-07-02T21:42:51Z","receivedAt":"2005-07-02T21:42:51Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Note that the fact that you use a common object store does not mean that \n> everything should be common.\n> \n> I still contend that tags and branches and things like that should be \n> personal. A \"gitforge\" thing should _not_ try to unify tags. Instead, give \n> people their own private area for keeping their own private references \n> (you can limit it to just a few kilobytes per person, so you might as well \n> just consider it to be part of their \"user information\" thing along with \n> whatever other preferences they have).\n> \n> Then, they call all share the objects, and there's never any confusion\n> about tags - everybody has their own tags, and you add a few simple\n> operations like \"copy user xxx's tag to my tag-space, and start a new \n> branch from that\".\n> \n> There're really no downsides. The only thing you need to have is some nice\n> tag-browser (and some simple permission model where developers can say\n> \"others can read my tag\" or \"this tag is visible only to me\" - the object \n> store may be shared, but if nobody can see your pointers into the object \n> store, you effectively have a totally private branch - which might be \n> what some people want).\n> \n> There's really never any reason to make tags global. Even in the case of\n> the kernel, people don't want to see a tag like \"v2.6.12\". They want to\n> see what _I_ tagged v2.6.12, so implicit in that whole thing is very much\n> that they want to see _my_ tags. Again, it's a _browsing_ issue, not a\n> \"tags should be global\" issue. They should be visible and easily \n> fetchable.\n> \n\nOK, so let me retell what I think I hear you say:\n\n- Store all the tags in the object store; they may conflict.\n- Let each source user have a set of refs, and provide a method for the \nend user to select which refs to get.\n\nIn other words, the only way (other than knowing what GPG keys to trust) \nto distinguish between your \"v2.6.12\" and J. Random Hacker's \"v2.6.12\" \nis that the former is referenced by *your* refs as opposed to JRH's \nrefs.  This also means the refs cannot be uniquely rebuilt from the \nobject storage.\n\n\t-hpa\n"},{"id":"5582","messageId":"42C70EEF.6050207@gmail.com","threadId":"1062","inReplyTo":"42C70A5B.9070606@zytor.com","subject":"Re: Tags","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2005-07-02T22:02:23Z","receivedAt":"2005-07-02T22:02:23Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"\n\nH. Peter Anvin wrote:\n...\n> \n> OK, so let me retell what I think I hear you say:\n> \n> - Store all the tags in the object store; they may conflict.\n> - Let each source user have a set of refs, and provide a method for the \n> end user to select which refs to get.\n> \n> In other words, the only way (other than knowing what GPG keys to trust) \n> to distinguish between your \"v2.6.12\" and J. Random Hacker's \"v2.6.12\" \n> is that the former is referenced by *your* refs as opposed to JRH's \n> refs.  This also means the refs cannot be uniquely rebuilt from the \n> object storage.\n\nWhy have tag objects at all?\n"},{"id":"5583","messageId":"20050702221434.GA23021@pasky.ji.cz","threadId":"1062","inReplyTo":"42C70A5B.9070606@zytor.com","subject":"Re: Tags","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2005-07-02T22:14:34Z","receivedAt":"2005-07-02T22:14:34Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Jul 02, 2005 at 11:42:51PM CEST, I got a letter\nwhere \"H. Peter Anvin\" <hpa@zytor.com> told me that...\n> Linus Torvalds wrote:\n> >\n> >Note that the fact that you use a common object store does not mean that \n> >everything should be common.\n\n\\o/ Finally I have some hope that we don't end up with something\nbraindead w.r.t. the tags... ;-)\n\n..snip..\n> OK, so let me retell what I think I hear you say:\n> \n> - Store all the tags in the object store; they may conflict.\n\nThey may have the same \"human-readable name\", but they will have a\ndifferent hash.\n\n> - Let each source user have a set of refs, and provide a method for the \n> end user to select which refs to get.\n> \n> In other words, the only way (other than knowing what GPG keys to trust) \n> to distinguish between your \"v2.6.12\" and J. Random Hacker's \"v2.6.12\" \n> is that the former is referenced by *your* refs as opposed to JRH's \n> refs.\n\nAfter all, this is the best way to distinguish it, isn't it? Just \"tag\nname\" without a name of the branch the tag concerns makes no sense -\nthat's the point I'm trying to get along. JRH's v2.6.12 wouldn't make\nmuch sense to you if you use Linus' v2.6.12, since the object JRH's\nv2.6.12 references simply may not be in the branch you use. Yes, JRH\ncould tag it somewhere in the common past, but that's kind of strange\nand is likely some private JRH's stuff. If Linus merged JRH, he will\ntake his v2.6.12 if it makes sense in his branch - so the decision\nis then up to the one who merges, which makes some sense too.\n\nFYI, I'll teach Cogito about the refs/tags/<branch>/<tag> later today\n(and totally offtopic, it already has some trivial cg-push now).\nIt will still fall back to refs/tags/<tag>.\n\n> This also means the refs cannot be uniquely rebuilt from the \n> object storage.\n\nWhy should they be, after all.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"5584","messageId":"Pine.LNX.4.58.0507021501450.8247@g5.osdl.org","threadId":"1062","inReplyTo":"42C70A5B.9070606@zytor.com","subject":"Re: Tags","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T22:17:07Z","receivedAt":"2005-07-02T22:17:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Jul 2005, H. Peter Anvin wrote:\n> \n> OK, so let me retell what I think I hear you say:\n> \n> - Store all the tags in the object store; they may conflict.\n\nNo. They cannot conflict.\n\nA git \"tag object\" cannmot conflict in any way. It is just a generic \n\"pointer object\", and like all other objects, it is defined by its \ncontents, and there are no \"conflicts\". If two people have exactly the \nsame pointer, they'll just have the same object - that's not a conflict, \nthat's just a fact of life with content-addressable filesystems.\n\nThe git \"tag object\" contains a suggested symbolic name, but that actually \nhas no meaning except as being informational. So for example:\n\n\t[torvalds@g5 linux]$ git-cat-file tag v2.6.12\n\tobject 9ee1c939d1cb936b1f98e8d81aeffab57bae46ab\n\ttype commit\n\ttag v2.6.12\n\t\n\tThis is the final 2.6.12 release\n\t-----BEGIN PGP SIGNATURE-----\n\tVersion: GnuPG v1.2.4 (GNU/Linux)\n\t\n\tiD8DBQBCsykyF3YsRnbiHLsRAvPNAJ482tCZwuxp/bJRz7Q98MHlN83TpACdHr37\n\to6X/3T+vm8K3bf3driRr34c=\n\t=sBHn\n\t-----END PGP SIGNATURE-----\n\nhere the \"symbolic name\" is \"v2.6.12\", but that's purely informational, \nand nothing at all cares if a million people have made their own tags that \nhave that same tag-name. The git _object_ is:\n\n\t[torvalds@g5 linux]$ git-rev-parse v2.6.12\n\t26791a8bcf0e6d33f43aef7682bdb555236d56de\n\nand that object name is going to be unique (modulo hash collissions)\n\n> - Let each source user have a set of refs, and provide a method for the \n> end user to select which refs to get.\n\nRight. Let users have any damn refs they want. They may be refs to tags\nobjects, but they may just be direct refs to the commit. The tag object \nreally has no meaning to git, except it allows signing. That's really the \n_only_ thing a tag object does: it introduces trust. There's no other \nreason to ever use one, really.\n\nAnd a \"tag ref\" thing is really nothing more (and nothing less) than a\nbranch. It's a 41-byte filename, although if you actually were to have a\n\"gitforge\" deamon, it could also be just the raw 20-byte SHA1 in a\ndatabase. Let people have their own refs, and have some good way to create\nthem and delete them, and copy them from others (and refer to other\npeoples refs - one common usage might be \"I want to merge with that other\nusers ref 'xyzzy'\".\n\nNote that the .git/refs/tags/xxx files are _literally_ treated exactly the \nsame as the same files under \"heads\". Or under \"mydir\". Git really doesn't \ncare, it's purely syntactic sugar. To git, a ref is a ref is a ref. It \njust refers to an object, and it's nothing more than a way to specify some \nrandom SHA1 at any time.\n\n> In other words, the only way (other than knowing what GPG keys to trust) \n> to distinguish between your \"v2.6.12\" and J. Random Hacker's \"v2.6.12\" \n> is that the former is referenced by *your* refs as opposed to JRH's \n> refs.  This also means the refs cannot be uniquely rebuilt from the \n> object storage.\n\nRight. All the refs are personal and \"fleeting\" - some refs are actively\nchanged all the time (branch refs - aka \"heads\" - get updated when you\nupdate the branch). Tags are really the same way in all technical ways,\nand the only real difference between a \"branch ref\" and a \"tag ref\" is\nyour _expectation_ of them - one you expect to be mostly stable, the other\nyou expect to be updated with development. _Technically_ there's no\ndifference between the two, though.\n\n(And you might also change tag contents occasionally. One reason might be\na bug and you decide to re-tag something else. But a more common reason\nmight be because you want to have tags like \"latest\" that don't actually\nupdate with development, but they update with some other event, like a\nrelease event or some automated test cycle completion or something like\nthat. So tags aren't _immutable_ even from an expectation standpoint, \nit's just that they tend to change _less_).\n\nNow, from tag _objects_ (as opposed to tag refs) you _can_ build them if\nsomebody created a tag object, and you have the signature so that you can\nre-associate the tag-name with the person. But you should consider that a\npretty heavy and unusual case. The normal case is that you just want to\nback up peoples refs. They're like a part of a personal \".gitrc\": you\ncould equally well think of them as \"these are my shorthands, because I\ndon't want to talk about 40-digit hex numbers all the time\". It's nothing\nmore than a personal address book, really.\n\n\t\tLinus\n"},{"id":"5585","messageId":"Pine.LNX.4.58.0507021517220.8247@g5.osdl.org","threadId":"1062","inReplyTo":"42C70EEF.6050207@gmail.com","subject":"Re: Tags","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-02T22:20:36Z","receivedAt":"2005-07-02T22:20:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Jul 2005, A Large Angry SCM wrote:\n> \n> Why have tag objects at all?\n\nTrust.\n\nNone of git itself normally has any \"trust\". The SHA1 means that the \n_integrity_ of the archive is ensured, but for some things (notably \nreleases), you want to have something else. That's the \"tag object\".\n\nAnd I really should probably have called them something else. _I_\npersonally tend to want to have a 1:1 relationship between my \"tag\nreferences\" (ie the 20-byte SHA1 pointer) and my \"tag objects\", but that's\nbecause my releases are things that I envision people may actually want to\nverify are mine.\n\nIn many cases, you'd never use a \"tag object\", and the \"tag reference\" \nwould just point directly to a commit, with no extra indirect object.\n\n\t\tLinus\n"},{"id":"5586","messageId":"20050702223219.GA21865@delft.aura.cs.cmu.edu","threadId":"1062","inReplyTo":"20050702203805.GB19206@delft.aura.cs.cmu.edu","subject":"Re: Tags","fromName":"Jan Harkes","fromEmail":"jaharkes@cs.cmu.edu","sentAt":"2005-07-02T22:32:19Z","receivedAt":"2005-07-02T22:32:19Z","isPatch":false,"sender":{"key":"jaharkes@cs.cmu.edu","avatar":"https://gravatar.com/avatar/cf95aecd150ca8ef33d6edc337ac4bb9e13aa4246fc3679257d578c7fddc1633?d=mp&s=160"},"body":"On Sat, Jul 02, 2005 at 04:38:06PM -0400, Jan Harkes wrote:\n> - Then you run the following script (untested)\n\nOk, I tested it and it was pretty broken, I assumed that git-fetch-script\naccepted the same arguments as git-pull-script.\n\nHere is one that actually seems to work.\n\nJan\n\n\n#!/bin/sh\n#\n# combine per-user private trees into a single repository.\n# assumes that user repositories are stored as \"$repos/<user>/<tree>.git\"\n#\nglobal=global.git\nrepos=/path/to/user/repositories\n\nexport GIT_DIR=\"$global\"\n\n# create global repository if it doesn't exist\ngit-init-db\n\nfor tree in $(cd \"$repos\" && find . -name '*.git' -prune | sed 'sX./XX')\ndo\n    root=\"$repos/$tree\"\n    for ref in $(cd \"$root\" && find refs -type f)  ; do\n\techo Synchronizing $tree\n\tgit fetch \"$root\" \"$ref\"\n\n\ttype=$(echo \"$ref\" | sed -ne 'sX^refs/\\([^/]*\\)/.*$X\\1Xp')\n\tname=$(echo \"$ref\" | sed -ne 'sX^refs/[^/]*/\\(.*\\)$X\\1Xp')\n\tdest=\"$GIT_DIR/refs/$type/$tree/$name\"\n\tmkdir -p $(dirname \"$dest\")\n\tcat \"$GIT_DIR/FETCH_HEAD\" > \"$dest\"\n    done\ndone\n"},{"id":"5587","messageId":"42C727FC.3030900@gmail.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507021517220.8247@g5.osdl.org","subject":"Re: Tags","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2005-07-02T23:49:16Z","receivedAt":"2005-07-02T23:49:16Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> \n> On Sat, 2 Jul 2005, A Large Angry SCM wrote:\n>>Why have tag objects at all?\n> \n> Trust.\n> \n> None of git itself normally has any \"trust\". The SHA1 means that the \n> _integrity_ of the archive is ensured, but for some things (notably \n> releases), you want to have something else. That's the \"tag object\".\n> \n\nBut can't the commit object do this just as well by signing the commit text?\n\n> And I really should probably have called them something else. _I_\n> personally tend to want to have a 1:1 relationship between my \"tag\n> references\" (ie the 20-byte SHA1 pointer) and my \"tag objects\", but that's\n> because my releases are things that I envision people may actually want to\n> verify are mine.\n> \n\nYour tendency is to use tag objects as a permanent, public label of some \nstate. Signing the commit text or the email stating that commit \n${COMMIT_SHA} would work just as well for verification purposes. Or even \na blob object containing the signed text \"${COMMIT_SHA} is vX.X.X.X\". \nEither way, you'd still need some kind of external reference to find the \nobject.\n\n> In many cases, you'd never use a \"tag object\", and the \"tag reference\" \n> would just point directly to a commit, with no extra indirect object.\n\nTag refs, like head refs and branches, are all just (temporary) \nnotational shorthand to make using the tools easier.\n\nThe problem with the Borg repository is not the objects but the object \nrefs. Isn't that just a namespace problem?\n"},{"id":"5588","messageId":"42C72B83.6030904@gmail.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507021501450.8247@g5.osdl.org","subject":"Re: Tags","fromName":"Dan Holmsand","fromEmail":"holmsand@gmail.com","sentAt":"2005-07-03T00:04:19Z","receivedAt":"2005-07-03T00:04:19Z","isPatch":false,"sender":{"key":"holmsand@gmail.com","avatar":"https://gravatar.com/avatar/5c722084bafd85e754a02efad01fe69107eb6f393253c49232c5c9f7faa974df?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> And a \"tag ref\" thing is really nothing more (and nothing less) than a\n> branch. \n\nI'm guessing that this is the root of the confusion here. To you, and to \ngit, a tag is just a another branch. And a tag object is pretty much a \nspecialized commit object, that can't have children and only one parent.\n\nBut people seem to *expect* tags to be connected somehow to a specific \nrepository. Or, rather, to a specific branch.\n\nThat's why people want e.g. cogito to get \"all the tags\" from \ntorvalds/linux-2.6.git when they cg-pull.\n\n From git's point of view, that doesn't really make any sense; it's like \nsaying that you should pull all the branches from a specific branch. But \nfrom a practical point of view, it *does* make sense if you hold the \nview that tags are connected to a branch, and that you should be able to \ndiff against v2.6.12 as soon as you've pulled the latest head.\n\nSo why not add tags to the branch itself?\n\nIt should be pretty straightforward: just make git look for tag refs in, \nsay, a .gittags tree in the current HEAD. The whole thing would pretty \nmuch as if you've symlinked .git/refs/tags to .gittags in the current \nworking tree, except that tag refs would have to be read directly from \nthe repository.\n\nThat way, tag refs could be handled pretty much just like any other \ngit-managed file: they can be added, deleted, changed, merged, \ncommitted, etc. We could track their history, and see who tagged what \nand when.\n\nAnd tags could easily be signed and contain arbitrary text, just like \nthe present day tag objects, as long as they start with a sha1 ref.\n\nThis way, a git branch could have public, shared tags, with a minimum of \nhassle. No special-casing needed for storage or transfer.\n\nAnd there would be no room for conflicting tag names (but you could \neasily use the same name in different branches, just as any file can \ndiffer in content between two branches).\n\nIt might be useful, though, to add some syntax for \"tag in a specific \nbranch\", say <branch-name>@<tag-name>.\n\nThe present tagging mechanism should be kept. It is useful for private \ntagging, and may be useful for signalling that \"this is a branch that is \nunlikely to change\".\n\nSo, am I missing something obvious here?\n\n/dan\n"},{"id":"5589","messageId":"Pine.LNX.4.58.0507021656250.8247@g5.osdl.org","threadId":"1062","inReplyTo":"42C727FC.3030900@gmail.com","subject":"Re: Tags","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-03T00:17:43Z","receivedAt":"2005-07-03T00:17:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 2 Jul 2005, A Large Angry SCM wrote:\n>\n> Linus Torvalds wrote:\n> > \n> > None of git itself normally has any \"trust\". The SHA1 means that the \n> > _integrity_ of the archive is ensured, but for some things (notably \n> > releases), you want to have something else. That's the \"tag object\".\n> > \n> \n> But can't the commit object do this just as well by signing the commit text?\n\nYes and no.\n\nTechnically yes, absolutely, you could add a signature to the commit text.\n\nHowever, that's just wrong for several reasons:\n\nFirst off, the signing is not necessarily done by the person committing\nsomething. Think of any paperwork: the person that signs the paperwork is \nnot necessarily the same person that _wrote_ the paperwork. A signature is \na \"witness\".\n\nFor an example of this, look at the signatures that we've had for a long \ntime on kernel.org: check out the files like \"patch-2.6.8.1.sign\". That's \na signature, but it's not a signature by _me_. It's kernel.org signing the \nthing so that downstream people can verify things.\n\nAnd it would be not only wrong, but literally _impossible_ for me to do it \nin the commit. I don't have (or want to have) the kernel.org private key. \nThat's not what the signature is about. kernel.org is signing that \"this \nis what I got, and what I passed on\". It's not signing that \"this is what \nI wrote\".\n\nIn a lot of systems, you tag something good after it has passed a\nregression test. Ie the _tag_ may happen days or even weeks after the\ncommit has been done.\n\nSo any system that signs commits directly is doing something _wrong_. \n\nSecondly, you can say that you trust other things. In git, you can tag \nindividual blobs, and you can tag individual trees. For an example of \nwhere it makes sense to tag (sign) individual file versions, we've \nactually had things like ISDN drivers (or firmware) that passed some telco \nverification suite, and in certain countries it used to be that you \nweren't legally supposed to use hadrware that hadn't passed that suite. In \ncases like that, you could sign the particular version of the driver, and \nsay \"this one is good\".\n\n(Yeah, those laws are happily going away, but I think the ISDN people in \ngermany actually ended up doing exactly that, except they obviously didn't \nuse git signatures. I think they had a list of file+md5sum).\n\nFinally, it's a tools issue. It's wrong to mix up the notion of committing \nand signing in the same thing, because that just complicates a tool that \nhas to be able to do both. Now you can have a nice graphical commit tool, \nand it doesn't need to know about public keys etc to be useful - you can \nuse another tool to do the signing.\n\nSmall is beautiful, but \"independent\" is even more so.\n\n> Your tendency is to use tag objects as a permanent, public label of some \n> state. Signing the commit text or the email stating that commit \n> ${COMMIT_SHA} would work just as well for verification purposes.\n\nWell, according to that logic, you'd never need signatures at all - you \ncan always keep them totally outside the system.\n\nBut if they are totally outside the system, then you have to have some\nother mechanism to track them, and you can never trust a git archive on\nits own. My goal with the tag objects was that you can just get my git\narchive, and the archive is _inherently_ trustworthy, because if you care,\nyou can verify it without any external input at all (except you need to\nknow my public key, of course, but that's not a tools issue any more,\nthat's about how signatures work).\n\nSo by having tag objects, I can just have refs to them, and anything that \ncan fetch a ref (which implies _any_ kind of \"pull\" functionality) can get \nit. No special cases. No crap.\n\nDo one thing, and do it well. Git does objects with relationships. That's \nreally what git is all about, and the \"tag object\" fits very well into \nthat mentality.\n\n\t\tLinus\n"},{"id":"5610","messageId":"42C86811.9030906@qualitycode.com","threadId":"1062","inReplyTo":"42C72B83.6030904@gmail.com","subject":"Re: Tags","fromName":"Kevin Smith","fromEmail":"yarcs@qualitycode.com","sentAt":"2005-07-03T22:34:57Z","receivedAt":"2005-07-03T22:34:57Z","isPatch":false,"sender":{"key":"yarcs@qualitycode.com","avatar":null},"body":"Dan Holmsand wrote:\n> So why not add tags to the branch itself?\n> \n> It should be pretty straightforward: just make git look for tag refs in, \n> say, a .gittags tree in the current HEAD. The whole thing would pretty \n> much as if you've symlinked .git/refs/tags to .gittags in the current \n> working tree, except that tag refs would have to be read directly from \n> the repository.\n> \n> That way, tag refs could be handled pretty much just like any other \n> git-managed file: they can be added, deleted, changed, merged, \n> committed, etc. We could track their history, and see who tagged what \n> and when.\n\nSounds like the way mercurial handles tags. It really seemed weird to me \nat first, but the more I think about it, the more it makes sense. Even \nmore so after reading this thread :-)\n\n   http://www.serpentine.com/mercurial/index.cgi?Tag\n\nKevin\n"},{"id":"5665","messageId":"m1vf3p2yks.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507021501450.8247@g5.osdl.org","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-05T13:04:51Z","receivedAt":"2005-07-05T13:04:51Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> (And you might also change tag contents occasionally. One reason might be\n> a bug and you decide to re-tag something else. But a more common reason\n> might be because you want to have tags like \"latest\" that don't actually\n> update with development, but they update with some other event, like a\n> release event or some automated test cycle completion or something like\n> that. So tags aren't _immutable_ even from an expectation standpoint, \n> it's just that they tend to change _less_).\n\nCould you include the person who generated the tag and the time the\ntag was generated in the tag object?\n\nFor a tag like \"latest\" it would help quite a bit if you could actually\nfind out which was the latest version of it :)\n\nEric\n"},{"id":"5669","messageId":"Pine.LNX.4.21.0507051155580.30848-100000@iabervon.org","threadId":"1062","inReplyTo":"m1vf3p2yks.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-07-05T16:21:09Z","receivedAt":"2005-07-05T16:21:09Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 5 Jul 2005, Eric W. Biederman wrote:\n\n> Could you include the person who generated the tag and the time the\n> tag was generated in the tag object?\n> \n> For a tag like \"latest\" it would help quite a bit if you could actually\n> find out which was the latest version of it :)\n\nActually, what you really want here is to put in refs/tags/latest the hash\nof the tag whose \"tag\" field is v2.6.13-rc1 (or whatever it is). Having a\ntag with the \"tag\" field of \"latest\" would be a bit silly, because the\nobject will probably stay in circulation long after it's no longer\ntrue. And the object itself would tell you that it was the latest version\nwhen it was created (but isn't every version?). That's why you want the\n_tag_ to say something useful about the version (maybe \"v2.6.12\", maybe\njust \"tested\"), and the _ref_ to tell you it's the latest.\n\nThe fact that lots of tags get refs named with their contents is just due\nto tags only getting used for a small portion of their possible uses. This\nonly happens when the feature you'd look something up under is a feature\nwhich is persistent.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5671","messageId":"m1br5hywde.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"Pine.LNX.4.21.0507051155580.30848-100000@iabervon.org","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-05T17:51:25Z","receivedAt":"2005-07-05T17:51:25Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Daniel Barkalow <barkalow@iabervon.org> writes:\n\n> On Tue, 5 Jul 2005, Eric W. Biederman wrote:\n>\n>> Could you include the person who generated the tag and the time the\n>> tag was generated in the tag object?\n>> \n>> For a tag like \"latest\" it would help quite a bit if you could actually\n>> find out which was the latest version of it :)\n>\n> The fact that lots of tags get refs named with their contents is just due\n> to tags only getting used for a small portion of their possible uses. This\n> only happens when the feature you'd look something up under is a feature\n> which is persistent.\n\nTrue but if you can you will get multiple tags with the\nsame suggested name.  So you need so way to find the one you\ncare about.\n\nEither a date or it's position in the tree, are all you have\nto go on.\n\nI picked on latest as that is an extreme example that had already\nbeen mentioned.\n\nEric\n"},{"id":"5672","messageId":"Pine.LNX.4.58.0507051132530.3570@g5.osdl.org","threadId":"1062","inReplyTo":"m1br5hywde.fsf@ebiederm.dsl.xmission.com","subject":"Re: Tags","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-07-05T18:33:52Z","receivedAt":"2005-07-05T18:33:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 5 Jul 2005, Eric W. Biederman wrote:\n> \n> True but if you can you will get multiple tags with the\n> same suggested name.  So you need so way to find the one you\n> care about.\n\nI do agree that it would make sense to have a \"tagger\" field with the same \nsemantics as the \"committer\" in a commit (including all the same fields: \nreal name, email, and date).\n\n\t\tLinus\n"},{"id":"5673","messageId":"7vy88ldpml.fsf@assigned-by-dhcp.cox.net","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507051132530.3570@g5.osdl.org","subject":"Re: Tags","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-07-05T19:22:42Z","receivedAt":"2005-07-05T19:22:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> On Tue, 5 Jul 2005, Eric W. Biederman wrote:\n>> \n>> True but if you can you will get multiple tags with the\n>> same suggested name.  So you need so way to find the one you\n>> care about.\n\nLT> I do agree that it would make sense to have a \"tagger\" field with the same \nLT> semantics as the \"committer\" in a commit (including all the same fields: \nLT> real name, email, and date).\n\nWhile we are talking about changing tag object format/fields,\nI've wondered if we would want to be able to associate more than\none objects with a single tag (i.e. have more than one \"object\"\nlines just like commits can have more than one \"parent\" lines).\nI admit that it would not be a \"tag\" anymore, rather, it would\nbe a \"bag\".\n\nI wanted to have something like this in the past for some reason\nI do not exactly remember anymore, but basically it was to\nrecord \"here is the list of related objects.\"\n\nI could fake it with a multi-parent commit with a commit message\nif all I want to include are commits with a single blob, but\nthat is (1) abusing the commit to record something that is not\neven a merge, and (2) the tree associated with that commit would\nnot mean anything.\n"},{"id":"5729","messageId":"pan.2005.07.06.18.03.57.445719@smurf.noris.de","threadId":"1062","inReplyTo":"7vy88ldpml.fsf@assigned-by-dhcp.cox.net","subject":"Re: Tags","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2005-07-06T18:04:03Z","receivedAt":"2005-07-06T18:04:03Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, Junio C Hamano wrote:\n\n> I wanted to have something like this in the past for some reason\n> I do not exactly remember anymore, but basically it was to\n> record \"here is the list of related objects.\"\n\nOne use I'd have for that is regression testing -- collect all IDs in one\nbag and then say \"gitk bad ^good\".\n\nOTOH, I dunno whether the core tools really need to understand that.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nIf at first you don't succeed, you must be a programmer.\n"},{"id":"5754","messageId":"m13bqrwaux.fsf@ebiederm.dsl.xmission.com","threadId":"1062","inReplyTo":"Pine.LNX.4.58.0507051132530.3570@g5.osdl.org","subject":"Re: Tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2005-07-07T03:31:18Z","receivedAt":"2005-07-07T03:31:18Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Tue, 5 Jul 2005, Eric W. Biederman wrote:\n>> \n>> True but if you can you will get multiple tags with the\n>> same suggested name.  So you need so way to find the one you\n>> care about.\n>\n> I do agree that it would make sense to have a \"tagger\" field with the same \n> semantics as the \"committer\" in a commit (including all the same fields: \n> real name, email, and date).\n\nOk here is a patch that implements it.\n\nI don't know how robust my code to get the defaults of tagger\nemail address and especially tagger name are but basically it\nworks.\n\nIn addition I added a message when git-tag-script is waiting\nfor you to type the tag message so people aren't confused.\n\nAnd of course I modified git-mktag to check that the tagger\nfield is present.\n\nNow git-pull-script just needs to be tweaked to optionally\nadd tags in the update into .git/refs/tags :)   Using git-fsck-cache\nto find tags is doable but it slows down as your archive grows.\n\nEric\n\n\ndiff --git a/date.c b/date.c\ndiff --git a/git-tag-script b/git-tag-script\n--- a/git-tag-script\n+++ b/git-tag-script\n@@ -1,12 +1,30 @@\n #!/bin/sh\n # Copyright (c) 2005 Linus Torvalds\n \n+usage() {\n+\techo 'git tag <tag name> [<sha1>]'\n+\texit 1\n+}\n+\n : ${GIT_DIR=.git}\n+if [ ! -d \"$GIT_DIR\" ]; then\n+\techo Not a git directory 1>&2\n+\texit 1\n+fi\n+\n+if [ $# -gt 2 -o $# -lt 1 ]; then\n+\tusage\n+fi\n \n object=${2:-$(cat \"$GIT_DIR\"/HEAD)}\n type=$(git-cat-file -t $object) || exit 1\n-( echo -e \"object $object\\ntype $type\\ntag $1\\n\"; cat ) > .tmp-tag\n+tagger_name=${GIT_COMMITTER_NAME:-$(sed -n -e \"s/^$(whoami):[^:]*:[^:]*:[^:]*:\\([^:,]*\\).*:.*$/\\1/p\" <  /etc/passwd)}\n+tagger_email=${GIT_COMMITTER_EMAIL:-\"$(whoami)@$(hostname --fqdn)\"}\n+tagger_date=$(date -d \"${GIT_COMMITTER_DATE:-$(date -R)}\" +\"%s %z\") || exit 1\n+echo \"Enter tag message now. ^D when finished\"\n+( echo -e \"object $object\\ntype $type\\ntag $1\\ntagger $tagger_name <$tagger_email> $tagger_date\\n\"; cat) > .tmp-tag\n rm -f .tmp-tag.asc\n gpg -bsa .tmp-tag && cat .tmp-tag.asc >> .tmp-tag\n-git-mktag < .tmp-tag\n-#rm .tmp-tag .tmp-tag.sig\n+exit 1\n+./git-mktag < .tmp-tag\n+rm -f .tmp-tag .tmp-tag.sig\ndiff --git a/mktag.c b/mktag.c\n--- a/mktag.c\n+++ b/mktag.c\n@@ -42,7 +42,7 @@ static int verify_tag(char *buffer, unsi\n \tint typelen;\n \tchar type[20];\n \tunsigned char sha1[20];\n-\tconst char *object, *type_line, *tag_line;\n+\tconst char *object, *type_line, *tag_line, *tagger_line;\n \n \tif (size < 64 || size > MAXSIZE-1)\n \t\treturn -1;\n@@ -91,6 +91,11 @@ static int verify_tag(char *buffer, unsi\n \t\t\tcontinue;\n \t\treturn -1;\n \t}\n+\t/* Verify the tagger line */\n+\ttagger_line = tag_line;\n+\n+\tif (memcmp(tagger_line, \"tagger \", 7) || (tagger_line[7] == '\\n'))\n+\t\treturn -1;\n \n \t/* The actual stuff afterwards we don't care about.. */\n \treturn 0;\n@@ -119,7 +124,7 @@ int main(int argc, char **argv)\n \t\tsize += ret;\n \t}\n \n-\t// Verify it for some basic sanity: it needs to start with \"object <sha1>\\ntype \"\n+\t// Verify it for some basic sanity: it needs to start with \"object <sha1>\\ntype\\ntagger \"\n \tif (verify_tag(buffer, size) < 0)\n \t\tdie(\"invalid tag signature file\");\n \n"}]}