{"thread":{"id":"180","subject":"on when to checksum","startedAt":"2005-04-20T22:25:07Z","lastAt":"2005-05-02T19:57:34Z","messageCount":8,"participants":["Tom Lord","Linus Torvalds","Andrew Timberlake-Newell"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1044","messageId":"200504202225.PAA15992@emf.net","threadId":"180","inReplyTo":null,"subject":"on when to checksum","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-20T22:25:07Z","receivedAt":"2005-04-20T22:25:07Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n\nLinus, \n\nI think you have made a mistake by moving the sha1 checksum from the\nzipped form to the inflated form.  Here is why:\n\nWhat you have set in motion with `git' is an ad-hoc p2p network for\nsharing filesystem trees -- a global distributed filesystem.  I\nbelieve your starter here has a good chance of taking off to be much,\nmuch larger than just a tool for the kernel.\n\nA subset of your work: blobs and blob databaes, has much wider application\nthan just sharing trees:  Those parts of `git' can form a very solid \nfoundation for many other applications as well.   To the extent `git'\nsucceeds in the context of the kernel, it will be invested in and\nextended and generalized --- and the kernel project will benefit.\nSo don't ignore those wider applications even though they are not your\nfocus today: they will generate investment that feeds back to your project.\n\nYour `git' is silent on transports and mirroring of blob databases --\ntasks for scripting, sure -- but those elements won't be far behind.\n\nEventually, slinging around blobs as atomic elements\nof payloads will become very common.\n\nThe blob handle (aka \"address\")/payload model of a blob db is very\nclean and simple.   In a network of nodes speaking to one and other\nby exchanging blobs, I forsee a prominent need for intermediate\nnodes that process blobs \"blindly\" and as quickly as possible.\n\nBlob compression is mostly goofy if regarded just as a way to \nsave on (diminishingly cheap) disk space but it is mostly \nsane if regarded as a way to cut the cost of network bandwidth\nroughly in half.\n\nMust intermediate nodes inflate the payloads passing through them\nor which they cache just to validate them?   That's not a desirable otucome\nfor many obvious reasonhs.\n\nThere *are* concerns about checksumming zips: it is necessary to nail\ndown the zip process and make sure it is absolutely and permanently\ndeterministic for this application.   But *that* is the problem to \nsolve, not avoid by moving what the checksum refers to.\n\nThanks,\n-t\n"},{"id":"1048","messageId":"Pine.LNX.4.58.0504201539180.6467@ppc970.osdl.org","threadId":"180","inReplyTo":"200504202225.PAA15992@emf.net","subject":"Re: on when to checksum","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-20T22:41:47Z","receivedAt":"2005-04-20T22:41:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 20 Apr 2005, Tom Lord wrote:\n> \n> I think you have made a mistake by moving the sha1 checksum from the\n> zipped form to the inflated form.  Here is why:\n\nI'd have agreed with you (and I did, violently) if it wasn't for the\nperformance issues. It makes a huge difference for write-tree, and to me,\nclearly performance _does_ matter.\n\nFractions of seconds may not sound like a lot, but they add up. I work \nwith 200-patch series myself all the time, so I'm very sensitive to a 0.3 \nsecond difference in performance.\n\n\t\tLinus\n"},{"id":"1051","messageId":"200504202252.PAA16837@emf.net","threadId":"180","inReplyTo":"Pine.LNX.4.58.0504201539180.6467@ppc970.osdl.org","subject":"Re: on when to checksum","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-20T22:52:33Z","receivedAt":"2005-04-20T22:52:33Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n   From: Linus Torvalds <torvalds@osdl.org>\n\n   On Wed, 20 Apr 2005, Tom Lord wrote:\n   > \n   > I think you have made a mistake by moving the sha1 checksum from the\n   > zipped form to the inflated form.  Here is why:\n\n   I'd have agreed with you (and I did, violently) if it wasn't for the\n   performance issues. It makes a huge difference for write-tree, and to me,\n   clearly performance _does_ matter.\n\n   Fractions of seconds may not sound like a lot, but they add up. I work \n   with 200-patch series myself all the time, so I'm very sensitive to a 0.3 \n   second difference in performance.\n\n\nHow many times per day do you invoke `write-tree' and why?\n\nIt takes a large multiple of `0.3s' to get me to take you seriously\non this point.\n\nI have long harbored the suspician that your perceived bandwidth\nimplies that you process a lot of patches unread or barely read --\nimplying that your day-to-day bitslingling could/should largely be\nhandled by an Arch-style patch-queue-manager (a script).\n\n-t\n"},{"id":"1052","messageId":"Pine.LNX.4.58.0504201601130.6467@ppc970.osdl.org","threadId":"180","inReplyTo":"200504202252.PAA16837@emf.net","subject":"Re: on when to checksum","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-20T23:07:11Z","receivedAt":"2005-04-20T23:07:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 20 Apr 2005, Tom Lord wrote:\n> \n> How many times per day do you invoke `write-tree' and why?\n\nEvery single commit does a write-tree, so when I merge with Andrew, it's \nusually a series of 100-250 of them in a row.\n\n(Actually, _usualyl_ it's smaller series, but it's the big series that can\nbe painful enough to matter).\n\n> It takes a large multiple of `0.3s' to get me to take you seriously\n> on this point.\n\nThe thing is, I don't \"trickle\" things in. That would be horribly \ninefficient for me. So I go over the patches, make a mbox, and do them all \nin one go. And then they need to happen _fast_. If it takes 20 minutes, I \ngo away for coffee or something, and then if something didn't apply \nhalf-way through, I will have lost my \"context\".\n\nThat's why I want things instant. Not because I have huge daily throughput \nissues, but I have huge _latency_ issues. \n\nI considered doing a \"two-level\" thing, where I first did the stuff in a\nlight-weigth patch manager, and then batched things up in the background\nfor the real thing. But the fact is, I don't think it's needed. Not the\nway git performs now. If I can apply a hundred patches in a minute or two,\nI have not \"lost the context\" if it turns out that there is some silly\nglitch with one of them.\n\n\t\tLinus\n"},{"id":"1061","messageId":"200504202339.QAA17753@emf.net","threadId":"180","inReplyTo":"Pine.LNX.4.58.0504201601130.6467@ppc970.osdl.org","subject":"Re: on when to checksum","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-20T23:39:39Z","receivedAt":"2005-04-20T23:39:39Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n(I'll have to study/think about that for a while before a proper\nreply.  Tomorrow, probably.)\n\nThanks,\n-t\n\n"},{"id":"1148","messageId":"002a01c54692$a723adb0$9b11a8c0@allianceoneinc.com","threadId":"180","inReplyTo":"200504202225.PAA15992@emf.net","subject":"RE: on when to checksum","fromName":"Andrew Timberlake-Newell","fromEmail":"andrew.timberlake-newell@allianceoneinc.com","sentAt":"2005-04-21T16:53:30Z","receivedAt":"2005-04-21T16:53:30Z","isPatch":false,"sender":{"key":"andrew.timberlake-newell@allianceoneinc.com","avatar":null},"body":"Tom Lord graced us with:\n> I think you have made a mistake by moving the sha1 checksum from the\n> zipped form to the inflated form.  Here is why:\n> \n> What you have set in motion with `git' is an ad-hoc p2p network for\n> sharing filesystem trees -- a global distributed filesystem.  I\n> believe your starter here has a good chance of taking off to be much,\n> much larger than just a tool for the kernel.\n\nThis might rather be a call for a git derivative.\n\nAs Linus has already mentioned in this thread, git is optimized for his need\nfor local speed.  But while sacrificing local speed for network speed would\nbreak git by stepping away from the git philosophy, a gitling with a\ndifferent philosophy but making use of gitish techniques could make that\nchange without being broken even though git itself can't.\n\n\n"},{"id":"2393","messageId":"200505021921.MAA26977@emf.net","threadId":"180","inReplyTo":"Pine.LNX.4.58.0504201601130.6467@ppc970.osdl.org","subject":"Re: on when to checksum","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-05-02T19:21:50Z","receivedAt":"2005-05-02T19:21:50Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n  The thing is, I don't \"trickle\" things in. That would be horribly \n  inefficient for me. So I go over the patches, make a mbox, and do them all \n  in one go. And then they need to happen _fast_. If it takes 20 minutes, I \n  go away for coffee or something, and then if something didn't apply \n  half-way through, I will have lost my \"context\".\n\n  That's why I want things instant. Not because I have huge daily throughput \n  issues, but I have huge _latency_ issues. \n\nI'm curious about what is the value of the \"batch\" nature of that\nproces?\n\nPresumably most patches apply cleanly and most or orthogonal (order\nindependent).   I'm sure that there are frequently interesting exceptions\nbut am I generally right about \"most\" here?\n\nSo, if I understand, you review each change before stuffing it in a\nmailbox, then you apply all the patches in that mailbox in batch.\nIn the majority of cases, the buffering of changes in the mailbox\nadds nothing.\n\nWhy isn't that more automated: when you approve a change, it could be\napplied at once, in the background.  If conflictless, it can be committed,\ntested, whatever.  If conflicting, *then* the change can be buffered\nup for you to look at.   Explicit declarations from programmers or \ntext-based computations about dependencies among the patches can help\nimprove the queue management in more complicated cases.\n\nIn other words, a more asynchronous process might save you time *and*\npay off by reserving more of your attention for areas where it's \nreally needed.\n\n-t\n\n"},{"id":"2397","messageId":"Pine.LNX.4.58.0505021250450.3594@ppc970.osdl.org","threadId":"180","inReplyTo":"200505021921.MAA26977@emf.net","subject":"Re: on when to checksum","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-02T19:57:34Z","receivedAt":"2005-05-02T19:57:34Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Tom Lord wrote:\n> \n> I'm curious about what is the value of the \"batch\" nature of that\n> proces?\n\nMy time.\n\nI don't know about other people, but I don't multitask. I do one thing, \nand that's it. I don't move my mouse around. I sit in my mail reader, and \nI read email. I don't read one email, switch to another window, apply it, \nswithc back, read the next email etc etc.\n\nIn fact, I claim that anybody who works that way is going to have an IQ of \nabout 15 points lower than somebody who batches things up. Just because \nyou end up losing your context, and that effectively makes you stupid. \n\nConcentration is a wonderful thing, but it _requires_ that you do things \nin a concentrated manner.\n\n> So, if I understand, you review each change before stuffing it in a\n> mailbox, then you apply all the patches in that mailbox in batch.\n> In the majority of cases, the buffering of changes in the mailbox\n> adds nothing.\n\nI read email, and while reading email I save the interesting ones off to\nanother mbox (I call mine \"doit\"). They get saved off for \"later perusal\".\n\nI do a first-order review at that stage, and in fact, 95% of the time, \nwhat goes into the \"doit\" folder _will_ get applied. Not 100%, though, \nexactly because at this stage I just read email and work in a mail-reader: \nI don't usually even look at the actual kernel sources that a patch \ninvolves. In particular, sometimes it turns out that the patch wasn't \nagainst my version at all, but against a -mm tree, and I just don't even \nworry about technical details at that stage.\n\nStage #2 is going through the \"doit\" folder at some later date (maybe a \ncouple of times a day), and going through it one more time. Maybe not that \nmuch more \"carefully\", but with a different intent - now I actually check \nsign-offs, add my own, and check out the actual problems in the source \ntree if needed.\n\nStage #3 is actually applying it.\n\n_Each_ stage culls out bad things.\n\nAnd I _really_ don't bounce between stages.\n\n> In other words, a more asynchronous process might save you time *and*\n> pay off by reserving more of your attention for areas where it's \n> really needed.\n\nIt's not asynchronous. It's batched in different stages so that I can \nwork better. And latency matters.\n\n\t\tLinus\n"}]}