{"thread":{"id":"27818","subject":"Git commit generation numbers","startedAt":"2011-07-14T18:24:27Z","lastAt":"2011-07-21T06:29:33Z","messageCount":47,"participants":["Linus Torvalds","Jeff King","Jakub Narebski","Ted Ts'o","Junio C Hamano","Geert Bosch","Long, Martin","Drew Northup","Shawn Pearce","Tony Luck","Christian Couder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"171352","messageId":"CA+55aFxZq1e8u7kXu1rNDy2UPgP3uOyC5y2j7idKSZ_4eL=bWw@mail.gmail.com","threadId":"27818","inReplyTo":null,"subject":"Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T18:24:27Z","receivedAt":"2011-07-14T18:24:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"Ok, so I see that the old discussion about generation numbers has resurfaced.\n\nAnd I have to say, with six years of git use, I think it's not a\ncoincidence that the notion of generation numbers has come up several\ntimes over the years: I think the lack of them is literally the only\nreal design mistake we have.\n\nAnd I absolutely *detest* the generation number cache thing I see on the list.\n\nMaybe I missed the discussion that actually added them to the commits\n(I don't read the git mailing list regularly any more) but I think\nit's a mistake to add an external cache to work around the fact that I\ndidn't add the generation numbers originally.\n\nSo I think we should just add the generation numbers now. We can make\nthe rule be that if a commit doesn't have a generation number, we end\nup having to compute it (with no real need for caching). Yes, it's\nexpensive. But it's going to be a *lot* less expensive over time as\npeople start using a git version that adds the generation numbers to\ncommits.\n\nAnd we can easily mix this - there's no \"flag-day\" issues. Old\nversions of git will ignore the generation number and generate new\ncommits that doesn't have it. New versions of git will generate them,\nand use them. And once the project starts having generation numbers in\nsome commits, the \"generating them\" part will get cheaper over time.\n\nI'll send out a patch that admittedly does not have much testing as a\nreply to this one. It ends up being really simple. Of course, maybe\nit's simple because I did something incredibly stupid, but please take\na look.\n\n                                 Linus\n"},{"id":"171356","messageId":"20110714183710.GA26820@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFxZq1e8u7kXu1rNDy2UPgP3uOyC5y2j7idKSZ_4eL=bWw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T18:37:10Z","receivedAt":"2011-07-14T18:37:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 11:24:27AM -0700, Linus Torvalds wrote:\n\n> And I have to say, with six years of git use, I think it's not a\n> coincidence that the notion of generation numbers has come up several\n> times over the years: I think the lack of them is literally the only\n> real design mistake we have.\n\nAgreed.\n\n> And I absolutely *detest* the generation number cache thing I see on\n> the list.\n\nI'd love to have in-commit generation numbers. I'm just not sure we can\nget the speeds we want without caching them for existing commits.\n\n> Maybe I missed the discussion that actually added them to the commits\n> (I don't read the git mailing list regularly any more) but I think\n> it's a mistake to add an external cache to work around the fact that I\n> didn't add the generation numbers originally.\n> \n> So I think we should just add the generation numbers now. We can make\n> the rule be that if a commit doesn't have a generation number, we end\n> up having to compute it (with no real need for caching). Yes, it's\n> expensive. But it's going to be a *lot* less expensive over time as\n> people start using a git version that adds the generation numbers to\n> commits.\n\nI'm not sure that is the best plan. Calculating generation numbers\ninvolves going to all roots. So once you have to find any generation\nnumber, it's going to be expensive, no matter how many recent commits\nhave generation numbers already in them (but it won't get _more_\nexpensive as more commits are added; you'll always be traversing from\nthe commit in question down to the roots).\n\nAs we add new commits with generation numbers, we won't need to do a\ncalculation to get their numbers. But if you are doing something like\n\"tag --contains\", you are going to want to know the generation number of\nold tags (otherwise, you can't know whether your cutoff might hit them\nor not). IOW, even if we add generation numbers _today_, every \"tag\n--contains\" in linux-2.6 is going to end up traversing from v3.0-rc7\ndown to the roots to get its generation number (v3.0-rc8 would get an\nembedded generation, of course).\n\nSo if you aren't going to cache generation numbers, then you might as\nwell write your traversal algorithm to assume you don't know them for\nold commits. Because calculating them needs to touch every ancestor, and\nthat's probably equivalent to the worst-case for your algorithm.\n\n\nThere's also one other issue with generation numbers. How do you handle\ngrafts and object-replacement refs?  If you graft history, your embedded\ngeneration numbers will all be junk, and you can't trust them.\n\n-Peff\n"},{"id":"171357","messageId":"CA+55aFwuK+krTA4OcnYhLXtKM5HQ1yuPK+J_vC-5R7AthrHWbg@mail.gmail.com","threadId":"27818","inReplyTo":"20110714183710.GA26820@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T18:47:45Z","receivedAt":"2011-07-14T18:47:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Jul 14, 2011 at 11:37 AM, Jeff King <peff@peff.net> wrote:\n>\n> I'd love to have in-commit generation numbers. I'm just not sure we can\n> get the speeds we want without caching them for existing commits.\n\nSo my argument would be that we'd simply be much better off fixing the\nfundamental data structure (which we can), and let it become the\nlong-term solution.\n\nNow, if *may* turn out that we'd want to have some cache for\ngeneration numbers in commits that don't have them, but I absolutely\nthink that that should be a \"add-on\" rather than anything fundamental.\nFor example, if we just merge the \"add generation numbers to the\ncommit object\" logic first, then the \"cache\" case never really needs\nto care about us generating new commits. They simply won't need the\ncache.\n\nAlso, I suspect that the cache could easily be done as a *small* and\n*incomplete* cache, ie you don't need to cache all commits, it would\nbe sufficient to cache a few hundred spread-out commits, and just know\nthat \"from any commit, the cached commit will be quickly reachable\".\n\n> I'm not sure that is the best plan. Calculating generation numbers\n> involves going to all roots. So once you have to find any generation\n> number, it's going to be expensive, no matter how many recent commits\n> have generation numbers already in them (but it won't get _more_\n> expensive as more commits are added; you'll always be traversing from\n> the commit in question down to the roots).\n\nIt only ends up being expensive if the commit has parents that don't\nhave generation numbers.\n\nThat's a fairly short-term problem. For the kernel, for example,\nbasically no development happens on a base that is older than one or\ntwo releases. So if I (and Greg, with the stable tree) start using my\npatch, within a couple of weeks, pretty much all development would\nhave a generation number in its history.\n\nSure, sometimes I'd merge from people who based their tree on\nsomething old, and I'd end up calculating it all. But it would get\nprogressively rarer.\n\n> As we add new commits with generation numbers, we won't need to do a\n> calculation to get their numbers. But if you are doing something like\n> \"tag --contains\", you are going to want to know the generation number of\n> old tags (otherwise, you can't know whether your cutoff might hit them\n> or not). IOW, even if we add generation numbers _today_, every \"tag\n> --contains\" in linux-2.6 is going to end up traversing from v3.0-rc7\n> down to the roots to get its generation number (v3.0-rc8 would get an\n> embedded generation, of course).\n\nSo that could easily be handled by caching. In fact, I suspect that\nyou could make the cache no associate with a commit ID, but be\nassociated with the tags and heads. But again, then the cache would be\na \"secondary\" issue, not something fundamental.\n\n> So if you aren't going to cache generation numbers, then you might as\n> well write your traversal algorithm to assume you don't know them for\n> old commits.\n\nBut that's how our algorithms are *already* written.\n\nSo why not have that as the fallback? You get the advantage of\ngeneration numbers only with modern things, but those are the ones you\nactually tend to use.\n\nMerge bases are *very* seldom historical, for example.\n\n                     Linus\n"},{"id":"171359","messageId":"CA+55aFy5Kr1oHgvQ0m4m+4zJVSTU-QPc_a-cbv=tDDMG0u_-2Q@mail.gmail.com","threadId":"27818","inReplyTo":"20110714183710.GA26820@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T18:52:01Z","receivedAt":"2011-07-14T18:52:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Jul 14, 2011 at 11:37 AM, Jeff King <peff@peff.net> wrote:\n>\n> There's also one other issue with generation numbers. How do you handle\n> grafts and object-replacement refs?  If you graft history, your embedded\n> generation numbers will all be junk, and you can't trust them.\n\nSo I don't think this is a real problem in practice.\n\nGrafts are already unreliable. You cannot sanely merge over a graft,\nand it has nothing to do with generation numbers.\n\nI'm actually sorry that we ever did grafting. It's fundamentally\nbroken, and can actually destroy your repository (by hiding real\nparents and then causing the commits to get garbage collected). So I\ndon't think grafting should be used as an argument for or against\nanything - it's a hack that breaks some fundamental git database\nconstraints.\n\n                             Linus\n"},{"id":"171362","messageId":"CA+55aFzvib7QF-J3fBj2brcQifXGqoeK1t7vfx6pcJmJAEO0dw@mail.gmail.com","threadId":"27818","inReplyTo":"CA+55aFwuK+krTA4OcnYhLXtKM5HQ1yuPK+J_vC-5R7AthrHWbg@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T18:55:39Z","receivedAt":"2011-07-14T18:55:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Jul 14, 2011 at 11:47 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> Also, I suspect that the cache could easily be done as a *small* and\n> *incomplete* cache, ie you don't need to cache all commits, it would\n> be sufficient to cache a few hundred spread-out commits, and just know\n> that \"from any commit, the cached commit will be quickly reachable\".\n\nPut another way: we could do the cache not as a real dynamic entity,\nbut as something that gets generated at \"git clone\" time or when\nre-packing.\n\nI'm actually much more nervous about a cache being inconsistent than I\nwould be about having generation numbers in the tree. The latter we\ncan (and should - but my patch didn't) add a fsck test for, and then\nyou would never get into some situation where there's some really\nsubtle issue with merge base calculation due to a corrupt cache.\n\n                    Linus\n"},{"id":"171364","messageId":"m3tyaoadfs.fsf@localhost.localdomain","threadId":"27818","inReplyTo":"CA+55aFy5Kr1oHgvQ0m4m+4zJVSTU-QPc_a-cbv=tDDMG0u_-2Q@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-07-14T19:08:34Z","receivedAt":"2011-07-14T19:08:34Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, Jul 14, 2011 at 11:37 AM, Jeff King <peff@peff.net> wrote:\n> >\n> > There's also one other issue with generation numbers. How do you handle\n> > grafts and object-replacement refs?  If you graft history, your embedded\n> > generation numbers will all be junk, and you can't trust them.\n> \n> So I don't think this is a real problem in practice.\n> \n> Grafts are already unreliable. You cannot sanely merge over a graft,\n> and it has nothing to do with generation numbers.\n> \n> I'm actually sorry that we ever did grafting. It's fundamentally\n> broken, and can actually destroy your repository (by hiding real\n> parents and then causing the commits to get garbage collected). So I\n> don't think grafting should be used as an argument for or against\n> anything - it's a hack that breaks some fundamental git database\n> constraints.\n\nWhat about object-replacement refs (i.e. \"git replace\" and refs/replace/)?\n\nThis is modern replacement for grafts mechanism, which is safe against\ngarbage collecting, and contrary to grafts it is transferable (as a ref).\n\nWith replacement objects (e.g. to repair some fragment of history to\nmake it bisectable - I think that was original idea behind introducing\ngit-replace, or instead of grafts to join with historical repository -\nIIRC the reason why grafts mechanism was created) you can also have\ninvalid generation numbers if they are stored in commit headers.  With\ngeneration cache we can simply invaliate it if grafts or replacements\nchange...\n\nP.S. grafts are quite useful when doing history surgery.  Create\ngrafts, check history, use git-filter-branch to make new DAG\npermanent, remove grafts.\n\nP.P.S. What about \"grafts lite\", i.e. shallow clone?  With generation\ncache we can invalidate it when depth changes...\n\n-- \nJakub Narębski\nPoland\n"},{"id":"171365","messageId":"20110714190844.GA26918@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFwuK+krTA4OcnYhLXtKM5HQ1yuPK+J_vC-5R7AthrHWbg@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T19:08:44Z","receivedAt":"2011-07-14T19:08:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 11:47:45AM -0700, Linus Torvalds wrote:\n\n> On Thu, Jul 14, 2011 at 11:37 AM, Jeff King <peff@peff.net> wrote:\n> >\n> > I'd love to have in-commit generation numbers. I'm just not sure we can\n> > get the speeds we want without caching them for existing commits.\n> \n> So my argument would be that we'd simply be much better off fixing the\n> fundamental data structure (which we can), and let it become the\n> long-term solution.\n> \n> Now, if *may* turn out that we'd want to have some cache for\n> generation numbers in commits that don't have them, but I absolutely\n> think that that should be a \"add-on\" rather than anything fundamental.\n> For example, if we just merge the \"add generation numbers to the\n> commit object\" logic first, then the \"cache\" case never really needs\n> to care about us generating new commits. They simply won't need the\n> cache.\n\nSure, I'd be fine with that (modulo the graft issue, which you don't\nseem to care about). I half-toyed with making an extra \"add generation\nnumbers to commit header\" on top of my series, but I wanted to first\nprove that generation numbers actually could yield speedups.\n\n> Also, I suspect that the cache could easily be done as a *small* and\n> *incomplete* cache, ie you don't need to cache all commits, it would\n> be sufficient to cache a few hundred spread-out commits, and just know\n> that \"from any commit, the cached commit will be quickly reachable\".\n\nYeah, that would work. Is it worth the trouble? Your cache size is still\nO(n). And you still have the complexity of _having_ a cache.  Yes, the\nsize is 1/100th of what it was (dropping from 6M to 600K on linux-2.6).\nBut you're also going to spend more time calculating. I think you'd have\nto measure to see how it performs in practice.\n\n> It only ends up being expensive if the commit has parents that don't\n> have generation numbers.\n> \n> That's a fairly short-term problem. For the kernel, for example,\n> basically no development happens on a base that is older than one or\n> two releases. So if I (and Greg, with the stable tree) start using my\n> patch, within a couple of weeks, pretty much all development would\n> have a generation number in its history.\n\nSure, that makes generation during commit-time cheaper, and eventually\nthe cost just goes away. I'm more concerned that it won't actually speed\nup algorithms where you look at old commits, which was the whole point\nin the first place.\n\n> > As we add new commits with generation numbers, we won't need to do a\n> > calculation to get their numbers. But if you are doing something like\n> > \"tag --contains\", you are going to want to know the generation number of\n> > old tags (otherwise, you can't know whether your cutoff might hit them\n> > or not). IOW, even if we add generation numbers _today_, every \"tag\n> > --contains\" in linux-2.6 is going to end up traversing from v3.0-rc7\n> > down to the roots to get its generation number (v3.0-rc8 would get an\n> > embedded generation, of course).\n> \n> So that could easily be handled by caching. In fact, I suspect that\n> you could make the cache no associate with a commit ID, but be\n> associated with the tags and heads. But again, then the cache would be\n> a \"secondary\" issue, not something fundamental.\n\nYeah, you could do that. And it would handle \"tag --contains\" and\n\"branch --contains\" (the latter doesn't even really need a cache; as the\nbranch tips move, they will get new commits with generation numbers). I\nsuspect we could get faster topo-sorting and possibly faster merge-base\ncalculation out of generation numbers, too.  But that won't happen if we\nonly have generation numbers for a handful of specific commits.\n\n> > So if you aren't going to cache generation numbers, then you might as\n> > well write your traversal algorithm to assume you don't know them for\n> > old commits.\n> \n> But that's how our algorithms are *already* written.\n\nSort of. We tend to rely on commit timestamps as a proxy for generation\nnumbers. But in the face of clock skew, git will give wrong answers\n(e.g., Ted posted some examples of name-rev giving wrong answers near\nsome skew in linux-2.6).\n\nIf we aren't going to go whole-hog on generation numbers, I'm much more\ntempted to simply keep using commit timestamps. It's easy to build a\ncache of commits with bogus timestamps (which I've already posted a\npatch for) if you want to better accuracy at the cost of more\ncomplexity. And as time progresses, you tend to ask about commits near\nthe skewed ones less often (and hopefully lessons learned from seeing\nhow the skew occurred will help us prevent them from reocurring in new\ncommits).\n\n-Peff\n"},{"id":"171367","messageId":"20110714191252.GB26918@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFzvib7QF-J3fBj2brcQifXGqoeK1t7vfx6pcJmJAEO0dw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T19:12:52Z","receivedAt":"2011-07-14T19:12:52Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 11:55:39AM -0700, Linus Torvalds wrote:\n\n> I'm actually much more nervous about a cache being inconsistent than I\n> would be about having generation numbers in the tree. The latter we\n> can (and should - but my patch didn't) add a fsck test for, and then\n> you would never get into some situation where there's some really\n> subtle issue with merge base calculation due to a corrupt cache.\n\nInteresting. I'm nervous about that, too, which is why I _favor_ the\ncache. Because we calculate the cache ourselves, we know its accurate\naccording to the parent pointers. If we find a bug, we fix it and bump\nthe cache version, which forces it to regenerate.\n\nContrast that with a bogus generation number that makes its way into an\nactual commit object. That's there for eternity, just like the commit\ntimestamp skew we already have. I find it much less likely to happen\nthan skew in the commit timestamp, if only because generations are a\ndirt-simple concept. But it is a case where there is duplicated\ninformation in the actual DAG, and if that information doesn't match up\nwe are screwed.\n\n-Peff\n"},{"id":"171370","messageId":"CA+55aFx=ACnVBGU8_9wa=9xTbxVoOWKnsqfmBvzq7qzOeMGSNA@mail.gmail.com","threadId":"27818","inReplyTo":"20110714190844.GA26918@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T19:23:31Z","receivedAt":"2011-07-14T19:23:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Jul 14, 2011 at 12:08 PM, Jeff King <peff@peff.net> wrote:\n>\n> If we aren't going to go whole-hog on generation numbers, I'm much more\n> tempted to simply keep using commit timestamps.\n\nSure. I think it's entirely reasonable to say that the issue basically\nboils down to one git question: \"can commit X be an ancestor of commit\nY\" (as a way to basically limit certain algorithms from having to walk\nall the way down). We've used commit dates for it, and realistically\nit really has worked very well. But it was always a broken heuristic.\n\nSo yes, I personally see generation counters as a way to do the commit\ndate comparisons right. And it would be perfectly fine to just say \"if\nthere are no generation numbers, we'll use the datestamps instead, and\nknow that they could be incorrect\".\n\nThat \"use the datestamps\" fallback thing may well involve all the\nheuristics we already do (ie check for the stamps looking sane, and\nnot trusting just one individual one).\n\n                           Linus\n"},{"id":"171372","messageId":"20110714194638.GE8453@thunk.org","threadId":"27818","inReplyTo":"CA+55aFzvib7QF-J3fBj2brcQifXGqoeK1t7vfx6pcJmJAEO0dw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2011-07-14T19:46:38Z","receivedAt":"2011-07-14T19:46:38Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Jul 14, 2011 at 11:55:39AM -0700, Linus Torvalds wrote:\n> On Thu, Jul 14, 2011 at 11:47 AM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n> >\n> > Also, I suspect that the cache could easily be done as a *small* and\n> > *incomplete* cache, ie you don't need to cache all commits, it would\n> > be sufficient to cache a few hundred spread-out commits, and just know\n> > that \"from any commit, the cached commit will be quickly reachable\".\n> \n> Put another way: we could do the cache not as a real dynamic entity,\n> but as something that gets generated at \"git clone\" time or when\n> re-packing.\n\nWould it be considered evil if we put the generation number in the\npack, but not consider it part of the formal object (i.e., it would be\njust a cache, but one that wouldn't change once the pack was created)?\n\n       \t      \t      \t   \t    \t   \t- Ted\n"},{"id":"171373","messageId":"CA+55aFzuQnfo1iywnp-WAajMHe2+6_HOM85aw0bS+p0xv5RyhA@mail.gmail.com","threadId":"27818","inReplyTo":"20110714194638.GE8453@thunk.org","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T19:51:39Z","receivedAt":"2011-07-14T19:51:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Jul 14, 2011 at 12:46 PM, Ted Ts'o <tytso@mit.edu> wrote:\n>\n> Would it be considered evil if we put the generation number in the\n> pack, but not consider it part of the formal object (i.e., it would be\n> just a cache, but one that wouldn't change once the pack was created)?\n\nThat would actually be a major change to data structures, and would\nrequire some serious surgery and be hard to support in a\nbackwards-compatible way (think different git versions accessing the\nsame repository).\n\nMuch bigger patch than the one I did.\n\nSo it sounds like it would work - and it would probably be a simple\nmatter of just incrementing the pack version number if you just say\n\"cannot access the pack with old versions\" - but I think it's a really\nfragile approach.\n\n                  Linus\n"},{"id":"171374","messageId":"20110714200144.GE26918@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFx=ACnVBGU8_9wa=9xTbxVoOWKnsqfmBvzq7qzOeMGSNA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T20:01:44Z","receivedAt":"2011-07-14T20:01:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 12:23:31PM -0700, Linus Torvalds wrote:\n\n> On Thu, Jul 14, 2011 at 12:08 PM, Jeff King <peff@peff.net> wrote:\n> >\n> > If we aren't going to go whole-hog on generation numbers, I'm much more\n> > tempted to simply keep using commit timestamps.\n> \n> Sure. I think it's entirely reasonable to say that the issue basically\n> boils down to one git question: \"can commit X be an ancestor of commit\n> Y\" (as a way to basically limit certain algorithms from having to walk\n> all the way down). We've used commit dates for it, and realistically\n> it really has worked very well. But it was always a broken heuristic.\n\nYeah, I agree with that.\n\n> So yes, I personally see generation counters as a way to do the commit\n> date comparisons right. And it would be perfectly fine to just say \"if\n> there are no generation numbers, we'll use the datestamps instead, and\n> know that they could be incorrect\".\n\nIn that case, is it really worth adding generation numbers to the cache?\nBecause they _can_ be wrong, too. I suspect they will be wrong less\noften than commit timestamps, if only because they're dirt simple to\ncalculate. But all it takes is some crappy porcelain doing:\n\n  git cat-file commit $foo |\n  munge_the_parents |\n  git hash-object -t commit --stdin -w\n\nto give us a bogus object. Sure, we can catch it via fsck. But we could\nalso catch commit timestamp skew via fsck just as easily.\n\n> That \"use the datestamps\" fallback thing may well involve all the\n> heuristics we already do (ie check for the stamps looking sane, and\n> not trusting just one individual one).\n\nThose aren't foolproof, of course. I asked people a few months ago to\nrun my skew-detection program on various repos, and some repos have long\nruns of skew (think somebody with a bad clock or a bogus program doing a\nwhole series). But they're fast and work OK in practice. We should apply\nthem more consistently (name-rev, for example, will tolerate a day of\nskew, but will not look past a single commit).\n\nAnd if people really want to be thorough, we can mark the skewed commits\nin a cache during \"git gc\" for them (or they can just say \"for this\ntraversal, I want to be thorough; turn off timestamp cutoffs\").\n\n\nOut of curiosity, what don't you like about the generation cache? The\nidea of using external storage? Generating it on the fly? The particular\nimplementation is too slow or crappy?\n\n-Peff\n"},{"id":"171376","messageId":"20110714200717.GF26918@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFzuQnfo1iywnp-WAajMHe2+6_HOM85aw0bS+p0xv5RyhA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T20:07:17Z","receivedAt":"2011-07-14T20:07:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 12:51:39PM -0700, Linus Torvalds wrote:\n\n> On Thu, Jul 14, 2011 at 12:46 PM, Ted Ts'o <tytso@mit.edu> wrote:\n> >\n> > Would it be considered evil if we put the generation number in the\n> > pack, but not consider it part of the formal object (i.e., it would be\n> > just a cache, but one that wouldn't change once the pack was created)?\n> \n> That would actually be a major change to data structures, and would\n> require some serious surgery and be hard to support in a\n> backwards-compatible way (think different git versions accessing the\n> same repository).\n\nIf we put it in the index, but not the pack, then it wouldn't be any\nmore painful than pack index v2. I don't recall there being huge fallout\nfrom that; we just gave a reasonable deprecation period before switching\nit on as the default.\n\nI'm not sure it is much less crappy than having the cache in a separate\nfile. It does take less space, since the pack index already contains all\nof the sha1s. But if we don't like the on-the-fly writing of what was in\nmy series, it would not be hard to generate the same cache during\npack-index time. Not having it in a separate file makes it hard to\ninvalidate the cache when the graph changes (due to grafts or replace\nrefs). But maybe we don't care about that. Or maybe it's OK to tell the\nuser to manually rebuild the pack index if they tweak those features.\n\n-Peff\n"},{"id":"171377","messageId":"20110714200817.GF8453@thunk.org","threadId":"27818","inReplyTo":"CA+55aFzuQnfo1iywnp-WAajMHe2+6_HOM85aw0bS+p0xv5RyhA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2011-07-14T20:08:17Z","receivedAt":"2011-07-14T20:08:17Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Jul 14, 2011 at 12:51:39PM -0700, Linus Torvalds wrote:\n> \n> So it sounds like it would work - and it would probably be a simple\n> matter of just incrementing the pack version number if you just say\n> \"cannot access the pack with old versions\" - but I think it's a really\n> fragile approach.\n\nSo if we ever change the pack format again, it's something to think\nabout adding, but probably not worth it on its own...\n\nWhat if we simply have a cache file per pack, which again is generated\nwhen the pack is first received or generated, but is otherwise not\ndynamic?  It's an extra file which is icky, but it would keep things\nsimpler.\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"171378","messageId":"69e0ad24-32b7-4e14-9492-6d0c3d653adf@email.android.com","threadId":"27818","inReplyTo":"20110714200144.GE26918@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-14T20:19:51Z","receivedAt":"2011-07-14T20:19:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nJeff King <peff@peff.net> wrote:\n>\n>Out of curiosity, what don't you like about the generation cache?\n\nThe thing I hate about it is very fundamental: I think it's a hack around a basic git design mistake. And it's a mistake we have known about for a long time.\n\nNow, I don't think it's a *fatal* mistake, but I do find it very broken to basically say \"we made a mistake in the original commit design, and instead of fixing it we create a separate workaround for it\".\n\nTHAT I find distasteful. My reaction is that if we're going to add generation numbers, then were should just do it the way we should have done them originally, rather than as some separate hack.\n\nSee? That's why I wouldn't have any problem with adding a separate cache on top of it, if it's really required, but I would hope that it isn't really needed.\n\nSo a cache in itself is not necessarily wrong. But leaving the original design mistake in place IS.\n\nAnd fixing it really ended up being a very tiny patch, no?\n\n     Linus\n"},{"id":"171379","messageId":"7vmxgg38xz.fsf@alter.siamese.dyndns.org","threadId":"27818","inReplyTo":"20110714183710.GA26820@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-07-14T20:26:32Z","receivedAt":"2011-07-14T20:26:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> There's also one other issue with generation numbers. How do you handle\n> grafts and object-replacement refs?  If you graft history, your embedded\n> generation numbers will all be junk, and you can't trust them.\n\nBy the way, I doubt your \"invalidate and recompute generation cache when\nreplacement changes\" would really work when we consider object transfer\n(which is the whole point of deprecating graft with object replacement\nmechanism). For the purpose of connectivity check during object transfer,\nwe deliberately _ignore_ the object replacements, so you would at least\nwant to have an ability to show the generation number according to the\n\"true\" history recorded in commits (which can come from Linus's in-commit\ngeneration number once everybody migrates) and the generation number that\ntakes grafts and replacements into account (for which we cannot depend on\nin-commit record).\n"},{"id":"171380","messageId":"20110714203141.GA28548@sigill.intra.peff.net","threadId":"27818","inReplyTo":"69e0ad24-32b7-4e14-9492-6d0c3d653adf@email.android.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T20:31:41Z","receivedAt":"2011-07-14T20:31:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 01:19:51PM -0700, Linus Torvalds wrote:\n\n> >Out of curiosity, what don't you like about the generation cache?\n> \n> The thing I hate about it is very fundamental: I think it's a hack\n> around a basic git design mistake. And it's a mistake we have known\n> about for a long time.\n> \n> Now, I don't think it's a *fatal* mistake, but I do find it very\n> broken to basically say \"we made a mistake in the original commit\n> design, and instead of fixing it we create a separate workaround for\n> it\".\n> \n> THAT I find distasteful. My reaction is that if we're going to add\n> generation numbers, then were should just do it the way we should have\n> done them originally, rather than as some separate hack.\n> \n> See? That's why I wouldn't have any problem with adding a separate\n> cache on top of it, if it's really required, but I would hope that it\n> isn't really needed.\n> \n> So a cache in itself is not necessarily wrong. But leaving the\n> original design mistake in place IS.\n\nThanks, that makes some sense to me.\n\nHowever, I'm not 100% convinced leaving generation numbers out was a\nmistake. The git philosophy seems always to have been to keep the\nminimal required information in the DAG. And I think that has served us\nwell, because we're not saddled with cruft that seemed like a good idea\nearly on, but isn't.\n\nGeneration numbers are _completely_ redundant with the actual structure\nof history represented by the parent pointers. Having them in there is\nnot about giving git more information that it doesn't have, but about\nbeing a cheap place to stuff a value that is a little expensive to\ncalculate.\n\nAnd so that seems a bit hack-ish to me.\n\nI liken it somewhat to the \"don't store renames\" debate. We don't want\nto crystallize forever in the history whatever crappy rename-detection\nalgorithm is done at the time of commit. We put the minimum amount of\ninformation in the DAG, and it's the runtime's responsibility to get the\nanswer.\n\nI think the decision is a little more gray with generation numbers,\nbecause it's not about \"you got this information with a wrong and crappy\nalgorithm\" like it might be with rename detection, but rather \"we're\nsticking this redundant number in the commit object, and we assume that\nit will always be useful enough to future algorithms to merit being\nhere\".\n\n> And fixing it really ended up being a very tiny patch, no?\n\nWell, yes. But it also doesn't yield a 100-fold speedup in \"git tag\n--contains\" for existing repositories. So it's not quite a full\nsolution.\n\n-Peff\n"},{"id":"171381","messageId":"20110714204122.GB28548@sigill.intra.peff.net","threadId":"27818","inReplyTo":"7vmxgg38xz.fsf@alter.siamese.dyndns.org","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-14T20:41:22Z","receivedAt":"2011-07-14T20:41:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 01:26:32PM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > There's also one other issue with generation numbers. How do you handle\n> > grafts and object-replacement refs?  If you graft history, your embedded\n> > generation numbers will all be junk, and you can't trust them.\n> \n> By the way, I doubt your \"invalidate and recompute generation cache when\n> replacement changes\" would really work when we consider object transfer\n> (which is the whole point of deprecating graft with object replacement\n> mechanism). For the purpose of connectivity check during object transfer,\n> we deliberately _ignore_ the object replacements, so you would at least\n> want to have an ability to show the generation number according to the\n> \"true\" history recorded in commits (which can come from Linus's in-commit\n> generation number once everybody migrates) and the generation number that\n> takes grafts and replacements into account (for which we cannot depend on\n> in-commit record).\n\nIt should actually work in that scenario, at least with replace refs,\nbut the performance is suboptimal. The copy of git doing the object\ntransfer will turn off read_replace_refs, our validity token will\nnot match, we will see that our cache is no longer valid, and\nregenerate it. Another run with replace-refs turned on will do the same\nthing in reverse. Even two programs running simultaneously will still be\ncorrect, because the cache is replaced atomically.\n\nHowever, there are two issues:\n\n  1. I don't think grafts have a \"respect grafts\" flag in the same way;\n     I haven't looked at how the packing code decides not to respect\n     them, but the \"stir graft info into the checksum\" data should use\n     the same check.\n\n  2. If you do a lot of object transfer, you will ping-pong back and\n     forth between cache versions, which is inefficient. It would\n     probably be better to store the cache that is valid under condition\n     $SHA1 as:\n\n       .git/cache/generations/$SHA1\n\n     In most cases, you would have a single file (i.e., you are not\n     using replace refs at all). But if you did, then you keep two\n     separate caches, one for the view from replace-refs, and one for\n     the standard view.\n\nIf we ignore replace refs and grafts, as Linus suggested, and always\nstore the true generation number, then we could generate it at pack time\n(and even put it in the pack index if we want to deal with a version\nbump there).\n\n-Peff\n"},{"id":"171386","messageId":"7vaacg35zd.fsf@alter.siamese.dyndns.org","threadId":"27818","inReplyTo":"20110714204122.GB28548@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-07-14T21:30:30Z","receivedAt":"2011-07-14T21:30:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> It should actually work in that scenario, at least with replace refs,...\n> regenerate it. Another run ...\n\nI know; that is what I called \"doubt it would really work\". Having to\nregenerate twice does not count as working.\n\n> However, there are two issues:\n>\n>   1. I don't think grafts have a \"respect grafts\" flag in the same way;\n>      I haven't looked at how the packing code decides not to respect\n>      them, but the \"stir graft info into the checksum\" data should use\n>      the same check.\n\nI do not think graft and object transfer meshes well at all, so I wouldn't\nworry about it.\n"},{"id":"171392","messageId":"CA+55aFyDzr+SfgSzWMr9pQuQUXTw9mcjZ-00NZof74PKZzbGPA@mail.gmail.com","threadId":"27818","inReplyTo":"20110714203141.GA28548@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T01:19:30Z","receivedAt":"2011-07-15T01:19:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Jul 14, 2011 at 1:31 PM, Jeff King <peff@peff.net> wrote:\n>\n> However, I'm not 100% convinced leaving generation numbers out was a\n> mistake. The git philosophy seems always to have been to keep the\n> minimal required information in the DAG.\n\nYes.\n\nAnd until I saw the patches trying to add generation numbers, I didn't\nreally try to push adding generation numbers to commits (although it\nactually came up as early as July 2005, so the \"let's use generation\nnumbers in commits\" thing is *really* old).\n\nIn other words, I do agree that we should strive for minimal required\ninformation.\n\nBut dammit, if you start using generation numbers, then they *are*\nrequired information. The fact that you then hide them in some\nunarchitected random file doesn't change anything! It just makes it\nugly and random, for chrissake!\n\nI really don't understand your logic that says that the cache is\nsomehow cleaner. It's a random hack! It's saying \"we don't have it in\nthe main data structure, so let's add it to some other one instead,\nand now we have a consistency and cache generation problem instead\".\n\nJust look at the size of the patches in question. Your caching patches\nare bigger and more complicated. Sure, part of it is that your series\nadds the code to _use_ the generation number, but look purely at the\ncode to maintain them.\n\nWhy do you think the odd separate cache is somehow better than just\ndoing it right? Seriously? If we require the generation numbers, then\nthey have *become* that minimal information that we should save!\n\n And I think that has served us\n> well, because we're not saddled with cruft that seemed like a good idea\n> early on, but isn't.\n\nAgain - we discussed adding generation numbers about 6 years ago. We\nclearly *should* have done it. Instead, we went with the hacky \"let's\nuse commit time\", that everybody really knew was technically wrong,\nand was a hack, but avoided the need.\n\nNow, six years later, you clearly are saying that we need the\ngeneration numbers, but then you go off and try to say that they\nshould be in some secondary non-architected random collection of data\nstructures that isn't covered by the security and maintenance\nguarantees that the core git objects are.\n\nDammit, one of the things that makes git special is that the data\nstructures are NOT random odd ad-hoc files. There is a design to them.\n\n> Generation numbers are _completely_ redundant with the actual structure\n> of history represented by the parent pointers.\n\nNot true. That's only true if you add \".. if you parse the whole\nhistory\" to that statement.\n\nAnd we've *never* parsed the whole history, because it's just too\nexpensive and doesn't scale. So right now we depend on commit dates\nwith a few hacks.\n\nSo no, generation numbers are not at all redundant. They are\nfundamental. It's why we had this discussion six years ago.\n\n> And so that seems a bit hack-ish to me.\n\nUm? If you feel that way, then why the hell are you pushing your EVEN\nMORE HACKISH CACHE PATCHES?\n\nThat's what this really boils down to. I think that if we have a value\nthat we need, then it should be recorded. In the data structures. Not\nin some random other location that isn't part of the real git data\nstructures.\n\nWe don't do caches in git, because we don't NEED to. Sure, gitk has\nit's hacky cache, but that's not core functionality.\n\nI think it's a sign of good design that we can do a \"find .git\" and\nexplain every single file, and show that it's all core functionality\n(again, with the exception of \"gitk.cache\", and I suspect that's\nbecause gitk is a script, not because of any really fundamental data\nissues), and explain it.\n\nI think the *cache* is a hell of a lot more hacky than just doing it right.\n\n> I liken it somewhat to the \"don't store renames\" debate.\n\nThat's total and utter bullshit.\n\nStoring renames is *wrong*. I've explained a million times why it's\nwrong. Doing it is a disaster. I know. I've used systems that did it.\nIt's crap. It's fundamentally information that is actively misleading\nand WRONG. It's not even that you can do rename detection at run-time,\nit's that you *HAVE* to do rename detection at run-time, because doing\nit at commit time is simply utterly and fundamentally *wrong*.\n\nJust look at \"git blame -C\" to remind yourself why rename information is wrong.\n\nBut even more importantly, look at git merges. Look at how git has\ngotten merging right since pretty much day #1, and has absolutely no\nissues with files that got generated two different ways. Look at every\nSCM that tries to do rename detection, and look at how THEY CANNOT DO\nMERGES RIGHT.\n\nIt's that simple. Rename detection is not about avoiding \"redundant\ndata\". It's about doing the right thing.\n\n                          Linus\n"},{"id":"171393","messageId":"186BDF84-7AE3-4E0F-8F6D-AA89A60C972C@adacore.com","threadId":"27818","inReplyTo":"CA+55aFyDzr+SfgSzWMr9pQuQUXTw9mcjZ-00NZof74PKZzbGPA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2011-07-15T02:41:57Z","receivedAt":"2011-07-15T02:41:57Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\nOn Jul 14, 2011, at 21:19, Linus Torvalds wrote:\n> But dammit, if you start using generation numbers, then they *are*\n> required information. The fact that you then hide them in some\n> unarchitected random file doesn't change anything! It just makes it\n> ugly and random, for chrissake!\n\nGeneration numbers never will be required information, because we\ncan always compute them. These numbers are really much more similar \nto other pack index information than anything else.\n\n<aside>\nSometimes I wish we'd have general \"depth\" information for each\nSHA1, which would be the maximum number of steps in the DAG to reach\na leaf. This way, if we want to do something like \"git log\ndrivers/net/slip.c\", we don't have to bother reading the majority\nof trees that have a depth less than two. The depth can also be used\nas a limiter for \"contains\" operations, where we want to see if\ncommit X contains commit Y: depth (X) has to be at least depth (Y).\n\nHowever, any such notion, wether generation or depth or whatever\nelse we'll think of tomorrow, is something particular to a certain\nimplementation of git. It does not add anything to the information\nwe stored.\n</aside>\n\nI don't think my commit should have a different SHA1 from yours,\nbecause your tree has a more generation numbers than mine.\n\nThe beauty and genius of GIT is that it just takes the minimum\namount of data needed to uniquely identify the information to be\nstored, and stores that in a UNIQUE format. By allowing generation\nnumbers to either be present or absent, that's all broken.\n\nIt's like computing the SHA1 of compressed data: it doesn't depend\non the data we store, just about the particular representation we\nchoose. Fortunately we have done away with the first mistake.\n\nSo, if you're going to add generation numbers, there has to be a\nflag day, after which generation numbers are required everywhere. \nOf course it would be possible to recognize \"old style\" commits \nand convert them on the fly, but that is true for pretty much \nany format change. However, adding redundant information seems \nlike a poor excuse for having a flag day.\n\nStoring generation data in pack indices on the other hand makes\nperfect sense: when we generate these indices, we do complete\ntraversals and have all required information trivially at hand.  We\ncan never have that many loose objects, so lack of generation\ninformation there isn't a big deal. By storing generation information\nin the index, we can be sure it is consistent with the data contained\nin the pack, so there are no cache invalidation issues.\n\nI know I must have missed some stupid and obvious reason why\nthis is all wrong, I just don't quite see it yet.\n\n  -Geert\n"},{"id":"171397","messageId":"20110715074656.GA31301@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFyDzr+SfgSzWMr9pQuQUXTw9mcjZ-00NZof74PKZzbGPA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-15T07:46:56Z","receivedAt":"2011-07-15T07:46:56Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 14, 2011 at 06:19:30PM -0700, Linus Torvalds wrote:\n\n> Yes.\n> \n> And until I saw the patches trying to add generation numbers, I didn't\n> really try to push adding generation numbers to commits (although it\n> actually came up as early as July 2005, so the \"let's use generation\n> numbers in commits\" thing is *really* old).\n> \n> In other words, I do agree that we should strive for minimal required\n> information.\n> \n> But dammit, if you start using generation numbers, then they *are*\n> required information. The fact that you then hide them in some\n> unarchitected random file doesn't change anything! It just makes it\n> ugly and random, for chrissake!\n\nSo you don't see a difference between storing the information directly\nin the commit object, where it affects the sha1 of the commit, and\ncalculating and storing it somewhere else? That is what seems ungit to\nme. You aren't adding new information to the DAG (note I said \"DAG\" and\nnot commit) that is not already there, but you are changing the ids of\ncommits in the DAG.\n\nI'm not saying that's a reason to ultimately reject the idea of putting\ngeneration numbers in commit objects. But it is a reason to give us\npause and figure out if there are other solutions, because it will be\nthe first time such redundant information has been added. And that's\nwhat I've been trying to do during this discussion with you: work out\nwhat the options are and evaluate them.\n\n> I really don't understand your logic that says that the cache is\n> somehow cleaner. It's a random hack! It's saying \"we don't have it in\n> the main data structure, so let's add it to some other one instead,\n> and now we have a consistency and cache generation problem instead\".\n\nAre packfiles unclean, or a random hack? How about pack indices? What\nabout Nico's and Shawn's ideas for a packv4 that would gain efficiency\nby storing objects not in their whole format, but in a way that would\nmake tree examination faster (but would be able to restore the whole\nobjects byte for byte)?\n\nThose things rely on the idea that the git DAG is a data model that we\npresent to the user, but that we're allowed to do things behind the\nscenes to make things faster. We're allowed to make an index of offsets\nof objects in the packfile for faster lookup. Why are we not allowed to\nuse an index for other object data if it will speed up our local\nalgorithms?\n\nAgain, I'm not saying that the patches I posted are necessarily the\nanswer. Maybe my cache implementation sucks. Maybe the value should go\ninto a pack index instead. Maybe the whole idea is stupid. But I don't\nthink it's worth rejecting out-of-hand the idea that the generation\nnumber might be stored outside of the commit object. I do think it's\nworth talking about what the actual downsides are, as compared to other\noptions.\n\nFor example, you mentioned there a consistency problem in the paragraph\nabove. What is it?\n\nIf you mean the problem with refs/replace, then yes, that is an open\nproblem to be solved (though not a hard one, as I already mentioned a\nsolution elsewhere). But is that problem better or worse with this\nsolution versus an embedded generation number? It seems to me that an\nembedded generation number is even worse.\n\n> Just look at the size of the patches in question. Your caching patches\n> are bigger and more complicated. Sure, part of it is that your series\n> adds the code to _use_ the generation number, but look purely at the\n> code to maintain them.\n\nIt's 300 lines of code. That can also be used to store arbitrary\nmeta-information for commits. I've already achieved significant speedups\nin some workflows by caching patch-id calculations, which would reuse\nthis code.  And I certainly don't think _that_ should go into the commit\nobject.\n\nYes, it's more complex than simply adding a generation number to the\ncommit header. But simply adding a generation number does not actually\ngive the 100-fold speedup I'm seeing. So again, I'm not interested in\nrejecting solutions out of hand; I'm interested in things like: is the\ncomplexity of the cache worth this speedup? What other options do we\nhave, and what speedup do they provide? Do we care enough about this\nspeedup to even bother?\n\n> Why do you think the odd separate cache is somehow better than just\n> doing it right? Seriously? If we require the generation numbers, then\n> they have *become* that minimal information that we should save!\n\nWhat do you do when generation numbers don't match the DAG represented\nby the parent pointers? Are you proposing to just ignore it? I'm not\nasking that question adversarially; ignoring may be the sane thing to\ndo, and we say \"generation numbers are to be trusted, even if they don't\nmatch parent pointers\".\n\n> Now, six years later, you clearly are saying that we need the\n> generation numbers, but then you go off and try to say that they\n> should be in some secondary non-architected random collection of data\n> structures that isn't covered by the security and maintenance\n> guarantees that the core git objects are.\n\nI don't think I said we clearly need them. I said we can get speedups by\nusing them, and I showed some patches. I _also_ posted patches showing\nhow to accomplish similar speedups using timestamps. Note that all of my\npatches started with \"RFC\". I am trying to figure out which is the best\nway to proceed.\n\nAnd why _would_ they need to be covered by the security and maintenance\nguarantees of core objects? You can trivially calculate them from the\ncore objects. Are pack indices also a \"secondary non-architected random\ncollection of data structures\"?\n\n> Dammit, one of the things that makes git special is that the data\n> structures are NOT random odd ad-hoc files. There is a design to them.\n\nThere is just as much documentation and design for the new file format I\nadded as there is for pack indices (in fact, they're quite similar in\ndesign). I really see them at the same level: something we calculate to\nspeed up some algorithms, but something we could regenerate at any time\nif we felt like.\n\n> > And so that seems a bit hack-ish to me.\n> \n> Um? If you feel that way, then why the hell are you pushing your EVEN\n> MORE HACKISH CACHE PATCHES?\n\nPlease, there is really no need to shout. And I find it quite silly that\nyou would refer to me as \"pushing\" these patches when they have been\nclearly listed as RFC, and everything I have posted in the nearby\nthreads has been about comparing different strategies (with patches and\ntimings for some of those other strategies!).\n\n> We don't do caches in git, because we don't NEED to. Sure, gitk has\n> it's hacky cache, but that's not core functionality.\n\nI'm sorry to tell you that there is already a cache for external\nconversion of diffs for blobs. And that I have a patch series which\nmakes \"git cherry\" much more pleasant to use by caching patch ids.\n\nDo we \"need\" those? No, of course not. Git works just fine without them,\nalbeit a bit slower. But is it sometimes worth making a space-time\ntradeoff to make some algorithms faster? I think it sometimes is,\ndepending on the space and time factors, and the complexity of the\nstorage (e.g., consistency problems with caching).\n\n> I think it's a sign of good design that we can do a \"find .git\" and\n> explain every single file, and show that it's all core functionality\n> (again, with the exception of \"gitk.cache\", and I suspect that's\n> because gitk is a script, not because of any really fundamental data\n> issues), and explain it.\n\nWould it make you happier if we stored the generation data in the pack\nindex when we index the packs?\n\n> I think the *cache* is a hell of a lot more hacky than just doing it right.\n\nYou still haven't explained how we would \"do it right\" and get the same\nspeedups. When I responded to your initial email, your answers were\nalong the lines of \"we could cache fewer things\". If your position is to\ndamn the speedup, the cache is not worth the complexity, I can buy that.\nIf your position is that the complexity is not worth it, and we are\nbetter off to keep using timestamps, I can buy that. If your position is\nthat you can find a clever way, using only the generation numbers in\nnewly created commits, to get similar speedups in \"git {tag,branch}\n--contains\", I'd love to hear it.\n\n> > I liken it somewhat to the \"don't store renames\" debate.\n> \n> That's total and utter bullshit.\n> \n> Storing renames is *wrong*. I've explained a million times why it's\n> wrong. Doing it is a disaster. I know. I've used systems that did it.\n> It's crap. It's fundamentally information that is actively misleading\n> and WRONG. It's not even that you can do rename detection at run-time,\n> it's that you *HAVE* to do rename detection at run-time, because doing\n> it at commit time is simply utterly and fundamentally *wrong*.\n\nYes, I am well aware that stored renames are wrong for merging. The\nproblem is that they not a function of a tree state (which is what a\ncommit stores), but rather of the difference between two states. So when\nyou diff the commit's state with some other arbitrary merge-base, any\nrenames recorded at commit time would be worthless.\n\nBut consider another case. Each time I run \"git log -M --raw\", I compute\nthe same renames over and over. Let's say I have a case in which this is\nannoyingly slow, and want to speed it up. The state of a particular\ncommit and the state of its parents are invariants for a particular sha1\ncommit id; this is a fundamental property of git, as you well know. So\nfor a given rename-detection algorithm (and any parameters it has), the\nset of renames between the states will also be an invariant.\n\nNow imagine I create a persistent cache mapping the commit sha1 for some\nsane default set of rename algorithm parameters to a set of rename\npairs. My annoyingly slow \"log -M\" is now faster, and I'm happier.\n\nI think you encounter a similar set of questions here as you do with the\nconcept of a generation header. If the information is an invariant for a\nparticular commit sha1, can we and should we store it in the commit\nobject? Is the speedup worth the complexity of a cache? What are the\ncircumstances under which the cache is not applicable, and how often do\nthey come up? Can we accurately detect when the cache is not applicable?\n\nAnd that is why I compared it to the idea of storing renames. Please\nnote that I did _not_ say they were exactly the same situation, or that\nthe answers to one set of questions were the same as the answers to\nanother.\n\n-Peff\n"},{"id":"171399","messageId":"m3fwm7aox1.fsf@localhost.localdomain","threadId":"27818","inReplyTo":"CA+55aFyDzr+SfgSzWMr9pQuQUXTw9mcjZ-00NZof74PKZzbGPA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-07-15T09:12:43Z","receivedAt":"2011-07-15T09:12:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n> On Thu, Jul 14, 2011 at 1:31 PM, Jeff King <peff@peff.net> wrote:\n> >\n> > However, I'm not 100% convinced leaving generation numbers out was a\n> > mistake. The git philosophy seems always to have been to keep the\n> > minimal required information in the DAG.\n> \n> Yes.\n> \n> And until I saw the patches trying to add generation numbers, I didn't\n> really try to push adding generation numbers to commits (although it\n> actually came up as early as July 2005, so the \"let's use generation\n> numbers in commits\" thing is *really* old).\n> \n> In other words, I do agree that we should strive for minimal required\n> information.\n> \n> But dammit, if you start using generation numbers, then they *are*\n> required information. The fact that you then hide them in some\n> unarchitected random file doesn't change anything! It just makes it\n> ugly and random, for chrissake!\n> \n> I really don't understand your logic that says that the cache is\n> somehow cleaner. It's a random hack! It's saying \"we don't have it in\n> the main data structure, so let's add it to some other one instead,\n> and now we have a consistency and cache generation problem instead\".\n\nYou store redundant information, one that is used to speed up\ncalculations, in a cache.\n \n[...]\n> > Generation numbers are _completely_ redundant with the actual structure\n> > of history represented by the parent pointers.\n\nWhat is more important the perceived structure of history can change\nby three mechanisms:\n\n * grafts\n * replace objects\n * shallow clone\n\nI can understand that you don't want to worry about grafts - they are\na terrible hack.  We can simply turn off using generation numbers\nstored in commit if they are present.\n\nThe problem with shallow clones is only at beginning, when some of\ncommits in shallow repository does not have generation numbers.  You\ncannot simply calculate generation number for a new commit in such\ncase.\n\nBut what about REPLACE OBJECTS?  If one for example use \"git replace\"\non root commit to join contemporary repository with historical\nrepository... this is not addressed in your emails.\n\n\nAnd let's not forget the fact that we need cache for old commits which\ndon't have yet generation number in a commit.\n\n\nBTW. you are not fair comparing size of code.  \n\nFirst, some of Peff code is about _using_ generation numbers, which\nwill be needed regardless of whether generation numbers are stored in\ncache or packfile index, or whether they are embedded in commit\nobjects.\n\nSecond, with generation number commit header you need to write fsck\ncode, and have to consider size of this yet-to-be-written code.\n\n[...]\n> > I liken it somewhat to the \"don't store renames\" debate.\n> \n> That's total and utter bullshit.\n\nI think Peff meant here that if you make mistakes in calculating\nrename info or generation number, and have incorrect information\nstored in commit object, you are f**ked.\n \n> Storing renames is *wrong*. I've explained a million times why it's\n> wrong. Doing it is a disaster. I know. I've used systems that did it.\n> It's crap. It's fundamentally information that is actively misleading\n> and WRONG. It's not even that you can do rename detection at run-time,\n> it's that you *HAVE* to do rename detection at run-time, because doing\n> it at commit time is simply utterly and fundamentally *wrong*.\n> \n> Just look at \"git blame -C\" to remind yourself why rename information is wrong.\n\nAlso doing full code movement and copying detection (that is what \"git\nblame -C\" does) rather than simplistic whole-file rename detection is\npretty much impossible at commit time.\n\nNb. most SCMs that use path-id based rename tracking require that user\nexplicitly marks renames using \"scm move\" or \"scm rename\" (well,\nMercurial has a tool for rename detection before commit, \"hg\naddremove\").  But asking user to mark code movements is simply\ninfeasible.\n \n> But even more importantly, look at git merges. Look at how git has\n> gotten merging right since pretty much day #1, and has absolutely no\n> issues with files that got generated two different ways. Look at every\n> SCM that tries to do rename detection, and look at how THEY CANNOT DO\n> MERGES RIGHT.\n> \n> It's that simple. Rename detection is not about avoiding \"redundant\n> data\". It's about doing the right thing.\n\nWell, rename tracking supporters say that heuristic rename detection\ncan be wrong.\n\n\nBy the way, what happened to \"wholesame directory rename detection\"\npatches?  Without them in the situation where one side renamed\ndirectory, and other created new file in said directory git on merge\ncreates file in re-created old name of directory...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"171400","messageId":"CANfMb_-ZxGGzpKDnhG46HK+DZ1UN+_kxccKuSrZtO41N0EFy6Q@mail.gmail.com","threadId":"27818","inReplyTo":"m3fwm7aox1.fsf@localhost.localdomain","subject":"Re: Git commit generation numbers","fromName":"Long, Martin","fromEmail":"martin@longhome.co.uk","sentAt":"2011-07-15T09:17:03Z","receivedAt":"2011-07-15T09:17:03Z","isPatch":false,"sender":{"key":"martin@longhome.co.uk","avatar":"https://gravatar.com/avatar/24daa6058f5db96c0590b65c055bc115a5f4a9ee9dcb67a8b9cce6c6469eac85?d=mp&s=160"},"body":"I strongly agree with Linus that the cache should not form part of the\nsolution to this problem, but could maybe be a later add-on, which\nimproved performance.\n\nThere is a possible improvement, which may remove the need for the\ncache. It doesn't solve the issue of broken numbers, but I think the\nkey to that is just to ensure the traversal algorithm is\ndeterministic, stable, and immutable.\n\nFirstly, I presume the generation number would not form part of the\nSHA1 calculation? No? Cool.\n\nWhen calculating a generation number by doing a traversal, would it\nnot be possible to update some, or all, commit objects touched, with\ntheir generation numbers. Again, this would be expensive, but there\nwould possibly be even quicker gains than Linus's original proposal to\njust add numbers to the new commit.\n\nA compromise might be to only update some commits - notably those with\n2 or more parents, so that both parents don't need to be traversed,\nand possibly every nth commit (to give regular checkpoints that can be\nutilised when traversing a branch). I would suggest commits with 2\nchildren for the latter, but with my limited knowledge of the\nimplementation, I understand that Is more difficult to find.\nObviously, these numbers would only be pegged locally, and wouldn't by\nsynced on push, as they already exist on the far end. However, it\ncould be possible to run a process on a bare repo to shoot through and\npeg commits, then at least new clones will be \"well pegged\"\n\nMartin Long\nUK\n"},{"id":"171404","messageId":"CANfMb_-cfAWBECGcUqQA3JCObRF+dSsx_Z2iCigYeKMdh7J7Zg@mail.gmail.com","threadId":"27818","inReplyTo":"CANfMb_-ZxGGzpKDnhG46HK+DZ1UN+_kxccKuSrZtO41N0EFy6Q@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Long, Martin","fromEmail":"martin@longhome.co.uk","sentAt":"2011-07-15T15:33:31Z","receivedAt":"2011-07-15T15:33:31Z","isPatch":false,"sender":{"key":"martin@longhome.co.uk","avatar":"https://gravatar.com/avatar/24daa6058f5db96c0590b65c055bc115a5f4a9ee9dcb67a8b9cce6c6469eac85?d=mp&s=160"},"body":">\n> Firstly, I presume the generation number would not form part of the\n> SHA1 calculation? No? Cool.\n\nI suspect this may be where my suggestion falls down. Though I suspect\nthere is a case for object metadata which doesn't form part of the\nSHA. Would generation number tampering be a concern?\n\nCaching offers the ability to store that metadata, to provide the same\nperformance gain, but maintain the integrity of the SHA chain.\nHowever, it does still leave the generation number liable to\ntampering, meaning a generic non-SHA metadata solution might be\nbetter.\n\nTBH, there are few situations where historical generations are useful\n- finding gen numbers of tags is one of them. Most cases are going to\nbe for new commits, and in that case, a few new commits at the tip of\neach branch will very quickly reduce the number of traversals. What\nuse case would really create enough traversals that it should be a\nperformance concern?\n"},{"id":"171407","messageId":"CA+55aFzS3KDNvKt-dXvYpuAQwFwD3+GCj8y8bRQCycPvrynT8Q@mail.gmail.com","threadId":"27818","inReplyTo":"20110715074656.GA31301@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T16:10:48Z","receivedAt":"2011-07-15T16:10:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 12:46 AM, Jeff King <peff@peff.net> wrote:\n>\n> So you don't see a difference between storing the information directly\n> in the commit object, where it affects the sha1 of the commit, and\n> calculating and storing it somewhere else?\n\nSure, I see the difference. And I think it's uglier to have two\ndifferent places for required information.\n\n>  That is what seems ungit to\n> me. You aren't adding new information to the DAG (note I said \"DAG\" and\n> not commit) that is not already there, but you are changing the ids of\n> commits in the DAG.\n\nUmm. It's redundant, but so what? We have tons of redundant\ninformation in there already. Those commits are very explicitly using\na 40-byte ASCII representation of the 20-byte SHA1 names. The very\noriginal deeper object structure is also redundant: we repeat the\nobject size in the object itself, even though it's part of the\nimplicit object format itself.\n\nWe also very purposefully repeat the type of the object there, even\nthough the type is basically always redundant (in fact, the core git\nfunctions require you to give the type of the object as part of the\nlookup, and will error out if the SHA1 points to the wrong type). That\nwas one of my original design decisions, exactly because I wanted the\nredundancy for verification.\n\nRedundancy isn't a problem. It's a source of sanity checking.\n\nI'm not seeing why you are harping on it.\n\nI think it's much worse to have the same information in two different\nplaces where it can cause inconsistencies that are hard to see and may\nnot be repeatable. If git ever finds the wrong merge base (because,\nsay, the generation numbers are wrong), I want it to be a *repeatable*\nthing. I want to be able to repeat on the git mailing list \"hey, guys,\nlook at what happens when I try to merge commits ABC and XYZ\". If you\ngo \"yeah, it works for me\", then that is bad.\n\nWhat I tried very hard to do in the git data structures is to make\nthem (a) immutable (so the DAG could never have two-way links, for\nexample) and (b) \"simple\".\n\nRight now, we do *have* a \"generation number\". It's just that it's\nvery easy to corrupt even by mistake. It's called \"committer date\". We\ncould improve on it.\n\n> Are packfiles unclean, or a random hack? How about pack indices?\n\nNo. Neither of them are unclean or random. The original git design was\nvery much about thinking of the object space as a \"filesystem\". Now,\nthe original object layout actually used the native OS filesystem, and\nI naively thought that would be ok. Using  aspecialized filesystem\ninstead doesn't really change anything. It's not fundamentally\ndifferent from the difference between running git on ext3 or btrfs or\nnfs or whatever. In fact, I think we've had more filesystem-related\nbugs wrt NFS than we've had with pack-files.\n\nThe pack indices are actually kind of ugly - and I would have\npreferred having them in the same file instead of having the worry of\nconsistency across two different files. They *are* the kind of thing\nthat could cause local inconsistency, but they are fairly simple, and\nthey have some serious protection in them (ie they aren't just SHA1'd\nin themselves, they contain a SHA1 of the pack-file they index in them\nto make sure that any inconsistency is findable). Again, that's\n\"redundancy\". But I consider the packfile/index to be just a\nfilesystem. It really fundamentally *is* that.\n\nPartly for that reason, I do think that if the generation count was\nembedded in the pack-file, that would not be an \"ugly\" decision. The\npack-files have definitely become \"core git data structures\", and are\nmore than just a local filesystem representation of the objects:\nthey're obviously also the data transport method, even if the rules\nthere are slightly different (no index, thank god, and incomplete\n\"thin\" packs).\n\nThat said, I don't think a generation count necessarily \"fits\" in the\npack-file. They are designed to be incremental, so it's not very\nnatural there. But I do think it would be conceptually prettier to\nhave the \"depth of commit\" be part of the \"filesystem\" data than to\nhave it as a separate ad-hoc cache.\n\n> Those things rely on the idea that the git DAG is a data model that we\n> present to the user, but that we're allowed to do things behind the\n> scenes to make things faster.\n\n.. and that is relevant to this discussion exactly *how*?\n\nIt's not. It's totally irrelevant. I certainly would never walk away\nfrom the DAG model. It's a fundamental git decision, and it's the\ncorrect one.\n\nBut it all boils down to one simple issue: we should have added\ngeneration counts back in 2005. It's likely the *one* data format\ndecision that I regret. Using commit dates was wrong. Everybody knew\nit was wrong, but we ended up going with it just to keep the format\nconstant.\n\nIf I had realized how small the patch was to add generation counters,\nand that it wouldn't have broken backwards compatibility (ie fsck\ndoesn't start complaining). I would have done it originally, instead\nof all the crazy hacks we did for commit date verification.\n\nAnd that is what this discussion fundamentally boils down to for me.\n\nIf we should have fixed it in the original specification, we damn well\nshould fix it today. It's been \"ignorable\" because it's just not been\nimportant enough. But if git now adds a fundamental cache for them,\nthen that information is clearly no longer \"not important enough\".\n\n                               Linus\n"},{"id":"171408","messageId":"1310746531.19224.32.camel@drew-northup.unet.maine.edu","threadId":"27818","inReplyTo":"CANfMb_-cfAWBECGcUqQA3JCObRF+dSsx_Z2iCigYeKMdh7J7Zg@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Drew Northup","fromEmail":"drew.northup@maine.edu","sentAt":"2011-07-15T16:15:31Z","receivedAt":"2011-07-15T16:15:31Z","isPatch":false,"sender":{"key":"drew.northup@maine.edu","avatar":"https://avatars.githubusercontent.com/u/18331571?v=4"},"body":"\nOn Fri, 2011-07-15 at 16:33 +0100, Long, Martin wrote:\n> >\n> > Firstly, I presume the generation number would not form part of the\n> > SHA1 calculation? No? Cool.\n> \n> I suspect this may be where my suggestion falls down. Though I suspect\n> there is a case for object metadata which doesn't form part of the\n> SHA. Would generation number tampering be a concern?\n\nIf you take Jeff's perspective on the purpose of generation numbers\n(representing metadata about the DAG in a more readily-available format)\nthen \"tampering\" is not really a concern as the metadata is merely local\n(to the running instance of Git) ephemera that we can cache between runs\nfor the sake of efficiency. Linus' perspective on generation numbers\nseems to be of a more hard and fast type of data.\n\nSo, are we really talking about [corpus] generation numbers (used to\ndescribe the state of the DAG in the way one describes his known family\ntree) or are we talking about _revision_numbers_ (used to describe the\ncommit, as Subversion does)? I think we've got two (or more) groups\ntalking about different things (and aims) and trying to use the same\nwords to do so.\n\n> Caching offers the ability to store that metadata, to provide the same\n> performance gain, but maintain the integrity of the SHA chain.\n> However, it does still leave the generation number liable to\n> tampering, meaning a generic non-SHA metadata solution might be\n> better.\n\nI'm not sure where you are going with this. I wouldn't think \"tampering\"\nwith _current_DAG-based ephemera would do much other than create a\nperformance hit. If you are really talking about a static\n_revision_number_ then that belongs in the commit, where it cannot be\nchanged (and may be completely meaningless when taken out of context, as\nSVN revision numbers are). What such a number may entail is probably up\nfor discussion, but perhaps in a different thread.\n\n> TBH, there are few situations where historical generations are useful\n> - finding gen numbers of tags is one of them. Most cases are going to\n> be for new commits, and in that case, a few new commits at the tip of\n> each branch will very quickly reduce the number of traversals. What\n> use case would really create enough traversals that it should be a\n> performance concern?\n\nThe answer to this is found in a previous thread\nhttp://article.gmane.org/gmane.comp.version-control.git/176807\n\n(remember, generation number vs. revision number...)\n\nAlso, please don't cull the CC list! (Added Geert Bosch)\n\n-- \n-Drew Northup\n________________________________________________\n\"As opposed to vegetable or mineral error?\"\n-John Pescatore, SANS NewsBites Vol. 12 Num. 59\n"},{"id":"171409","messageId":"CAJo=hJtuxNLhSjn_sDJxG7xu5k2wbJ_QLf_n+Z1E=o2AndAuJQ@mail.gmail.com","threadId":"27818","inReplyTo":"CA+55aFzS3KDNvKt-dXvYpuAQwFwD3+GCj8y8bRQCycPvrynT8Q@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-07-15T16:18:46Z","receivedAt":"2011-07-15T16:18:46Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Fri, Jul 15, 2011 at 09:10, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> Right now, we do *have* a \"generation number\". It's just that it's\n> very easy to corrupt even by mistake. It's called \"committer date\". We\n> could improve on it.\n...\n> If I had realized how small the patch was to add generation counters,\n> and that it wouldn't have broken backwards compatibility (ie fsck\n> doesn't start complaining). I would have done it originally, instead\n> of all the crazy hacks we did for commit date verification.\n\nWhat about going forward making the requirement that a new commit must\nhave a committer date whose date is >= the maximum date of its\nparents?\n\nWe could also add a check during fast-forward merges to refuse to\nperform the merge if the incoming commit has a committer date too far\nforward in the future (e.g. more than 5 minutes). If you pull from a\nmoron whose system clock is set such that the committer date isn't a\nproxy for generation number, Git would just refuse the merge, and you\ncould ask them to fix their objects.\n\n-- \nShawn.\n"},{"id":"171410","messageId":"CA+55aFw_XjWm+4XwsN6CRJnsrcEu5YEChOHSHN51UUBN6PynWw@mail.gmail.com","threadId":"27818","inReplyTo":"CAJo=hJtuxNLhSjn_sDJxG7xu5k2wbJ_QLf_n+Z1E=o2AndAuJQ@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T16:44:21Z","receivedAt":"2011-07-15T16:44:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 9:18 AM, Shawn Pearce <spearce@spearce.org> wrote:\n>\n> What about going forward making the requirement that a new commit must\n> have a committer date whose date is >= the maximum date of its\n> parents?\n\nSo you suggest just making commit dates be the generation number.\n\nI'd be ok with that. It's basically what we've been doing for the last\nsix years.\n\nBut in that case, we shouldn't be doing the generation count cache either.\n\nBtw, I do agree that we probably should add a warning for the case\n(\"your clock is wrong - your commit date is before the commit date of\nyour parents\") and maybe require the use of \"-f\" or something to\noverride it. That would certainly be a good thing quite independently\nof anything else. So regardless of generation counts, it's probably\nworth it.\n\nBut if you think commit date is good enough for generation counts -\nand I'm not arguing against it - then please tell me why you would\nthen want to have a separate generation count cache.\n\nSo I would like to repeat: I think our commit-date based hack has been\npretty successful. We've lived with it for years and years. Even the\n\"let's try to fix it by adding slop\" code is from three years ago\n(commit 7d004199d1), which means that for three years we never really\nsaw any serious problems. I forget what problem we actually did see -\nI have this dim memory of it being Ted that had problems with a merge\nbecause git picked a crap merge base, but that may just be my\nAlzheimer's speaking.\n\nObviously there are cases where we miss some merge base and it doesn't\nreally end up mattering, so we may well have a *ton* of commits that\nhave bad dates, but they just haven't affected us enough for us to\ncare. That's fine too - I dislike how our algorithm isn't truly\nreliable, but at the same time I think we're so robust that it all\nworks regardless.\n\nSo I think it's ugly and fairly hacky, but it has worked well enough\nin practice. I dislike our commit dates, but I don't _hate_ them. I do\nthink it was a mistake, but not one I'm especially ashamed of.\n\nSo why do I dislike the generation count cache so much? I dislike it\nexactly because\n\n  \"if the commit date isn't good enough, then dammit, we should have\njust added a generation count\".\n\nAnd if we should have added it six years ago, then we should add it\ntoday. Not say \"oh, we made a mistake six years ago, let's work around\nthe mistake instead of fixing it\".\n\nThat's really what it boils down to. Let's not paper over a mistake.\nEither we need the generation depth or we don't. And if we do need it,\nwe should replace the date-based hackery with it (where \"replace\" may\nwell be \"still fall back on our traditional date-based hackery in the\nabsense of generation counters\").\n\nBut if we decide that we don't really need generation counters AT ALL,\nand can just continue with the commit date hack, then I'm personally\nok with that too.\n\nSo to me, it's a \"either or\" situation. Either the commit dates are\ngood enough, or we should add generation counts to the commits.\n\nBut in *neither* case is it ok to do some external cache to work around it.\n\n                          Linus\n"},{"id":"171414","messageId":"20110715184211.GH8453@thunk.org","threadId":"27818","inReplyTo":"CA+55aFw_XjWm+4XwsN6CRJnsrcEu5YEChOHSHN51UUBN6PynWw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Ted Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2011-07-15T18:42:11Z","receivedAt":"2011-07-15T18:42:11Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Jul 15, 2011 at 09:44:21AM -0700, Linus Torvalds wrote:\n> So I would like to repeat: I think our commit-date based hack has been\n> pretty successful. We've lived with it for years and years. Even the\n> \"let's try to fix it by adding slop\" code is from three years ago\n> (commit 7d004199d1), which means that for three years we never really\n> saw any serious problems. I forget what problem we actually did see -\n> I have this dim memory of it being Ted that had problems with a merge\n> because git picked a crap merge base, but that may just be my\n> Alzheimer's speaking.\n\nMy original main issue was simply that \"git tag --contains\" and \"git\nbranch --contains\" was either (a) incorrect, or (b) slower than\npopping up gitk and pulling the information out of the GUI.  The\nreason for (b) is because of gitk.cache.\n\nMaybe the answer then is creating a command-line tool (it doesn't have to\nbe in \"core\" of git) which just pulls the dammned information out of\ngitk.cache....\n\n(Yes, it's gross, but I'm not worrying about the long-term\narchitecture of git or anything high-falutin' like that.  I'm just a\npoor dumb user who just wants git tag --contains and git branch\n--contains to be fast and accurate...)\n\n\t\t\t\t\t\t- Ted\n"},{"id":"171415","messageId":"CA+8MBbJNZdkpOhA5Kke0VqUA9qCFdzfEP5cWPTMF3eUfDsGRiQ@mail.gmail.com","threadId":"27818","inReplyTo":"CA+55aFw_XjWm+4XwsN6CRJnsrcEu5YEChOHSHN51UUBN6PynWw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Tony Luck","fromEmail":"tony.luck@intel.com","sentAt":"2011-07-15T18:46:38Z","receivedAt":"2011-07-15T18:46:38Z","isPatch":false,"sender":{"key":"tony.luck@intel.com","avatar":"https://avatars.githubusercontent.com/u/5446021?v=4"},"body":"On Fri, Jul 15, 2011 at 9:44 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> Btw, I do agree that we probably should add a warning for the case\n> (\"your clock is wrong - your commit date is before the commit date of\n> your parents\") and maybe require the use of \"-f\" or something to\n> override it. That would certainly be a good thing quite independently\n> of anything else. So regardless of generation counts, it's probably\n> worth it.\n\nWhat if my clock is wrong in the opposite direction - set to some time\nout in 2025.\nIt would pass the check you propose and let the commit go in - but would\ncause problems for everyone if that tree was pulled into upstream.\n\nYou'd also want a check in pull(merge) that none of the commits being\nadded were in the future (as defined by the time on your machine).\n\n-Tony\n"},{"id":"171417","messageId":"CA+55aFzeN9jdzU1RBVZ6UDGUT0YvNkzpycKG_mT-z+dt_3GnMw@mail.gmail.com","threadId":"27818","inReplyTo":"CA+8MBbJNZdkpOhA5Kke0VqUA9qCFdzfEP5cWPTMF3eUfDsGRiQ@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T18:58:09Z","receivedAt":"2011-07-15T18:58:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 11:46 AM, Tony Luck <tony.luck@intel.com> wrote:\n>\n> What if my clock is wrong in the opposite direction - set to some time\n> out in 2025.\n> It would pass the check you propose and let the commit go in - but would\n> cause problems for everyone if that tree was pulled into upstream.\n\nI think Shawn suggested that we just notice it at merge time.\n\nBut yes, it's why (a) I'd suggest we have a \"-f\" to override and (b) I\ndo think that generation counts are a better idea. You could still\nscrew them up, but it would be due to an outright bug or malicious\nbehavior, rather than simple incompetence on the part of a user.\n\nIncompetent users (where \"date on the machine set to the wrong\ncentury\" is just _one_ sign of incompetence) are something git should\npretty much take for granted. It may not be the common case, but it's\ncertainly something we should design for and take into account.\n\nIn contrast, if somebody *wants* to screw his repository up by\nre-writing objects with \"git hash-object\" etc, be my guest. We should\njust make sure fsck catches anything serious.\n\nSo I would suggest checking the date regardless of any generation\ncount issues, because it would possibly find badly configured machines\nthat should be fixed. The same way we complain when we find no name.\n\nWhether it should then be a correctness issue or not is kind of separate.\n\n> You'd also want a check in pull(merge) that none of the commits being\n> added were in the future (as defined by the time on your machine).\n\nI don't think you need to care about \"none of the commits\", just\nmaking sure the tip is reasonable. That would not only be expensive,\nand not what we normally do (we show the diff against endpoints, not\nall changes, etc). It would also cause problems for \"fixed\"\nrepositories (ie anything that has historical dates that are wrong,\nbut are ok now).\n\n                  Linus\n"},{"id":"171418","messageId":"CA+55aFyWTxE0Pc=+i7ZsSbATgEG+gDta6880Bq83wBk1OXTgVQ@mail.gmail.com","threadId":"27818","inReplyTo":"20110715184211.GH8453@thunk.org","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T19:00:52Z","receivedAt":"2011-07-15T19:00:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 11:42 AM, Ted Ts'o <tytso@mit.edu> wrote:\n>\n> My original main issue was simply that \"git tag --contains\" and \"git\n> branch --contains\" was either (a) incorrect, or (b) slower than\n> popping up gitk and pulling the information out of the GUI.  The\n> reason for (b) is because of gitk.cache.\n\nWith \"original issue\" I actually meant the case that caused us to add\nthe \"slop\" commit (7d004199d1). But I was too lazy to try to find the\narchives from March 2008..\n\n                  Linus\n"},{"id":"171425","messageId":"20110715194807.GA356@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFzS3KDNvKt-dXvYpuAQwFwD3+GCj8y8bRQCycPvrynT8Q@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-15T19:48:07Z","receivedAt":"2011-07-15T19:48:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jul 15, 2011 at 09:10:48AM -0700, Linus Torvalds wrote:\n\n> I think it's much worse to have the same information in two different\n> places where it can cause inconsistencies that are hard to see and may\n> not be repeatable. If git ever finds the wrong merge base (because,\n> say, the generation numbers are wrong), I want it to be a *repeatable*\n> thing. I want to be able to repeat on the git mailing list \"hey, guys,\n> look at what happens when I try to merge commits ABC and XYZ\". If you\n> go \"yeah, it works for me\", then that is bad.\n\nHaving the information in two different places is my concern, too. And I\nthink the fundamental difference between putting it inside or outside\nthe commit sha1 (where outside encompasses putting it in a cache, in the\npack-index, or whatever), is that I see the commit sha1 as somehow more\n\"definitive\". That is, it is the sole data we pass from repo to repo\nduring pushes and pulls, and it is the thing that is consistency-checked\nby hashes.\n\nSo if there is an inconsistency between what the parent pointers\nrepresent, and what the generation number in \"outside\" storage says,\nthen the outside storage is wrong, and the parent pointers are the right\nanswer. It becomes a lot more fuzzy to me if there is an inconsistency\nbetween what the parent pointers represent, and what the generation\nnumber says.\n\nHow should that situation be handled? Should fsck check for it and\ncomplain? Should we just ignore it, even though it may cause our\ntraversal algorithms to be inaccurate? Like clock skew, there's not much\nthat can be done if the commits are published.\n\nThose are serious questions that I think should be considered if we are\ngoing to put a generation header into the commit object, and I haven't\nseen answers for them yet.\n\n> Partly for that reason, I do think that if the generation count was\n> embedded in the pack-file, that would not be an \"ugly\" decision. The\n> pack-files have definitely become \"core git data structures\", and are\n> more than just a local filesystem representation of the objects:\n> they're obviously also the data transport method, even if the rules\n> there are slightly different (no index, thank god, and incomplete\n> \"thin\" packs).\n> \n> That said, I don't think a generation count necessarily \"fits\" in the\n> pack-file. They are designed to be incremental, so it's not very\n> natural there. But I do think it would be conceptually prettier to\n> have the \"depth of commit\" be part of the \"filesystem\" data than to\n> have it as a separate ad-hoc cache.\n\nSure, I would be fine with that. When you say \"packfile\", do you mean\nthe the general concept, as in it could go in the pack index as opposed\nto the packfile itself? Or specifically in the packfile? The latter\nseems a lot more problematic to me in terms of implementation.\n\n> > Those things rely on the idea that the git DAG is a data model that we\n> > present to the user, but that we're allowed to do things behind the\n> > scenes to make things faster.\n> \n> .. and that is relevant to this discussion exactly *how*?\n\nBecause keeping the generation information outside of the DAG keeps the\nmodel we present to the user simple (and not just the user; the\ninformation that we present to other programs), but lets git still use\nthe information without calculating it from scratch each time. Just like\nwe present the data as a DAG of loose objects via things like \"git\ncat-file\", even though the underlying storage inside a packfile may be\nvery different. I just don't see those two ideas as fundamentally\ndifferent.\n\n> It's not. It's totally irrelevant. I certainly would never walk away\n> from the DAG model. It's a fundamental git decision, and it's the\n> correct one.\n\nOf course not. I never suggested we should.\n\n> And that is what this discussion fundamentally boils down to for me.\n> \n> If we should have fixed it in the original specification, we damn well\n> should fix it today. It's been \"ignorable\" because it's just not been\n> important enough. But if git now adds a fundamental cache for them,\n> then that information is clearly no longer \"not important enough\".\n\nOK, so let's say we add generation headers to each commit. What happens\nnext? Are we going to convert algorithms that use timestamps to use\ncommit generations? How are we going to handle performance issues when\ndealing with older parts of history that don't have generations?\n\nAgain, those are serious questions that need answered. I respect that\nyou think the lack of a generation header is a design decision that\nshould be corrected. As I said before, I'm not 100% sure I agree, but\nnor do I completely disagree (and I think it largely boils down to a\nphilosophical distinction, which I think you will agree should take a\nbackseat to real, practical concerns). But it's not 2005, and we have a\nton of history without generation numbers. So adding them now is only\none piece of the puzzle.\n\nWhat's your solution for the rest of it?\n\n-Peff\n"},{"id":"171428","messageId":"20110715200713.GA969@sigill.intra.peff.net","threadId":"27818","inReplyTo":"20110715194807.GA356@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-15T20:07:13Z","receivedAt":"2011-07-15T20:07:13Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jul 15, 2011 at 03:48:07PM -0400, Jeff King wrote:\n\n> OK, so let's say we add generation headers to each commit. What happens\n> next? Are we going to convert algorithms that use timestamps to use\n> commit generations? How are we going to handle performance issues when\n> dealing with older parts of history that don't have generations?\n> \n> Again, those are serious questions that need answered. I respect that\n> you think the lack of a generation header is a design decision that\n> should be corrected. As I said before, I'm not 100% sure I agree, but\n> nor do I completely disagree (and I think it largely boils down to a\n> philosophical distinction, which I think you will agree should take a\n> backseat to real, practical concerns). But it's not 2005, and we have a\n> ton of history without generation numbers. So adding them now is only\n> one piece of the puzzle.\n> \n> What's your solution for the rest of it?\n\nI just read some of your later emails to others in the thread. It seems\nlike your answer is \"assume the timestamp-based limiting is good enough\nfor old history\".\n\nI'm OK with that. It obviously falls down in a few specific situations,\nbut certainly has not been an unbearable problem for the past 5 years.\n\n-Peff\n"},{"id":"171441","messageId":"CA+55aFx0KyAZRsy7gZ3Z4woWC-uWcLu11gcUrR+9MJR5NOSkrA@mail.gmail.com","threadId":"27818","inReplyTo":"20110715194807.GA356@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T21:17:26Z","receivedAt":"2011-07-15T21:17:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 12:48 PM, Jeff King <peff@peff.net> wrote:\n>\n> Having the information in two different places is my concern, too. And I\n> think the fundamental difference between putting it inside or outside\n> the commit sha1 (where outside encompasses putting it in a cache, in the\n> pack-index, or whatever), is that I see the commit sha1 as somehow more\n> \"definitive\". That is, it is the sole data we pass from repo to repo\n> during pushes and pulls, and it is the thing that is consistency-checked\n> by hashes.\n\nSure. That is also the data that is the same for everybody.\n\nThat's a big deal, in the sense that it's the only thing we should\nrely on if we want consistent behavior. Immediately if core\nfunctionality starts using any other data, behavior becomes \"local\".\nAnd I think that's really *really* dangerous.\n\nSure, we have \"local behavior\" in a lot of small details. We very much\nintentionally have it in the ref-logs, and since branches and tags are\nlocal we also have it in things like \"--decorate\", which obviously\ndepends on exactly which local refs you have.\n\nWe also have local behavior in things like .git/config etc files, so\ngit can behave very differently for different people even with what is\notherwise an identical repository.\n\nSo local behavior is good and expected for some things. We *want* it\nfor things like colorization decisions, we want it for aliases, and we\nwant it for branch naming.\n\nBut really core behavior shouldn't depend on local information. I\nthink it would be wrong if something like a merge base decision would\nbe based on any local information.\n\nFor example, should it matter whether something is packed or not? I\nreally don't think so. That's a pretty random implementation detail,\nand if we get different end results because some commit happens to be\npacked, vs not packed (because, say, we'd be hiding generation\ninformation in the pack) that would be wrong.\n\nNow, we do have things like merge resolution caches etc (which\nobviously do save and use local information again), but I think that's\npretty well clarified.\n\n> So if there is an inconsistency between what the parent pointers\n> represent, and what the generation number in \"outside\" storage says,\n> then the outside storage is wrong, and the parent pointers are the right\n> answer. It becomes a lot more fuzzy to me if there is an inconsistency\n> between what the parent pointers represent, and what the generation\n> number says.\n\nSo I really don't see why you harp on that. If the generation counters\nare in the objects THEY BY DEFINITION CANNOT BE INCONSISTENT.\n\nThat's a big issue.\n\nSure, they may be LYING, but that's a different thing entirely. They\nwill be lying to everybody consistently. There would never be any\nquestion about what the generation number of a commit is.\n\nSee what I'm trying to say? There's no way that they would cause\ndifferent behavior for different people. Everything is 100%\nconsistent.\n\nThe exact same thing is true of commit dates, btw. They may be\nconfused as hell, and they may cause us to do bad things when we\ntraverse the history, but different clocks on different machines will\nstill not cause git to act differently on different machines. There's\nno possibility of inconsistency.\n\n(Of course, different *versions* of git may traverse the history\ndifferently, since we've changed the heuristics over time. So we do\nhave that kind of inconsistent behavior, where we give different\nresults from different versions of git).\n\nAnd btw, having \"incorrect\" data in the git objects is not the end of\nthe world. You can generate merge commits that simply have the wrong\nparents. That will be confusing as hell to the user, and it will make\nfuture merges not work very well, but it's a bug in the archive, and\nthat's \"ok\". The developers may not be very happy about it. In fact,\nafaik we've had a few cases like that in the kernel tree, because\nearly git had bugs where it would not properly forget parents after a\nfailed merge. Most of them are ARM-related, because the ARM tree was\none of the first users of git (outside of me, but I had fewer issues\nwith what happens when things go wrong).\n\nSo I would not be *too* shocked if we'd end up with \"odd\" generation\ncounts due to some odd bug. It sounds unlikely, but my point is that\nthat is not at all what I'd *worry* about.\n\n> How should that situation be handled? Should fsck check for it and\n> complain? Should we just ignore it, even though it may cause our\n> traversal algorithms to be inaccurate? Like clock skew, there's not much\n> that can be done if the commits are published.\n\nRight. I simply think it's not a big deal.\n\nIOW, if we would rely on generation counts instead of clock dates,\nmaybe the generation counts would have occasional problems too, but I\nsuspect they'd be *much* rarer than time-based issues, because at\nleast the generation count is a well-defined number rather than a\nrandom thing we pick out of emails and badly maintained machines.\n\nThat said, I'm not 100% sure at all that we want generation numbers at\nall. Their use is pretty limited. If we had had them from the\nbeginning, I think we would simply have replaced the date-based commit\nlist sorting with a generation-number-based one, and it should have\nbeen possible to guarantee that we never output a parent before the\ncommit in rev-parse.\n\nAs it is, I have to admit that looking at it, I shudder at changing\nthe current date-based logic and replacing it with a \"date or\ngeneration number\".\n\nThe date-based one, despite all its fuzziness and not being very well\ndefined (\"Global clock in a distributed system? You're a moron\") and\nup being a *nice* heuristic for certain human interaction. So it's not\na wonderful solution from a technical standpoint, but it does have (I\nthink) some nice UI advantages.\n\n(For an example of that: using \"--topo-sort\" for revision history may\nbe a very good thing technically, but even if it wasn't for the fact\nthat it's more expensive, I think that our largely time-based default\norder for \"git log\" in many ways is a better interface for humans. Of\ncourse, when mixed with actually giving a history graph, that changes,\nbecause then you want the \"related\" commits to group together, rather\nthan by time. So I think it's just basically a fuzzy area, without any\nclear hard rules - which is probably why using that fuzzy timestamp\nworks so well in practice)\n\n> Those are serious questions that I think should be considered if we are\n> going to put a generation header into the commit object, and I haven't\n> seen answers for them yet.\n\nI do agree that the really *big* question is \"do we even need it at\nall\". I do like perhaps just tightening the commit timestamp rules.\nBecause I do think they would probably work very well for the\n\"contains\" problem too.\n\nWith the exact same fuzzy downsides, of course. Timestamps aren't\nperfect, and they need that annoying fuzz factor thing.\n\n>> That said, I don't think a generation count necessarily \"fits\" in the\n>> pack-file. They are designed to be incremental, so it's not very\n>> natural there. But I do think it would be conceptually prettier to\n>> have the \"depth of commit\" be part of the \"filesystem\" data than to\n>> have it as a separate ad-hoc cache.\n>\n> Sure, I would be fine with that. When you say \"packfile\", do you mean\n> the the general concept, as in it could go in the pack index as opposed\n> to the packfile itself? Or specifically in the packfile? The latter\n> seems a lot more problematic to me in terms of implementation.\n\nI was thinking the \"general\" issue - it might make most sense to put\nthem in the index.\n>> If we should have fixed it in the original specification, we damn well\n>> should fix it today. It's been \"ignorable\" because it's just not been\n>> important enough. But if git now adds a fundamental cache for them,\n>> then that information is clearly no longer \"not important enough\".\n>\n> OK, so let's say we add generation headers to each commit. What happens\n> next? Are we going to convert algorithms that use timestamps to use\n> commit generations? How are we going to handle performance issues when\n> dealing with older parts of history that don't have generations?\n\nSo I do think the _initial_ question need to be the other way around:\ndo we have to have generation numbers at all?\n\nI think it's likely a design misfeature not to have them, but\nconsidering that we don't, and have been able to make do without for\nso long, I'm also perfectly willing to believe that we could speed up\n\"contains\" dramatically with the same kind of (crazy and inexact)\ntricks we use for merge bases.\n\n(Looking at a profile, a third - and the top entry - of the \"git tag\n--contains\" profile cost is just in \"clear_commit_marks()\" - not doing\nany real work, rather *undoing* the work in order to re-do things. So\nit's entirely possible that the real issue is simply that\n\"in_merge_bases()\" is badly done, and we could speed things up a lot\nindependently of anything else).\n\nFor example, for the \"git tag --contains\" thing, what's the\nperformance effect of just skipping tags that are much older than the\ncommit we ask for?\n\n                Linus\n"},{"id":"171444","messageId":"20110715215416.GB2117@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFx0KyAZRsy7gZ3Z4woWC-uWcLu11gcUrR+9MJR5NOSkrA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-15T21:54:16Z","receivedAt":"2011-07-15T21:54:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jul 15, 2011 at 02:17:26PM -0700, Linus Torvalds wrote:\n\n> > Having the information in two different places is my concern, too. And I\n> > think the fundamental difference between putting it inside or outside\n> > the commit sha1 (where outside encompasses putting it in a cache, in the\n> > pack-index, or whatever), is that I see the commit sha1 as somehow more\n> > \"definitive\". That is, it is the sole data we pass from repo to repo\n> > during pushes and pulls, and it is the thing that is consistency-checked\n> > by hashes.\n> \n> Sure. That is also the data that is the same for everybody.\n> \n> That's a big deal, in the sense that it's the only thing we should\n> rely on if we want consistent behavior. Immediately if core\n> functionality starts using any other data, behavior becomes \"local\".\n> And I think that's really *really* dangerous.\n\nYes, I see your argument. I just don't think it's all that big a deal,\nbecause the information is so easily derived from data that _is_ the\nsame for everybody (and when you _do_ want it to be different locally,\nbecause you are grafting, that is easy to do).\n\nBut I think at this point we have both said all there is to say. There\nis no actual data to be brought forth in this argument, and we obviously\ndisagree on this point. So I think we may have to agree to disagree.\n\nAnd as I said before, I am willing to concede generation numbers in the\ncommit header. But we need the rest of the solution, too.\n\n> So I really don't see why you harp on that. If the generation counters\n> are in the objects THEY BY DEFINITION CANNOT BE INCONSISTENT.\n> \n> That's a big issue.\n> \n> Sure, they may be LYING, but that's a different thing entirely. They\n> will be lying to everybody consistently. There would never be any\n> question about what the generation number of a commit is.\n>\n> See what I'm trying to say? There's no way that they would cause\n> different behavior for different people. Everything is 100%\n> consistent.\n\nRead my email again. I am clearly talking about inconsistency between\ntwo data items in the sha1-checked DAG itself. You then proceed to yell\nat me that they are not inconsistent, now talking about inconsistency\nbetween different people with the same DAG, but different caches. In\nother words, you are talking about an entirely different type of\ninconsistency. And then you proceed to say that the generation numbers\nmay be lying, which is _exactly_ what I meant when I said inconsistency.\n\nI don't mind arguing with you, even if I think you use capital letters\ntoo frequently; but when you do use them, please take care that I really\nam being a bonehead, and it is not you misrepresenting what I said.\n\nAs to lying (aka inconsistency between items within the DAG), you say:\n\n> And btw, having \"incorrect\" data in the git objects is not the end of\n> the world. You can generate merge commits that simply have the wrong\n> parents. That will be confusing as hell to the user, and it will make\n> future merges not work very well, but it's a bug in the archive, and\n> that's \"ok\". The developers may not be very happy about it. In fact,\n> afaik we've had a few cases like that in the kernel tree, because\n> early git had bugs where it would not properly forget parents after a\n> failed merge. Most of them are ARM-related, because the ARM tree was\n> one of the first users of git (outside of me, but I had fewer issues\n> with what happens when things go wrong).\n\nNo, it's not the end of the world. I just think it's worse than the\npossibility of inconsistency between two users' idea of the graph,\nbecause the bug stays with you for all of history, instead of getting\nfixed with a new version of git.\n\n> That said, I'm not 100% sure at all that we want generation numbers at\n> all. Their use is pretty limited. If we had had them from the\n> beginning, I think we would simply have replaced the date-based commit\n> list sorting with a generation-number-based one, and it should have\n> been possible to guarantee that we never output a parent before the\n> commit in rev-parse.\n> \n> As it is, I have to admit that looking at it, I shudder at changing\n> the current date-based logic and replacing it with a \"date or\n> generation number\".\n> \n> The date-based one, despite all its fuzziness and not being very well\n> defined (\"Global clock in a distributed system? You're a moron\") and\n> up being a *nice* heuristic for certain human interaction. So it's not\n> a wonderful solution from a technical standpoint, but it does have (I\n> think) some nice UI advantages.\n\nThat is the conclusion I am coming to, also. I don't find the external\ncache as odious as you obviously do. But that was why I posted the\npatches with an RFC tag. I wanted to see how painful people found the\nconcept. But if it's too ugly a concept, I think the path of least\nresistance is just making timestamps suck less (by using more consistent\nand robust skew avoidance[1] in our various algorithms, and by perhaps\ntaking more care to notify the user of skew early, before commits are\npublished).\n\nAnd then we don't really need generation numbers anymore.  As elegant as\nthey might have been if they were there from day one, it's just not\nworth the hassle of maintaining the dual solution.\n\n[1] We use \"N slop commits\" in some places and \"allow 86400 seconds of\n    skew\" in other places. We should probably use both, and apply them\n    consistently.\n\n> > Those are serious questions that I think should be considered if we are\n> > going to put a generation header into the commit object, and I haven't\n> > seen answers for them yet.\n> \n> I do agree that the really *big* question is \"do we even need it at\n> all\". I do like perhaps just tightening the commit timestamp rules.\n> Because I do think they would probably work very well for the\n> \"contains\" problem too.\n> \n> With the exact same fuzzy downsides, of course. Timestamps aren't\n> perfect, and they need that annoying fuzz factor thing.\n\nYeah. But in practice, that fuzz is really easy to implement, has worked\npretty well so far, and doesn't actually hurt performance measurably,\nbecause skew is rare, and a constant, small timestamp tends to equate to\na constant, small number of commits.\n\n> > Sure, I would be fine with that. When you say \"packfile\", do you mean\n> > the the general concept, as in it could go in the pack index as opposed\n> > to the packfile itself? Or specifically in the packfile? The latter\n> > seems a lot more problematic to me in terms of implementation.\n> \n> I was thinking the \"general\" issue - it might make most sense to put\n> them in the index.\n\nIf we were to go the cache route, I think I am leaning that way, too, if\nonly because we don't duplicate the 20-byte sha1 per commit, which keeps\nour I/O down.\n\n> > OK, so let's say we add generation headers to each commit. What happens\n> > next? Are we going to convert algorithms that use timestamps to use\n> > commit generations? How are we going to handle performance issues when\n> > dealing with older parts of history that don't have generations?\n> \n> So I do think the _initial_ question need to be the other way around:\n> do we have to have generation numbers at all?\n\nNo, we don't need them. My \"contains\" patches were already implemented\nusing timestamps, and it's pretty fast. They fall down only in the face\nlying timestamps (i.e., skew). The whole reason to switch to generation\nheaders was that we could assume they would be correct, and our\nalgorithms using them would be more likely to be correct.\n\nAnd I do think a generation header would be more likely to be correct\nthan a timestamp, if only because timestamps are harder to get right.\n\n> I think it's likely a design misfeature not to have them, but\n> considering that we don't, and have been able to make do without for\n> so long, I'm also perfectly willing to believe that we could speed up\n> \"contains\" dramatically with the same kind of (crazy and inexact)\n> tricks we use for merge bases.\n\nAlready done. I can point you to the patches if you want.\n\n> For example, for the \"git tag --contains\" thing, what's the\n> performance effect of just skipping tags that are much older than the\n> commit we ask for?\n\nIt's as fast as using generations. See these two patches:\n\n  http://article.gmane.org/gmane.comp.version-control.git/150261\n\n  http://article.gmane.org/gmane.comp.version-control.git/150262\n\n-Peff\n"},{"id":"171447","messageId":"CA+55aFzE-okH9gaEyuSFdorK-7v3odpsk65ZTqCMHFz80n65ug@mail.gmail.com","threadId":"27818","inReplyTo":"CA+55aFx0KyAZRsy7gZ3Z4woWC-uWcLu11gcUrR+9MJR5NOSkrA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T23:10:23Z","receivedAt":"2011-07-15T23:10:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 2:17 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> For example, for the \"git tag --contains\" thing, what's the\n> performance effect of just skipping tags that are much older than the\n> commit we ask for?\n\nHmm.\n\nMaybe there is something seriously wrong with this trivial patch, but\nit gave the right results for the test-cases I threw at it, and passes\nthe tests.\n\nBefore:\n\n   [torvalds@i5 linux]$ time git tag --contains v2.6.24 > correct\n\n   real\t0m7.548s\n   user\t0m7.344s\n   sys\t0m0.116s\n\nAfter:\n\n   [torvalds@i5 linux]$ time ~/git/git tag --contains v2.6.24 > date-cut-off\n\n   real\t0m0.161s\n   user\t0m0.140s\n   sys\t0m0.016s\n\nand 'correct' and 'date-cut-off' both give the same answer.\n\nThe date-based \"slop\" thing is (at least *meant* to be - note the lack\nof any extensive testing) \"at least five consecutive commits that have\ndates that are more than five days off\".\n\nSomebody should double-check my logic. Maybe I'm doing something\nstupid. Because that's a *big* difference.\n\n                     Linus\n\n\n commit.c |   42 +++++++++++++++++++++++++++++++++++++++++-\n 1 files changed, 41 insertions(+), 1 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex ac337c7d7dc1..0d33c33a6520 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -737,16 +737,56 @@ struct commit_list *get_merge_bases(struct commit *one, struct commit *two,\n \treturn get_merge_bases_many(one, 1, &two, cleanup);\n }\n \n+#define VISITED (1 << 16)\n+\n+static int is_recursive_descendant(struct commit *commit, struct commit *target)\n+{\n+\tint slop = 5;\n+\tparse_commit(target);\n+\tfor (;;) {\n+\t\tstruct commit_list *parents;\n+\t\tif (commit == target)\n+\t\t\treturn 1;\n+\t\tif (commit->object.flags & VISITED)\n+\t\t\treturn 0;\n+\t\tcommit->object.flags |= VISITED;\n+\t\tparse_commit(commit);\n+\t\tif (commit->date + 5*24*60*60 < target->date) {\n+\t\t\tif (--slop <= 0)\n+\t\t\t\treturn 0;\n+\t\t} else\n+\t\t\tslop = 5;\n+\t\tparents = commit->parents;\n+\t\tif (!parents)\n+\t\t\treturn 0;\n+\t\tcommit = parents->item;\n+\t\tparents = parents->next;\n+\t\twhile (parents) {\n+\t\t\tif (is_recursive_descendant(parents->item, target))\n+\t\t\t\treturn 1;\n+\t\t\tparents = parents->next;\n+\t\t}\n+\t}\n+}\n+\n+static int is_descendant(struct commit *commit, struct commit *target)\n+{\n+\tint ret = is_recursive_descendant(commit, target);\n+\tclear_commit_marks(commit, VISITED);\n+\treturn ret;\n+}\n+\n int is_descendant_of(struct commit *commit, struct commit_list *with_commit)\n {\n \tif (!with_commit)\n \t\treturn 1;\n+\n \twhile (with_commit) {\n \t\tstruct commit *other;\n \n \t\tother = with_commit->item;\n \t\twith_commit = with_commit->next;\n-\t\tif (in_merge_bases(other, &commit, 1))\n+\t\tif (is_descendant(commit, other))\n \t\t\treturn 1;\n \t}\n \treturn 0;\n"},{"id":"171448","messageId":"CA+55aFwpVoqK7TaG0R3JJO07eOyWQ9pR1sHUGBQt0kmM0vk2bw@mail.gmail.com","threadId":"27818","inReplyTo":"CA+55aFzE-okH9gaEyuSFdorK-7v3odpsk65ZTqCMHFz80n65ug@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T23:16:20Z","receivedAt":"2011-07-15T23:16:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 15, 2011 at 4:10 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> Maybe there is something seriously wrong with this trivial patch, but\n> it gave the right results for the test-cases I threw at it, and passes\n> the tests.\n>\n> Before:\n\nI have fewer branches than tags, but I get something similar for \"git\nbranch --contains\":\n\n  [torvalds@i5 linux]$ time git branch --contains v2.6.12 | sha1sum\n  9d4224eec98ec7b0bcd5331dfa5badb9ef1fd510  -\n\n  real\t0m4.205s\n  user\t0m4.112s\n  sys\t0m0.084s\n  [torvalds@i5 linux]$ time ~/git/git branch --contains v2.6.12 | sha1sum\n  9d4224eec98ec7b0bcd5331dfa5badb9ef1fd510  -\n\n  real\t0m0.112s\n  user\t0m0.100s\n  sys\t0m0.008s\n\nie identical results, except one took 4.2s and with the patch it took 0.1s.\n\nThis is all hot-cache, of course, and on a fast machine.\n\n                    Linus\n"},{"id":"171449","messageId":"CA+55aFwSgYYaQ8gTQmCw6SNMyr-bz5rJPr0o9xoog-1aCqb5rA@mail.gmail.com","threadId":"27818","inReplyTo":"CA+55aFwpVoqK7TaG0R3JJO07eOyWQ9pR1sHUGBQt0kmM0vk2bw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-15T23:36:40Z","receivedAt":"2011-07-15T23:36:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"And one last comment:\n\nOn Fri, Jul 15, 2011 at 4:16 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> I have fewer branches than tags, but I get something similar for \"git\n> branch --contains\":\n\nThe time-based heuristic does seem to be important. If I just remove\nit, I get increasingly long times for things that aren't contained in\nmy branches.\n\nAnd in fact, I think that is why the code used the merge-base helper\nfunctions - not because it wanted merge bases, but because the merge\nbase stuff will work from either end until it decides things aren't\nrelevant any more. Because *without* the time-based heuristics, the\ntrivial \"is this a descendant\" algorithm ends up working very badly\nfor the case where the target doesn't exist in the branches. Examples\nof NOT having a date-based cut-off, but just doing the straightforward\n(non-merge-base) ancestry walk:\n\n  time ~/git/git branch --contains v2.6.12\n  real\t0m0.113s\n\n  [torvalds@i5 linux]$ time ~/git/git branch --contains v2.6.39\n  real\t0m3.691s\n\nand what ends up happening is that in the latter case, every branch\nwalks all the way to the root and checks every commit (walking all the\nmerges too). While in the first case, it's very quick because it will\nfind that particular commit when it walk straight backwards (so it\ndoesn't even have to do a lot of recursion - the first branch that\nhits that commit will be a success), so it won't have to look at all\nthe side ways of getting there.\n\nOf course, the above particular difference happens to be due to the\n\"depth-first\" implementation working well for the thing I am searching\nfor.  But it does show that the date-based cut-off matters due to\ntraversal issues like that.\n\n                          Linus\n"},{"id":"171452","messageId":"20110716004034.GA32230@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFzE-okH9gaEyuSFdorK-7v3odpsk65ZTqCMHFz80n65ug@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-16T00:40:34Z","receivedAt":"2011-07-16T00:40:34Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jul 15, 2011 at 04:10:23PM -0700, Linus Torvalds wrote:\n\n> On Fri, Jul 15, 2011 at 2:17 PM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n> >\n> > For example, for the \"git tag --contains\" thing, what's the\n> > performance effect of just skipping tags that are much older than the\n> > commit we ask for?\n> \n> Hmm.\n> \n> Maybe there is something seriously wrong with this trivial patch, but\n> it gave the right results for the test-cases I threw at it, and passes\n> the tests.\n> \n> Before:\n> \n>    [torvalds@i5 linux]$ time git tag --contains v2.6.24 > correct\n> \n>    real\t0m7.548s\n>    user\t0m7.344s\n>    sys\t0m0.116s\n> \n> After:\n> \n>    [torvalds@i5 linux]$ time ~/git/git tag --contains v2.6.24 > date-cut-off\n> \n>    real\t0m0.161s\n>    user\t0m0.140s\n>    sys\t0m0.016s\n> \n> and 'correct' and 'date-cut-off' both give the same answer.\n\nWithout even looking carefully at your patches for any minor mistakes, I\ncan tell you that the speedup you're seeing is approximately right.\nBecause it's almost exactly the same optimization I made in my\ntimestamp-based patches (links to which I sent you earlier today).\n\nHowever, you can make it even faster. The \"tag --contains\" code will ask\n\"is_descendant_of\" repeatedly for the same set of \"want\" commits. So you\nend up traversing some parts of the graph over and over. My patches\nshare the marks over a set of contains traversals, so you only ever\ntouch each commit once. And that's what my patches do.\n\nWith yours, on my box:\n\n  $ time git tag --contains HEAD~1000 >/dev/null\n  real    0m0.113s\n  user    0m0.104s\n  sys     0m0.008s\n\nand mine:\n\n  $ time git tag --contains HEAD~1000 >/dev/null\n  real    0m0.035s\n  user    0m0.020s\n  sys     0m0.012s\n\nI suspect you can make the difference even more prominent by having more\ntags, or by having multiple \"want\" commits.\n\n-Peff\n"},{"id":"171453","messageId":"20110716004232.GB32230@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CA+55aFwSgYYaQ8gTQmCw6SNMyr-bz5rJPr0o9xoog-1aCqb5rA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-16T00:42:32Z","receivedAt":"2011-07-16T00:42:32Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jul 15, 2011 at 04:36:40PM -0700, Linus Torvalds wrote:\n\n> On Fri, Jul 15, 2011 at 4:16 PM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n> >\n> > I have fewer branches than tags, but I get something similar for \"git\n> > branch --contains\":\n> \n> The time-based heuristic does seem to be important. If I just remove\n> it, I get increasingly long times for things that aren't contained in\n> my branches.\n> \n> And in fact, I think that is why the code used the merge-base helper\n> functions - not because it wanted merge bases, but because the merge\n> base stuff will work from either end until it decides things aren't\n> relevant any more. Because *without* the time-based heuristics, the\n> trivial \"is this a descendant\" algorithm ends up working very badly\n> for the case where the target doesn't exist in the branches. Examples\n> of NOT having a date-based cut-off, but just doing the straightforward\n> (non-merge-base) ancestry walk:\n> \n>   time ~/git/git branch --contains v2.6.12\n>   real\t0m0.113s\n> \n>   [torvalds@i5 linux]$ time ~/git/git branch --contains v2.6.39\n>   real\t0m3.691s\n\nYes, exactly. That is why my first patch (which goes to a recursive\nsearch), takes about the same amount of time as \"git rev-list --all\"\n(and I suspect your 3.691s above is similar). And then the second one\ndrops that again to .03s.\n\nI think you are simply recreating the strategy and timings I have posted\nseveral times now.\n\n-Peff\n"},{"id":"171457","messageId":"CAP8UFD3p8rv9BoPkTYSr_qRztKhWmmHgjHi0pZ6gN9YzkSX0Jw@mail.gmail.com","threadId":"27818","inReplyTo":"20110715184211.GH8453@thunk.org","subject":"Re: Git commit generation numbers","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2011-07-16T09:16:45Z","receivedAt":"2011-07-16T09:16:45Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Jul 15, 2011 at 8:42 PM, Ted Ts'o <tytso@mit.edu> wrote:\n> On Fri, Jul 15, 2011 at 09:44:21AM -0700, Linus Torvalds wrote:\n>> So I would like to repeat: I think our commit-date based hack has been\n>> pretty successful. We've lived with it for years and years. Even the\n>> \"let's try to fix it by adding slop\" code is from three years ago\n>> (commit 7d004199d1), which means that for three years we never really\n>> saw any serious problems. I forget what problem we actually did see -\n>> I have this dim memory of it being Ted that had problems with a merge\n>> because git picked a crap merge base, but that may just be my\n>> Alzheimer's speaking.\n>\n> My original main issue was simply that \"git tag --contains\" and \"git\n> branch --contains\" was either (a) incorrect, or (b) slower than\n> popping up gitk and pulling the information out of the GUI.  The\n> reason for (b) is because of gitk.cache.\n>\n> Maybe the answer then is creating a command-line tool (it doesn't have to\n> be in \"core\" of git) which just pulls the dammned information out of\n> gitk.cache....\n>\n> (Yes, it's gross, but I'm not worrying about the long-term\n> architecture of git or anything high-falutin' like that.  I'm just a\n> poor dumb user who just wants git tag --contains and git branch\n> --contains to be fast and accurate...)\n\nIf  \"git tag --contains\" and \"git branch --contains\" give incorrect\nanswers because the commiter date is wrong in some commits, then why\nnot use \"git replace\" to \"change\" the commiter date in the commits\nthat have a wrong date? Is it because you don't want to use \"git\nreplace\", or because there is no script to do it automatically, or is\nthere another reason?\n\nThanks,\nChristian.\n"},{"id":"171538","messageId":"20110718034106.GB2468@sigill.intra.peff.net","threadId":"27818","inReplyTo":"CAP8UFD3p8rv9BoPkTYSr_qRztKhWmmHgjHi0pZ6gN9YzkSX0Jw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-18T03:41:06Z","receivedAt":"2011-07-18T03:41:06Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Jul 16, 2011 at 11:16:45AM +0200, Christian Couder wrote:\n\n> If  \"git tag --contains\" and \"git branch --contains\" give incorrect\n> answers because the commiter date is wrong in some commits, then why\n> not use \"git replace\" to \"change\" the commiter date in the commits\n> that have a wrong date? Is it because you don't want to use \"git\n> replace\", or because there is no script to do it automatically, or is\n> there another reason?\n\nThat would work. There are a few tricky things, though:\n\n  1. Most commits have less than 100 skewed commits. But some have many\n     (e.g., thousands in the mesa repo). How well does git cope with\n     large numbers of replace refs, performance-wise?\n\n  2. Declaring which commits are skewed is actually tricky. You can find\n     a commit whose timestamp is less than the timestamp of one of its\n     ancestors. But you don't know whether it is skewed, or the\n     ancestor.\n\n     If you are implementing a list of commits whose timestamps\n     shouldn't be used for traversal cutoff, it doesn't really matter\n     who is _right_; you just care about whether the timestamps are\n     strictly increasing from that point.\n\n     But once you start replacing commits, you need to put in a\n     reasonable value for the timestamp. So you may well be replacing a\n     perfectly valid commit with one that has bogus, skewed information\n     in the commit timestamp.\n\n  3. Any value you put in is actually going to be a lie during things\n     like \"git log --pretty=raw\". That may be OK. But it is letting an\n     optimization meant to make traversal fast and accurate bleed into\n     the actual data we show the user.\n\n  4. Sometimes we need to do traversals on the real objects (e.g.,\n     because we are doing upload-pack). To get the benefit, those\n     traversals would presumably need to look at both the original\n     object and the replacement, use the timestamp from the replacement\n     for traversal, but otherwise use the original object.\n\n-Peff\n"},{"id":"171637","messageId":"201107190614.38431.chriscool@tuxfamily.org","threadId":"27818","inReplyTo":"20110718034106.GB2468@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Christian Couder","fromEmail":"chriscool@tuxfamily.org","sentAt":"2011-07-19T04:14:38Z","receivedAt":"2011-07-19T04:14:38Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Monday 18 July 2011 05:41:06 Jeff King wrote:\n> On Sat, Jul 16, 2011 at 11:16:45AM +0200, Christian Couder wrote:\n> > If  \"git tag --contains\" and \"git branch --contains\" give incorrect\n> > answers because the commiter date is wrong in some commits, then why\n> > not use \"git replace\" to \"change\" the commiter date in the commits\n> > that have a wrong date? Is it because you don't want to use \"git\n> > replace\", or because there is no script to do it automatically, or is\n> > there another reason?\n> \n> That would work. There are a few tricky things, though:\n> \n>   1. Most commits have less than 100 skewed commits. But some have many\n>      (e.g., thousands in the mesa repo). How well does git cope with\n>      large numbers of replace refs, performance-wise?\n\nIf it did not cope well, it should be possible to improve the performance.\n\nAnyway, another way to fix the problem with \"git replace\" could be to create \nbranches with commits that have a fixed commiter date and then to use \"git \nreplace\" only to connect these branches to the graph.\n\nFor example if you have this:\n\nA - B - X1 - X2 - X3 - C - D\n\nwhere X1, X2 and X3 are skewed, then you can create this:\n\nA - B - X1 - X2 - X3 - C - D\n         \\ Y1 - Y2 - Y3\n\nwhere Y1, Y2, Y3 are the same as X1, X2, X3 except they are not skewed.\nThen you only need to do \"git replace X3 Y3\" so you create only one replace \nref.\n\n> \n>   2. Declaring which commits are skewed is actually tricky. You can find\n>      a commit whose timestamp is less than the timestamp of one of its\n>      ancestors. But you don't know whether it is skewed, or the\n>      ancestor.\n> \n>      If you are implementing a list of commits whose timestamps\n>      shouldn't be used for traversal cutoff, it doesn't really matter\n>      who is _right_; you just care about whether the timestamps are\n>      strictly increasing from that point.\n> \n>      But once you start replacing commits, you need to put in a\n>      reasonable value for the timestamp. So you may well be replacing a\n>      perfectly valid commit with one that has bogus, skewed information\n>      in the commit timestamp.\n\nPerhaps but with \"git replace\" you can choose to create new replace refs and \ndeprecate the old replace refs to fix this where you got it wrong.\n\nIt would be easier to do that if \"git replace\" supported sub directories like \n\"refs/replace/clock-skew/ted-july-2011/\", so you could manage the replace refs \nmore easily.\n\nFor example you could create new refs in \"refs/replace/clock-skew/ted-\njuly-2011-2/\" if you found a better fix. And then use these new refs instead of \nthose in \"refs/replace/clock-skew/ted-july-2011/\".\n\n>   3. Any value you put in is actually going to be a lie during things\n>      like \"git log --pretty=raw\". That may be OK. But it is letting an\n>      optimization meant to make traversal fast and accurate bleed into\n>      the actual data we show the user.\n\nWith replace refs, the user could choose the \"lies\" told to him/her by \nselecting the replace refs or set of replace refs that are used.\n\nAs commits are immutable, when they are created with bad data, the best we can \ndo is let the user choose if they want to see the original or another \"fixed\" \nversion. Because the original will always be \"true\" in a way.\n\n>   4. Sometimes we need to do traversals on the real objects (e.g.,\n>      because we are doing upload-pack). To get the benefit, those\n>      traversals would presumably need to look at both the original\n>      object and the replacement, use the timestamp from the replacement\n>      for traversal, but otherwise use the original object.\n\nYeah, or maybe when we do traversals on real objects we could afford not to \nrely on commiter date or some other \"fragile\" data.\n\nThanks,\nChristian.\n"},{"id":"171698","messageId":"20110719200022.GB3957@sigill.intra.peff.net","threadId":"27818","inReplyTo":"201107190614.38431.chriscool@tuxfamily.org","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-19T20:00:22Z","receivedAt":"2011-07-19T20:00:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jul 19, 2011 at 06:14:38AM +0200, Christian Couder wrote:\n\n> >      But once you start replacing commits, you need to put in a\n> >      reasonable value for the timestamp. So you may well be replacing a\n> >      perfectly valid commit with one that has bogus, skewed information\n> >      in the commit timestamp.\n> \n> Perhaps but with \"git replace\" you can choose to create new replace refs and \n> deprecate the old replace refs to fix this where you got it wrong.\n> \n> It would be easier to do that if \"git replace\" supported sub directories like \n> \"refs/replace/clock-skew/ted-july-2011/\", so you could manage the replace refs \n> more easily.\n\nI think all of the arguments I cut from your email are reasonable, but\nthe crux of the issue comes down to this point.\n\nIf you are interested in actually correcting the skew, then yes, replace\nrefs are a good solution. But doing so is going to involve somebody\nlooking at the commits and deciding which ones are wrong, and what they\nshould be. And maybe that's a good thing to do for people who really\ncare about cleaning history.\n\nBut for something like \"speed up revision traversal by assuming commit\ntimestamps are roughly increasing\", we want something very automated,\nand what is needs to say is much weaker (not \"this is what this commit\n_should_ say\", but rather \"this commit might be right, but it is not a\ngood point for cutting off a traversal\"). So that's a much easier\nproblem, and it's easy to do in an automated way.\n\nSo I think while you could use replace refs to handle this issue, it is\nnot always going to be the right solution, and there is room for\nsomething simpler (and weaker).\n\n-Peff\n"},{"id":"171783","messageId":"CAP8UFD2duHW7MtLVrnjE9UMyU8bzx2xbJFzot5R2CThon8Dr3w@mail.gmail.com","threadId":"27818","inReplyTo":"20110719200022.GB3957@sigill.intra.peff.net","subject":"Re: Git commit generation numbers","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2011-07-21T06:29:33Z","receivedAt":"2011-07-21T06:29:33Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Tue, Jul 19, 2011 at 10:00 PM, Jeff King <peff@peff.net> wrote:\n> On Tue, Jul 19, 2011 at 06:14:38AM +0200, Christian Couder wrote:\n>\n>> Perhaps but with \"git replace\" you can choose to create new replace refs and\n>> deprecate the old replace refs to fix this where you got it wrong.\n>>\n>> It would be easier to do that if \"git replace\" supported sub directories like\n>> \"refs/replace/clock-skew/ted-july-2011/\", so you could manage the replace refs\n>> more easily.\n>\n> I think all of the arguments I cut from your email are reasonable, but\n> the crux of the issue comes down to this point.\n>\n> If you are interested in actually correcting the skew, then yes, replace\n> refs are a good solution. But doing so is going to involve somebody\n> looking at the commits and deciding which ones are wrong, and what they\n> should be.\n\nI think that we can help the user a lot to find the skew, and then to\ndecide which commits are wrong, and then to fix the skew even if the\nfix we suggest is far from being perfect.\n\n> And maybe that's a good thing to do for people who really\n> care about cleaning history.\n\nYeah, so maybe at one point we will want to help these people even if\nwe have implemented automatic generation numbers. Then this means that\nautomated generation numbers are useful only if:\n\n1) there are commits with skews\n2) the heuristics to deal with some skew don't work\n3) the user is too lazy to use the help we (can) provide to fix the skews\n\nI think that we can probably find heuristics that will deal with at\nleast 95% of the cases. For example we could perhaps decide that we\ndon't cut off a traversal until the date difference is greater than 5\ndays.\n\nThen in the hopefully few cases where there are really big skews that\nwon't be caught by our heuristics, (but that we can automatically\ndetect when fetching or commiting,) we can perhaps afford to ask the\nuser to do a small analysis to properly fix the skew.\n\nI mean that at one point when things are too weird it is ok and\nperhaps even a good thing to involve the user.\n\n> But for something like \"speed up revision traversal by assuming commit\n> timestamps are roughly increasing\", we want something very automated,\n> and what is needs to say is much weaker (not \"this is what this commit\n> _should_ say\", but rather \"this commit might be right, but it is not a\n> good point for cutting off a traversal\"). So that's a much easier\n> problem, and it's easy to do in an automated way.\n\nYeah, generation numbers look like an easy thing to do. And yeah,\nbeing automated is great too. But it does not mean it is the right\nthing to do. (Or perhaps we could have them but not save them in any\ncache, nor in the commit object.)\n\n> So I think while you could use replace refs to handle this issue, it is\n> not always going to be the right solution, and there is room for\n> something simpler (and weaker).\n\nYou know, replace refs can be used to fix or improve a lot of things\nlike bad authors, clock skews, bisecting on a fixed up history,\nworking on a larger or smaller repository than the original, and so\non. And of course for each of these problems you may find another\nsolution tailored to the problem at hand that will seem simpler or\neasier. But in the end if you develop all these other solutions you\nwill have developed a lot of stuff that will be harder to maintain,\nless generic, more complex and so on, that properly developed replace\nrefs.\n\nThanks,\nChristian.\n"}]}