{"thread":{"id":"55341","subject":"Distinguishing FF vs non-FF updates in the reflog?","startedAt":"2021-03-17T20:07:14Z","lastAt":"2021-03-26T07:44:24Z","messageCount":19,"participants":["Han-Wen Nienhuys","Martin Fick","Jeff King","Ævar Arnfjörð Bjarmason","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"419545","messageId":"CAFQ2z_MefCwiWdhs0buJv5Zok+nsgaOvUCcsSnfm_PP0WozZKA@mail.gmail.com","threadId":"55341","inReplyTo":null,"subject":"Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-17T20:06:06Z","receivedAt":"2021-03-17T20:07:14Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"Hi there,\n\nI'm working on some extensions to Gerrit for which it would be very\nbeneficial if we could tell from the reflog if an update is a\nfast-forward or not: if we find a SHA1 in the reflog, and see there\nwere only FF updates since, we can be sure that the SHA1 is reachable\nfrom the branch, without having to open packfiles and decode commits.\n\nFor the reftable format, I think we could store this easily by\nintroducing more record types. Today we have 0 = deletion, 1 = update,\nand we could add 2 = FF update, 3 = non-FF update.\n\nHowever, the textual reflog format doesn't easily allow for this.\nHowever, we might add a convention, eg. have the message start with\n'FF' or 'NFF' depending on the nature of the update.\n\nDoes this make sense, and if yes is it worth proposing a change?\n\nthanks,\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419583","messageId":"5359503.g8GvsOHjsp@mfick-lnx","threadId":"55341","inReplyTo":"CAFQ2z_MefCwiWdhs0buJv5Zok+nsgaOvUCcsSnfm_PP0WozZKA@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-03-17T21:21:52Z","receivedAt":"2021-03-17T21:27:51Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Wednesday, March 17, 2021 9:06:06 PM MDT Han-Wen Nienhuys wrote:\n> I'm working on some extensions to Gerrit for which it would be very\n> beneficial if we could tell from the reflog if an update is a\n> fast-forward or not: if we find a SHA1 in the reflog, and see there\n> were only FF updates since, we can be sure that the SHA1 is reachable\n> from the branch, without having to open packfiles and decode commits.\n\nI don't think this would be reliable.\n\n1) Not all updates make it to the reflogs\n2) Reflogs can be edited or mucked with\n3) On NFS reflogs can outright be wrong even when used properly as their are \ncaching issues. We specifically have seen entries that appear to be FFs that \nwere not.\n\nI believe that today git can do very fast reachability checks without opening \npack files by using some of its indexes (bitmap code or https://git-scm.com/\ndocs/commit-graph ?). It probably makes sense to add this ability to jgit if \nthat is what you need?\n\n-Martin \n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"419620","messageId":"CAFQ2z_MavgAGDyJzc9-+j6zTDODP7hCdPHtB5dyx-reLMSLX3Q@mail.gmail.com","threadId":"55341","inReplyTo":"5359503.g8GvsOHjsp@mfick-lnx","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-18T08:58:56Z","receivedAt":"2021-03-18T08:59:44Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Wed, Mar 17, 2021 at 10:22 PM Martin Fick <mfick@codeaurora.org> wrote:\n>\n> On Wednesday, March 17, 2021 9:06:06 PM MDT Han-Wen Nienhuys wrote:\n> > I'm working on some extensions to Gerrit for which it would be very\n> > beneficial if we could tell from the reflog if an update is a\n> > fast-forward or not: if we find a SHA1 in the reflog, and see there\n> > were only FF updates since, we can be sure that the SHA1 is reachable\n> > from the branch, without having to open packfiles and decode commits.\n>\n> I don't think this would be reliable.\n>\n> 1) Not all updates make it to the reflogs\n> 2) Reflogs can be edited or mucked with\n> 3) On NFS reflogs can outright be wrong even when used properly as their are\n> caching issues. We specifically have seen entries that appear to be FFs that\n> were not.\n\nCan you tell a little more about 3) ? SInce we don't annotate non-FF\nvs FF today, what does \"appear to be FFs\" mean?\n\nBut you are right: since the reflog for a branch is in a different\nfile from the branch head, there is no way to do an update to both of\nthem at the same time. I guess this will have to be a reftable-only\nfeature.\n\n> I believe that today git can do very fast reachability checks without opening\n> pack files by using some of its indexes (bitmap code or https://git-scm.com/\n> docs/commit-graph ?). It probably makes sense to add this ability to jgit if\n> that is what you need?\n\nThe bitmaps are generated by GC, and you can't GC all the time. JGit\nhas support for bitmaps, and its support actually predates C-Git's\nsupport for it. (It was added to JGit by Colby Ranger who worked in\nShawn's team).\n\nI expect that the commit graph doesn't work for my intended use-case.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\nRegistergericht und -nummer: Hamburg, HRB 86891\nSitz der Gesellschaft: Hamburg\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419647","messageId":"YFOrbjeunnVfQNRC@coredump.intra.peff.net","threadId":"55341","inReplyTo":"CAFQ2z_MavgAGDyJzc9-+j6zTDODP7hCdPHtB5dyx-reLMSLX3Q@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-03-18T19:35:10Z","receivedAt":"2021-03-18T19:35:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 18, 2021 at 09:58:56AM +0100, Han-Wen Nienhuys wrote:\n\n> > 1) Not all updates make it to the reflogs\n> > 2) Reflogs can be edited or mucked with\n> > 3) On NFS reflogs can outright be wrong even when used properly as their are\n> > caching issues. We specifically have seen entries that appear to be FFs that\n> > were not.\n> \n> Can you tell a little more about 3) ? SInce we don't annotate non-FF\n> vs FF today, what does \"appear to be FFs\" mean?\n> \n> But you are right: since the reflog for a branch is in a different\n> file from the branch head, there is no way to do an update to both of\n> them at the same time. I guess this will have to be a reftable-only\n> feature.\n\nEach individual reflog entry (in the branch reflog and the HEAD reflog)\nshould still be consistent, though. They give the \"before\" and \"after\"\nobject ids, and the ff-ness is an immutable property of those commit\nids.\n\n> > I believe that today git can do very fast reachability checks without opening\n> > pack files by using some of its indexes (bitmap code or https://git-scm.com/\n> > docs/commit-graph ?). It probably makes sense to add this ability to jgit if\n> > that is what you need?\n> \n> The bitmaps are generated by GC, and you can't GC all the time. JGit\n> has support for bitmaps, and its support actually predates C-Git's\n> support for it. (It was added to JGit by Colby Ranger who worked in\n> Shawn's team).\n\nBitmaps can help with these checks, but we don't actually look at them\nin most of the algorithms one might use for computing ancestry. One of\nthe reasons for that is that they often backfire as an optimization,\nbecause:\n\n  - as you note, they are often not up to date because they require a\n    repack. So they won't help when asking about very recently added\n    commits (which people tend to ask about more than ancient ones).\n\n  - the bitmap file format doesn't have any index. So a reader has to\n    scan the whole thing upon opening to decide which commits have\n    bitmaps.\n\nFor several years we had a patch at GitHub that checked for bitmaps\nduring \"--contains\" traversals. Even though it did sometimes backfire,\nit was enough of a net win to be worth keeping, compared to actually\nopening commit objects to follow their parent pointers. But with\ncommit-graphs, it was a strict loss, and we stopped using it entirely\nlast year. (We do still look at bitmaps for our branch ahead/behind\nchecks using a custom patch; I'm suspicious of its performance for the\nsame reasons, but we haven't dug carefully into it).\n\nBut...\n\n> I expect that the commit graph doesn't work for my intended use-case.\n\n...I think commit-graphs are a big win here. They are more often kept up\nto date, because they can be generated incrementally with effort\nproportional to the number of new commits. And they make a big\ndifference if the traversal has to cover a lot of commits. E.g., here's\nthe most extreme case in git.git, checking ancestry of the oldest\ncommit:\n\n  $ time git merge-base --is-ancestor e83c5163316f89bfbde7d9ab23ca2e25604af290 HEAD; echo $?\n\n  real\t0m0.014s\n  user\t0m0.008s\n  sys\t0m0.005s\n  0\n\n  $ time git -c core.commitgraph=false merge-base --is-ancestor e83c5163316f89bfbde7d9ab23ca2e25604af290 HEAD; echo $?\n\n  real\t0m0.398s\n  user\t0m0.369s\n  sys\t0m0.028s\n  0\n\nOf course most results won't be so dramatic, because they wouldn't have\nto traverse many commits in the first place (so they are already pretty\nfast with or without the commit-graph).  But that 14ms should be an\nupper bound for this repo. And naturally that scales with the number of\ncommits; in linux.git it's 43ms, compared to 8.7s without commit-graphs).\n\n-Peff\n"},{"id":"419648","messageId":"YFOuT6L0dsrCGTBk@coredump.intra.peff.net","threadId":"55341","inReplyTo":"CAFQ2z_MefCwiWdhs0buJv5Zok+nsgaOvUCcsSnfm_PP0WozZKA@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-03-18T19:47:27Z","receivedAt":"2021-03-18T19:48:15Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 17, 2021 at 09:06:06PM +0100, Han-Wen Nienhuys wrote:\n\n> I'm working on some extensions to Gerrit for which it would be very\n> beneficial if we could tell from the reflog if an update is a\n> fast-forward or not: if we find a SHA1 in the reflog, and see there\n> were only FF updates since, we can be sure that the SHA1 is reachable\n> from the branch, without having to open packfiles and decode commits.\n\nI left some numbers in another part of the thread, but IMHO performance\nisn't that compelling a reason to do this these days, if you are using\ncommit-graphs.\n\nJust walking the reflog might be _slightly_ faster, though not\nnecessarily (it depends on whether the depth of the object graph or the\ndepth of the reflog chain is deeper). It might matter more if you are\nusing a more exotic storage scheme, where switching from accessing\nreflogs to objects implies extra round-trips to a server (e.g., custom\nstorage backends with JGit; I don't know the state of the art in what\nGoogle is doing there).\n\n> For the reftable format, I think we could store this easily by\n> introducing more record types. Today we have 0 = deletion, 1 = update,\n> and we could add 2 = FF update, 3 = non-FF update.\n> \n> However, the textual reflog format doesn't easily allow for this.\n> However, we might add a convention, eg. have the message start with\n> 'FF' or 'NFF' depending on the nature of the update.\n> \n> Does this make sense, and if yes is it worth proposing a change?\n\nAt GitHub we do something similar. We don't generally use reflogs much\nat all, but we keep a custom \"audit log\": a single append-only file that\nrecords every ref update in the repository. And its format just happens\nto be one reflog entry per line, prefixed by the updated ref.\n\nAnd there we do generally annotate the FF-ness of an update by stuffing\nit into the free-form message field (in fact, we shove in a small JSON\nobject, so we record multiple fields like the pushing id, IP, etc).\n\nBut the main goal there isn't performance (and in fact we don't\ngenerally consult it for anything outside of debugging). The reason we\nrecord FF-ness is for later debugging or analysis. We don't prune from\nthe audit log, and we don't consider it for reachability when we prune\nobjects (since otherwise you'd never be able to prune anything!). So the\nobjects sometimes aren't available later to compute, but we still want\nto know if the user did a force-push, etc.\n\nI don't think that really applies to regular reflogs, because they do\nimply reachability (and they are not great for later analysis, because\nwe may selectively expire unreachable entries).\n\n-Peff\n"},{"id":"419663","messageId":"4400050.5IlZNYTcJN@mfick-lnx","threadId":"55341","inReplyTo":"CAFQ2z_MavgAGDyJzc9-+j6zTDODP7hCdPHtB5dyx-reLMSLX3Q@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-03-18T22:24:03Z","receivedAt":"2021-03-18T22:29:54Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Thursday, March 18, 2021 9:58:56 AM MDT Han-Wen Nienhuys wrote:\n> On Wed, Mar 17, 2021 at 10:22 PM Martin Fick <mfick@codeaurora.org> wrote:\n> > On Wednesday, March 17, 2021 9:06:06 PM MDT Han-Wen Nienhuys wrote:\n> > > I'm working on some extensions to Gerrit for which it would be very\n> > > beneficial if we could tell from the reflog if an update is a\n> > > fast-forward or not: if we find a SHA1 in the reflog, and see there\n> > > were only FF updates since, we can be sure that the SHA1 is reachable\n> > > from the branch, without having to open packfiles and decode commits.\n> > \n> > I don't think this would be reliable.\n> > \n> > 1) Not all updates make it to the reflogs\n> > 2) Reflogs can be edited or mucked with\n> > 3) On NFS reflogs can outright be wrong even when used properly as their\n> > are caching issues. We specifically have seen entries that appear to be\n> > FFs that were not.\n> \n> Can you tell a little more about 3) ? SInce we don't annotate non-FF\n> vs FF today, what does \"appear to be FFs\" mean?\n\nTo be honest I don't recall for sure, but I will describe what I think has \nhappened. I think that we have seen a server(A) update a branch from\nC1 to C2A, and then later another server(B) update the same branch from C1 to \nC2B. Obviously the move from C2A to C2B is not a FF, but that move is not what \nis recorded. Each of those updates was a FF when viewed as separate entries, \nbut if you look at both lines you can see that the second entry does not start \nwhere the first one left off. This would be detectable, it would appear as if \nthere was a missed entry that did a rewind from C2A to C1, but that rewind \npresumably actually came in as part of the second update from server (B)  as \nit had a cached version of the branch and believed it still pointed to C1 when \nit made its update,\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"419665","messageId":"21556546.pz8f0XPdgJ@mfick-lnx","threadId":"55341","inReplyTo":"CAFQ2z_MavgAGDyJzc9-+j6zTDODP7hCdPHtB5dyx-reLMSLX3Q@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-03-18T22:31:24Z","receivedAt":"2021-03-18T22:37:45Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Thursday, March 18, 2021 9:58:56 AM MDT Han-Wen Nienhuys wrote:\n> The bitmaps are generated by GC, and you can't GC all the time. \n\nI believe that I recently saw an effort to make this incremental, perhaps \nrelated to the geometric repacking series? If that were the case, you could gc \nmuch more often cheaply. Perhaps it could be something done on every upload at \nsome point the way that reflog effectively does on every update?\n\nAs Peff pointed out though, commit-graphs would still be a better way to go \nanyway since that is closer to their intent,\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"419667","messageId":"YFPaMj9PVfXlp4bL@coredump.intra.peff.net","threadId":"55341","inReplyTo":"21556546.pz8f0XPdgJ@mfick-lnx","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-03-18T22:54:42Z","receivedAt":"2021-03-18T22:55:21Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 18, 2021 at 04:31:24PM -0600, Martin Fick wrote:\n\n> On Thursday, March 18, 2021 9:58:56 AM MDT Han-Wen Nienhuys wrote:\n> > The bitmaps are generated by GC, and you can't GC all the time. \n> \n> I believe that I recently saw an effort to make this incremental, perhaps \n> related to the geometric repacking series? If that were the case, you could gc \n> much more often cheaply. Perhaps it could be something done on every upload at \n> some point the way that reflog effectively does on every update?\n\nThat geometric repacking work is leading up to having a bitmap for a\nmulti-pack-index. Which will make them _cheaper_, but still not\nespecially cheap (because we've reordered the objects corresponding to\neach bit, and also because our writing process still does a lot of\nO(nr_commits) work).\n\nIn the very long run, I think the way out would be to stop using pack or\nmidx ordering as the basis of the bitmap, and instead have a stable\nobject ordering that can be appended to. That would allow true\nincremental generation of the bitmaps (leaving old ones in place, and\njust adding a new ones to represent new commits). But that's such a big\ndeparture from the status quo that having a midx bitmap seemed like a\nmore attainable middle ground in the meantime.\n\n-Peff\n"},{"id":"419946","messageId":"CAFQ2z_NaXC-h5U3v2JtrFU8rGNHstTSN3CDoazUiUUk522sW7A@mail.gmail.com","threadId":"55341","inReplyTo":"4400050.5IlZNYTcJN@mfick-lnx","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-22T12:31:25Z","receivedAt":"2021-03-22T12:33:06Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Thu, Mar 18, 2021 at 11:24 PM Martin Fick <mfick@codeaurora.org> wrote:\n>\n> On Thursday, March 18, 2021 9:58:56 AM MDT Han-Wen Nienhuys wrote:\n> > On Wed, Mar 17, 2021 at 10:22 PM Martin Fick <mfick@codeaurora.org> wrote:\n> > > On Wednesday, March 17, 2021 9:06:06 PM MDT Han-Wen Nienhuys wrote:\n> > > > I'm working on some extensions to Gerrit for which it would be very\n> > > > beneficial if we could tell from the reflog if an update is a\n> > > > fast-forward or not: if we find a SHA1 in the reflog, and see there\n> > > > were only FF updates since, we can be sure that the SHA1 is reachable\n> > > > from the branch, without having to open packfiles and decode commits.\n> > >\n> > > I don't think this would be reliable.\n> > >\n> > > 1) Not all updates make it to the reflogs\n> > > 2) Reflogs can be edited or mucked with\n> > > 3) On NFS reflogs can outright be wrong even when used properly as their\n> > > are caching issues. We specifically have seen entries that appear to be\n> > > FFs that were not.\n> >\n> > Can you tell a little more about 3) ? SInce we don't annotate non-FF\n> > vs FF today, what does \"appear to be FFs\" mean?\n>\n> To be honest I don't recall for sure, but I will describe what I think has\n> happened. I think that we have seen a server(A) update a branch from\n> C1 to C2A, and then later another server(B) update the same branch from C1 to\n> C2B. Obviously the move from C2A to C2B is not a FF, but that move is not what\n> is recorded. Each of those updates was a FF when viewed as separate entries,\n\nI think those would fail with the way that Gerrit uses JGit, because\nC1 -> C2B would fail with LOCK_ERROR. I guess there are code paths in\nGit (?) that will execute force-push without checking if the update is\nFF or not.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419948","messageId":"87eeg7qpyr.fsf@evledraar.gmail.com","threadId":"55341","inReplyTo":"CAFQ2z_MefCwiWdhs0buJv5Zok+nsgaOvUCcsSnfm_PP0WozZKA@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-03-22T13:26:36Z","receivedAt":"2021-03-22T13:27:17Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Mar 17 2021, Han-Wen Nienhuys wrote:\n\n> Hi there,\n>\n> I'm working on some extensions to Gerrit for which it would be very\n> beneficial if we could tell from the reflog if an update is a\n> fast-forward or not: if we find a SHA1 in the reflog, and see there\n> were only FF updates since, we can be sure that the SHA1 is reachable\n> from the branch, without having to open packfiles and decode commits.\n>\n> For the reftable format, I think we could store this easily by\n> introducing more record types. [snip].\n\nAside from what others have mentioned here, you're talking about the\nlog_type field are you not? I.e.:\nhttps://googlers.googlesource.com/sop/jgit/+/reftable/Documentation/technical/reftable.md#log-block-format\n\nHas that \"log_type = 0x0\" tombstone proven to be a worthwhile\noptimization past the stash case mention there (which is presumably not\nrelevant to the vast majority of Google's use-cases).\n\nI.e. it's redundant to looking at the record and seeing if new_id =\nZERO_OID.\n\nSimilarly can't ff v.s. non-ff be deduced unambiguously by looking ahead\nto the next record, and seeing if the current record's \"old_id\" matches\nthat of the last record's \"new_id\". If it does it's a FF, if not it's a\nnon-FF (or a create/delete).\n\nI'm not arguing that a quicker lookup isn't needed, I'm just trying to\ndig at what \"beneficial\" here is. The format is ordered, and the common\ncase is that the page we have in memory has the last record.\n\nWhat sort of case are we talking about where not unpacking the log_data\nsegment is making a difference?\n\n> However, the textual reflog format doesn't easily allow for this.\n> However, we might add a convention, eg. have the message start with\n> 'FF' or 'NFF' depending on the nature of the update.\n\nMaybe a bit ugly, but a \"..\" and \"...\" prefix would at least be\nconsistent with \"fetch\" output. Or e.g. \"commit:\" and \"+commit:\" for ff\nand non-ff (and we could make it \"\\t commit:\" v.s. \"\\t+commit:\"\nv.s. current \"\\tcommit:\" to distinguish all three in the current\ntext-based format. Per \"OUTPUT\" in git-fetch(1).\n\n> [Ævar: snipped from earlier] Today we have 0 = deletion, 1 = update,\n> and we could add 2 = FF update, 3 = non-FF update.\n\nI've written log table implementations (a site table in a RDBMS) for git\n(one table for refs) which had:\n\n    create, ff, non-ff, delete\n\nI wonder if that quad-state would be useful for reftable too, with this\nproposed change you'd still need to unpack the record and see if the\nold_id is ZERO_OID to check if it's a creation, would you not?\n\nI also wonder if it couldn't be:\n\n    0 = deletion, 1 = non-ff-update, 2 = ff-update, 4 = creation\n\nSo the format wouldn't forever carry the historical wart of this not\nhaving been considered from the beginning.\n\nIt would mean that the few current reftable users (just Google?) would\nhave to look at the record to see if it's *really* a non-ff-update, but\npresumably they need to do so now for ff v.s. non-ff, so they're no\nworse off than they are now.\n\nThen when those users know they're on a version that distinguishes these\nthey can hard rely on 1 not being a \"ff for sure\", not a \"maybe\" status\nfor new updates. Presumably they either don't care about ancient reflog\nrecords, or a one-off migration of rewriting the records for older\nentries could be done.\n\nAlso between my [1] and this proposal we have at least a reftable v1.01\nin the wild (the filename locking behavior change discussed in [1]), and\nthis would make it v1.02, but the only up-to-date spec is for v1.00 (and\nmaybe JGit has other changes I haven't tracked).\n\nThat [1] change is minor, but still, a spec change.\n\nSo just a *poke* that having some version where the spec is kept\nup-to-date with that and this change if it happens would be very useful,\nespecially if the reftable-in-git.git lands one of these days.\n\n1. https://lore.kernel.org/git/87k0tzulf1.fsf@evledraar.gmail.com/\n"},{"id":"419951","messageId":"CAFQ2z_N8tCMZG62rNSY=HoRGuKnfk1W-Y_GOXz3SeaZO6=cWWA@mail.gmail.com","threadId":"55341","inReplyTo":"YFOuT6L0dsrCGTBk@coredump.intra.peff.net","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-22T14:40:46Z","receivedAt":"2021-03-22T14:41:51Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Thu, Mar 18, 2021 at 8:47 PM Jeff King <peff@peff.net> wrote:\n> > I'm working on some extensions to Gerrit for which it would be very\n> > beneficial if we could tell from the reflog if an update is a\n> > fast-forward or not: if we find a SHA1 in the reflog, and see there\n> > were only FF updates since, we can be sure that the SHA1 is reachable\n> > from the branch, without having to open packfiles and decode commits.\n>\n> I left some numbers in another part of the thread, but IMHO performance\n> isn't that compelling a reason to do this these days, if you are using\n> commit-graphs.\n>\n> Just walking the reflog might be _slightly_ faster, though not\n> necessarily (it depends on whether the depth of the object graph or the\n> depth of the reflog chain is deeper). It might matter more if you are\n> using a more exotic storage scheme, where switching from accessing\n> reflogs to objects implies extra round-trips to a server (e.g., custom\n> storage backends with JGit; I don't know the state of the art in what\n> Google is doing there).\n\nJGit doesn't currently support commit-graph, so it's hard to predict\nwhat performance will be like, but isn't commit-graph is keyed by\nSHA1? That makes it hard to do caching, especially when considering\nlarge repositories.\n\nAFAIU, commit-graph would help speed up reachability checks, by being\nable to shortcut cases where the commit number proves that some commit\nis not ancestor of the other, but you still have to do a revwalk to\nconclusively prove reachability.\n\nIn our storage system, the revwalk runs on top of packfile data that\nmust be faulted-in (slow!) from datacenter-wide storage. It's made\nworse because we don't support midx yet.\n\nThe application that I'm thinking of providing a way for automation to\ndeal with lagging replicas. This could be done by specifying a\n\n  X-Need-GitRef: $repositoryname~$refname~$SHA1\n\nheader on Gerrit requests, that specify that the given $SHA1 must have\nbeen in a recent ref update, and be reachable from $refname. The\nreflog has this information organized in a form that suited very well\nto answering these questions quickly (assuming the reflog is annotated\nsuch that we can distinguish FF and non-FF updates)\n\n\n\n> > Does this make sense, and if yes is it worth proposing a change?\n>\n> At GitHub we do something similar. We don't generally use reflogs much\n> at all, but we keep a custom \"audit log\": a single append-only file that\n> records every ref update in the repository. And its format just happens\n> to be one reflog entry per line, prefixed by the updated ref.\n\nThe interest of having a standard/convention in Git would be to not\nrequire reftable for folks that want to use this feature.\n\n> And there we do generally annotate the FF-ness of an update by stuffing\n> it into the free-form message field (in fact, we shove in a small JSON\n> object, so we record multiple fields like the pushing id, IP, etc).\n>\n> But the main goal there isn't performance (and in fact we don't\n> generally consult it for anything outside of debugging). The reason we\n> record FF-ness is for later debugging or analysis. We don't prune from\n> the audit log, and we don't consider it for reachability when we prune\n> objects (since otherwise you'd never be able to prune anything!). So the\n> objects sometimes aren't available later to compute, but we still want\n> to know if the user did a force-push, etc.\n\nWe store reflogs in a global database table, which has this kind of\ninformation, but the Google-specific format is harder to make work\nwith Gerrit, which is open source.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\nRegistergericht und -nummer: Hamburg, HRB 86891\nSitz der Gesellschaft: Hamburg\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419952","messageId":"CAFQ2z_NSh3XxjGx56r=xBP2WBk7ggUjh4rXSb5ivPtkS_r4iBQ@mail.gmail.com","threadId":"55341","inReplyTo":"87eeg7qpyr.fsf@evledraar.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-22T14:59:40Z","receivedAt":"2021-03-22T15:00:54Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Mon, Mar 22, 2021 at 2:26 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> > I'm working on some extensions to Gerrit for which it would be very\n> > beneficial if we could tell from the reflog if an update is a\n> > fast-forward or not: if we find a SHA1 in the reflog, and see there\n> > were only FF updates since, we can be sure that the SHA1 is reachable\n> > from the branch, without having to open packfiles and decode commits.\n> >\n> > For the reftable format, I think we could store this easily by\n> > introducing more record types. [snip].\n>\n> Aside from what others have mentioned here, you're talking about the\n> log_type field are you not? I.e.:\n> https://googlers.googlesource.com/sop/jgit/+/reftable/Documentation/technical/reftable.md#log-block-format\n\nCorrect.\n\n> Has that \"log_type = 0x0\" tombstone proven to be a worthwhile\n> optimization past the stash case mention there (which is presumably not\n> relevant to the vast majority of Google's use-cases).\n\nI've never really understood the log_type=0x0 use case. I think it was\nadded solely to cater for a use case in CGit's stash command.\n\n> I.e. it's redundant to looking at the record and seeing if new_id =\n> ZERO_OID.\n>\n> Similarly can't ff v.s. non-ff be deduced unambiguously by looking ahead\n> to the next record, and seeing if the current record's \"old_id\" matches\n> that of the last record's \"new_id\". If it does it's a FF, if not it's a\n> non-FF (or a create/delete).\n\nI don't see how that will tell you FF vs non-FF-ness.  Both an FF\nupdate and a non-FF  update look like 'new_oid = 20-random-bytes'.\nBarring further info, you have to lookup the commit object for those\nbytes, and then walk back to see if you pass old_oid.\n\nAFAICT, a correct sequence of ref updates (FF or not) always has\nprev.new_oid = current.old_oid.\n\n> > [Ævar: snipped from earlier] Today we have 0 = deletion, 1 = update,\n> > and we could add 2 = FF update, 3 = non-FF update.\n>\n> I've written log table implementations (a site table in a RDBMS) for git\n> (one table for refs) which had:\n>\n>     create, ff, non-ff, delete\n>\n> I wonder if that quad-state would be useful for reftable too, with this\n> proposed change you'd still need to unpack the record and see if the\n> old_id is ZERO_OID to check if it's a creation, would you not?\n\nDelete & create are handled with ZERO_OID.\n\nThe reftable format makes it so that you have to decode a record in\norder to read past it (there is no size framing the table entry\nlevel), so there is no big performance advantage in encoding this\ninformation in the log_type. You merely use a log_type bit rather than\na 20 byte raw ID. Since log records are zlib compressed anyway, it\nprobably also makes no space difference.\n\n> I also wonder if it couldn't be:\n>\n>     0 = deletion, 1 = non-ff-update, 2 = ff-update, 4 = creation\n>\n> So the format wouldn't forever carry the historical wart of this not\n> having been considered from the beginning.\n\nIf you do it like this, you will force that all implementations to\nhave to compute whether a (forced) update is a FF or not. I don't know\nif that is a problem. A 'maybe non-FF' value would be useful. Perhaps\nwe could even do simply\n\n     0 = deletion, 1 = maybe-ff-update, 2 = guaranteed-ff-update\n\n> It would mean that the few current reftable users (just Google?) would\n> have to look at the record to see if it's *really* a non-ff-update, but\n> presumably they need to do so now for ff v.s. non-ff, so they're no\n> worse off than they are now.\n\nAt Google, we currently don't record log records in reftable yet. From\nour perspective, we could probably change the standard 'in place'.\nJGit has supported reftable since Nov 2019, but I'm unaware of users;\nI did hear about GerritForge wanting to try it out in production this\nyear.\n\n> Then when those users know they're on a version that distinguishes these\n> they can hard rely on 1 not being a \"ff for sure\", not a \"maybe\" status\n> for new updates. Presumably they either don't care about ancient reflog\n> records, or a one-off migration of rewriting the records for older\n> entries could be done.\n>\n> Also between my [1] and this proposal we have at least a reftable v1.01\n> in the wild (the filename locking behavior change discussed in [1]), and\n> this would make it v1.02, but the only up-to-date spec is for v1.00 (and\n> maybe JGit has other changes I haven't tracked).\n\nThe file locking update has been added to the standard,\nhttps://github.com/git/git/pull/951.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419953","messageId":"87blbbqju3.fsf@evledraar.gmail.com","threadId":"55341","inReplyTo":"CAFQ2z_NSh3XxjGx56r=xBP2WBk7ggUjh4rXSb5ivPtkS_r4iBQ@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-03-22T15:39:00Z","receivedAt":"2021-03-22T15:39:56Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 22 2021, Han-Wen Nienhuys wrote:\n\n> On Mon, Mar 22, 2021 at 2:26 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>> > I'm working on some extensions to Gerrit for which it would be very\n>> > beneficial if we could tell from the reflog if an update is a\n>> > fast-forward or not: if we find a SHA1 in the reflog, and see there\n>> > were only FF updates since, we can be sure that the SHA1 is reachable\n>> > from the branch, without having to open packfiles and decode commits.\n>> >\n>> > For the reftable format, I think we could store this easily by\n>> > introducing more record types. [snip].\n>>\n>> Aside from what others have mentioned here, you're talking about the\n>> log_type field are you not? I.e.:\n>> https://googlers.googlesource.com/sop/jgit/+/reftable/Documentation/technical/reftable.md#log-block-format\n>\n> Correct.\n>\n>> Has that \"log_type = 0x0\" tombstone proven to be a worthwhile\n>> optimization past the stash case mention there (which is presumably not\n>> relevant to the vast majority of Google's use-cases).\n>\n> I've never really understood the log_type=0x0 use case. I think it was\n> added solely to cater for a use case in CGit's stash command.\n>\n>> I.e. it's redundant to looking at the record and seeing if new_id =\n>> ZERO_OID.\n>>\n>> Similarly can't ff v.s. non-ff be deduced unambiguously by looking ahead\n>> to the next record, and seeing if the current record's \"old_id\" matches\n>> that of the last record's \"new_id\". If it does it's a FF, if not it's a\n>> non-FF (or a create/delete).\n>\n> I don't see how that will tell you FF vs non-FF-ness.  Both an FF\n> update and a non-FF  update look like 'new_oid = 20-random-bytes'.\n> Barring further info, you have to lookup the commit object for those\n> bytes, and then walk back to see if you pass old_oid.\n\nBecause both the next reflog format and reftable's have the update-ref\n<newvalue> <oldvalue> for each update. So:\n\n    $ cut -d ' ' -f1-2 .git/logs/refs/remotes/origin/master | head -n 2\n    1c52ecf4ba0f4f7af72775695fee653f50737c71 ba2aa15129e59f248d8cdd30404bc78b5178f61d\n    ba2aa15129e59f248d8cdd30404bc78b5178f61d 6d3ef5b467eccd2769f1aa1c555d317d3c8dc707\n\nWe can know with !strcmp(rows[0][1], rows[1][0]) whether the latest\nupdate is a ff or non-ff (it is). As opposed to:\n\n    $ cut -d ' ' -f1-2 .git/logs/refs/remotes/origin/seen | head -n 2\n    8b26a41f4bf7e7c0f097cb91012d08fe8ae30e7f 0f3a981cbd5be5f97e9504ab770cd88f988fe820\n    0f3a981cbd5be5f97e9504ab770cd88f988fe820 fdd019edfe6bc60d0100d5751c41e4f6ad28a2ef\n\nWhere the rows[0][1] value is not the same as rows[1][0].\n\nWhat you can't do is find whether any given commit(s) in those ranges\nare part of the FF or not, since you just have the start/end points.\n\nBut I'm vaguely paranoid that we're talking past one another here and\nI've misunderstood what you want...\n\n> AFAICT, a correct sequence of ref updates (FF or not) always has\n> prev.new_oid = current.old_oid.\n>\n>> > [Ævar: snipped from earlier] Today we have 0 = deletion, 1 = update,\n>> > and we could add 2 = FF update, 3 = non-FF update.\n>>\n>> I've written log table implementations (a site table in a RDBMS) for git\n>> (one table for refs) which had:\n>>\n>>     create, ff, non-ff, delete\n>>\n>> I wonder if that quad-state would be useful for reftable too, with this\n>> proposed change you'd still need to unpack the record and see if the\n>> old_id is ZERO_OID to check if it's a creation, would you not?\n>\n> Delete & create are handled with ZERO_OID.\n>\n> The reftable format makes it so that you have to decode a record in\n> order to read past it (there is no size framing the table entry\n> level), so there is no big performance advantage in encoding this\n> information in the log_type. You merely use a log_type bit rather than\n> a 20 byte raw ID. Since log records are zlib compressed anyway, it\n> probably also makes no space difference.\n>\n>> I also wonder if it couldn't be:\n>>\n>>     0 = deletion, 1 = non-ff-update, 2 = ff-update, 4 = creation\n>>\n>> So the format wouldn't forever carry the historical wart of this not\n>> having been considered from the beginning.\n>\n> If you do it like this, you will force that all implementations to\n> have to compute whether a (forced) update is a FF or not. I don't know\n> if that is a problem. A 'maybe non-FF' value would be useful. Perhaps\n> we could even do simply\n>\n>      0 = deletion, 1 = maybe-ff-update, 2 = guaranteed-ff-update\n\nThe point is that nobody relies on this now, but also that you'd of\ncourse want 1 = i-know-for-sure-non-ff-update, 2 =\ni-know-for-sure-ff-update.\n\nBut since logs are transitory (aren't they also expired in reftable,\ncan't see that from skimming hte spec) and this is just a helpful\nside-index I think a one-time migration for the tiny minority of\nreftable users who'd care would be preferrable to the vast majority of\nfuture reftable users (if it ever lands in git.git) having to deal with\nthis (albeit small) special-case forever.\n\n>> It would mean that the few current reftable users (just Google?) would\n>> have to look at the record to see if it's *really* a non-ff-update, but\n>> presumably they need to do so now for ff v.s. non-ff, so they're no\n>> worse off than they are now.\n>\n> At Google, we currently don't record log records in reftable yet. From\n> our perspective, we could probably change the standard 'in place'.\n> JGit has supported reftable since Nov 2019, but I'm unaware of users;\n> I did hear about GerritForge wanting to try it out in production this\n> year.\n\nAh, so not even Google's using it, the backcompat is just for good\nmeasure...\n\n>> Then when those users know they're on a version that distinguishes these\n>> they can hard rely on 1 not being a \"ff for sure\", not a \"maybe\" status\n>> for new updates. Presumably they either don't care about ancient reflog\n>> records, or a one-off migration of rewriting the records for older\n>> entries could be done.\n>>\n>> Also between my [1] and this proposal we have at least a reftable v1.01\n>> in the wild (the filename locking behavior change discussed in [1]), and\n>> this would make it v1.02, but the only up-to-date spec is for v1.00 (and\n>> maybe JGit has other changes I haven't tracked).\n>\n> The file locking update has been added to the standard,\n> https://github.com/git/git/pull/951.\n\nAh yes. Now I remember. I managed to miss/not recall that. Thanks.\n"},{"id":"419954","messageId":"CAFQ2z_ML8s0Gk4Zmg+2mxzkfP1AbL=zkeUG0yKEtoege7it-vA@mail.gmail.com","threadId":"55341","inReplyTo":"87blbbqju3.fsf@evledraar.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-22T15:56:32Z","receivedAt":"2021-03-22T15:57:34Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Mon, Mar 22, 2021 at 4:39 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> We can know with !strcmp(rows[0][1], rows[1][0]) whether the latest\n> update is a ff or non-ff (it is). As opposed to:\n>\n>     $ cut -d ' ' -f1-2 .git/logs/refs/remotes/origin/seen | head -n 2\n>     8b26a41f4bf7e7c0f097cb91012d08fe8ae30e7f 0f3a981cbd5be5f97e9504ab770cd88f988fe820\n>     0f3a981cbd5be5f97e9504ab770cd88f988fe820 fdd019edfe6bc60d0100d5751c41e4f6ad28a2ef\n>\n> Where the rows[0][1] value is not the same as rows[1][0].\n\nI'm confused.\n\nrows[0][1] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\nrows[1][0] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n\nthey are the same. I don't understand your argument.\n\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\nRegistergericht und -nummer: Hamburg, HRB 86891\nSitz der Gesellschaft: Hamburg\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419957","messageId":"878s6fqgze.fsf@evledraar.gmail.com","threadId":"55341","inReplyTo":"CAFQ2z_ML8s0Gk4Zmg+2mxzkfP1AbL=zkeUG0yKEtoege7it-vA@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-03-22T16:40:37Z","receivedAt":"2021-03-22T16:41:43Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 22 2021, Han-Wen Nienhuys wrote:\n\n> On Mon, Mar 22, 2021 at 4:39 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>> We can know with !strcmp(rows[0][1], rows[1][0]) whether the latest\n>> update is a ff or non-ff (it is). As opposed to:\n>>\n>>     $ cut -d ' ' -f1-2 .git/logs/refs/remotes/origin/seen | head -n 2\n>>     8b26a41f4bf7e7c0f097cb91012d08fe8ae30e7f 0f3a981cbd5be5f97e9504ab770cd88f988fe820\n>>     0f3a981cbd5be5f97e9504ab770cd88f988fe820 fdd019edfe6bc60d0100d5751c41e4f6ad28a2ef\n>>\n>> Where the rows[0][1] value is not the same as rows[1][0].\n>\n> I'm confused.\n>\n> rows[0][1] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n> rows[1][0] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n>\n> they are the same. I don't understand your argument.\n\nSorry, I mean same = ff update, not the same = non-ff. So I flipped\nthose around in describing it.\n\nBut in any case, the point is that you can reliably tell from a log of\nupdates whether individual updates are ff or non-ff, as long as you:\n\n1. Look at two records but not one, or rather the ff-ness of the update\n   you're looking at now depends on <oldvalue> of that update matching\n   the <newvalue> of the update before that.\n\n2. The log is guaranteed to be in the same order as the update-ref\n   calls, and not to be pruned in the middle.\n\n3. Don't care about the FF-ness of the first entry in the reflog\n   (e.g. prune it at N entries or 2 weeks, you can't see if the oldest\n   entry you have is a FF or not)\n\nSo as I noted I've materialized this in a RDBMS before as something like\nENUM('delete', 'create', 'ff', 'non-ff'), but that was:\n\nA. In the delete/create case as much about the convenience of\n   selecting/grouping as anything else, easier to select operation =\n   \"delete\" than {old,new} = \"<copy paste or calefully type exactly 40 x\n   '0'>\"\n\nB. In SQL it's a PITA both typing and index-wise to get information\n   about a record based on a \"previous\" record when the relation between\n   the two is a a UNIQUE INDEX and AUTO-INCREMENT id. I.e. for an\n   AUTO_INCREMENT UNIQUE INDEX of repo,refname,id to get the \"last\"\n   record you need to select where repo && refname is the same, and id <\n   current_id LIMIT 1.\n\nC. You're accessing the data ad-hoc & manually via SQL, not some\n   Git-specific query language/wrapper.\n\nBut in the case of reftable I'm wondering what the use-case is, since\npresumably whatever API now serves up the reflog from it can just as\nwell look at one more record and fill in a matrialized \"ff\" or \"non-ff\"\nfield for the current record based on that.\n\nAlso: If you're adding more values a genuinely useful value to log (that\nI've seen custom logging for) is \"was this a forced push? || update-ref\nwithout the <oldvalue>?\", which is *not* the same as \"not a\nfast-forward\". I.e. it's the difference between \"update this with this\n<oldvalue>\" (ff or non-ff) v.s. \"I don't care what the oldvalue is, just\nupdate it\" (f orce or not).\n\n(To complicate matters and make them even more harder to explain, the\nwhole \"force with lease\" thing is not actually a \"force push\" in that\nsense, it's just an update-ref with an oldvalue you expect to result in\na non-ff update).\n"},{"id":"419960","messageId":"CAFQ2z_OVbcxZtb-hS1DZ3X7FpdQD25W0VqtGLXdkQ4bjgCPS-w@mail.gmail.com","threadId":"55341","inReplyTo":"878s6fqgze.fsf@evledraar.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2021-03-22T17:12:42Z","receivedAt":"2021-03-22T17:13:42Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Mon, Mar 22, 2021 at 5:40 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, Mar 22 2021, Han-Wen Nienhuys wrote:\n>\n> > On Mon, Mar 22, 2021 at 4:39 PM Ævar Arnfjörð Bjarmason\n> > <avarab@gmail.com> wrote:\n> >> We can know with !strcmp(rows[0][1], rows[1][0]) whether the latest\n> >> update is a ff or non-ff (it is). As opposed to:\n> >>\n> >>     $ cut -d ' ' -f1-2 .git/logs/refs/remotes/origin/seen | head -n 2\n> >>     8b26a41f4bf7e7c0f097cb91012d08fe8ae30e7f 0f3a981cbd5be5f97e9504ab770cd88f988fe820\n> >>     0f3a981cbd5be5f97e9504ab770cd88f988fe820 fdd019edfe6bc60d0100d5751c41e4f6ad28a2ef\n> >>\n> >> Where the rows[0][1] value is not the same as rows[1][0].\n> >\n> > I'm confused.\n> >\n> > rows[0][1] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n> > rows[1][0] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n> >\n> > they are the same. I don't understand your argument.\n>\n> Sorry, I mean same = ff update, not the same = non-ff. So I flipped\n> those around in describing it.\n>\n> But in any case, the point is that you can reliably tell from a log of\n> updates whether individual updates are ff or non-ff, as long as you:\n\n?\n\n\n$ git checkout -b non-ff origin/master\nPrevious HEAD position was 59f4325222 Add \"test-tool dump-reftable\" command.\nBranch 'non-ff' set up to track remote branch 'master' from 'origin'.\nSwitched to a new branch 'non-ff'\n\n$ git reset --hard HEAD^\nHEAD is now at 56a57652ef Sync with Git 2.30.2 for CVE-2021-21300\n\n$ cat .git/logs/refs/heads/non-ff\n0000000000000000000000000000000000000000\n13d7ab6b5d7929825b626f050b62a11241ea4945 Han-Wen Nienhuys\n<hanwen@google.com> 1616432938 +0100 branch: Created from\norigin/master\n13d7ab6b5d7929825b626f050b62a11241ea4945\n56a57652ef8e4ca2f108a8719b8caeed5e153c95 Han-Wen Nienhuys\n<hanwen@google.com> 1616432948 +0100 reset: moving to HEAD^\n\nHow can I tell from this reflog that \"moving to HEAD^\" is a non-FF update?\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Halimah DeLaine Prado\n"},{"id":"419962","messageId":"3334455.nVGR68FYPA@mfick-lnx","threadId":"55341","inReplyTo":"CAFQ2z_NaXC-h5U3v2JtrFU8rGNHstTSN3CDoazUiUUk522sW7A@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2021-03-22T17:45:04Z","receivedAt":"2021-03-22T17:46:29Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, March 22, 2021 1:31:25 PM MDT Han-Wen Nienhuys wrote:\n> On Thu, Mar 18, 2021 at 11:24 PM Martin Fick <mfick@codeaurora.org> wrote:\n> > On Thursday, March 18, 2021 9:58:56 AM MDT Han-Wen Nienhuys wrote:\n> > > On Wed, Mar 17, 2021 at 10:22 PM Martin Fick <mfick@codeaurora.org> \nwrote:\n> > > > On Wednesday, March 17, 2021 9:06:06 PM MDT Han-Wen Nienhuys wrote:\n> > > > > I'm working on some extensions to Gerrit for which it would be very\n> > > > > beneficial if we could tell from the reflog if an update is a\n> > > > > fast-forward or not: if we find a SHA1 in the reflog, and see there\n> > > > > were only FF updates since, we can be sure that the SHA1 is\n> > > > > reachable\n> > > > > from the branch, without having to open packfiles and decode\n> > > > > commits.\n> > > > \n> > > > I don't think this would be reliable.\n> > > > \n> > > > 1) Not all updates make it to the reflogs\n> > > > 2) Reflogs can be edited or mucked with\n> > > > 3) On NFS reflogs can outright be wrong even when used properly as\n> > > > their\n> > > > are caching issues. We specifically have seen entries that appear to\n> > > > be\n> > > > FFs that were not.\n> > > \n> > > Can you tell a little more about 3) ? SInce we don't annotate non-FF\n> > > vs FF today, what does \"appear to be FFs\" mean?\n> > \n> > To be honest I don't recall for sure, but I will describe what I think has\n> > happened. I think that we have seen a server(A) update a branch from\n> > C1 to C2A, and then later another server(B) update the same branch from C1\n> > to C2B. Obviously the move from C2A to C2B is not a FF, but that move is\n> > not what is recorded. Each of those updates was a FF when viewed as\n> > separate entries,\n> I think those would fail with the way that Gerrit uses JGit, because\n> C1 -> C2B would fail with LOCK_ERROR.\n\nIf jgit knew that the branch no longer pointed to C1, then yes it should fail \nwith LOCK_ERROR. However this situation is believed to arise due to jgit's \ncaching which could make it think that the branch still points to C1 even \nthought it has already been advanced to C2A! :(\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"419968","messageId":"xmqqh7l36nol.fsf@gitster.g","threadId":"55341","inReplyTo":"878s6fqgze.fsf@evledraar.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-03-22T18:36:10Z","receivedAt":"2021-03-22T18:37:02Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> I'm confused.\n>>\n>> rows[0][1] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n>> rows[1][0] == \"0f3a981cbd5be5f97e9504ab770cd88f988fe820\"\n>>\n>> they are the same. I don't understand your argument.\n>\n> Sorry, I mean same = ff update, not the same = non-ff. So I flipped\n> those around in describing it.\n\nI am confused too.  Are you tacking something else, a gap in a run\nof reflog entries?  If I go from commit A to B to C, the first log\nentry would record the transtion from A->B, and the second entry\nwould record the transition from B->C, and the lack of gap does not\nsay anything about the relationship between A and B, or B and C.  A\ncan be, and does not have to be, an ancestor of B, and B can be, and\ndoes not have to be, an ancestor of C.  Hopping from A to B to C would\nleave the same pair of reflog records and I do not think you can tell\nthe reachability among A and B and C from them.\n\n"},{"id":"420239","messageId":"YF2Qr/ClD4a3jCUQ@coredump.intra.peff.net","threadId":"55341","inReplyTo":"CAFQ2z_N8tCMZG62rNSY=HoRGuKnfk1W-Y_GOXz3SeaZO6=cWWA@mail.gmail.com","subject":"Re: Distinguishing FF vs non-FF updates in the reflog?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-03-26T07:43:43Z","receivedAt":"2021-03-26T07:44:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 22, 2021 at 03:40:46PM +0100, Han-Wen Nienhuys wrote:\n\n> > I left some numbers in another part of the thread, but IMHO performance\n> > isn't that compelling a reason to do this these days, if you are using\n> > commit-graphs.\n> >\n> > Just walking the reflog might be _slightly_ faster, though not\n> > necessarily (it depends on whether the depth of the object graph or the\n> > depth of the reflog chain is deeper). It might matter more if you are\n> > using a more exotic storage scheme, where switching from accessing\n> > reflogs to objects implies extra round-trips to a server (e.g., custom\n> > storage backends with JGit; I don't know the state of the art in what\n> > Google is doing there).\n> \n> JGit doesn't currently support commit-graph, so it's hard to predict\n> what performance will be like, but isn't commit-graph is keyed by\n> SHA1? That makes it hard to do caching, especially when considering\n> large repositories.\n\nYes, it's keyed by sha1. It's essentially replacing \"inflate the commit\nobject and parse it\" with \"here are the parsed values as mmap-able\n32-bit integer fields\" (there's some other stuff with generation\nnumbers, too, but the main speedup is simply that accessing each commit\nis orders of magnitude cheaper).\n\nIt caches well, because those properties of the commit are immutable.\nBut if you meant \"when pulling data from the commit-graph file, is it\nfriendly to block cache\", then no, it's not linear. You'd binary search\nwithin it to find each commit, just as you would a pack .idx (and just\nlike a .idx, I'd expect a system that is pulling data from a network\nsource to want to grab the whole commit-graph file. They tend to be much\nsmaller than the main .idx for a given repo).\n\n> AFAIU, commit-graph would help speed up reachability checks, by being\n> able to shortcut cases where the commit number proves that some commit\n> is not ancestor of the other, but you still have to do a revwalk to\n> conclusively prove reachability.\n\nRight. You'll still walk a lot of the commits, but you'll do so much\nfaster (the generation numbers can also help prune some uninteresting\nside paths, but again, I think the main value for this operation is just\ngetting the parent info much faster).\n\n-Peff\n"}]}