{"thread":{"id":"52980","subject":"Notes from Git Contributor Summit, Los Angeles (April 5, 2020)","startedAt":"2020-03-12T03:55:31Z","lastAt":"2020-05-19T20:10:44Z","messageCount":106,"participants":["James Ramsay","Emily Shaffer","Derrick Stolee","Konstantin Ryabitsev","Jonathan Nieder","Junio C Hamano","Jeff King","Eric Wong","Jakub Narebski","Damien Robert","Konstantin Tokarev","Elijah Newren","Phillip Susi","Philip Oakley","Philippe Blain","Christian Couder","Taylor Blau","Eric Sunshine","Phillip Wood","Josh Steadmon","brian m. carlson","Noam Soloveichik","nbelakovski@gmail.com"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"393102","messageId":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","threadId":"52980","inReplyTo":null,"subject":"Notes from Git Contributor Summit, Los Angeles (April 5, 2020)","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T03:55:21Z","receivedAt":"2020-03-12T03:55:31Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"It was great to see everyone at the Contributor Summit last week, in \nperson and virtually.\n\nParticular thanks go to Peff for facilitating, and to GitHub for \norganizing the logistics of the meeting place and food. Thank you!\n\nOn the day, the topics below were discussed:\n\n1. Ref table (8 votes)\n2. Hooks in the future (7 votes)\n3. Obliterate (6 votes)\n4. Sparse checkout (5 votes)\n5. Partial Clone (6 votes)\n6. GC strategies (6 votes)\n7. Background operations/maintenance (4 votes)\n8. Push performance (4 votes)\n9. Obsolescence markers and evolve (4 votes)\n10. Expel ‘git shell’? (3 votes)\n11. GPL enforcement (3 votes)\n12. Test harness improvements (3 votes)\n13. Cross implementation test suite (3 votes)\n14. Aspects of merge-ort: cool, or crimes against humanity? (2 votes)\n15. Reachability checks (2 votes)\n16. “I want a reviewer” (2 votes)\n17. Security (2 votes)\n\nNotes were taken in the linked Google Doc, but for those who’d rather \nread the notes here, I’ll also send the notes as replies to this \nmessage.\n\nhttps://docs.google.com/document/d/15a_MPnKaEPbC92a4jhprlHvkyirDh2CtTtgOxNbnIbA/edit#heading=h.vvhyp0oa4hhz\n\nRegards,\nJames\n"},{"id":"393103","messageId":"1B71B54C-E000-4CEB-8AC6-3DB86E96E31A@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 1/17] Reftable","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T03:56:10Z","receivedAt":"2020-03-12T03:56:16Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. In case you’re not aware what it is. It was introduced in JGit. \n???Prefix table??\n\n2. Gerrit team likes to get this in cgit\n\n3. From the Stump the Experts yesterday, the question was “If you \ncould go back and change anything what would it be?”: Loose refs can \ncause difficulties. So it would be nice to make reftables a first-class \ncitizen. There are issues with OSes with case-insensitive filesystems. \nReftables can help with this.\n\n4. Stolee: contributing an entire copy of the source of a library \nelsewhere as one patch makes it hard to review, and doesn’t feel like \na contribution to Git.\n\n5. Brian: agree. Is it an external library that needs to be pulled in \nevery time a new version added in JGit.\n\n6. Edward: having it as external library moves the maintenance burden\n\n7. Jonathan N: example of xdiff, we have a copy, Mercurial has a copy, \nand they have been patched in different ways. Can we separate these \nconcerns? One: patches that can be reviewed separately. Two: licensing. \nThree: ongoing maintenance approach.\n\n8. Peff: benefits of external library are clear. What is the maintenance \nburden of not maintaining this in the core git tree. More concerned \nabout niceties in Git that aren’t in other libraries, like strbufs and \ndata structures. Lowest common denominator isn’t ideal. Can this cost \nbe mitigated?\n\n9. Ed: I have the same concerns. We also have strbufs, but they are not \nthe same. We also might run into licensing issues.\n\n10. Stolee: also cross platform compatibility… It might not perform \nwell on different platforms.\nPeff: It feels to me there are a lot of hairy filesystem details \nreftables need to do.\n\n11. Brian: Atomic renames have issues on Windows.\n\n12. Jonathan N: Han-Wen wanted a more substantial review, and we just \nprovided one (actionable for\n\n13. Jonathan: write a summary email to Han-Wen)\n\n14. Brian: (inaudible) Having a reftable library would be interesting to \ntest SHA256 changes.\n\n15. Stolee: would be nice to have tests regarding case-sensitivity & \ndirectory/file conflicts\n\n16. Ed: wait, are we loosening the restriction?\n\n17. Peff: no, for backwards-compatibility we cannot. Would love to get \nrid of that restriction, though.\n\n18. Jonathan N: Immediate benefit wrt D/F conflicts is being able to \nkeep reflogs for deleted branches\n"},{"id":"393104","messageId":"0D7F1872-7614-46D6-BB55-6FEAA79F1FE6@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 2/17] Hooks in the future","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T03:56:53Z","receivedAt":"2020-03-12T03:56:59Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Emily: hooks in the config file. Sent a read only patch, but didn’t \nget much traction. Add a new header in the config file, then have prefix \nnumber so that security scan could be configured at system level that \nwould run last, and then hook could also be configured at the project \nlevel.\n\n2. Peff: Having hooks in the config would be nice. But don’t do it at \n`hooks.prereceive`, but use a subconfig like `hooks.prereceive.command` \nso it’s possible to add options later on.\n\n3. Brian: sometimes the need to overridden, ordering works for me. For \nGit LFS it would be helpful to have a pre-push hook for LFS, and a \ndifferent one for something else. Want flexibility about finding and \ndiscovering hooks.\n\n4. Emily: if you want to specify a hook that is in the work tree, then \nit has to be configured after cloning.\n\n5. Jonathan: It’s better to start with something low complexity as \nlong as it can be extended/changed later. If there's room to tweak it \nover time then I'm not too worried about the initial version being \nperfect — we can make mistakes and learn from them. A challenge will \nbe how hooks interact. Analogy to the challenges of stacked union \nfilesystems and security modules in Linux. Analogy to sequence number \nallocation for unit scripts\n\n6. CB: Declare dependencies instead of a sequence number? In theory \nindependent hooks can also run in parallel.\n\n7. Peff: Maybe that’s something to not worry about from the start. \nLike, how many hooks do you expect to run anyway.\n\n8. Christian: At booking.com they use a lot of hooks, and they also sent \npatches to the mailing list to improve that.\n\n9. Emily: In-tree hooks?\n\n10. Brian: You can do `git cat-file <ref> | sh` to run a hook.\n\n11. Brandon: Is it possible to globally to disable all hooks locally? It \nmight be a security concern. Or is it something we might want to add?\n\n12. Peff: No it’s not.\n"},{"id":"393105","messageId":"5B2FEA46-A12F-4DE7-A184-E8856EF66248@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 3/17] Obliterate","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T03:57:24Z","receivedAt":"2020-03-12T03:57:35Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Jonathan N: sometimes people accidentally add a big file they don’t \nneed. Have to use BFG and it’s a pain. Next time, maybe you just deal \nwith it and ignore. This happened to Chrome. Some huge blob that was in \nthe repo, should no longer be in the repo, but don’t want to rewrite \nthe history. Other use cases are confidential information, like \npassword, credit card number etc. Initial reactions: it’s already out \nthere, rotate. Second reaction: if it’s a toxic blob it needs to be \nremoved everywhere! What if someone taught kernel repo to\n\n2. James: I’ve been in a lot of meetings with customers where they \nmentioned it’s not possible to rotate the information that was leaked \ninto the repo\n\n3. Demetr: How far back do we allow to go to obliterate?\n\n4. Jonathan N: there are indeed horrible real-world examples where \nthings to be obliterated are from a long time ago.\n\n5. James: real cost to changing object ids: Git and tools interacting \nwith it really assume that history is immutable.\n\n6. Elijah: replace refs helps, but not supported by hosts like GitHub \netc\n\n     a. Stolee: breaks commit graph because of generation numbers.\n     b. Replace refs for blobs, then special packfile, there were edge \ncases.\n\n7. Demetr: Backward compatibility, wouldn’t custom handling be \nproblematic for old clients.\n\n8. Jeff H: can we introduce a new type of object -- a \"revoked blob\" if \nyou will that burns the original one but also holds the original SHA in \nthe ODB ??\n\n9. Peff: what would this mean for signatures? New opportunity to forge \nsignatures.\n\n10. Jonathan N: if a new entity, this means you’ve changed the content \nwhich we want to avoid. Maybe a list of revoked blobs. If fsck notices \nmissing, it should be happy. Protocol support, if someone tries to \ninclude a patch with it, just ignore it. Not great. Improvement would be \nto send a list of things I deliberately didn’t send. Could also \ncommunicate blobs to be deleted, but ignore that for v1. Learn from \nMercurial who have a very complicated signed revocation mechanism.\n\n11. Brian: the remote can’t be trusted, ala leftpad maintainer could \ndo something malicious causing repo to become invalid.\n\n12. Jonathan N: main scenario I’m considering is trusted company \nremote.\n\n13. Terry: partial clone and solve large files. Maybe the server could \nhandle it by converting normal clone into partial, and then handle the \nerror if someone asks for that blob.\n\n14. Jakub: one idea would simply be to treat this as a missing blob in a \npartial clone\n\n15. Michael Haggerty: does this only apply to blobs? (Peff: no, commit \nmessages can contain sensitive information; Johannes: trees contain file \nnames which also can contain sensitive information)\n\n16. Jonathan N: partial clone is not a solution for the desire to get \nrid of the blob on the server side.\n"},{"id":"393106","messageId":"577F08D9-719A-4C66-B135-DC62E68AC941@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 4/17] Sparse checkout","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T03:58:06Z","receivedAt":"2020-03-12T03:58:12Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Stolee: built in! We’re making improvements! We’re looking at UX \n- add, remove, state. Harder things like grep needs everything, but not \nexpected.\n\n2. Elijah: we do have a strongly worded warning, so we could just change \nit.\n\n3. Jonathan N: I like both modes.\n\n4. Terry: how is GC wired up? If I change a cone, will it be reclaimed?\n\n5. Stolee: GC doesn’t remove reachable objects. Haven’t found people \nneed to do this, unless they accidentally rehydrated something massive \nthey didn’t really need. Day to day work doesn’t introduce too much.\n\n6. Terry: Android devs have massive special machines. Constantly running \nout of disk space.\n\n7. Stolee: more of a partial clone feature, than a sparse checkout \nfeature. If I checkout three branches, go offline, I don’t want GC to \nclean things that I had downloaded.\n\n8. Jonathan N: switching between Word and Powerpoint. Would it be useful \nto attach cone to branch rather than repo.\n\n9. Stolee: Office team is building some kind of magic to automatically \ndetect from branch.\n\n10. Brian: can use reflog maybe. Prune based on that? People who run out \nof disk space could have shorter reflog.\n\n11. Elijah: biggest problem people run into doing a rebase/pull, hit \nconflicts, then they need to update sparsity patterns, which they \ncan’t do because there are conflicts. Working on a patch.\n\n12. Stolee: Office scoper tool would automatically recalculate \ndependencies and update sparsity config so that they can build.\n\n13. During break, Minh brought up an idea that we could use in-tree data \nto manage the dependency chain: The tree could contain files that \ncontain directory names, and users use config to specify the list of \nthose files to use for the sparse-checkout definition. When Git updates \nthe working directory and those files change, the sparse-checkout can be \nupdated to include the union of those directories. Stolee will look into \nhow this could work and whether this works for existing customers.\n"},{"id":"393107","messageId":"58425B78-7C6D-41CD-92AE-434D0A58F968@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 5/17] Partial Clone","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:00:54Z","receivedAt":"2020-03-12T04:01:00Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Stolee: what is the status, who is deploying it, what issues need to \nbe handled? Example, downloading tags. Hard to highly recommend it.\n\n2. Taylor: we deployed it. No activity except for internal testing. Some \nmore activity, but no crashes. Have been dragging our feet. Chicken egg, \ncan’t deploy it because the client may not work, but hoping to hear \nabout problems.\n\n3. ZJ: dark launched for a mission critical repos. Internal questions \nfrom CI team, not sure about performance. Build farm hitting it hard \nwith different filter specs.\n\n4. Taylor: we have patches we are promising to the list. Blob none and \nlimit, for now, but add them incrementally. Bitmap patches are on the \nlist\n\n5. James: I’ve been talking to customers who have high interest in \nthis. But they are hesitant. Do people have similar situations, like \nshallow clones?\n\n6. Jonathan N: we’re not using it en masse with server farms (see \nTerry’s stats). Performance issues with catchup, long periods of \ndownloading and no progress. Missing progress display means a user waits \nand gets worried. On server side, reachability check can be expensive, \nin part because enumerating refs is expensive.\n\n7. Peff: client experience sucks with N+1 situations. If the server \noperator side is tolerable, that way it’s easier to move the client \nside forward. By default, v2 just serves them up, no reachability check. \nNot sure if we’ll do that forever. Often have to inflate blobs that \nare deltas, and then delta compression which is not needed.\n\n8. Stolee: Jonathan built a batch download when changing trees. Possible \nto improve by sending haves.\n\n9. Jonathan N: if you’re in the blob none filter, and say I have a \ncommit, I might not actually have what the server expects.\n\n10. Peff: could enumerate blobs\n\n12. Demetr: Partial clones are dangerous for DoS attacks\n\n12. Jonathan: JGit forbids most filters that can't use bitmaps.\n\n13. Peff: just blob filters? Yes, so far.\n\n14. Jonathan: as far as the client experience goes, we’re not batching \noften enough and not showing progress on catch-up fetches. Any other UX \nissues?\n\n15. Jeff: no, those two are what I meant.\n\n16. James: another question for git service providers: Is it a \nreplacement for LFS?\n\n17. Brian: some files can compress, others don’t. Repacking can blow \nup if you try to compress something that can’t be compressed. How do \nwe identify which objects we compress, and which we don’t.\n\n18. Jonathan N: if you see something already compressed, tell zlib to do \npassthrough compression.\n\n19. Taylor: two problems - which projects do you want to quarantine, \nwhere do you put them. CDN offloading would be nice.\n\n20. Stolee: reachability bitmaps are tied to a single packfile. Becomes \nmore and more expensive. Even just having them in another file requires \na lot of work.\n\n21. Taylor: we’re looking at some heuristics so that some parts of the \npack can just be moved over verbatim.\n\n22. Peff: I see three problems: multi pack lookups, bitmaps,\n\n23. Jonathan N: we never generate on the fly deltas\n\n24. Peff: there are pathological cases.\n\n25. Terry: we are seeing 89k partial clones per day. Majority is clone. \nShallow clone equivalent.\n\n26. Peff: why? Is it better?\n\n27. Jonathan N: initial clone is about the same as shallow. One reason \nwe encourage, if you do a follow up, with shallow clone it is expensive \nfor the server.\n\n28. Stolee: if you persist the previous shallow clone, it is much much \ncheaper to do incremental fetch.\n\n29. Terry: JGit has enough shallow clone bugs that we often just send \neverything. Make shallow clone obsolete\n\n30. Jonathan N: Jenkins style CI, option for shallow clone. Want to run \ndiff or git describe, have to turn it off. Partial clone is simpler.\n\n31. Minh: could the server force the client to partial clone?\n\n32. Brian: risks, working on an airplane. I don’t want to do any kind \nof fetch operation on poor connection. Could be good for CI, but don’t \nwant to break things for humans.\n\n33. Jonathan N: if I am going to get on an airplane, is there a way to \nfill it in the background. There are workarounds, like run `git show` \nwhich needs everything.\n\n34. Elijah: I want to fetch a bunch more stuff, but don’t fetch \nanymore, throw an error rather than hanging.\n\n35. Jonathan: filter blob:none is people's first experience of the \nfeature. Make it a first class ui concept, present a user oriented UI \nlike git sparse-checkout?\n\n36. Taylor: It looks like it’s simple to use, but there’s a lot to \ndo to actually use it. And Scalar is doing that for you.\n\n37. James: Some of our customers would be interested to have a feature \nthat pushes down configuration to all the users. It would give them LFS \nby default, without the end-users doing something.\n\n38. Jonathan: We considered enabling a global config at Google. For \nexample for 1+GB files.\n"},{"id":"393108","messageId":"B459AEB1-BF77-44C9-B06A-4B96C2E22287@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 6/17] GC strategies","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:01:54Z","receivedAt":"2020-03-12T04:02:00Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Jonathan N: Git has a flexible packfile format. Compared to CVS where \nthings are stored as deltas against the next revision of the same file. \nGC can be a huge operation if it’s not done regularly. \"git gc\" makes \none huge pack. Better amortized behavior to have multiple packs with \nexponentially increasing size and combine them when needed (Martin \nDick's exproll).\n\n2. Jonathan N: There are also unreachable objects to take care about. GC \ncan/should delete them. But at the same time someone else might be \ncreating history that still needs those objects. To give objects a grace \nperiod, we turn the unused objects into loose objects and look at the \ncreation time. But alternatively there’s the proposal to move these \nunreachable objects into a packfile for all these objects. But this can \nbe a problem for older git clients, because they might not know the pack \nis garbage and might move objects across packs. See the hash function \ntransition doc for details.\n\n3. Terry: JGit has these unreachable garbage packs\n\n4. Peff: You want to solve this loose objects explosion problem?\n\n5. Peff: what if you reference an object in the garbage pack from an \nobject in a non-garbage pack?\n\n6. Jonathan N: At GC time the object from the garbage pack is copied to \na non-garbage pack. Basically rescue it from the garbage. It only saves \nthe referenced objects, not the whole garbage pack.\n\n7. Jonathan N: It has been running in production for >2 years.\n\n8. Peff: There are so many non-atomic operations that can happen. And \nraces can happen.\n\n9. Jonathan N: If you find races, please comment on the JGit change that \ndescribes the algorithm. Happens-before relation and grace period.\n\n10. `git gc --prune-now` should no longer create loose objects first, \nbefore just deleting them.\n"},{"id":"393109","messageId":"35FDF767-CE0C-4C51-88A8-12965CD2D4FF@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 7/17] Background operations/maintenance","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:02:48Z","receivedAt":"2020-03-12T04:02:54Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Stolee: Are we interested in having a background process doing \nthings?\n\n2. Emily: There are a lot of different ways to do this. Even only \nlooking at Linux there are different ways.\n\n3. Stolee: Without looking at how? What background operations would we \nlike to have?\n\n4. Emily: Is it a good candidate for `git undo`? To keep track of what \nthe user was doing and to make it possible to roll back?\n\n5. Brian: It can run into scalability issues. Also there might be repos \non my disk that never change and don’t need background processing. At \nGitHub we do maintenance based on the number of pushes.\n\n6. Stolee: Kind of maintenance will differ from client and server, \ninterests are different. For Scalar we have this one process looking at \nall repos and will do operations on them.\n\n7. Peff: On server-side you’ll have millions of repos and even one \nprocess looking at all processes have impact on the system. Most hosting \nproviders already have services taking care of this, so I think this \nfeature is only interesting for client-side.\n\n8. Brian: We should be careful. For example I’m constantly creating \ntest repos in /tmp.\n\n9. Stolee: Thanks for the input, we’ll do research and come back to \nthis.\n"},{"id":"393110","messageId":"E13E6610-10BA-4394-A8BE-AE91D621A0E8@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 8/17] Push performance","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:03:44Z","receivedAt":"2020-03-12T04:03:49Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Terry: Chrome has 500MB file pushed up. Using Gerrit, feature work \nbecomes stale over a few days, then push. For a few months pushes would \npush gigabytes of data.\n\n2. Stolee: where we do the tree walk, we are doing it from the merge \nbase.\nJonathan N: Minh rescued us by advertising more .have refs to avoid it \nbeing pushed. In protocol V2 for push there are 3 major changes \nproposed: one, abbreviating ref advertisement; two, adding negotiation; \nthree, push to fast moving ref if you don’t care if its a fast \nforward. Are there other cases?\n\n3. Minh: performance on reachability. Would help to know what branch you \nare pushing.\n\n4. Peff: I might be pushing a random sha, without a branch.\n\n5. Brian: I’ve seen cases with 80k refs, we tried then to send minimal \namounts of objects. We spend a lot of time negotiating, to eventually \nonly send 4 objects. It’s not very efficient, you could just spend \nless time on that and send a few more objects.\n\n6. Minh: can we invert the pattern? Just send the new thing, and then \nthe server says give me more.\n\n7. Peff: You’ll get N+1 issues.\n\n8. Jonathan N: I like Jeff Hostetler’s idea in Zoom chat. You can look \nat the branch and see when the author changes and use that as a crude \nheuristic to ask the server if they have that commit.\n"},{"id":"393111","messageId":"9CE46D29-4BCD-4E95-B2DA-939EA10D7934@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 9/17] Obsolescence markers and evolve","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:04:44Z","receivedAt":"2020-03-12T04:04:49Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Brandon: I thought it would be interesting to have a similar feature \nas Mercurial has. Mercurial evolve will help you do a big rebase commit \nby commit. Giving you more insights how commits change over time.\n\n2. Peff: This has been discussed a lot of time on the list already.\n\n3. Jonathan N: It will help with Googlers productivity, but it’s \nsmaller compared to other performance fixes.\n\n4. Brian: It’s a great feature and I would like to have it, but I’m \nnot sure it gives enough value to someone to sit down and implement it.\n\n5. Emily: Is it a good candidate for GSoC?\n\n6. Brian: If we have a good design.\n\n7. Stolee: It should be easier to use than interactive rebase.\n\n8. Stolee: It would be nice to have instead of fixup commits I would \nsend to you new commits which mark your original commits are obsolete.\n"},{"id":"393112","messageId":"AF7A56C0-FDDF-476B-B7B2-F58CE6115353@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 10/17] Expel ‘git shell’?","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:05:57Z","receivedAt":"2020-03-12T04:06:03Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Jonathan N: Cannot use safely on its own. So why do we still have it?\n\n2. Jonathan N: It’s not an interactive shell, it’s a login shell. To \ngive the user only access to a git repo.\n\n3. Jonathan N: Gitolite is the only sensible thing that uses git-shell. \nIf this is the only good use-case? So can we donate it to them?\n\n4. Peff: If it’s a tool for security, but no one is using it, so \nit’s dangerous to have it around. It’s mostly stand-alone, so it \nshould be possible.\n"},{"id":"393113","messageId":"9D605928-2112-4EDB-85CB-806E4F895449@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 11/17] GPL enforcement","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:07:41Z","receivedAt":"2020-03-12T04:07:47Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Peff: Hypothetically if a company would not distribute the sources of \na modified version of git they ship. What should we do about it? Should \nwe take legal actions and make them aware that they are doing something \nthey should not do? Making sure they also treat other projects better.\n\n2. Brandon: Would we gather together with other projects affected by \nthis company?\n\n3. Peff: How hard do we want to take it on this company?\n\n4. Ed: I’m also bothered by this. And they just send me a tar ball. At \nMicrosoft we are aggressive about doing this right.\n\n5. Jonathan N: Can we make it really easy to comply, e.g. by making a \nbuild target that contains everything?\n\n6. ZJ: They could just push to GitHub. That would be fine.\n\n7. Peff: I brought this up, because we were made aware of this by the \nConservancy. So I wanted to hear how people are feeling about it.\n\n8. Brian: A more aggressive approach would be appropriate if we have \nmade them aware of the issue and they decided to not comply on purpose.\n\n9. Peff: Code change doesn’t matter, whether it’s a security fix or \nfeature. And I’m fine giving them a bit of lag time, like a day. Not a \nday.\n\n10. Peff: You’re not obliged to send the source code, but you should \nprovide the offer to share the source. In this case, they sent us a tar \nball, but the sources are not on their open-source. So they probably do \nnot yet apply.\n\n11. CB: But everyone on Mac can request to send you the source code. We \ncould release a form somewhere to give people an easy option to request \nthis.\n"},{"id":"393114","messageId":"A39F6554-8D11-4181-A615-C6562D851716@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 12/17] Test harness improvements","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:08:49Z","receivedAt":"2020-03-12T04:08:55Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Jonathan N: Test harness is an important part of the development \nprocess, shapes what kinds of tests people write. What can we improve?\n\n2. Peff: I love our test harness\n\n3. Brian: It’s amazing for integration tests. For our C code it’s a \nlot harder to do unit tests. We sometimes have a portability issue, \nabout POSIX shell vs others.\n\n4. ZJ: I like how it also acts as a piece of documentation.\n\n5. Jonathan N: If we had more unit tests: if I am working on refs, I \nmight like to run all tests related to that. And now we have this lack \nof dependency graph between this\n\n6. Peff: I’m super nervous about that. Tools like code coverage could \ndo this. But I’ve seen cases where all new tests are green, and tests \nin the area I expected they succeed. But at some far corner it seems to \nfail. So you’re optimizing for speed, might be losing in correctness. \nI’m biased because I can run all tests on my computer in 1 minute. But \nfor Windows this doesn’t seem to work that.\n\n7. Peff: We can spend time on speeding up things, making it better \nparallelized for example. I’ll send some patches out on this.\n\n8. Jonathan N: Really nice contribution to Git by David Barr, whose \nbackground was as a Java developer and thus the code was written in a \nJava way with clear API boundaries and unit tests.\n\n9. Brian: Yes if your function is doing too much, it should be split up \nmaking it possible to test the separate pieces and then have a function \nthat calls those and tests the end result.\n"},{"id":"393115","messageId":"66774F5B-E37F-4676-9274-0066EC38CC48@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 13/17] Cross implementation test suite","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:09:36Z","receivedAt":"2020-03-12T04:09:42Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Carlos: some aspects are under specified, or work in very specific \nways, but need agreement of correct behaviour. For example implementing \na command line tool that will have expectations, or expected repo state \nso another tool can generate the right output. For example libgit2 \nkeeping up with ignore rules. How does JGit handle this?\n\n2. Jonathan N: JGit has some tests of matching behavior which I do not \nlike. Invokes git-grep, generate patterns and compare output. Having \nnon-deterministic tests is not great. I like the idea of table driven \ntests, common data, but different manifestations of how you test those \nthings.\n\n3. Patrick: config formatted tests, need to write drivers for other \nprojects. Stopped because writing all the tests in this format was not \nfun. Basics work though. Spoke to Peff 2 years ago, likely easy to write \ndrivers for Git.\n\n4. Peff: already replaced tests with table driven, and prefer that. \nThere are table driven tests for attribute matching.\n\n5. Brian: valuable for LFS. Know attribute matching is not up to spec. \nCould benefit from the tests to help identify gaps. We are MIT licensed, \nso we can’t just drop them in, but we could import them in CI.\n\n6. Peff: make whatever is in Git as authority, add tables, and then \nthese can be used by other projects.\n\n7. Jonathan N: example is diff tests\n"},{"id":"393116","messageId":"84A85206-F4A7-4F36-A302-C3986D6AFF91@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 14/17] Aspects of merge-ort: cool, or crimes against humanity?","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:11:35Z","receivedAt":"2020-03-12T04:11:41Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Elijah: ORT stands for Ostensibly Recursive’s Twin. As a merge \nstrategy, just like you can call ‘git merge -s recursive’ you can \ncall ‘git merge -s ort’.  Git’s option parsing doesn’t require \nthe space after the ‘-s’.\n\n2. Major question is about performance & possible layering violations. \nMerge recursive calls unpacks trees to walks trees, then needs to get \nall file and directory names and so walks the trees on the right again, \nand then the trees on the left again. Then diff needs to walk sets of \ntrees twice, and then insert_stage_data() does a bunch more narrow tree \nwalks for rename detection. Lots of tree walking. Replaced that with two \ntree walks.\n\n3. Using traverse_trees() instead of unpack_trees(), and avoid the index \nentirely (not even touching or creating cache_entry’s), and building \nup information as I need. I’m not calling diffcore_std(), but instead \ndirectly calling diffcore_rename(). Is this horrifying? Or is it \njustified by the performance gains?\n\n4. Peff: both, some of it sounds like an improvement, but maybe there \nwere hidden benefits previously.\n\n5. Elijah: I write to a tree before I do anything.\n\n6. Peff: I like that. Seems like a clean up to me. We have written \nlibgit2-like code for merging server-side\n\n7. Elijah: I’ve been adding tests for the past few years, more to add, \nfeel good about it.\n\n8. Jonathan N: If you are using a lower-layer thing, I would not say \nyou’re not doing anything you shouldn’t. But if you docs say you \nshould not to use diffcore_rename(), you can update the docs to say that \nit’s fine to use it.\n\n9. Elijah: three places directly write tree objects. All have different \ndata structures they are writing from. Should I pull them out? But then \nmy data structure was also different, so I’d have a fourth.\n\n10. Peff: not worried because trees are simple. Worried about policy \nlogic. Can’t write a tree entry with a double slash. Want this to be \nenforced everywhere, but no idea how hard that would be to implement. \nNot about lines of code, but consistency of policy. Fearful that only \none place does it.\n\n11. Elijah: I know merge-ort checks this, but it’s not nearby, so it \ncould change.\n\n12. Peff: as bad as it is to round trip through the index, it may bypass \nquality checks, which you will need to manually implement.\n\n13. Elijah: usability side, with the tree I’ve created, I could have \n.git/AUTOMERGE. I have an old tree, a new tree, and a checkout can get \nme there. Fixed a whole bunch of bugs for sparsity and submodules.\n\n14. Elijah: If we use this to on-the-fly remerge as part of git-log in \norder to compare the merge commit to what the automatic merging would \nhave done, where/how should we write objects as we go?\n\n15. Jonathan N: can end up with proliferation of packs, would be nice to \nhave similar to fast import and have in memory store. Dream not to have \nloose files written ever.\n\n16. Peff: I like your dream. But fast import packs are bad. We assume \nthat packs are good, and thus need to use GC aggressively. This \nincreases pollution of that problem. I know about objects, but not \nwritten to disc, risk that you can write objects that are broken, but \ngit doesn’t know because git thinks it has the object but it’s only \nin memory. Log is conceptually a read operation, but this would create \nthe need for writes.\n\n17. Elijah: you could write into a temporary directory. Worried about \n`gc --auto` in the middle of my operation. If I write to a temp pack I \ncould potentially avoid it.\n\n18. Elijah: large files. Rename detection might not work efficiently OR \ncorrectly for sufficiently large files (binary or not). Limited bucket \nsize means that completely different files treated as renames when both \nare over 8MB. Should big files just not be compared?\n\n19. Peff: maybe we should fix the hash…\n\n20. Elijah: present situation is broken, maybe we can cheat in the short \nterm, and avoid fixing?\n\n21. Peff: seems more correct for now, but we’d need to document\n\n22. Elijah: checkout --overwrite-ignore flag. Should merge have the same \nflag.\n\n23. Jonathan N: gitignore original use case was build outputs which can \nbe regenerate. But then some people want to ignore `.hg` which is much \nmore precious.\n\n24. Peff: we can plumb it through later to other commands\n\n25. Brian: CI doesn’t really care. Moving between branches it would \ncomplain. For checkout and merge it makes sense to support just \ndestroying.\n"},{"id":"393117","messageId":"82E8EA09-EFC5-47BD-84C1-C4F5BC98580B@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 15/17] Reachability checks","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:13:24Z","receivedAt":"2020-03-12T04:13:29Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Jonathan N: seed the idea that it would be nice to hint the ref that \nyour commit might be reachable from to help the server avoid iterating \nover all refs. Also, any strategies for speeding up reachability checks?\n\n2. Demitr: reachability by user, or would you consider open to everyone?\n\n3. Stolee: we don’t do branch level security, but we do tailor ref \nlist to default, favorites and those you’ve pushed. There is also a \nfull endpoint.\n\n4. Brian: security model we have to have is that we assume everyone has \nread to everything. There are too many ways to attack. Useful for \nperformance reasons, but not sure reachability checks provide much \nbenefit. Don’t think it’s difficult to automate.\n\n5. Demitr: what about security issues\n\n6. Stolee: we’d say find another way.\n\n7. Terry: we have a mono repo, easier to test everything. JGit goes down \nto object level.\n\n8. Peff: Git doesn’t go down to that level, doesn’t validate haves.\n\n9. Jonathan: two lessons, no one except Gerrit cares strongly about \nthis; second if we like the model by branch permissions, worth making it \nwork well in Git to prevent distance between JGit and Git.\n\n10. Terry: can remove a branch very quickly and prevent new people \ngetting it\n\n11. Peff: don’t deny its usefulness, but the performance implication \nis concerning. Trying to keep objects private from determined attackers. \nBut pushing a malicious commit to Linux, a user can see it, and won’t \nunderstand reachability doesn’t imply endorsement.\n\n12. Jonathan: if Git has an easy cheap way to do it, people would use \nit.\n\n13. Peff: have flirted with it, but might have to open 50GB of \npackfiles, or bitmap has corner cases. There are some obvious ways to \nimprove, but a lot of work. V2 spec says you’re not allowed to check \nreachability.\n\n14. Jonathan N: nah, it says you don't advertise a capability describing \nwhether it is checking reachability.\n\n15. Peff: submodule, but then the commit disappears and becomes \nunreachable. How do you handle?\n\n16. Jonathan N: encourage folks to do fast forward only updates. In \nhooks instead of the git layer\n\n17. Peff: you might not know what ref has reachability to that commit. I \nlike the hint thing, if it’s just a hint.\n"},{"id":"393118","messageId":"6DAC1E49-9CA0-4074-867E-F22CD26C9FEB@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 16/17] “I want a reviewer”","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:14:25Z","receivedAt":"2020-03-12T04:14:30Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Jonathan N: An experience some folks have is, sending a patch and \nhearing nothing. That must mean patch is awesome. But then realize I \nneed to do something to get a review. In Git, people like Peff are good \nabout responding to newcomers. As an author it can be hard to excite \npeople enough to review your patch. Relatedly, you might get a review, \nbut it doesn’t give you the feedback you wanted. As a reviewer, you \nwant to help people to grow and make progress quickly, but it might not \nbe easy to identify patches where this will be possible.\n\n2. Emily: A few months ago we started doing code review book club. Git \ndevel IRC, and mailing list, could we be more public about these? I \nqueue my patch to list of things that have been idle and needs a review, \nthen a bot pops something off the list to increase attention for people \nto review?\n\n3. Jonathan Tan: during book club we discuss and review together. \nEveryone can benefit from review experience and expertise. Emily is \nhoping for similar knowledge transfer in the IRC channel.\n\n4. Brian: general case that patches don’t get lost. There is the git \ncontext script, but I am now a reviewer because I have touched \neverything for SHA256. But we are losing patches and bug reports because \nthings get missed. What tool would we use? How would we do it?\n\n5. Jonathan N: patchwork exists, need to learn how to use it :)\n\n6. Peff: this is all possible on the mailing list. I see things that \nlook interesting, and have a to do folder. If someone replies, I’ll \ntake it off the list. Once a week go through all the items. I like the \nbook club idea, instead of it being ad hoc, or by me, a group of a few \npeople review the list in the queue. You might want to use a separate \ntool, like IRC, but it would be good to have it bring it back to the \nmailing list as a summary. Public inbox could be better, but someone \nneeds to write it. Maybe nerd snipe Eric?\n\n7. Stolee: not just about doing reviews, but training reviewers.\n"},{"id":"393119","messageId":"A982100B-A7EB-400F-83EA-75F14D7260AF@jramsay.com.au","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"[TOPIC 17/17] Security","fromName":"James Ramsay","fromEmail":"james@jramsay.com.au","sentAt":"2020-03-12T04:16:04Z","receivedAt":"2020-03-12T04:16:11Z","isPatch":false,"sender":{"key":"james@jramsay.com.au","avatar":null},"body":"1. Demtr: what are people doing to prevent security issues? For example, \nnot allowing things into trees that would be problematic for various \nfilesystems.\n\n2. Jonathan N: transfer fsck objects by default, to validate at the \ntrust boundary (in case some code paths at use time are missing some \nvalidation)\n\n3. Peff: we have had buffer overflows, most are logic errors, and mostly \npaths related. Recently we’ve tightened up which paths are allowed. \nForbidding things that might be valid on Linux, but problems on Windows. \nCan’t catch everything though, because Windows is so so complex\n\n4. Stolee: I am fearful, and do not know all the rules.\n\n5. Peff: I don’t think it is possible.\n\n6. Demetr: only latin chars, numbers and a few other characters. Do not \nallow any special symbols.\n\n7. Brian: that’s going to break lots of existing projects. Some \nprojects have never been on Windows, and therefore people have no \nconcern about Windows. People checking files that are strange to \ndeliberately test strange files in their own software. If Windows has an \nAPI to test filepath, there is not much we can do to protect it. \nCompatibility is important.\n\n8. Peff: probably some cleanup needed, maybe can’t clone git.git. Some \npaths that are innocuous, are a problem in strange situations.\n\n9. Jonathan N: what in Git's design scares the crap out of you?\n\n10. ZJ: GitLab shells out for everything. We had injections. Now we have \na DSL to verify things. Looking at --end-of-options.\n\n11. Peff: C is terrifying. Rust rewrite please. Still have integer \noverflow risks. Tried to deal with it a few years ago, and found some \nmore a few months back. A happy story: OID array uses signed integer, \nbecause no-one has more than 2billion objects. Someone had 3billion \nobjects. Just the SHA1s are 60GB. Found it because it triggered overflow \nin st_add. As soon as they wrapped around, it crashed, preventing under \nallocation\n\n12. Jeff H: communication between processes\n\n13. <musical interlude>\n\n14. Peff: I feel good about where we read and write strings to each \nother. Maybe if we were using JSON encode/decode it might be easier to \nhandle obscure cases\n"},{"id":"393126","messageId":"20200312133127.GK212281@google.com","threadId":"52980","inReplyTo":"6DAC1E49-9CA0-4074-867E-F22CD26C9FEB@jramsay.com.au","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-03-12T13:31:27Z","receivedAt":"2020-03-12T13:31:36Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Thu, Mar 12, 2020 at 03:14:25PM +1100, James Ramsay wrote:\n> 5. Jonathan N: patchwork exists, need to learn how to use it :)\n\nWe've actually got a meeting with some Patchwork folks today - if\nanybody has a burning need they want filled via Patchwork, just say so,\nand we'll try to ask.\n\n - Emily\n"},{"id":"393127","messageId":"20200312141628.GL212281@google.com","threadId":"52980","inReplyTo":"0D7F1872-7614-46D6-BB55-6FEAA79F1FE6@jramsay.com.au","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-03-12T14:16:28Z","receivedAt":"2020-03-12T14:16:38Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Thu, Mar 12, 2020 at 02:56:53PM +1100, James Ramsay wrote:\n> 1. Emily: hooks in the config file. Sent a read only patch, but didn’t get\n> much traction. Add a new header in the config file, then have prefix number\n> so that security scan could be configured at system level that would run\n> last, and then hook could also be configured at the project level.\n> \n> 2. Peff: Having hooks in the config would be nice. But don’t do it at\n> `hooks.prereceive`, but use a subconfig like `hooks.prereceive.command` so\n> it’s possible to add options later on.\n> \n> 3. Brian: sometimes the need to overridden, ordering works for me. For Git\n> LFS it would be helpful to have a pre-push hook for LFS, and a different one\n> for something else. Want flexibility about finding and discovering hooks.\n> \n> 4. Emily: if you want to specify a hook that is in the work tree, then it\n> has to be configured after cloning.\n> \n> 5. Jonathan: It’s better to start with something low complexity as long as\n> it can be extended/changed later. If there's room to tweak it over time then\n> I'm not too worried about the initial version being perfect — we can make\n> mistakes and learn from them. A challenge will be how hooks interact.\n> Analogy to the challenges of stacked union filesystems and security modules\n> in Linux. Analogy to sequence number allocation for unit scripts\n> \n> 6. CB: Declare dependencies instead of a sequence number? In theory\n> independent hooks can also run in parallel.\n> \n> 7. Peff: Maybe that’s something to not worry about from the start. Like, how\n> many hooks do you expect to run anyway.\n> \n> 8. Christian: At booking.com they use a lot of hooks, and they also sent\n> patches to the mailing list to improve that.\n> \n> 9. Emily: In-tree hooks?\n> \n> 10. Brian: You can do `git cat-file <ref> | sh` to run a hook.\n> \n> 11. Brandon: Is it possible to globally to disable all hooks locally? It\n> might be a security concern. Or is it something we might want to add?\n> \n> 12. Peff: No it’s not.\n\nThanks for the notes, James.\n\nI came away with the understanding that we want the config hook to look\nsomething like this (barring misunderstanding of config file syntax,\nplus or minus naming quibbles):\n\n[hook \"/path/to/executable.sh\"]\n\tevent = pre-commit\n\nThe idea being that by using a subsection, we can extend the format\nlater much more easily, but by starting simply, we can start using it\nand see what we need or don't want. We can use config order to begin\nwith.\n\nThis means that we could do something like this:\n\n[hook \"/path/to/executable.sh\"]\n\tevent = pre-commit\n\torder = 123\n\tmustSucceed = false\n\tparallelizable = true\n\netc, etc as needed.\n\nBut I wonder if we also want to be able to do something like this:\n\n[hook \"/etc/git-secrets/git-secrets\"]\n\tevent = pre-commit\n\tevent = prepare-commit-msg\n\t...\n\nI guess the point is that we can choose to allow this, or not. I could\nsee there being some trouble if you wanted the execution order to work\ndifferently (e.g. run it first for pre-commit but last for\nprepare-commit-msg)...\n\nI think, though, that something like\nhook.pre-commit.\"path/to/executable.sh\" won't work. It doesn't seem like\nmultiple subsections are OK in config syntax, as far as I can see. I'd\nbe interested to know I'm wrong :)\n\nWill try and get some work on this soon, but honestly my hope is to get\nbugreport squared away first.\n\n - Emily\n"},{"id":"393128","messageId":"a6ef088a-0c68-58ec-ef5b-4f353f0b0299@gmail.com","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"Re: Notes from Git Contributor Summit, Los Angeles (April 5, 2020)","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-03-12T14:38:02Z","receivedAt":"2020-03-12T14:38:08Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 3/11/2020 11:55 PM, James Ramsay wrote:\n> It was great to see everyone at the Contributor Summit last week, in person and virtually.\n> \n> Particular thanks go to Peff for facilitating, and to GitHub for organizing the logistics of the meeting place and food. Thank you!\n\nThanks for taking excellent notes!\n\n> On the day, the topics below were discussed:\n> \n> 1. Ref table (8 votes)\n> 2. Hooks in the future (7 votes)\n> 3. Obliterate (6 votes)\n> 4. Sparse checkout (5 votes)\n> 5. Partial Clone (6 votes)\n> 6. GC strategies (6 votes)\n> 7. Background operations/maintenance (4 votes)\n> 8. Push performance (4 votes)\n> 9. Obsolescence markers and evolve (4 votes)\n> 10. Expel ‘git shell’? (3 votes)\n> 11. GPL enforcement (3 votes)\n> 12. Test harness improvements (3 votes)\n> 13. Cross implementation test suite (3 votes)\n> 14. Aspects of merge-ort: cool, or crimes against humanity? (2 votes)\n> 15. Reachability checks (2 votes)\n> 16. “I want a reviewer” (2 votes)\n> 17. Security (2 votes)\n\nWow, this split into separate emails was a fantastic idea to control the multi-threaded discussion. Kudos!\n\n-Stolee\n"},{"id":"393143","messageId":"20200312173134.bpflnl6n3w6mywlg@chatter.i7.local","threadId":"52980","inReplyTo":"20200312133127.GK212281@google.com","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-03-12T17:31:34Z","receivedAt":"2020-03-12T17:31:42Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, Mar 12, 2020 at 06:31:27AM -0700, Emily Shaffer wrote:\n> On Thu, Mar 12, 2020 at 03:14:25PM +1100, James Ramsay wrote:\n> > 5. Jonathan N: patchwork exists, need to learn how to use it :)\n> \n> We've actually got a meeting with some Patchwork folks today - if\n> anybody has a burning need they want filled via Patchwork, just say so,\n> and we'll try to ask.\n\nJust to highlight this -- a long while ago someone asked me to set up a \npatchwork instance for Git, but I believe they never used it:\n\nhttps://patchwork.kernel.org/project/git/list/\n\n-K\n"},{"id":"393147","messageId":"20200312174212.GA120942@google.com","threadId":"52980","inReplyTo":"20200312173134.bpflnl6n3w6mywlg@chatter.i7.local","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2020-03-12T17:42:12Z","receivedAt":"2020-03-12T17:42:17Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi!\n\nKonstantin Ryabitsev wrote:\n> On Thu, Mar 12, 2020 at 06:31:27AM -0700, Emily Shaffer wrote:\n\n>> We've actually got a meeting with some Patchwork folks today - if\n>> anybody has a burning need they want filled via Patchwork, just say so,\n>> and we'll try to ask.\n>\n> Just to highlight this -- a long while ago someone asked me to set up a\n> patchwork instance for Git, but I believe they never used it:\n>\n> https://patchwork.kernel.org/project/git/list/\n\nThat was me.  In fact, we are using it, but mostly read-only (similar\nto lore patchwork) so far.  I'm hoping we can learn more about how to\nautomatically close reviews when a patch has landed, assign delegates\nto reviews, set up bundles, etc and write some docs so it becomes\nuseful to more people.\n\nThanks,\nJonathan\n"},{"id":"393150","messageId":"20200312180012.bgjls4esndkg3iqf@chatter.i7.local","threadId":"52980","inReplyTo":"20200312174212.GA120942@google.com","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-03-12T18:00:12Z","receivedAt":"2020-03-12T18:00:17Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, Mar 12, 2020 at 10:42:12AM -0700, Jonathan Nieder wrote:\n> >> We've actually got a meeting with some Patchwork folks today - if\n> >> anybody has a burning need they want filled via Patchwork, just say so,\n> >> and we'll try to ask.\n> >\n> > Just to highlight this -- a long while ago someone asked me to set up a\n> > patchwork instance for Git, but I believe they never used it:\n> >\n> > https://patchwork.kernel.org/project/git/list/\n> \n> That was me.  In fact, we are using it, but mostly read-only (similar\n> to lore patchwork) so far.  I'm hoping we can learn more about how to\n> automatically close reviews when a patch has landed, assign delegates\n> to reviews, set up bundles, etc and write some docs so it becomes\n> useful to more people.\n\nFYI, I can set it up with git-patchwork-bot, which does some of the \nabove. You can read more here:\n\nhttps://korg.wiki.kernel.org/userdoc/patchwork#adding_patchwork-bot_integration\n\nIf that's something you would like to see, please send a request per \nthat doc.\n\n-K\n"},{"id":"393152","messageId":"20200312180651.yonj6wzuatur25w6@chatter.i7.local","threadId":"52980","inReplyTo":"5B2FEA46-A12F-4DE7-A184-E8856EF66248@jramsay.com.au","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2020-03-12T18:06:51Z","receivedAt":"2020-03-12T18:06:55Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, Mar 12, 2020 at 02:57:24PM +1100, James Ramsay wrote:\n> 8. Jeff H: can we introduce a new type of object -- a \"revoked blob\" if you\n> will that burns the original one but also holds the original SHA in the ODB\n> ??\n> \n> 9. Peff: what would this mean for signatures? New opportunity to forge\n> signatures.\n\nEasy, you just quickly find a collision for that blob's sha1 and put \nthat in place of the offending original. ;)\n\n(Fully tongue-in-cheek.)\n\n-K\n"},{"id":"393221","messageId":"xmqqeetwcf4k.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200312141628.GL212281@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-03-13T17:56:59Z","receivedAt":"2020-03-13T17:57:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n> This means that we could do something like this:\n>\n> [hook \"/path/to/executable.sh\"]\n> \tevent = pre-commit\n> \torder = 123\n> \tmustSucceed = false\n> \tparallelizable = true\n>\n> etc, etc as needed.\n\nYou can do\n\n    [hook \"pre-commit\"]\n\torder = 123\n\tpath = \"/path/to/executable.sh\"\n\n    [hook \"pre-commit\"]\n\torder = 234\n\tpath = \"/path/to/another-executable.sh\"\n\nas well, and using the second level for what hook the (sub)section\nis about, instead of \"we have this path that is used for a hook.\nWhat hook is it?\", feels (at least to me) more natural.\n\n> But I wonder if we also want to be able to do something like this:\n>\n> [hook \"/etc/git-secrets/git-secrets\"]\n> \tevent = pre-commit\n> \tevent = prepare-commit-msg\n\nOnce you start going this route, it no longer makes sense to give\npriority (you called it \"order\") to a path and have that same number\nused in contexts of different hooks.  Your git-secrets script may\nwant to be called early among pre-commit hooks but late among the\nprepare-commit-msg hooks, for example.\n\n> I think, though, that something like\n> hook.pre-commit.\"path/to/executable.sh\" won't work.\n\nThat is why Peff already suggested in the TOPIC notes to use\n\"command\" in the message you are responding to (I used \"path\" in the\nabove description).\n"},{"id":"393234","messageId":"20200313204738.GA557052@coredump.intra.peff.net","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"Re: Notes from Git Contributor Summit, Los Angeles (April 5, 2020)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-13T20:47:38Z","receivedAt":"2020-03-13T20:47:43Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 12, 2020 at 02:55:21PM +1100, James Ramsay wrote:\n\n> It was great to see everyone at the Contributor Summit last week, in person\n> and virtually.\n> \n> Particular thanks go to Peff for facilitating, and to GitHub for organizing\n> the logistics of the meeting place and food. Thank you!\n\nThanks very much to you and others who took these notes! It's nice to\nhave a more permanent record of these discussions.\n\n-Peff\n"},{"id":"393237","messageId":"20200313212531.GA22502@dcvr","threadId":"52980","inReplyTo":"6DAC1E49-9CA0-4074-867E-F22CD26C9FEB@jramsay.com.au","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Eric Wong","fromEmail":"e@yhbt.net","sentAt":"2020-03-13T21:25:31Z","receivedAt":"2020-03-13T21:25:33Z","isPatch":false,"sender":{"key":"e@yhbt.net","avatar":null},"body":"James Ramsay <james@jramsay.com.au> wrote:\n\nJames: first off, thank you for these accessible summaries for\nnon-JS users and those who could not attend for various reasons(*)\n\n<snip>\n\n> 6. Peff: this is all possible on the mailing list. I see things that look\n> interesting, and have a to do folder. If someone replies, I’ll take it off\n> the list. Once a week go through all the items. I like the book club idea,\n> instead of it being ad hoc, or by me, a group of a few people review the\n> list in the queue. You might want to use a separate tool, like IRC, but it\n> would be good to have it bring it back to the mailing list as a summary.\n> Public inbox could be better, but someone needs to write it. Maybe nerd\n> snipe Eric?\n\nWhat now? :o\n\nThere's a lot of things it could be better at, but a more\nconcrete idea of what you want would help.\n\nRight now I only have enough resources to do bugfixes along with scalability\nand performance improvements so more people can run it and keep\nit 100% reproducible and centralization resistant.\n\nI'm also planning on some local tooling along the lines of\nnotmuch/mairix which is NNTP/HTTPS-aware but not sure when I'll\nbe able to do that...\n\n\n(*) I stopped attending events over a decade ago for privacy reasons\n    (facial recognition, invasive airport searches, etc.)\n"},{"id":"393247","messageId":"20200315003654.GA711@dcvr","threadId":"52980","inReplyTo":"20200314172715.GA1178875@coredump.intra.peff.net","subject":"inbox indexing wishlist [was: [TOPIC 16/17] “I want a reviewer”]","fromName":"Eric Wong","fromEmail":"e@yhbt.net","sentAt":"2020-03-15T00:36:54Z","receivedAt":"2020-03-15T01:55:01Z","isPatch":false,"sender":{"key":"e@yhbt.net","avatar":null},"body":"Jeff King <peff@peff.net> wrote:\n> On Fri, Mar 13, 2020 at 09:25:31PM +0000, Eric Wong wrote:\n> \n> > > 6. Peff: this is all possible on the mailing list. I see things that look\n> > > interesting, and have a to do folder. If someone replies, I’ll take it off\n> > > the list. Once a week go through all the items. I like the book club idea,\n> > > instead of it being ad hoc, or by me, a group of a few people review the\n> > > list in the queue. You might want to use a separate tool, like IRC, but it\n> > > would be good to have it bring it back to the mailing list as a summary.\n> > > Public inbox could be better, but someone needs to write it. Maybe nerd\n> > > snipe Eric?\n> > \n> > What now? :o\n> > \n> > There's a lot of things it could be better at, but a more\n> > concrete idea of what you want would help.\n> \n> short answer: searching for threads that only one person participated in\n\n+Cc meta@public-inbox.org\n\nOK, something I've thought of doing anyways in the past...\n\n> The discussion here was around people finding useful things to do on the\n> list: triaging or fixing bugs, responding to questions, etc. And I said\n> my mechanism for doing that was to hold interesting-looking but\n> not-yet-responded-to mails in my git-list inbox, treating it like a todo\n> list, and then eventually:\n> \n>   1. I sweep through and spend time on each one.\n> \n>   2. I see that somebody else responded, and I drop it from my queue.\n> \n>   3. It ages out and I figure that it must not have been that important\n>      (I do this less individually, and more by occasionally declaring\n>      bankruptcy).\n> \n> That's easy for me because I use mutt, and I basically keep my own list\n> archive anyway. But it would probably be possible to use an existing\n> archive and just search for \"threads with only one author from the last\n> 7 days\". And people could sweep through that[1].\n> \n> You already allow date-based searches, so it would really just be adding\n> the \"thread has only one author\" search. It's conceptually simple, but\n> it might be hard to index (because of course it may change as messages\n> are added to the archive, though any updates are bounded to the set of\n> threads the new messages are in).\n\nExactly on being conceptually simple but requiring some deeper\nchanges to the way indexing works.  I'll have to think about it\na bit, but it should be doable without being too intrusive,\ninvasive or expensive for existing users.\n\n> But to be clear, I don't think you have any obligation here. I just\n> wondered if it might be interesting enough that you would implement it\n> for fun. :) As far as I'm concerned, if you never implemented another\n> feature for public-inbox, what you've done already has been a great\n> service to the community.\n\nThanks.  I'll keep that index change in mind and it should be\ndoable if I remain alive and society doesn't collapse...\n\n> [1] The obvious thing this lacks compared to my workflow is a way to\n>     mark threads as \"seen\" or \"not interesting\". But that implies\n>     per-user storage.\n\nYeah, that would be part of the local tools bit I've been\nthinking about (user labels such as \"important\", \"seen\",\n\"replied\", \"new\", \"ignore\", ... flags).\n"},{"id":"393253","messageId":"20200314172715.GA1178875@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200313212531.GA22502@dcvr","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-14T17:27:15Z","receivedAt":"2020-03-15T02:27:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Mar 13, 2020 at 09:25:31PM +0000, Eric Wong wrote:\n\n> > 6. Peff: this is all possible on the mailing list. I see things that look\n> > interesting, and have a to do folder. If someone replies, I’ll take it off\n> > the list. Once a week go through all the items. I like the book club idea,\n> > instead of it being ad hoc, or by me, a group of a few people review the\n> > list in the queue. You might want to use a separate tool, like IRC, but it\n> > would be good to have it bring it back to the mailing list as a summary.\n> > Public inbox could be better, but someone needs to write it. Maybe nerd\n> > snipe Eric?\n> \n> What now? :o\n> \n> There's a lot of things it could be better at, but a more\n> concrete idea of what you want would help.\n\nshort answer: searching for threads that only one person participated in\n\nThe discussion here was around people finding useful things to do on the\nlist: triaging or fixing bugs, responding to questions, etc. And I said\nmy mechanism for doing that was to hold interesting-looking but\nnot-yet-responded-to mails in my git-list inbox, treating it like a todo\nlist, and then eventually:\n\n  1. I sweep through and spend time on each one.\n\n  2. I see that somebody else responded, and I drop it from my queue.\n\n  3. It ages out and I figure that it must not have been that important\n     (I do this less individually, and more by occasionally declaring\n     bankruptcy).\n\nThat's easy for me because I use mutt, and I basically keep my own list\narchive anyway. But it would probably be possible to use an existing\narchive and just search for \"threads with only one author from the last\n7 days\". And people could sweep through that[1].\n\nYou already allow date-based searches, so it would really just be adding\nthe \"thread has only one author\" search. It's conceptually simple, but\nit might be hard to index (because of course it may change as messages\nare added to the archive, though any updates are bounded to the set of\nthreads the new messages are in).\n\nBut to be clear, I don't think you have any obligation here. I just\nwondered if it might be interesting enough that you would implement it\nfor fun. :) As far as I'm concerned, if you never implemented another\nfeature for public-inbox, what you've done already has been a great\nservice to the community.\n\n-Peff\n\n[1] The obvious thing this lacks compared to my workflow is a way to\n    mark threads as \"seen\" or \"not interesting\". But that implies\n    per-user storage.\n"},{"id":"393280","messageId":"86pndd794k.fsf@gmail.com","threadId":"52980","inReplyTo":"AC2EB721-2979-43FD-922D-C5076A57F24B@jramsay.com.au","subject":"Re: Notes from Git Contributor Summit, Los Angeles (April 5, 2020)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-03-15T18:42:19Z","receivedAt":"2020-03-15T18:42:25Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"James Ramsay\" <james@jramsay.com.au> writes:\n\n> It was great to see everyone at the Contributor Summit last week, in\n> person and virtually.\n>\n> Particular thanks go to Peff for facilitating, and to GitHub for\n> organizing the logistics of the meeting place and food. Thank you!\n>\n> On the day, the topics below were discussed:\n>\n> 1. Ref table (8 votes)\n> 2. Hooks in the future (7 votes)\n> 3. Obliterate (6 votes)\n> 4. Sparse checkout (5 votes)\n> 5. Partial Clone (6 votes)\n> 6. GC strategies (6 votes)\n> 7. Background operations/maintenance (4 votes)\n> 8. Push performance (4 votes)\n> 9. Obsolescence markers and evolve (4 votes)\n> 10. Expel ‘git shell’? (3 votes)\n> 11. GPL enforcement (3 votes)\n> 12. Test harness improvements (3 votes)\n> 13. Cross implementation test suite (3 votes)\n> 14. Aspects of merge-ort: cool, or crimes against humanity? (2 votes)\n> 15. Reachability checks (2 votes)\n> 16. “I want a reviewer” (2 votes)\n> 17. Security (2 votes)\n\nThank you very much for sending split writeup to the mailing list.\n\nOne question to all participating live (in person): how those topics\nwere proposed, and how they were voted for?  This was done before remote\naccess was turned on, I think.\n\nThanks in advance,\n-- \nJakub Narębski\n"},{"id":"393288","messageId":"20200315221940.bdgi5mluxuetq2lz@doriath","threadId":"52980","inReplyTo":"5B2FEA46-A12F-4DE7-A184-E8856EF66248@jramsay.com.au","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Damien Robert","fromEmail":"damien.olivier.robert@gmail.com","sentAt":"2020-03-15T22:19:40Z","receivedAt":"2020-03-15T23:13:27Z","isPatch":false,"sender":{"key":"damien.olivier.robert@gmail.com","avatar":null},"body":"From James Ramsay, Thu 12 Mar 2020 at 14:57:24 (+1100) :\n> 6. Elijah: replace refs helps, but not supported by hosts like GitHub etc\n>     a. Stolee: breaks commit graph because of generation numbers.\n>     b. Replace refs for blobs, then special packfile, there were edge cases.\n\nI am interested in more details on how to handle this using replace.\n\nMy situation: coworkers push big files by mistake, I don't want to rewrite\nhistory because they are not too well versed with git, but I want to keep\n*my* repo clean.\n\nPartial solution:\n- identify the large blobs (easy)\n- write a replace ref (easy):\n  $ git replace b5f74037bb91 $(git hash-object -w -t blob /dev/null)\n  and replace the file (if it is still in the repo) by an empty file.\n\nNow the pain points start:\n- first the index does not handle replace (I think), so the replaced file\n  appear as changed in git status, even through eg git diff shows nothing.\n\n=> Solution: configure .git/info/sparse-checkout\n\n- secondly, I want to remove the large blob from my repo.\n\nIdeally I'd like to repack everything but filter this blob, except that\nrepack does not understand --filter. So I need to use `git pack-objects`\ndirectly and then do the naming and clean up that repack usually does\nmanually, which is error prone.\n\nFurthermore, while `git pack-objects` accepts --filter, I can only filter on\nblob size, not blob oid. (there is filter=sparse:oid where I could reuse my\nsparse checkout file, but I would need to make a blob of it first). And if I\nhave one large file I want to keep, I cannot filter by blob size.\n\nAnother solution would be to use `git unpack-objects` to unpack all objects\n(except I would need to do that in an empty git dir), remove the blob, and\nthen repack everything.\n\nAm I missing a simpler solution?\n\n- finally, checkouting to a ref including the replaced (now missing) blob\n  gives error messages of the form:\nerror: invalid object 100644 b5f74037bb91c45606b233b0ad6aad86f8e3875e for 'Silverman-Height-NonTorsion.pdf'\n\nOn the one hand it is reassuring that git checks that the real object\n(rather than only the replaced object) is still there, on the other hand it\nwould be nice to ask git to completely forget about the original object\n(except fsck of course).\n\nThanks,\nDamien\n"},{"id":"393301","messageId":"3839451584363302@sas2-a098efd00d24.qloud-c.yandex.net","threadId":"52980","inReplyTo":"20200315221940.bdgi5mluxuetq2lz@doriath","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Konstantin Tokarev","fromEmail":"annulen@yandex.ru","sentAt":"2020-03-16T12:55:39Z","receivedAt":"2020-03-16T12:55:45Z","isPatch":false,"sender":{"key":"annulen@yandex.ru","avatar":null},"body":"\n\n16.03.2020, 02:13, \"Damien Robert\" <damien.olivier.robert@gmail.com>:\n> From James Ramsay, Thu 12 Mar 2020 at 14:57:24 (+1100) :\n>>  6. Elijah: replace refs helps, but not supported by hosts like GitHub etc\n>>      a. Stolee: breaks commit graph because of generation numbers.\n>>      b. Replace refs for blobs, then special packfile, there were edge cases.\n>\n> I am interested in more details on how to handle this using replace.\n>\n> My situation: coworkers push big files by mistake, I don't want to rewrite\n> history because they are not too well versed with git, but I want to keep\n> *my* repo clean.\n\nWouldn't it be better to prevent *them* from such mistakes, e.g. by using\npre-push review system like Gerrit?\n\n-- \nRegards,\nKonstantin\n\n"},{"id":"393304","messageId":"CABPp-BEnYTvakuP9nBi3Q_-mP3i7BJEvKofC3_4N8cO9JkF22Q@mail.gmail.com","threadId":"52980","inReplyTo":"20200315221940.bdgi5mluxuetq2lz@doriath","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2020-03-16T16:32:45Z","receivedAt":"2020-03-16T16:33:01Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sun, Mar 15, 2020 at 4:16 PM Damien Robert\n<damien.olivier.robert@gmail.com> wrote:\n>\n> From James Ramsay, Thu 12 Mar 2020 at 14:57:24 (+1100) :\n> > 6. Elijah: replace refs helps, but not supported by hosts like GitHub etc\n> >     a. Stolee: breaks commit graph because of generation numbers.\n> >     b. Replace refs for blobs, then special packfile, there were edge cases.\n>\n> I am interested in more details on how to handle this using replace.\n\nThis comment at the conference was in reference to how people rewrite\nhistory to remove the big blobs, but then run into issues because\nthere are many places outside of git that reference old commit IDs\n(wiki pages, old emails, issues/tickets, etc.) that are now broken.\n\nreplace refs can help in that situation, because replace refs can be\nused to not only replace existing objects with something else, they\ncan be used to replace non-existing objects with something else,\nessentially setting up an alias for an object.\n\ngit filter-repo uses this when it rewrites history to give you a way\nto access NEW commit hashes using OLD commit hashes, despite the old\ncommit hashes not being stored in the repository.  The old commit\nhashes are just replace refs that replace non-existing objects (at\nleast within the newly rewritten repo) that happen to match old commit\nhashes and map to the new commit hashes.  Unfortunately this isn't\nquite a perfect solution, there are still three known downsides:\n\n  * replace refs cannot be abbreviated, unlike real object ids.  Thus,\nif you have an abbreviated old commit hash, git won't recognize it in\nsuch a setup.\n  * commit-graph apparently assumes that the existence of replace refs\nimplies that commit objects in the repo have likely been replaced\n(even though that is not the case for this situation), and thus is\ndisabled when such refs are present.\n  * external GUI programs such as GitHub and Gerrit and likely others\ndo not honor replace refs, instead showing you some form of \"Not\nFound\" error.\n\n\nAs for using replace refs to attempt to alleviate problems without\nrewriting history, that's an even bigger can of worms and it doesn't\nsolve clone/fetch/gc/fsck nor the many other places you highlighted in\nyour email.\n"},{"id":"393321","messageId":"87lfo0881d.fsf@vps.thesusis.net","threadId":"52980","inReplyTo":"20200315221940.bdgi5mluxuetq2lz@doriath","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Phillip Susi","fromEmail":"phill@thesusis.net","sentAt":"2020-03-16T18:32:46Z","receivedAt":"2020-03-16T18:32:51Z","isPatch":false,"sender":{"key":"phill@thesusis.net","avatar":null},"body":"\nDamien Robert writes:\n\n> My situation: coworkers push big files by mistake, I don't want to rewrite\n> history because they are not too well versed with git, but I want to keep\n> *my* repo clean.\n>\n> Partial solution:\n> - identify the large blobs (easy)\n> - write a replace ref (easy):\n>   $ git replace b5f74037bb91 $(git hash-object -w -t blob /dev/null)\n>   and replace the file (if it is still in the repo) by an empty file.\n>\n> Now the pain points start:\n> - first the index does not handle replace (I think), so the replaced file\n>   appear as changed in git status, even through eg git diff shows nothing.\n\nInstead of replacing the blob with an empty file, why not replace the\ntree that references it with one that does not?  That way you won't have\nthe file in your checkout at all, and the index won't list it so status\nwon't show it as changed.\n\n"},{"id":"393327","messageId":"20200316193144.GD1073710@coredump.intra.peff.net","threadId":"52980","inReplyTo":"86pndd794k.fsf@gmail.com","subject":"Re: Notes from Git Contributor Summit, Los Angeles (April 5, 2020)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-16T19:31:44Z","receivedAt":"2020-03-16T19:31:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Mar 15, 2020 at 07:42:19PM +0100, Jakub Narebski wrote:\n\n> One question to all participating live (in person): how those topics\n> were proposed, and how they were voted for?  This was done before remote\n> access was turned on, I think.\n\nDuring breakfast people wrote topics no a whiteboard and people voted on\nthem by putting a tick-mark on the board. The topics and votes were\ntransferred to the Google Doc for notes. Next time I think we'll just go\nstraight to the online doc to save time and make things friendlier for\nremote folks (the whiteboard is a holdover from when we didn't have any\nremotes). And I'll make it clear that remote people are welcome to add\ntopics and vote via the doc.\n\n-Peff\n"},{"id":"393328","messageId":"5cab1530-f8b6-cef3-7b93-48fad410a160@iee.email","threadId":"52980","inReplyTo":"20200315221940.bdgi5mluxuetq2lz@doriath","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2020-03-16T20:01:13Z","receivedAt":"2020-03-16T20:01:21Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Damien, James, (and 4 other who voted for the topic)\n\nI had been thinking about 'missing' blobs for a long while as a earlier\n'partial clone' concept (unpublished)\n\nOn 15/03/2020 22:19, Damien Robert wrote:\n> From James Ramsay, Thu 12 Mar 2020 at 14:57:24 (+1100) :\n>> 6. Elijah: replace refs helps, but not supported by hosts like GitHub etc\n>>     a. Stolee: breaks commit graph because of generation numbers.\n>>     b. Replace refs for blobs, then special packfile, there were edge cases.\n> I am interested in more details on how to handle this using replace.\n>\n> My situation: coworkers push big files by mistake, I don't want to rewrite\n> history because they are not too well versed with git, but I want to keep\n> *my* repo clean.\n>\n> Partial solution:\n> - identify the large blobs (easy)\n> - write a replace ref (easy):\n>   $ git replace b5f74037bb91 $(git hash-object -w -t blob /dev/null)\n>   and replace the file (if it is still in the repo) by an empty file.\n\nHere, my idea was to create a deliberately malformed blob object that\nwould allow self reference to say \"this blob is deliberately missing\".\n(i.e. the same content would exist under two oids, one valid, one\ninvalid) The change would require extra code (more below).\n\nManaging the verification of the replacement is a bigger problem,\nespecially if already pushed to a server.\n> Now the pain points start:\n> - first the index does not handle replace (I think), so the replaced file\n>   appear as changed in git status, even through eg git diff shows nothing.\n>\n> => Solution: configure .git/info/sparse-checkout\n>\n> - secondly, I want to remove the large blob from my repo.\n>\n> Ideally I'd like to repack everything but filter this blob, except that\n> repack does not understand --filter. So I need to use `git pack-objects`\n> directly and then do the naming and clean up that repack usually does\n> manually, which is error prone.\n>\n> Furthermore, while `git pack-objects` accepts --filter, I can only filter on\n> blob size, not blob oid. (there is filter=sparse:oid where I could reuse my\n> sparse checkout file, but I would need to make a blob of it first). And if I\n> have one large file I want to keep, I cannot filter by blob size.\n>\n> Another solution would be to use `git unpack-objects` to unpack all objects\n> (except I would need to do that in an empty git dir), remove the blob, and\n> then repack everything.\n>\n> Am I missing a simpler solution?\n>\n> - finally, checkouting to a ref including the replaced (now missing) blob\n>   gives error messages of the form:\n> error: invalid object 100644 b5f74037bb91c45606b233b0ad6aad86f8e3875e for 'Silverman-Height-NonTorsion.pdf'\n>\n> On the one hand it is reassuring that git checks that the real object\n> (rather than only the replaced object) is still there, on the other hand it\n> would be nice to ask git to completely forget about the original object\n> (except fsck of course).\n>\n> Thanks,\n> Damien\nMy notes on the \"13. Obliterate\" ideas.\n\n1. If the object is in the wild & is dangerous : Stop: Failed: Damage\nlimitation.\n2. If the object is external, but still tame : Seek and recapture;\neither treat as internal, or treat as wild [1].\n3. The object is in captivity, even if distributed around an enclosure.\nProceed to vaccination [4].\n4. Create new blob object with exact content \"Git revoke: <oid>\" (or\nsimilar) This object includes the embedded object type coding as part of\nthe object format. This object is/becomes part of the git signature/oid\ncommit hierarchy. This should (ultimately) be on 'master' branch as it\nis the verifier for the obliteration.\n5. In the old revoked object <oid>, replace the object content (after\nzlib etc) with the same content as created in step 4. This deliberately\n_malformed_ object would normally cause fsck to barf. see [6]\n6. However here we/fsck would detect the length and prefix of the\n(barfed) object contents and so determine its oid (the oid of the\ncontent). This results in an oid equal to that found in 4. which can be\nlooked up and determined to be a self referral to this obliterated oid,\nso an fsck 'pass:obliterated' result is returned. This content could be\nactually be stored in any removed file if checked out!\n\nConsequences:\nPacks and other served object contents no longer contain the  revoked oid.\nHygiene/vaccination needs applied to other distributed recipients of the\nformer defective object.\n\nPossible attacks: Attacker removes other important commits/blobs/trees\nby adding a 'revoke' which propagates to other users: Separate the\nhygiene cycle from the initial server revocation.\n\nFor trees(?) and commits the message is \"revoke <oid> <use-oid>\". But\nwhere to 'hold' the commit & tree (maybe require that tree revoke is\ntreated as a commit revoke, so the the new tree is got for free). We\nstill need the new commit to be walked by fsck/gc, and the old oid\ncontents to be gc'd.\nFor a 'commit' revocation it (the new msg/trees/revision) could maybe be\na 2nd (or third parent after {0}) so a 'normal walk finds it, but\nprobably that's just a recipe for disaster.\nMaybe a revocation reflog that doesn't expire? or can be rebuilt (fsck\nwould extend it's lost/found to include a revoked list).\n\nThe new (XY) problem is now one of tying in the new revoked blob to the\n'old' commit/tree hierarchy which only handles tracked files! Maybe its\na .revoked file (like a .gitignore) which has a list of the old oids and\nhas actual blobs attached under a .revoked tree.\n\nAlso need to make sure that re-packing is done if the blob/tree/commit\nwas a delta-base at the point of obliteration. Also need to prompt the\nlocal user, just in case it's a spoof!. Plus need a way of 'sending' the\nrevocation. (and flag for what to do about a fetch pack containing a\nrevocation for which we have the original, esp if we have it as a pack\nthat will take a long time to recreate. Need a way of writing the\n'defective' object (more code).\n\nNewhash transition. When histories are rewritten, then the obliterated\nartefacts are truly removed. For new repos using the newhash then the\nrevocation mechanism is essentially the same other than extending the\nnominal size of the revocation objects.\n\nPerhaps use the 'submodule' commit object type (i.e 'stuff held\nelsewhere')  for the holder of the revoked ID (for commits & trees).\nThis could be locked into the history (details not fully thought\nthrough..).\n\nIf there is a design error within Git, its the lack of an 'after the\nfact' redaction mechanism (and how it is spread across branches and\ndistributed users/servers) - not easy.\n\nPhilip\n"},{"id":"393339","messageId":"73783776-C6B3-42FE-B8F7-2E43BC175B58@gmail.com","threadId":"52980","inReplyTo":"20200312133127.GK212281@google.com","subject":"Re: [TOPIC 16/17] “I want a reviewer”","fromName":"Philippe Blain","fromEmail":"levraiphilippeblain@gmail.com","sentAt":"2020-03-17T00:43:06Z","receivedAt":"2020-03-17T00:43:14Z","isPatch":false,"sender":{"key":"levraiphilippeblain@gmail.com","avatar":"https://avatars.githubusercontent.com/u/44212482?v=4"},"body":"Hi Emily,\n\n> Le 12 mars 2020 à 09:31, Emily Shaffer <emilyshaffer@google.com> a écrit :\n> \n> On Thu, Mar 12, 2020 at 03:14:25PM +1100, James Ramsay wrote:\n>> 5. Jonathan N: patchwork exists, need to learn how to use it :)\n> \n> We've actually got a meeting with some Patchwork folks today - if\n> anybody has a burning need they want filled via Patchwork, just say so,\n> and we'll try to ask.\n\nI just read this so I don't know if it's too late, but patchwork does not cope well with how Gitgitgadget uses the same email address for all submissions.\nI reported that here:  https://lore.kernel.org/git/75987318-A9A7-4235-8B1D-315B29B644E8@gmail.com/, but haven't opened an issue yet on patchwork's bug tracker.\nI'm not sure either if the best course of action is on the GGG or the patchwork side, though, as perJunio's suggestion in the above thread...\n\nPhilippe."},{"id":"393345","messageId":"CAP8UFD0wJo4onz0_Vw4-bcX1h61=J=ZiKfM-fMXLj4B9q0aveg@mail.gmail.com","threadId":"52980","inReplyTo":"58425B78-7C6D-41CD-92AE-434D0A58F968@jramsay.com.au","subject":"Allowing only blob filtering was: [TOPIC 5/17] Partial Clone","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2020-03-17T07:38:08Z","receivedAt":"2020-03-17T07:38:22Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi Taylor and Peff,\n\nOn Thu, Mar 12, 2020 at 5:01 AM James Ramsay <james@jramsay.com.au> wrote:\n>\n> 1. Stolee: what is the status, who is deploying it, what issues need to\n> be handled? Example, downloading tags. Hard to highly recommend it.\n>\n> 2. Taylor: we deployed it. No activity except for internal testing. Some\n> more activity, but no crashes. Have been dragging our feet. Chicken egg,\n> can’t deploy it because the client may not work, but hoping to hear\n> about problems.\n>\n> 3. ZJ: dark launched for a mission critical repos. Internal questions\n> from CI team, not sure about performance. Build farm hitting it hard\n> with different filter specs.\n>\n> 4. Taylor: we have patches we are promising to the list. Blob none and\n> limit, for now, but add them incrementally. Bitmap patches are on the\n> list\n\nWe (GitLab) would be interested in seeing the patches you already have\nthat only allow blob filtering.\n\nThanks,\nChristian.\n"},{"id":"393372","messageId":"cover.1584477196.git.me@ttaylorr.com","threadId":"52980","inReplyTo":"CAP8UFD0wJo4onz0_Vw4-bcX1h61=J=ZiKfM-fMXLj4B9q0aveg@mail.gmail.com","subject":"[RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-17T20:39:05Z","receivedAt":"2020-03-17T20:39:11Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Hi Christian,\n\nOf course, I would be happy to send along our patches. They are included\nin the series below, and correspond roughly to what we are running at\nGitHub. (For us, there have been a few more clean-ups and additional\npatches, but I squashed them into 2/2 below).\n\nThe approach is roughly that we have:\n\n  - 'uploadpack.filter.allow' -> specifying the default for unspecified\n    filter choices, itself defaulting to true in order to maintain\n    backwards compatibility, and\n\n  - 'uploadpack.filter.<filter>.allow' -> specifying whether or not each\n    filter kind is allowed or not. (Originally this was given as 'git\n    config uploadpack.filter=blob:none.allow true', but this '=' is\n    ambiguous to configuration given over '-c', which itself uses an '='\n    to separate keys from values.)\n\nI noted in the second patch that there is the unfortunate possibility of\nencountering a SIGPIPE when trying to write the ERR sideband back to a\nclient who requested a non-supported filter. Peff and I have had some\ndiscussion off-list about resurrecting SZEDZER's work which makes room\nin the buffer by reading one packet back from the client when the server\nencounters a SIGPIPE. It is for this reason that I am marking the series\nas 'RFC'.\n\nFor reference, our configuration at GitHub looks something like:\n\n  [uploadpack]\n    allowAnySHA1InWant = true\n    allowFilter = true\n  [uploadpack \"filter\"]\n    allow = false\n  [uploadpack \"filter.blob:limit\"]\n    allow = true\n  [uploadpack \"filter.blob:none\"]\n    allow = true\n\nwith a few irrelevant details elided for the purposes of the list :-).\n\nI'd be happy to take in any comments that you or others might have\nbefore dropping the 'RFC' status.\n\nTaylor Blau (2):\n  list_objects_filter_options: introduce 'list_object_filter_config_name'\n  upload-pack.c: allow banning certain object filter(s)\n\n Documentation/config/uploadpack.txt | 12 ++++++\n list-objects-filter-options.c       | 25 +++++++++++\n list-objects-filter-options.h       |  6 +++\n t/t5616-partial-clone.sh            | 23 ++++++++++\n upload-pack.c                       | 67 +++++++++++++++++++++++++++++\n 5 files changed, 133 insertions(+)\n\n--\n2.26.0.rc2.2.g888d9484cf\n"},{"id":"393373","messageId":"c75806d011b04f2ad7efbbec01613a2d0b1f570b.1584477196.git.me@ttaylorr.com","threadId":"52980","inReplyTo":"cover.1584477196.git.me@ttaylorr.com","subject":"[RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-17T20:39:09Z","receivedAt":"2020-03-17T20:39:13Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"In a subsequent commit, we will add configuration options that are\nspecific to each kind of object filter, in which case it is handy to\nhave a function that translates between 'enum\nlist_objects_filter_choice' and an appropriate configuration-friendly\nstring.\n---\n list-objects-filter-options.c | 25 +++++++++++++++++++++++++\n list-objects-filter-options.h |  6 ++++++\n 2 files changed, 31 insertions(+)\n\ndiff --git a/list-objects-filter-options.c b/list-objects-filter-options.c\nindex 256bcfbdfe..6b6aa0b3ec 100644\n--- a/list-objects-filter-options.c\n+++ b/list-objects-filter-options.c\n@@ -15,6 +15,31 @@ static int parse_combine_filter(\n \tconst char *arg,\n \tstruct strbuf *errbuf);\n \n+const char *list_object_filter_config_name(enum list_objects_filter_choice c)\n+{\n+\tswitch (c) {\n+\tcase LOFC_BLOB_NONE:\n+\t\treturn \"blob:none\";\n+\tcase LOFC_BLOB_LIMIT:\n+\t\treturn \"blob:limit\";\n+\tcase LOFC_TREE_DEPTH:\n+\t\treturn \"tree:depth\";\n+\tcase LOFC_SPARSE_OID:\n+\t\treturn \"sparse:oid\";\n+\tcase LOFC_COMBINE:\n+\t\treturn \"combine\";\n+\tcase LOFC_DISABLED:\n+\tcase LOFC__COUNT:\n+\t\t/*\n+\t\t * Include these to catch all enumerated values, but\n+\t\t * break to treat them as a bug. Any new values of this\n+\t\t * enum will cause a compiler error, as desired.\n+\t\t */\n+\t\tbreak;\n+\t}\n+\tBUG(\"list_object_filter_choice_name: invalid argument '%d'\", c);\n+}\n+\n /*\n  * Parse value of the argument to the \"filter\" keyword.\n  * On the command line this looks like:\ndiff --git a/list-objects-filter-options.h b/list-objects-filter-options.h\nindex 2ffb39222c..e5259e4ac6 100644\n--- a/list-objects-filter-options.h\n+++ b/list-objects-filter-options.h\n@@ -17,6 +17,12 @@ enum list_objects_filter_choice {\n \tLOFC__COUNT /* must be last */\n };\n \n+/*\n+ * Returns a configuration key suitable for describing the given object filter,\n+ * e.g.: \"blob:none\", \"combine\", etc.\n+ */\n+const char *list_object_filter_config_name(enum list_objects_filter_choice c);\n+\n struct list_objects_filter_options {\n \t/*\n \t * 'filter_spec' is the raw argument value given on the command line\n-- \n2.26.0.rc2.2.g888d9484cf\n\n"},{"id":"393374","messageId":"888d9484cf4130e90f451134c236a290a6c5e18d.1584477196.git.me@ttaylorr.com","threadId":"52980","inReplyTo":"cover.1584477196.git.me@ttaylorr.com","subject":"[RFC PATCH 2/2] upload-pack.c: allow banning certain object filter(s)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-17T20:39:11Z","receivedAt":"2020-03-17T20:39:16Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Git clients may ask the server for a partial set of objects, where the\nset of objects being requested is refined by one or more object filters.\nServer administrators can configure 'git upload-pack' to allow or ban\nthese filters by setting the 'uploadpack.allowFilter' variable to\n'true' or 'false', respectively.\n\nHowever, administrators using bitmaps may wish to allow certain kinds of\nobject filters, but ban others. Specifically, they may wish to allow\nobject filters that can be optimized by the use of bitmaps, while\nrejecting other object filters which aren't and represent a perceived\nperformance degradation (as well as an increased load factor on the\nserver).\n\nAllow configuring 'git upload-pack' to support object filters on a\ncase-by-case basis by introducing a new configuration variable and\nsection:\n\n  - 'uploadpack.filter.allow'\n\n  - 'uploadpack.filter.<kind>.allow'\n\nwhere '<kind>' may be one of 'blob:none', 'blob:limit', 'tree:depth',\nand so on. The additional '.' between 'filter' and '<kind>' is part of\nthe sub-section.\n\nSetting the second configuration variable for any valid value of\n'<kind>' explicitly allows or disallows restricting that kind of object\nfilter.\n\nIf a client requests the object filter <kind> and the respective\nconfiguration value is not set, 'git upload-pack' will default to the\nvalue of 'uploadpack.filter.allow', which itself defaults to 'true' to\nmaintain backwards compatibility. Note that this differs from\n'uploadpack.allowfilter', which controls whether or not the 'filter'\ncapability is advertised.\n\nNB: this introduces an unfortunate possibility that attempt to write the\nERR sideband will cause a SIGPIPE. This can be prevented by some of\nSZEDZER's previous work, but it is silenced in 't' for now.\n---\n Documentation/config/uploadpack.txt | 12 ++++++\n t/t5616-partial-clone.sh            | 23 ++++++++++\n upload-pack.c                       | 67 +++++++++++++++++++++++++++++\n 3 files changed, 102 insertions(+)\n\ndiff --git a/Documentation/config/uploadpack.txt b/Documentation/config/uploadpack.txt\nindex ed1c835695..6213bd619c 100644\n--- a/Documentation/config/uploadpack.txt\n+++ b/Documentation/config/uploadpack.txt\n@@ -57,6 +57,18 @@ uploadpack.allowFilter::\n \tIf this option is set, `upload-pack` will support partial\n \tclone and partial fetch object filtering.\n \n+uploadpack.filter.allow::\n+\tProvides a default value for unspecified object filters (see: the\n+\tbelow configuration variable).\n+\tDefaults to `true`.\n+\n+uploadpack.filter.<filter>.allow::\n+\tExplicitly allow or ban the object filter corresponding to `<filter>`,\n+\twhere `<filter>` may be one of: `blob:none`, `blob:limit`, `tree:depth`,\n+\t`sparse:oid`, or `combine`. If using combined filters, both `combine`\n+\tand all of the nested filter kinds must be allowed.\n+\tDefaults to `uploadpack.filter.allow`.\n+\n uploadpack.allowRefInWant::\n \tIf this option is set, `upload-pack` will support the `ref-in-want`\n \tfeature of the protocol version 2 `fetch` command.  This feature\ndiff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh\nindex 77bb91e976..ee1af9b682 100755\n--- a/t/t5616-partial-clone.sh\n+++ b/t/t5616-partial-clone.sh\n@@ -235,6 +235,29 @@ test_expect_success 'implicitly construct combine: filter with repeated flags' '\n \ttest_cmp unique_types.expected unique_types.actual\n '\n \n+test_expect_success 'upload-pack fails banned object filters' '\n+\t# Ensure that configuration keys are normalized by capitalizing\n+\t# \"blob:None\" below:\n+\ttest_config -C srv.bare uploadpack.filter.blob:None.allow false &&\n+\ttest_must_fail ok=sigpipe git clone --no-checkout --filter.blob:none \\\n+\t\t\"file://$(pwd)/srv.bare\" pc3\n+'\n+\n+test_expect_success 'upload-pack fails banned combine object filters' '\n+\ttest_config -C srv.bare uploadpack.filter.allow false &&\n+\ttest_config -C srv.bare uploadpack.filter.combine.allow true &&\n+\ttest_config -C srv.bare uploadpack.filter.tree:depth.allow true &&\n+\ttest_config -C srv.bare uploadpack.filter.blob:none.allow false &&\n+\ttest_must_fail ok=sigpipe git clone --no-checkout --filter=tree:1 \\\n+\t\t--filter=blob:none \"file://$(pwd)/srv.bare\" pc3\n+'\n+\n+test_expect_success 'upload-pack fails banned object filters with fallback' '\n+\ttest_config -C srv.bare uploadpack.filter.allow false &&\n+\ttest_must_fail ok=sigpipe git clone --no-checkout --filter=blob:none \\\n+\t\t\"file://$(pwd)/srv.bare\" pc3\n+'\n+\n test_expect_success 'partial clone fetches blobs pointed to by refs even if normally filtered out' '\n \trm -rf src dst &&\n \tgit init src &&\ndiff --git a/upload-pack.c b/upload-pack.c\nindex c53249cac1..81f2701f99 100644\n--- a/upload-pack.c\n+++ b/upload-pack.c\n@@ -69,6 +69,8 @@ static int filter_capability_requested;\n static int allow_filter;\n static int allow_ref_in_want;\n static struct list_objects_filter_options filter_options;\n+static struct string_list allowed_filters = STRING_LIST_INIT_DUP;\n+static int allow_filter_fallback = 1;\n \n static int allow_sideband_all;\n \n@@ -848,6 +850,45 @@ static int process_deepen_not(const char *line, struct string_list *deepen_not,\n \treturn 0;\n }\n \n+static int allows_filter_choice(enum list_objects_filter_choice c)\n+{\n+\tconst char *key = list_object_filter_config_name(c);\n+\tstruct string_list_item *item = string_list_lookup(&allowed_filters,\n+\t\t\t\t\t\t\t   key);\n+\tif (item)\n+\t\treturn (intptr_t) item->util;\n+\treturn allow_filter_fallback;\n+}\n+\n+static struct list_objects_filter_options *banned_filter(\n+\tstruct list_objects_filter_options *opts)\n+{\n+\tsize_t i;\n+\n+\tif (!allows_filter_choice(opts->choice))\n+\t\treturn opts;\n+\n+\tif (opts->choice == LOFC_COMBINE)\n+\t\tfor (i = 0; i < opts->sub_nr; i++) {\n+\t\t\tstruct list_objects_filter_options *sub = &opts->sub[i];\n+\t\t\tif (banned_filter(sub))\n+\t\t\t\treturn sub;\n+\t\t}\n+\treturn NULL;\n+}\n+\n+static void die_if_using_banned_filter(struct packet_writer *w,\n+\t\t\t\t       struct list_objects_filter_options *opts)\n+{\n+\tstruct list_objects_filter_options *banned = banned_filter(opts);\n+\tif (!banned)\n+\t\treturn;\n+\n+\tpacket_writer_error(w, _(\"filter '%s' not supported\\n\"),\n+\t\t\t    list_object_filter_config_name(banned->choice));\n+\tdie(_(\"git upload-pack: banned object filter requested\"));\n+}\n+\n static void receive_needs(struct packet_reader *reader, struct object_array *want_obj)\n {\n \tstruct object_array shallows = OBJECT_ARRAY_INIT;\n@@ -885,6 +926,7 @@ static void receive_needs(struct packet_reader *reader, struct object_array *wan\n \t\t\t\tdie(\"git upload-pack: filtering capability not negotiated\");\n \t\t\tlist_objects_filter_die_if_populated(&filter_options);\n \t\t\tparse_list_objects_filter(&filter_options, arg);\n+\t\t\tdie_if_using_banned_filter(&writer, &filter_options);\n \t\t\tcontinue;\n \t\t}\n \n@@ -1044,6 +1086,9 @@ static int find_symref(const char *refname, const struct object_id *oid,\n \n static int upload_pack_config(const char *var, const char *value, void *unused)\n {\n+\tconst char *sub, *key;\n+\tint sub_len;\n+\n \tif (!strcmp(\"uploadpack.allowtipsha1inwant\", var)) {\n \t\tif (git_config_bool(var, value))\n \t\t\tallow_unadvertised_object_request |= ALLOW_TIP_SHA1;\n@@ -1065,6 +1110,26 @@ static int upload_pack_config(const char *var, const char *value, void *unused)\n \t\t\tkeepalive = -1;\n \t} else if (!strcmp(\"uploadpack.allowfilter\", var)) {\n \t\tallow_filter = git_config_bool(var, value);\n+\t} else if (!parse_config_key(var, \"uploadpack\", &sub, &sub_len, &key) &&\n+\t\t   key && !strcmp(key, \"allow\")) {\n+\t\tif (sub && skip_prefix(sub, \"filter.\", &sub) && sub_len >= 7) {\n+\t\t\tstruct string_list_item *item;\n+\t\t\tchar *spec;\n+\n+\t\t\t/*\n+\t\t\t * normalize the filter, and chomp off '.allow' from the\n+\t\t\t * end\n+\t\t\t */\n+\t\t\tspec = xstrdup_tolower(sub);\n+\t\t\tspec[sub_len - 7] = 0;\n+\n+\t\t\titem = string_list_insert(&allowed_filters, spec);\n+\t\t\titem->util = (void *) (intptr_t) git_config_bool(var, value);\n+\n+\t\t\tfree(spec);\n+\t\t} else if (!strcmp(\"uploadpack.filter.allow\", var)) {\n+\t\t\tallow_filter_fallback = git_config_bool(var, value);\n+\t\t}\n \t} else if (!strcmp(\"uploadpack.allowrefinwant\", var)) {\n \t\tallow_ref_in_want = git_config_bool(var, value);\n \t} else if (!strcmp(\"uploadpack.allowsidebandall\", var)) {\n@@ -1308,6 +1373,8 @@ static void process_args(struct packet_reader *request,\n \t\tif (allow_filter && skip_prefix(arg, \"filter \", &p)) {\n \t\t\tlist_objects_filter_die_if_populated(&filter_options);\n \t\t\tparse_list_objects_filter(&filter_options, p);\n+\t\t\tdie_if_using_banned_filter(&data->writer,\n+\t\t\t\t\t\t   &filter_options);\n \t\t\tcontinue;\n \t\t}\n \n-- \n2.26.0.rc2.2.g888d9484cf\n"},{"id":"393375","messageId":"CAPig+cTVtv+uzzpoZ-BT=F=srdt1ewvgeBAAr9R+OUCYSov65A@mail.gmail.com","threadId":"52980","inReplyTo":"c75806d011b04f2ad7efbbec01613a2d0b1f570b.1584477196.git.me@ttaylorr.com","subject":"Re: [RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2020-03-17T20:53:44Z","receivedAt":"2020-03-17T20:53:59Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Tue, Mar 17, 2020 at 4:40 PM Taylor Blau <me@ttaylorr.com> wrote:\n> In a subsequent commit, we will add configuration options that are\n> specific to each kind of object filter, in which case it is handy to\n> have a function that translates between 'enum\n> list_objects_filter_choice' and an appropriate configuration-friendly\n> string.\n> ---\n\nMissing sign-off (but perhaps that's intentional since this is RFC).\n\n> diff --git a/list-objects-filter-options.c b/list-objects-filter-options.c\n> @@ -15,6 +15,31 @@ static int parse_combine_filter(\n> +const char *list_object_filter_config_name(enum list_objects_filter_choice c)\n> +{\n> +       switch (c) {\n> +       case LOFC_BLOB_NONE:\n> +               return \"blob:none\";\n> +       case LOFC_BLOB_LIMIT:\n> +               return \"blob:limit\";\n> +       case LOFC_TREE_DEPTH:\n> +               return \"tree:depth\";\n> +       case LOFC_SPARSE_OID:\n> +               return \"sparse:oid\";\n> +       case LOFC_COMBINE:\n> +               return \"combine\";\n> +       case LOFC_DISABLED:\n> +       case LOFC__COUNT:\n> +               /*\n> +                * Include these to catch all enumerated values, but\n> +                * break to treat them as a bug. Any new values of this\n> +                * enum will cause a compiler error, as desired.\n> +                */\n\nIn general, people will see a warning, not an error, unless they\nspecifically use -Werror (or such) to turn the warning into an error,\nso this statement is misleading. Also, while some compilers may\ncomplain, others may not. So, although the comment claims that we will\nnotice an unhandled enum constant at compile-time, that isn't\nnecessarily the case.\n\nMoreover, the comment itself, in is present form, is rather\nsuperfluous since its merely repeating what the BUG() invocation just\nbelow it already tells me. In fact, as a reader of this code, I would\nbe more interested in knowing why those two cases do not have string\nequivalents which are returned (although perhaps even that would be\nobvious to someone familiar with the code, hence the comment can\nprobably be dropped altogether).\n\n> +               break;\n> +       }\n> +       BUG(\"list_object_filter_choice_name: invalid argument '%d'\", c);\n> +}\n"},{"id":"393376","messageId":"CAPig+cRgnqmwCCjFV32K_ysawHBkJN_y6=Do_oKXjjpy0BSvUQ@mail.gmail.com","threadId":"52980","inReplyTo":"888d9484cf4130e90f451134c236a290a6c5e18d.1584477196.git.me@ttaylorr.com","subject":"Re: [RFC PATCH 2/2] upload-pack.c: allow banning certain object filter(s)","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2020-03-17T21:11:42Z","receivedAt":"2020-03-17T21:11:56Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Tue, Mar 17, 2020 at 4:40 PM Taylor Blau <me@ttaylorr.com> wrote:\n> NB: this introduces an unfortunate possibility that attempt to write the\n> ERR sideband will cause a SIGPIPE. This can be prevented by some of\n> SZEDZER's previous work, but it is silenced in 't' for now.\n\ns/SZEDZER/SZEDER/\n\n> diff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh\n> @@ -235,6 +235,29 @@ test_expect_success 'implicitly construct combine: filter with repeated flags' '\n> +test_expect_success 'upload-pack fails banned object filters' '\n> +       # Ensure that configuration keys are normalized by capitalizing\n> +       # \"blob:None\" below:\n> +       test_config -C srv.bare uploadpack.filter.blob:None.allow false &&\n\nI found the wording of the comment more confusing than clarifying.\nPerhaps rewriting it like this could help:\n\n    Test case-insensitivity by intentional use of \"blob:None\" rather than\n    \"blob:none\".\n\nor something.\n\n> +       test_must_fail ok=sigpipe git clone --no-checkout --filter.blob:none \\\n> +               \"file://$(pwd)/srv.bare\" pc3\n> +'\n"},{"id":"393387","messageId":"20200318100327.GA1227946@coredump.intra.peff.net","threadId":"52980","inReplyTo":"CAPig+cTVtv+uzzpoZ-BT=F=srdt1ewvgeBAAr9R+OUCYSov65A@mail.gmail.com","subject":"Re: [RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-18T10:03:27Z","receivedAt":"2020-03-18T10:03:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Mar 17, 2020 at 04:53:44PM -0400, Eric Sunshine wrote:\n\n> > +       case LOFC_DISABLED:\n> > +       case LOFC__COUNT:\n> > +               /*\n> > +                * Include these to catch all enumerated values, but\n> > +                * break to treat them as a bug. Any new values of this\n> > +                * enum will cause a compiler error, as desired.\n> > +                */\n> \n> In general, people will see a warning, not an error, unless they\n> specifically use -Werror (or such) to turn the warning into an error,\n> so this statement is misleading. Also, while some compilers may\n> complain, others may not. So, although the comment claims that we will\n> notice an unhandled enum constant at compile-time, that isn't\n> necessarily the case.\n\nYes, but that's the best we can do, isn't it?\n\nThere's sort of a meta-issue here which Taylor and I discussed off-list\nand which led to this comment. We quite often write switch statements\nover enums like this:\n\n  switch (foo)\n  case FOO_ONE:\n\t...do something...\n  case FOO_TWO:\n        ...something else...\n  default:\n\tBUG(\"I don't know what to do with %d\", foo);\n  }\n\nThat's reasonable and does the right thing at runtime if we ever hit\nthis case. But it has the unfortunate side effect that we lose any\n-Wswitch warning that could tell us at compile time that we're missing a\ncase. Not everybody would see such a warning, as you note, but\ndevelopers on gcc and clang generally would (it's part of -Wall).\n\nBut we can't just remove the default case. Even though enums don't\ngenerally take on other values, it's legal for them to do so. So we do\nwant to make sure we BUG() in that instance.\n\nThis is awkward to solve in the general case[1]. But because we're\nreturning in each case arm here, it's easy to just put the BUG() after\nthe switch. Anything that didn't return is unhandled, and we get the\nbest of both: -Wswitch warnings when we need to add a new filter type,\nand a BUG() in the off chance that we see an unexpected value.\n\nBut the cost is that we have to enumerate the set of values that are\ndefined but not handled here (LOFC__COUNT, for instance, isn't a real\nenum value but rather a placeholder to let other code know how many\nfilter types there are).\n\nSo...I dunno. Worth it as a general technique?\n\n-Peff\n\n[1] In the general case where you don't return, you have to somehow know\n    whether the value was actually handled or not (and BUG() if it\n    wasn't). Presumably by keeping a separate flag variable, which is\n    pretty ugly. -Wswitch-enum is supposed to deal with this by\n    requiring that you list all of the values even if you have a default\n    case. But it triggers in a lot of other places in the code that I\n    think would be made much harder to read by having to list out the\n    enumerated possibilities.\n"},{"id":"393388","messageId":"20200318101825.GB1227946@coredump.intra.peff.net","threadId":"52980","inReplyTo":"cover.1584477196.git.me@ttaylorr.com","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-18T10:18:25Z","receivedAt":"2020-03-18T10:18:27Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Mar 17, 2020 at 02:39:05PM -0600, Taylor Blau wrote:\n\n> Hi Christian,\n> \n> Of course, I would be happy to send along our patches. They are included\n> in the series below, and correspond roughly to what we are running at\n> GitHub. (For us, there have been a few more clean-ups and additional\n> patches, but I squashed them into 2/2 below).\n> \n> The approach is roughly that we have:\n> \n>   - 'uploadpack.filter.allow' -> specifying the default for unspecified\n>     filter choices, itself defaulting to true in order to maintain\n>     backwards compatibility, and\n> \n>   - 'uploadpack.filter.<filter>.allow' -> specifying whether or not each\n>     filter kind is allowed or not. (Originally this was given as 'git\n>     config uploadpack.filter=blob:none.allow true', but this '=' is\n>     ambiguous to configuration given over '-c', which itself uses an '='\n>     to separate keys from values.)\n\nOne thing that's a little ugly here is the embedded dot in the\nsubsection (i.e., \"filter.<filter>\"). It makes it look like a four-level\nkey, but really there is no such thing in Git.  But everything else we\ntried was even uglier.\n\nI think we want to declare a real subsection for each filter and not\njust \"uploadpack.filter.<filter>\". That gives us room to expand to other\nconfig options besides \"allow\" later on if we need to.\n\nWe don't want to claim \"uploadpack.allow\" and \"uploadpack.<filter>.allow\";\nthat's too generic.\n\nLikewise \"filter.allow\" is too generic.\n\nWe could do \"uploadpackfilter.allow\" and \"uploadpackfilter.<filter>.allow\",\nbut that's both ugly _and_ separates these options from the rest of\nuploadpack.*.\n\nWe could use a character besides \".\", which would reduce confusion. But\nwhat? Using colon is kind of ugly, because it's already syntactically\nsignificant in filter names, and you get:\n\n  uploadpack.filter:blob:none.allow\n\nWe tried equals, like:\n\n  uploadpack.filter=blob:none.allow\n\nbut there's an interesting side effect. Doing:\n\n  git -c uploadpack.filter=blob:none.allow=true upload-pack ...\n\ndoesn't work, because the \"-c\" parser ends the key at the first \"=\". As\nit should, because otherwise we'd get confused by an \"=\" in a value.\nThis is a failing of the \"-c\" syntax; it can't represent values with\n\"=\". Fixing it would be awkward, and I've never seen it come up in\npractice outside of this (you _could_ have a branch with a funny name\nand try to do \"git -c branch.my=funny=branch.remote=origin\" or\nsomething, but the lack of bug reports suggests nobody is that\nmasochistic).\n\nSo...maybe the extra dot is the last bad thing?\n\n> I noted in the second patch that there is the unfortunate possibility of\n> encountering a SIGPIPE when trying to write the ERR sideband back to a\n> client who requested a non-supported filter. Peff and I have had some\n> discussion off-list about resurrecting SZEDZER's work which makes room\n> in the buffer by reading one packet back from the client when the server\n> encounters a SIGPIPE. It is for this reason that I am marking the series\n> as 'RFC'.\n\nFor reference, the patch I was thinking of was this:\n\n  https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n\n-Peff\n"},{"id":"393390","messageId":"13dd0152-b20a-51e1-5940-5e4b67242e9b@iee.email","threadId":"52980","inReplyTo":"888d9484cf4130e90f451134c236a290a6c5e18d.1584477196.git.me@ttaylorr.com","subject":"Re: [RFC PATCH 2/2] upload-pack.c: allow banning certain object filter(s)","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2020-03-18T11:18:16Z","receivedAt":"2020-03-18T11:18:26Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi\nOn 17/03/2020 20:39, Taylor Blau wrote:\n> Git clients may ask the server for a partial set of objects, where the\n> set of objects being requested is refined by one or more object filters.\n> Server administrators can configure 'git upload-pack' to allow or ban\n> these filters by setting the 'uploadpack.allowFilter' variable to\n> 'true' or 'false', respectively.\n>\n> However, administrators using bitmaps may wish to allow certain kinds of\n> object filters, but ban others. Specifically, they may wish to allow\n> object filters that can be optimized by the use of bitmaps, while\n> rejecting other object filters which aren't and represent a perceived\n> performance degradation (as well as an increased load factor on the\n> server).\n>\n> Allow configuring 'git upload-pack' to support object filters on a\n> case-by-case basis by introducing a new configuration variable and\n> section:\n>\n>   - 'uploadpack.filter.allow'\n>\n>   - 'uploadpack.filter.<kind>.allow'\n>\n> where '<kind>' may be one of 'blob:none', 'blob:limit', 'tree:depth',\n> and so on. The additional '.' between 'filter' and '<kind>' is part of\n> the sub-section.\n>\n> Setting the second configuration variable for any valid value of\n> '<kind>' explicitly allows or disallows restricting that kind of object\n> filter.\n>\n> If a client requests the object filter <kind> and the respective\n> configuration value is not set, 'git upload-pack' will default to the\n> value of 'uploadpack.filter.allow', which itself defaults to 'true' to\n> maintain backwards compatibility. Note that this differs from\n> 'uploadpack.allowfilter', which controls whether or not the 'filter'\n> capability is advertised.\n>\n> NB: this introduces an unfortunate possibility that attempt to write the\n> ERR sideband will cause a SIGPIPE. This can be prevented by some of\n> SZEDZER's previous work, but it is silenced in 't' for now.\n> ---\n>  Documentation/config/uploadpack.txt | 12 ++++++\n>  t/t5616-partial-clone.sh            | 23 ++++++++++\n>  upload-pack.c                       | 67 +++++++++++++++++++++++++++++\n>  3 files changed, 102 insertions(+)\n>\n> diff --git a/Documentation/config/uploadpack.txt b/Documentation/config/uploadpack.txt\n> index ed1c835695..6213bd619c 100644\n> --- a/Documentation/config/uploadpack.txt\n> +++ b/Documentation/config/uploadpack.txt\n> @@ -57,6 +57,18 @@ uploadpack.allowFilter::\n>  \tIf this option is set, `upload-pack` will support partial\n>  \tclone and partial fetch object filtering.\n>  \n> +uploadpack.filter.allow::\n> +\tProvides a default value for unspecified object filters (see: the\n> +\tbelow configuration variable).\n> +\tDefaults to `true`.\n> +\n> +uploadpack.filter.<filter>.allow::\n> +\tExplicitly allow or ban the object filter corresponding to `<filter>`,\n> +\twhere `<filter>` may be one of: `blob:none`, `blob:limit`, `tree:depth`,\n> +\t`sparse:oid`, or `combine`. If using combined filters, both `combine`\n> +\tand all of the nested filter kinds must be allowed.\n\nDoesn't the man page at least need the part from the commit message \"The\nadditional '.' between 'filter' and '<kind>' is part of\nthe sub-section.\" as it's not a common mechanism (other comments not\nwithstanding)\n\nPhilip\n> +\tDefaults to `uploadpack.filter.allow`.\n> +\n>  uploadpack.allowRefInWant::\n>  \tIf this option is set, `upload-pack` will support the `ref-in-want`\n>  \tfeature of the protocol version 2 `fetch` command.  This feature\n> diff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh\n> index 77bb91e976..ee1af9b682 100755\n> --- a/t/t5616-partial-clone.sh\n> +++ b/t/t5616-partial-clone.sh\n> @@ -235,6 +235,29 @@ test_expect_success 'implicitly construct combine: filter with repeated flags' '\n>  \ttest_cmp unique_types.expected unique_types.actual\n>  '\n>  \n> +test_expect_success 'upload-pack fails banned object filters' '\n> +\t# Ensure that configuration keys are normalized by capitalizing\n> +\t# \"blob:None\" below:\n> +\ttest_config -C srv.bare uploadpack.filter.blob:None.allow false &&\n> +\ttest_must_fail ok=sigpipe git clone --no-checkout --filter.blob:none \\\n> +\t\t\"file://$(pwd)/srv.bare\" pc3\n> +'\n> +\n> +test_expect_success 'upload-pack fails banned combine object filters' '\n> +\ttest_config -C srv.bare uploadpack.filter.allow false &&\n> +\ttest_config -C srv.bare uploadpack.filter.combine.allow true &&\n> +\ttest_config -C srv.bare uploadpack.filter.tree:depth.allow true &&\n> +\ttest_config -C srv.bare uploadpack.filter.blob:none.allow false &&\n> +\ttest_must_fail ok=sigpipe git clone --no-checkout --filter=tree:1 \\\n> +\t\t--filter=blob:none \"file://$(pwd)/srv.bare\" pc3\n> +'\n> +\n> +test_expect_success 'upload-pack fails banned object filters with fallback' '\n> +\ttest_config -C srv.bare uploadpack.filter.allow false &&\n> +\ttest_must_fail ok=sigpipe git clone --no-checkout --filter=blob:none \\\n> +\t\t\"file://$(pwd)/srv.bare\" pc3\n> +'\n> +\n>  test_expect_success 'partial clone fetches blobs pointed to by refs even if normally filtered out' '\n>  \trm -rf src dst &&\n>  \tgit init src &&\n> diff --git a/upload-pack.c b/upload-pack.c\n> index c53249cac1..81f2701f99 100644\n> --- a/upload-pack.c\n> +++ b/upload-pack.c\n> @@ -69,6 +69,8 @@ static int filter_capability_requested;\n>  static int allow_filter;\n>  static int allow_ref_in_want;\n>  static struct list_objects_filter_options filter_options;\n> +static struct string_list allowed_filters = STRING_LIST_INIT_DUP;\n> +static int allow_filter_fallback = 1;\n>  \n>  static int allow_sideband_all;\n>  \n> @@ -848,6 +850,45 @@ static int process_deepen_not(const char *line, struct string_list *deepen_not,\n>  \treturn 0;\n>  }\n>  \n> +static int allows_filter_choice(enum list_objects_filter_choice c)\n> +{\n> +\tconst char *key = list_object_filter_config_name(c);\n> +\tstruct string_list_item *item = string_list_lookup(&allowed_filters,\n> +\t\t\t\t\t\t\t   key);\n> +\tif (item)\n> +\t\treturn (intptr_t) item->util;\n> +\treturn allow_filter_fallback;\n> +}\n> +\n> +static struct list_objects_filter_options *banned_filter(\n> +\tstruct list_objects_filter_options *opts)\n> +{\n> +\tsize_t i;\n> +\n> +\tif (!allows_filter_choice(opts->choice))\n> +\t\treturn opts;\n> +\n> +\tif (opts->choice == LOFC_COMBINE)\n> +\t\tfor (i = 0; i < opts->sub_nr; i++) {\n> +\t\t\tstruct list_objects_filter_options *sub = &opts->sub[i];\n> +\t\t\tif (banned_filter(sub))\n> +\t\t\t\treturn sub;\n> +\t\t}\n> +\treturn NULL;\n> +}\n> +\n> +static void die_if_using_banned_filter(struct packet_writer *w,\n> +\t\t\t\t       struct list_objects_filter_options *opts)\n> +{\n> +\tstruct list_objects_filter_options *banned = banned_filter(opts);\n> +\tif (!banned)\n> +\t\treturn;\n> +\n> +\tpacket_writer_error(w, _(\"filter '%s' not supported\\n\"),\n> +\t\t\t    list_object_filter_config_name(banned->choice));\n> +\tdie(_(\"git upload-pack: banned object filter requested\"));\n> +}\n> +\n>  static void receive_needs(struct packet_reader *reader, struct object_array *want_obj)\n>  {\n>  \tstruct object_array shallows = OBJECT_ARRAY_INIT;\n> @@ -885,6 +926,7 @@ static void receive_needs(struct packet_reader *reader, struct object_array *wan\n>  \t\t\t\tdie(\"git upload-pack: filtering capability not negotiated\");\n>  \t\t\tlist_objects_filter_die_if_populated(&filter_options);\n>  \t\t\tparse_list_objects_filter(&filter_options, arg);\n> +\t\t\tdie_if_using_banned_filter(&writer, &filter_options);\n>  \t\t\tcontinue;\n>  \t\t}\n>  \n> @@ -1044,6 +1086,9 @@ static int find_symref(const char *refname, const struct object_id *oid,\n>  \n>  static int upload_pack_config(const char *var, const char *value, void *unused)\n>  {\n> +\tconst char *sub, *key;\n> +\tint sub_len;\n> +\n>  \tif (!strcmp(\"uploadpack.allowtipsha1inwant\", var)) {\n>  \t\tif (git_config_bool(var, value))\n>  \t\t\tallow_unadvertised_object_request |= ALLOW_TIP_SHA1;\n> @@ -1065,6 +1110,26 @@ static int upload_pack_config(const char *var, const char *value, void *unused)\n>  \t\t\tkeepalive = -1;\n>  \t} else if (!strcmp(\"uploadpack.allowfilter\", var)) {\n>  \t\tallow_filter = git_config_bool(var, value);\n> +\t} else if (!parse_config_key(var, \"uploadpack\", &sub, &sub_len, &key) &&\n> +\t\t   key && !strcmp(key, \"allow\")) {\n> +\t\tif (sub && skip_prefix(sub, \"filter.\", &sub) && sub_len >= 7) {\n> +\t\t\tstruct string_list_item *item;\n> +\t\t\tchar *spec;\n> +\n> +\t\t\t/*\n> +\t\t\t * normalize the filter, and chomp off '.allow' from the\n> +\t\t\t * end\n> +\t\t\t */\n> +\t\t\tspec = xstrdup_tolower(sub);\n> +\t\t\tspec[sub_len - 7] = 0;\n> +\n> +\t\t\titem = string_list_insert(&allowed_filters, spec);\n> +\t\t\titem->util = (void *) (intptr_t) git_config_bool(var, value);\n> +\n> +\t\t\tfree(spec);\n> +\t\t} else if (!strcmp(\"uploadpack.filter.allow\", var)) {\n> +\t\t\tallow_filter_fallback = git_config_bool(var, value);\n> +\t\t}\n>  \t} else if (!strcmp(\"uploadpack.allowrefinwant\", var)) {\n>  \t\tallow_ref_in_want = git_config_bool(var, value);\n>  \t} else if (!strcmp(\"uploadpack.allowsidebandall\", var)) {\n> @@ -1308,6 +1373,8 @@ static void process_args(struct packet_reader *request,\n>  \t\tif (allow_filter && skip_prefix(arg, \"filter \", &p)) {\n>  \t\t\tlist_objects_filter_die_if_populated(&filter_options);\n>  \t\t\tparse_list_objects_filter(&filter_options, p);\n> +\t\t\tdie_if_using_banned_filter(&data->writer,\n> +\t\t\t\t\t\t   &filter_options);\n>  \t\t\tcontinue;\n>  \t\t}\n>  \n\n"},{"id":"393404","messageId":"xmqqtv2lfrk7.fsf_-_@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200318101825.GB1227946@coredump.intra.peff.net","subject":"Re*: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-03-18T18:26:00Z","receivedAt":"2020-03-18T18:26:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>>   - 'uploadpack.filter.<filter>.allow' -> specifying whether or not each\n>>     filter kind is allowed or not. (Originally this was given as 'git\n>>     config uploadpack.filter=blob:none.allow true', but this '=' is\n>>     ambiguous to configuration given over '-c', which itself uses an '='\n>>     to separate keys from values.)\n>\n> One thing that's a little ugly here is the embedded dot in the\n> subsection (i.e., \"filter.<filter>\"). It makes it look like a four-level\n> key, but really there is no such thing in Git.  But everything else we\n> tried was even uglier.\n\nI think this gives us the best arrangement by upfront forcing all\nthe configuration handers for \"<subcommand>.*.<token>\" namespace,\ncurrent and future, to use \"<group-prefix>\" before the unbounded set\nof user-specifiable values that affects the <subcommand> (which is\n\"uploadpack\").\n\nSo far, the configuration variables that needs to be grouped by\nunbounded set of user-specifiable values we supported happened to\nhave only one sensible such set for each <subcommand>, so we could\nget away without such <group-prefix> and it was perfectly OK to\nhave, say \"guitool.<name>.cmd\".\n\nSyntactically, the convention to always end such <group-prefix> with\na dot \".\" may look unusual, or once readers' eyes get used to them,\nmay look natural.  One tiny sad thing about it is that it cannot be\nmechanically enforced, but that is minor.\n\n> We could do \"uploadpackfilter.allow\" and \"uploadpackfilter.<filter>.allow\",\n> but that's both ugly _and_ separates these options from the rest of\n> uploadpack.*.\n\nThere is an existing instance of a configuration that affects\n<subcommand> that uses a different word after <subcommand>, which is\ncredentialCache.ignoreSIGHUP, and I tend to agree that it is ugly.\n\nBy the way, I noticed the following while I was studying the current\npractice, so before I forget...\n\n-- >8 --\nSubject: [PATCH] separate tar.* config to its own source file\n\nEven though there is only one configuration variable in the\nnamespace, it is not quite right to have tar.umask described\namong the variables for tag.* namespace.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n Documentation/config.txt     | 2 ++\n Documentation/config/tag.txt | 7 -------\n Documentation/config/tar.txt | 6 ++++++\n 3 files changed, 8 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 08b13ba72b..2450589a0e 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -447,6 +447,8 @@ include::config/submodule.txt[]\n \n include::config/tag.txt[]\n \n+include::config/tar.txt[]\n+\n include::config/trace2.txt[]\n \n include::config/transfer.txt[]\ndiff --git a/Documentation/config/tag.txt b/Documentation/config/tag.txt\nindex 6d9110d84c..5062a057ff 100644\n--- a/Documentation/config/tag.txt\n+++ b/Documentation/config/tag.txt\n@@ -15,10 +15,3 @@ tag.gpgSign::\n \tconvenient to use an agent to avoid typing your gpg passphrase\n \tseveral times. Note that this option doesn't affect tag signing\n \tbehavior enabled by \"-u <keyid>\" or \"--local-user=<keyid>\" options.\n-\n-tar.umask::\n-\tThis variable can be used to restrict the permission bits of\n-\ttar archive entries.  The default is 0002, which turns off the\n-\tworld write bit.  The special value \"user\" indicates that the\n-\tarchiving user's umask will be used instead.  See umask(2) and\n-\tlinkgit:git-archive[1].\ndiff --git a/Documentation/config/tar.txt b/Documentation/config/tar.txt\nnew file mode 100644\nindex 0000000000..de8ff48ea9\n--- /dev/null\n+++ b/Documentation/config/tar.txt\n@@ -0,0 +1,6 @@\n+tar.umask::\n+\tThis variable can be used to restrict the permission bits of\n+\ttar archive entries.  The default is 0002, which turns off the\n+\tworld write bit.  The special value \"user\" indicates that the\n+\tarchiving user's umask will be used instead.  See umask(2) and\n+\tlinkgit:git-archive[1].\n"},{"id":"393410","messageId":"xmqqlfnxfo3o.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200318100327.GA1227946@coredump.intra.peff.net","subject":"Re: [RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-03-18T19:40:43Z","receivedAt":"2020-03-18T19:40:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> But the cost is that we have to enumerate the set of values that are\n> defined but not handled here (LOFC__COUNT, for instance, isn't a real\n> enum value but rather a placeholder to let other code know how many\n> filter types there are).\n>\n> So...I dunno. Worth it as a general technique?\n\n\"This is a possible value in the enum we are switching on, so I\nwrite a case arm for it, but we do nothing for it here\" is OK, but\nif it were \"we do nothing for it here or anywhere\" (i.e. the maximum\nenum value defined as a sentinel), the resulting code would be ugly.\n\nI am not sure if the tradeoff is good to force such an ugliness on\nreaders' eyes to squelch the -Wswitch warnings.\n\nSo, I dunno.\n"},{"id":"393418","messageId":"20200318210538.GA31397@syl.local","threadId":"52980","inReplyTo":"CAPig+cTVtv+uzzpoZ-BT=F=srdt1ewvgeBAAr9R+OUCYSov65A@mail.gmail.com","subject":"Re: [RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-18T21:05:38Z","receivedAt":"2020-03-18T21:05:45Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Hi Eric,\n\nOn Tue, Mar 17, 2020 at 04:53:44PM -0400, Eric Sunshine wrote:\n> On Tue, Mar 17, 2020 at 4:40 PM Taylor Blau <me@ttaylorr.com> wrote:\n> > In a subsequent commit, we will add configuration options that are\n> > specific to each kind of object filter, in which case it is handy to\n> > have a function that translates between 'enum\n> > list_objects_filter_choice' and an appropriate configuration-friendly\n> > string.\n> > ---\n>\n> Missing sign-off (but perhaps that's intentional since this is RFC).\n\nYes, the missing sign-off (in this patch as well as 2/2) is intentional,\nsince this is an RFC. Sorry for not calling this out more clearly in my\ncover.\n\nThanks,\nTaylor\n"},{"id":"393419","messageId":"20200318211801.GC31397@syl.local","threadId":"52980","inReplyTo":"CAPig+cRgnqmwCCjFV32K_ysawHBkJN_y6=Do_oKXjjpy0BSvUQ@mail.gmail.com","subject":"Re: [RFC PATCH 2/2] upload-pack.c: allow banning certain object filter(s)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-18T21:18:01Z","receivedAt":"2020-03-18T21:18:05Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Mar 17, 2020 at 05:11:42PM -0400, Eric Sunshine wrote:\n> On Tue, Mar 17, 2020 at 4:40 PM Taylor Blau <me@ttaylorr.com> wrote:\n> > NB: this introduces an unfortunate possibility that attempt to write the\n> > ERR sideband will cause a SIGPIPE. This can be prevented by some of\n> > SZEDZER's previous work, but it is silenced in 't' for now.\n>\n> s/SZEDZER/SZEDER/\n\nThank you for pointing this out, and my apologies to SZEDER.\n\n> > diff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh\n> > @@ -235,6 +235,29 @@ test_expect_success 'implicitly construct combine: filter with repeated flags' '\n> > +test_expect_success 'upload-pack fails banned object filters' '\n> > +       # Ensure that configuration keys are normalized by capitalizing\n> > +       # \"blob:None\" below:\n> > +       test_config -C srv.bare uploadpack.filter.blob:None.allow false &&\n>\n> I found the wording of the comment more confusing than clarifying.\n> Perhaps rewriting it like this could help:\n>\n>     Test case-insensitivity by intentional use of \"blob:None\" rather than\n>     \"blob:none\".\n>\n> or something.\n\nSure, your suggestion does clarify things. I'll apply it to my fork.\n\n> > +       test_must_fail ok=sigpipe git clone --no-checkout --filter.blob:none \\\n> > +               \"file://$(pwd)/srv.bare\" pc3\n> > +'\n\nThanks,\nTaylor\n"},{"id":"393421","messageId":"20200318212011.GD31397@syl.local","threadId":"52980","inReplyTo":"13dd0152-b20a-51e1-5940-5e4b67242e9b@iee.email","subject":"Re: [RFC PATCH 2/2] upload-pack.c: allow banning certain object filter(s)","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-18T21:20:11Z","receivedAt":"2020-03-18T21:20:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Hi Philip,\n\nOn Wed, Mar 18, 2020 at 11:18:16AM +0000, Philip Oakley wrote:\n> Hi\n> On 17/03/2020 20:39, Taylor Blau wrote:\n> > Git clients may ask the server for a partial set of objects, where the\n> > set of objects being requested is refined by one or more object filters.\n> > Server administrators can configure 'git upload-pack' to allow or ban\n> > these filters by setting the 'uploadpack.allowFilter' variable to\n> > 'true' or 'false', respectively.\n> >\n> > However, administrators using bitmaps may wish to allow certain kinds of\n> > object filters, but ban others. Specifically, they may wish to allow\n> > object filters that can be optimized by the use of bitmaps, while\n> > rejecting other object filters which aren't and represent a perceived\n> > performance degradation (as well as an increased load factor on the\n> > server).\n> >\n> > Allow configuring 'git upload-pack' to support object filters on a\n> > case-by-case basis by introducing a new configuration variable and\n> > section:\n> >\n> >   - 'uploadpack.filter.allow'\n> >\n> >   - 'uploadpack.filter.<kind>.allow'\n> >\n> > where '<kind>' may be one of 'blob:none', 'blob:limit', 'tree:depth',\n> > and so on. The additional '.' between 'filter' and '<kind>' is part of\n> > the sub-section.\n> >\n> > Setting the second configuration variable for any valid value of\n> > '<kind>' explicitly allows or disallows restricting that kind of object\n> > filter.\n> >\n> > If a client requests the object filter <kind> and the respective\n> > configuration value is not set, 'git upload-pack' will default to the\n> > value of 'uploadpack.filter.allow', which itself defaults to 'true' to\n> > maintain backwards compatibility. Note that this differs from\n> > 'uploadpack.allowfilter', which controls whether or not the 'filter'\n> > capability is advertised.\n> >\n> > NB: this introduces an unfortunate possibility that attempt to write the\n> > ERR sideband will cause a SIGPIPE. This can be prevented by some of\n> > SZEDZER's previous work, but it is silenced in 't' for now.\n> > ---\n> >  Documentation/config/uploadpack.txt | 12 ++++++\n> >  t/t5616-partial-clone.sh            | 23 ++++++++++\n> >  upload-pack.c                       | 67 +++++++++++++++++++++++++++++\n> >  3 files changed, 102 insertions(+)\n> >\n> > diff --git a/Documentation/config/uploadpack.txt b/Documentation/config/uploadpack.txt\n> > index ed1c835695..6213bd619c 100644\n> > --- a/Documentation/config/uploadpack.txt\n> > +++ b/Documentation/config/uploadpack.txt\n> > @@ -57,6 +57,18 @@ uploadpack.allowFilter::\n> >  \tIf this option is set, `upload-pack` will support partial\n> >  \tclone and partial fetch object filtering.\n> >\n> > +uploadpack.filter.allow::\n> > +\tProvides a default value for unspecified object filters (see: the\n> > +\tbelow configuration variable).\n> > +\tDefaults to `true`.\n> > +\n> > +uploadpack.filter.<filter>.allow::\n> > +\tExplicitly allow or ban the object filter corresponding to `<filter>`,\n> > +\twhere `<filter>` may be one of: `blob:none`, `blob:limit`, `tree:depth`,\n> > +\t`sparse:oid`, or `combine`. If using combined filters, both `combine`\n> > +\tand all of the nested filter kinds must be allowed.\n>\n> Doesn't the man page at least need the part from the commit message \"The\n> additional '.' between 'filter' and '<kind>' is part of\n> the sub-section.\" as it's not a common mechanism (other comments not\n> withstanding)\n\nThanks, you're certainly right. I wrote the man pages back when the\nconfiguration was spelled:\n\n  $ git config uploadpack.filter=blob:none.allow true\n\nBut now that there is the extra '.', it's worth calling out here, too.\nI'll make sure that this is addressed based on the outcome of the\ndiscussion below when these patches hit non-RFC status.\n\n> Philip\n\nThanks,\nTaylor\n"},{"id":"393422","messageId":"20200318212818.GE31397@syl.local","threadId":"52980","inReplyTo":"20200318101825.GB1227946@coredump.intra.peff.net","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-03-18T21:28:18Z","receivedAt":"2020-03-18T21:28:26Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Wed, Mar 18, 2020 at 06:18:25AM -0400, Jeff King wrote:\n> On Tue, Mar 17, 2020 at 02:39:05PM -0600, Taylor Blau wrote:\n>\n> > Hi Christian,\n> >\n> > Of course, I would be happy to send along our patches. They are included\n> > in the series below, and correspond roughly to what we are running at\n> > GitHub. (For us, there have been a few more clean-ups and additional\n> > patches, but I squashed them into 2/2 below).\n> >\n> > The approach is roughly that we have:\n> >\n> >   - 'uploadpack.filter.allow' -> specifying the default for unspecified\n> >     filter choices, itself defaulting to true in order to maintain\n> >     backwards compatibility, and\n> >\n> >   - 'uploadpack.filter.<filter>.allow' -> specifying whether or not each\n> >     filter kind is allowed or not. (Originally this was given as 'git\n> >     config uploadpack.filter=blob:none.allow true', but this '=' is\n> >     ambiguous to configuration given over '-c', which itself uses an '='\n> >     to separate keys from values.)\n>\n> One thing that's a little ugly here is the embedded dot in the\n> subsection (i.e., \"filter.<filter>\"). It makes it look like a four-level\n> key, but really there is no such thing in Git.  But everything else we\n> tried was even uglier.\n>\n> I think we want to declare a real subsection for each filter and not\n> just \"uploadpack.filter.<filter>\". That gives us room to expand to other\n> config options besides \"allow\" later on if we need to.\n>\n> We don't want to claim \"uploadpack.allow\" and \"uploadpack.<filter>.allow\";\n> that's too generic.\n>\n> Likewise \"filter.allow\" is too generic.\n\nI wonder. A multi-valued 'uploadpack.filter.allow' *might* solve some\nproblems, but the more I turn it over in my head, the more that I think\nthat it's creating more headaches for us than it's removing.\n\nOn the pro's side, is that we could have this be a multi-valued key\nwhere each value is the name of an allowed filter. I guess that would\nsolve the subsection-naming problem, but it is admittedly generic, not\nto mention the fact that we already *use* this key to specify a default\nvalue for missing 'uploadpack.filter.<filter>.allow' values. For that\nreason, it seems like a non-starter to me.\n\n> We could do \"uploadpackfilter.allow\" and \"uploadpackfilter.<filter>.allow\",\n> but that's both ugly _and_ separates these options from the rest of\n> uploadpack.*.\n>\n> We could use a character besides \".\", which would reduce confusion. But\n> what? Using colon is kind of ugly, because it's already syntactically\n> significant in filter names, and you get:\n>\n>   uploadpack.filter:blob:none.allow\n>\n> We tried equals, like:\n>\n>   uploadpack.filter=blob:none.allow\n>\n> but there's an interesting side effect. Doing:\n>\n>   git -c uploadpack.filter=blob:none.allow=true upload-pack ...\n>\n> doesn't work, because the \"-c\" parser ends the key at the first \"=\". As\n> it should, because otherwise we'd get confused by an \"=\" in a value.\n> This is a failing of the \"-c\" syntax; it can't represent values with\n> \"=\". Fixing it would be awkward, and I've never seen it come up in\n> practice outside of this (you _could_ have a branch with a funny name\n> and try to do \"git -c branch.my=funny=branch.remote=origin\" or\n> something, but the lack of bug reports suggests nobody is that\n> masochistic).\n\nThanks for adding some more detail to this decision.\n\nAnother thing we could do is just simply use a different character. It\nmay be a little odd, but it keeps the filter-related variables in their\nown sub-section, allowing us to add more configuration sub-variables in\nthe future. I guess that calling it something like:\n\n  $ git config uploadpack.filter@blob:none.allow <true|false>\n\nis a little strange (i.e., why '@' over '#'? There's certainly no\nprecedent here that I can think of...), but maybe it is slightly\nless-weird than a pseudo-four-level key.\n\n> So...maybe the extra dot is the last bad thing?\n>\n> > I noted in the second patch that there is the unfortunate possibility of\n> > encountering a SIGPIPE when trying to write the ERR sideband back to a\n> > client who requested a non-supported filter. Peff and I have had some\n> > discussion off-list about resurrecting SZEDZER's work which makes room\n> > in the buffer by reading one packet back from the client when the server\n> > encounters a SIGPIPE. It is for this reason that I am marking the series\n> > as 'RFC'.\n>\n> For reference, the patch I was thinking of was this:\n>\n>   https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n\nThanks.\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"393423","messageId":"CAPig+cTWb55K70v1MahHbTi12F5Zi6stKc1vjY2=9jSvEm7jww@mail.gmail.com","threadId":"52980","inReplyTo":"20200318100327.GA1227946@coredump.intra.peff.net","subject":"Re: [RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2020-03-18T22:38:49Z","receivedAt":"2020-03-18T22:39:07Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Mar 18, 2020 at 6:03 AM Jeff King <peff@peff.net> wrote:\n> On Tue, Mar 17, 2020 at 04:53:44PM -0400, Eric Sunshine wrote:\n> > > +       case LOFC__COUNT:\n> > > +               /*\n> > > +                * Include these to catch all enumerated values, but\n> > > +                * break to treat them as a bug. Any new values of this\n> > > +                * enum will cause a compiler error, as desired.\n> > > +                */\n> >\n> > In general, people will see a warning, not an error, unless they\n> > specifically use -Werror (or such) to turn the warning into an error,\n> > so this statement is misleading. Also, while some compilers may\n> > complain, others may not. So, although the comment claims that we will\n> > notice an unhandled enum constant at compile-time, that isn't\n> > necessarily the case.\n>\n> Yes, but that's the best we can do, isn't it?\n\nTo be clear, I wasn't questioning the code structure at all. I was\nspecifically referring to the comment talking about \"error\" when it\nshould say \"warning\" or \"possible warning\".\n\nMoreover, normally, we use comments to highlight something in the code\nwhich is not obvious or straightforward, so I was questioning whether\nthis comment is even helpful since the code seems reasonably clear.\nAnd...\n\n> But we can't just remove the default case. Even though enums don't\n> generally take on other values, it's legal for them to do so. So we do\n> want to make sure we BUG() in that instance.\n>\n> This is awkward to solve in the general case[1]. But because we're\n> returning in each case arm here, it's easy to just put the BUG() after\n> the switch. Anything that didn't return is unhandled, and we get the\n> best of both: -Wswitch warnings when we need to add a new filter type,\n> and a BUG() in the off chance that we see an unexpected value.\n>\n> So...I dunno. Worth it as a general technique?\n\n...if this is or will become an idiom we want in this codebase, then\nit would be silly to write an explanatory comment every place we\nemploy it. Instead, a document such as CodingGuidelines would likely\nbe a better fit for such knowledge.\n"},{"id":"393424","messageId":"xmqq4kulffps.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200318212818.GE31397@syl.local","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-03-18T22:41:51Z","receivedAt":"2020-03-18T22:41:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Taylor Blau <me@ttaylorr.com> writes:\n\n>> We tried equals, like:\n>>\n>>   uploadpack.filter=blob:none.allow\n>>\n>> but there's an interesting side effect. Doing:\n>>\n>>   git -c uploadpack.filter=blob:none.allow=true upload-pack ...\n>>\n>> doesn't work, because the \"-c\" parser ends the key at the first \"=\". As\n>> it should, because otherwise we'd get confused by an \"=\" in a value.\n>> This is a failing of the \"-c\" syntax; it can't represent values with\n>> \"=\". \n\ns/value/key/ I presume ;-)\n"},{"id":"393474","messageId":"20200319170350.GA4075823@coredump.intra.peff.net","threadId":"52980","inReplyTo":"xmqqtv2lfrk7.fsf_-_@gitster.c.googlers.com","subject":"Re: Re*: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-19T17:03:50Z","receivedAt":"2020-03-19T17:03:53Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 18, 2020 at 11:26:00AM -0700, Junio C Hamano wrote:\n\n> > One thing that's a little ugly here is the embedded dot in the\n> > subsection (i.e., \"filter.<filter>\"). It makes it look like a four-level\n> > key, but really there is no such thing in Git.  But everything else we\n> > tried was even uglier.\n> \n> I think this gives us the best arrangement by upfront forcing all\n> the configuration handers for \"<subcommand>.*.<token>\" namespace,\n> current and future, to use \"<group-prefix>\" before the unbounded set\n> of user-specifiable values that affects the <subcommand> (which is\n> \"uploadpack\").\n> \n> So far, the configuration variables that needs to be grouped by\n> unbounded set of user-specifiable values we supported happened to\n> have only one sensible such set for each <subcommand>, so we could\n> get away without such <group-prefix> and it was perfectly OK to\n> have, say \"guitool.<name>.cmd\".\n\nYeah. We have often just split those out into a separate hierarchy from\n<subcommand> E.g., tar.<format>.command, which is really feeding the\ngit-archive command. We could do that here, too, but I wasn't sure of a\ngood name (this really is upload-pack specific, though I guess in theory\nother commands could grow a need to look at or restrict \"remote object\nfilters\").\n\n> Syntactically, the convention to always end such <group-prefix> with\n> a dot \".\" may look unusual, or once readers' eyes get used to them,\n> may look natural.  One tiny sad thing about it is that it cannot be\n> mechanically enforced, but that is minor.\n\nThe biggest downside to implying a 4-level key is that the\ncase-sensitivity rules may be different. I.e., you can say:\n\n  UploadPack.filter.blob:none.Allow\n\nbut not:\n\n  UploadPack.Filter.blob:none.Allow\n\nSince \"filter\" is part of the subsection, it's case sensitive. We could\nmatch it case-insensitively in upload_pack_config(), but it would crop\nup in other laces (e.g., \"git config --unset\" would still care).\n\n> > We could do \"uploadpackfilter.allow\" and \"uploadpackfilter.<filter>.allow\",\n> > but that's both ugly _and_ separates these options from the rest of\n> > uploadpack.*.\n> \n> There is an existing instance of a configuration that affects\n> <subcommand> that uses a different word after <subcommand>, which is\n> credentialCache.ignoreSIGHUP, and I tend to agree that it is ugly.\n\nI don't think that's what's going on here. It affects only the\ncredential-cache subcommand, but we avoid hyphens in our key names.\nSo it really is the subcommand; it's just that the name is a superset of\nanother command name. :)\n\n> By the way, I noticed the following while I was studying the current\n> practice, so before I forget...\n> \n> -- >8 --\n> Subject: [PATCH] separate tar.* config to its own source file\n> \n> Even though there is only one configuration variable in the\n> namespace, it is not quite right to have tar.umask described\n> among the variables for tag.* namespace.\n\nYeah, this is definitely an improvement. But I was surprised that\ntar.<format>.* wasn't covered here. It is documented in git-archive.\nProbably worth moving or duplicating it in git-config.\n\n-Peff\n"},{"id":"393475","messageId":"20200319170954.GB4075823@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200318212818.GE31397@syl.local","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-19T17:09:54Z","receivedAt":"2020-03-19T17:09:56Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 18, 2020 at 03:28:18PM -0600, Taylor Blau wrote:\n\n> I wonder. A multi-valued 'uploadpack.filter.allow' *might* solve some\n> problems, but the more I turn it over in my head, the more that I think\n> that it's creating more headaches for us than it's removing.\n\nIMHO we should avoid multi-valued keys when there's not a compelling\nreason. There are a lot of corner cases they introduce (e.g., there's no\nstandard way to override them rather than adding to the list).\n\n> Another thing we could do is just simply use a different character. It\n> may be a little odd, but it keeps the filter-related variables in their\n> own sub-section, allowing us to add more configuration sub-variables in\n> the future. I guess that calling it something like:\n> \n>   $ git config uploadpack.filter@blob:none.allow <true|false>\n> \n> is a little strange (i.e., why '@' over '#'? There's certainly no\n> precedent here that I can think of...), but maybe it is slightly\n> less-weird than a pseudo-four-level key.\n\nI guess it's subjective, but the \"@\" just feels odd because it's\nassociated with so many other meanings. Likewise \"#\".\n\n-Peff\n"},{"id":"393476","messageId":"20200319171020.GC4075823@coredump.intra.peff.net","threadId":"52980","inReplyTo":"xmqq4kulffps.fsf@gitster.c.googlers.com","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-19T17:10:20Z","receivedAt":"2020-03-19T17:10:23Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 18, 2020 at 03:41:51PM -0700, Junio C Hamano wrote:\n\n> Taylor Blau <me@ttaylorr.com> writes:\n> \n> >> We tried equals, like:\n> >>\n> >>   uploadpack.filter=blob:none.allow\n> >>\n> >> but there's an interesting side effect. Doing:\n> >>\n> >>   git -c uploadpack.filter=blob:none.allow=true upload-pack ...\n> >>\n> >> doesn't work, because the \"-c\" parser ends the key at the first \"=\". As\n> >> it should, because otherwise we'd get confused by an \"=\" in a value.\n> >> This is a failing of the \"-c\" syntax; it can't represent values with\n> >> \"=\". \n> \n> s/value/key/ I presume ;-)\n\nYes. :)\n\n-Peff\n"},{"id":"393477","messageId":"20200319171502.GD4075823@coredump.intra.peff.net","threadId":"52980","inReplyTo":"CAPig+cTWb55K70v1MahHbTi12F5Zi6stKc1vjY2=9jSvEm7jww@mail.gmail.com","subject":"Re: [RFC PATCH 1/2] list_objects_filter_options: introduce 'list_object_filter_config_name'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-03-19T17:15:02Z","receivedAt":"2020-03-19T17:15:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 18, 2020 at 06:38:49PM -0400, Eric Sunshine wrote:\n\n> To be clear, I wasn't questioning the code structure at all. I was\n> specifically referring to the comment talking about \"error\" when it\n> should say \"warning\" or \"possible warning\".\n> \n> Moreover, normally, we use comments to highlight something in the code\n> which is not obvious or straightforward, so I was questioning whether\n> this comment is even helpful since the code seems reasonably clear.\n> And...\n\nOK, I agree with all that. :)\n\n> ...if this is or will become an idiom we want in this codebase, then\n> it would be silly to write an explanatory comment every place we\n> employ it. Instead, a document such as CodingGuidelines would likely\n> be a better fit for such knowledge.\n\nYeah, that makes sense. If we do use this technique, though, we'll have\nto explicitly list \"case\" lines for the enum values which are meant to\nbreak out to the BUG(). And there it _is_ worth commenting on \"yes, we\nknow about this value but it is not handled here because...\". Which is\nwhat you asked for in your original message. :)\n\nSomething like:\n\n  switch (c) {\n  case LOFC_BLOB_NONE:\n\treturn \"blob:none\":\n  ..etc...\n  case LOFC__COUNT:\n\t/* not a real filter type; just a marker for counting the number */\n\tbreak;\n  case LOFC_DISABLED:\n\t/* we have no name for \"no filter at all\" */\n\tbreak;\n  }\n  BUG(...);\n\n-Peff\n"},{"id":"394124","messageId":"20200326222720.bkqkhalof6s2srfq@doriath","threadId":"52980","inReplyTo":"3839451584363302@sas2-a098efd00d24.qloud-c.yandex.net","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Damien Robert","fromEmail":"damien.olivier.robert@gmail.com","sentAt":"2020-03-26T22:27:20Z","receivedAt":"2020-03-26T22:27:27Z","isPatch":false,"sender":{"key":"damien.olivier.robert@gmail.com","avatar":null},"body":"From Konstantin Tokarev, Mon 16 Mar 2020 at 15:55:39 (+0300) :\n> > My situation: coworkers push big files by mistake, I don't want to rewrite\n> > history because they are not too well versed with git, but I want to keep\n> > *my* repo clean.\n\n> Wouldn't it be better to prevent *them* from such mistakes, e.g. by using\n> pre-push review system like Gerrit?\n\nSo my coworkers are mathematicians, and not all of them are comfortable\nwith dvcs, and I already have a hard time convincing them to use git rather\nthan dropbox. I take it upon myself to make it as easy as possible to use\ngit (by telling them to push to a different branch when there is a conflict\nso that I can handle the conflict myself).\n\nI don't think Gerrit is a solution there...\n"},{"id":"394125","messageId":"20200326223056.csfa4vbuir5lab5b@doriath","threadId":"52980","inReplyTo":"CABPp-BEnYTvakuP9nBi3Q_-mP3i7BJEvKofC3_4N8cO9JkF22Q@mail.gmail.com","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Damien Robert","fromEmail":"damien.olivier.robert@gmail.com","sentAt":"2020-03-26T22:30:56Z","receivedAt":"2020-03-26T22:31:03Z","isPatch":false,"sender":{"key":"damien.olivier.robert@gmail.com","avatar":null},"body":"From Elijah Newren, Mon 16 Mar 2020 at 09:32:45 (-0700) :\n> > I am interested in more details on how to handle this using replace.\n\n> This comment at the conference was in reference to how people rewrite\n> history to remove the big blobs, but then run into issues because\n> there are many places outside of git that reference old commit IDs\n> (wiki pages, old emails, issues/tickets, etc.) that are now broken.\n[...]\n\nInteresting, thanks for the context!\n\n> As for using replace refs to attempt to alleviate problems without\n> rewriting history, that's an even bigger can of worms and it doesn't\n> solve clone/fetch/gc/fsck nor the many other places you highlighted in\n> your email.\n\nI agreed, but one part that makes it easier in my context is that I don't\nneed to distribute the replaced references, I just need them for myself.\nThis alleviate a lot of problems already, and as I outlined in my email the\ncombination of replace ref and sparse checkout is almost enough.\n\n-- \nDamien Robert\nhttp://www.normalesup.org/~robert/pro\n"},{"id":"394127","messageId":"20200326223749.ej4a3z7n37s5lw6r@doriath","threadId":"52980","inReplyTo":"87lfo0881d.fsf@vps.thesusis.net","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"Damien Robert","fromEmail":"damien.olivier.robert@gmail.com","sentAt":"2020-03-26T22:37:49Z","receivedAt":"2020-03-26T22:37:54Z","isPatch":false,"sender":{"key":"damien.olivier.robert@gmail.com","avatar":null},"body":"From Phillip Susi, Mon 16 Mar 2020 at 14:32:46 (-0400) :\n> Instead of replacing the blob with an empty file, why not replace the\n> tree that references it with one that does not?  That way you won't have\n> the file in your checkout at all, and the index won't list it so status\n> won't show it as changed.\n\nThat's an interesting solution, but it only works if the tree itself does\nnot change.\n\n- This is the case when these large objects were uploaded and then removed\n  in another commit [*].\n  But in this case I won't checkout (usually) back to this tree anyway, so the\n  error due to the missing blob is not a big problem.\n\n- When my coauthors use git as a dropbox alternative where they upload big\n  pdf files (rather than only source code or .tex files), they also want to\n  keep them there. If they had uploaded these files to the special literature/\n  folder I had made for them, I could just replace the literature/ tree by\n  an empty one, but they managed to upload them in the root folder which is\n  subject to change unfortunately, and it would be annoying to make a new\n  replace ref each time.\n\n[*] by myself usually, for instance when people commit spurious tex\ngenerated files like eg *.synctex, despite my .gitignore (I don't know how\nthey manage this...)\n\n-- \nDamien Robert\nhttp://www.normalesup.org/~robert/pro\n"},{"id":"395023","messageId":"20200407230132.GD137962@google.com","threadId":"52980","inReplyTo":"xmqqeetwcf4k.fsf@gitster.c.googlers.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-07T23:01:32Z","receivedAt":"2020-04-07T23:01:39Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Mar 13, 2020 at 10:56:59AM -0700, Junio C Hamano wrote:\n> Emily Shaffer <emilyshaffer@google.com> writes:\n\nPhew - now that git-bugreport looks to be on the path to 'next' I get to\nwork on hooks again :) Forgive me for the late reply.\n\n> \n> > This means that we could do something like this:\n> >\n> > [hook \"/path/to/executable.sh\"]\n> > \tevent = pre-commit\n> > \torder = 123\n> > \tmustSucceed = false\n> > \tparallelizable = true\n> >\n> > etc, etc as needed.\n> \n> You can do\n> \n>     [hook \"pre-commit\"]\n> \torder = 123\n> \tpath = \"/path/to/executable.sh\"\n> \n>     [hook \"pre-commit\"]\n> \torder = 234\n> \tpath = \"/path/to/another-executable.sh\"\n> \n> as well, and using the second level for what hook the (sub)section\n> is about, instead of \"we have this path that is used for a hook.\n> What hook is it?\", feels (at least to me) more natural.\n\nYeah, I see what you mean, and it's true I misread the notes and\nmisremembered Peff's suggestion. I was reworking my RFC patch some\ntoday, and noticed that the following two configs:\n\nA.gitconfig:\n  [hook \"pre-commit\"]\n    command = \"/path/to/executable.sh\"\n    option = foo\n  [hook \"pre-commit\"]\n    command = \"/path/to/another-executable.sh\"\n\nB.gitconfig:\n  [hook \"pre-commit\"]\n    command = \"/path/to/executable.sh\"\n  [hook \"pre-commit\"]\n    option = foo\n    command = \"/path/to/another-executable.sh\"\n\nare indistinguishable during the config parse - both show up during the\nconfig callback looking like:\n\n  value = \"hook.pre-commit.command\"; var = \"/path/to/executable.sh\"\n  value = \"hook.pre-commit.option\"; var = \"foo\"\n  value = \"hook.pre-commit.command\"; var = \"/path/to/another-executable.sh\"\n\nI didn't see anything to get around this in the config parser library;\nif I missed it I'd love to know.\n\nUsing the hook path as the subsection still doesn't help, of course. I\nthink the only way I see around it is to require a specific value at the\nbeginning of each hook config section, e.g. \"each hook entry must begin\nwith 'command'\"; that means that the config parser callback can look\nsomething like:\n\n  parse section, subsection, key\n  if section.subsection = \"hook.pre-commit\":\n    if key = \"command\":\n      add a new hook to the hook list\n    else:\n      operate on the tail of the hook list\n\nThe price of this is poor user experience for those handcrafting their\nown hook configs, but I don't think it's poorer than carefully spelling\nout \"123:~/my-hook-path.sh:whatever:other:options\" or something. I'll\nadd that I had planned to teach 'git-hook' to write and modify config\nfiles for the user with an interactive-rebase-like UI, so a brittle\nconfig layout might not be the end of the world.\n\nOr, I suppose, we could teach the config parser how to understand\n\"structlike\" configs like this where repeated header entries need to be\ncollated together. That seems to be contrary to the semantics of the\nconfig file right now, though, and it looks like it'd require a rework\nof the config_set implementation: today config_set_element looks like\n\n  struct config_set_element {\n          struct hashmap_entry ent;\n          char *key; /* \"hook.pre-commit.command\" */\n          struct string_list value_list; /* \"/path/to/executable.sh\"\n\t                                  * \"path/to/another-executable.sh\"\n\t\t\t\t\t  */\n  };\n\nI'm not very keen on the idea of changing the way configs are stored for\neveryone, although if folks are unsatisfied with the way it is now and\nwant to do that, I guess it's an option. But it's certainly more\noverhead than my earlier suggestion.\n\nThoughts?\n\n - Emily\n"},{"id":"395028","messageId":"20200407235116.GE137962@google.com","threadId":"52980","inReplyTo":"20200407230132.GD137962@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-07T23:51:16Z","receivedAt":"2020-04-07T23:51:27Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Tue, Apr 07, 2020 at 04:01:32PM -0700, Emily Shaffer wrote:\n> Thoughts?\n\nJonathan Nieder and I discussed this a little bit offline, and he\nsuggested another thought:\n\n[hook \"unique-name\"]\n  pre-commit = ~/path-to-hook.sh args-for-precommit\n  pre-push = ~/path-to-hook.sh\n  order = 001\n\nThen, in another config:\n\nhook.unique-name.pre-push-order = 123\n\nor,\n\nhook.unique-name.enable = false\nhook.unique-name.pre-commit-enable = true\n\nTo pick it apart a little more:\n\n - Let's give each logical action a unique name, e.g. \"git-secrets\".\n - Users can sign up for a certain event by providing the command to\n   run, e.g. `hook.git-secrets.pre-commit = git-secrets pre-commit`.\n - Users can set up defaults for the logical action, e.g.\n   `hook.git-secrets.before = gerrit` (where \"gerrit\" is the unique name\n   for another logical action), and then change it on a per-hook basis\n   e.g. `hook.git-secrets.pre-commit-before = clang-tidy`\n\nThere's some benefit:\n - We don't have to kludge something new (multiple sections with the\n   same name, but logically disparate) into the config semantics where\n   it doesn't really fit.\n - Users could, for example, turn off all \"git-secrets\" invocations in a\n   repo without knowing which hooks it's attached to, e.g.\n   `hook.git-secrets.enable = false`\n - We still have the option to add and remove parameters like 'order' or\n   'before'/'after' or 'parallelizable' or etc., on a per-hook basis or\n   for all flavors of a logical action such as \"git-secrets\"\n - It may be easier for a config-authoring iteration of 'git-hook' to\n   modify existing configs than it would be if the ordering of config\n   entries is vital.\n\nOne drawback I can think of is that these unique names could be either\ndifficult to autogenerate and guarantee uniqueness, or difficult for\nhumans to parse. I'd have to rethink the UI for writing or editing with\ngit-hook (rather than editing the config by hand), although I think with\nthe mood shifting away from configs looking like\n\"hook.pre-commit=123:~/path-to-thing.sh\" my UI mockups are all invalid\nanyways :)\n\nWe also considered something like:\n\n[hook \"git-secrets-pre-commit\"]\n  command = ~/path-to-hooks.sh args-for-precommit\n  order = 001\n\n[hook \"git-secrets-pre-push\"]\n  comand = ~/path-to-hook.sh\n  order = 123\n\nbut concluded that it's more verbose without adding additional value\nover the earlier proposal above. This syntax can achieve a subset of\ngoals but is missing extra value like an easy path to disable all hooks\nin that logical action, or nice defaults when you don't expect the order\nor parallelism to change.\n\nDefinitely interested in hearing more ideas :)\n\n - Emily\n"},{"id":"395033","messageId":"xmqq4ktures1.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200407235116.GE137962@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-04-08T00:40:14Z","receivedAt":"2020-04-08T00:40:24Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n> [hook \"unique-name\"]\n>   pre-commit = ~/path-to-hook.sh args-for-precommit\n>   pre-push = ~/path-to-hook.sh\n>   order = 001\n>\n> Then, in another config:\n>\n> hook.unique-name.pre-push-order = 123\n>\n> or,\n>\n> hook.unique-name.enable = false\n> hook.unique-name.pre-commit-enable = true\n>\n> To pick it apart a little more:\n>\n>  - Let's give each logical action a unique name, e.g. \"git-secrets\".\n>  - Users can sign up for a certain event by providing the command to\n>    run, e.g. `hook.git-secrets.pre-commit = git-secrets pre-commit`.\n>  - Users can set up defaults for the logical action, e.g.\n>    `hook.git-secrets.before = gerrit` (where \"gerrit\" is the unique name\n>    for another logical action), and then change it on a per-hook basis\n>    e.g. `hook.git-secrets.pre-commit-before = clang-tidy`\n\nSorry, but the description and the tokens used in there are so\ndetached from the current reality that I am having a hard time\ntrying to even guess what you two were talking about.  \n\nFor example, how would I express that I am using program X as my\n'push-to-checkout' hook in a way consistent with the above\ndescription?  Would \"push\" correspond to your \"git-secrets\" and\n\"checkout\" to your \"pre-commit\", or would these be placed where you\nwrote \"unique-name\"?\n\n\n\n"},{"id":"395034","messageId":"20200408010904.GF137962@google.com","threadId":"52980","inReplyTo":"xmqq4ktures1.fsf@gitster.c.googlers.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-08T01:09:04Z","receivedAt":"2020-04-08T01:09:15Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Tue, Apr 07, 2020 at 05:40:14PM -0700, Junio C Hamano wrote:\n> Emily Shaffer <emilyshaffer@google.com> writes:\n> \n> > [hook \"unique-name\"]\n> >   pre-commit = ~/path-to-hook.sh args-for-precommit\n> >   pre-push = ~/path-to-hook.sh\n> >   order = 001\n> >\n> > Then, in another config:\n> >\n> > hook.unique-name.pre-push-order = 123\n> >\n> > or,\n> >\n> > hook.unique-name.enable = false\n> > hook.unique-name.pre-commit-enable = true\n> >\n> > To pick it apart a little more:\n> >\n> >  - Let's give each logical action a unique name, e.g. \"git-secrets\".\n> >  - Users can sign up for a certain event by providing the command to\n> >    run, e.g. `hook.git-secrets.pre-commit = git-secrets pre-commit`.\n> >  - Users can set up defaults for the logical action, e.g.\n> >    `hook.git-secrets.before = gerrit` (where \"gerrit\" is the unique name\n> >    for another logical action), and then change it on a per-hook basis\n> >    e.g. `hook.git-secrets.pre-commit-before = clang-tidy`\n> \n> Sorry, but the description and the tokens used in there are so\n> detached from the current reality that I am having a hard time\n> trying to even guess what you two were talking about.  \n\nAck, sorry about that. Point taken.\n\n> \n> For example, how would I express that I am using program X as my\n> 'push-to-checkout' hook in a way consistent with the above\n> description?  Would \"push\" correspond to your \"git-secrets\" and\n> \"checkout\" to your \"pre-commit\", or would these be placed where you\n> wrote \"unique-name\"?\n\nIf you are using program X, which lives at /bin/x, and you want to use\nit as your push-to-checkout hook:\n\n[hook \"x\"]\n  push-to-checkout = /bin/x\n\n\"unique-name\" is unique and arbitrary, which is why I mentioned it could\nbe either difficult to machine-generate or difficult to human-read.\n\n - Emily\n"},{"id":"395246","messageId":"20200410213146.GA2075494@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200407235116.GE137962@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-10T21:31:46Z","receivedAt":"2020-04-10T21:31:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 07, 2020 at 04:51:16PM -0700, Emily Shaffer wrote:\n\n> On Tue, Apr 07, 2020 at 04:01:32PM -0700, Emily Shaffer wrote:\n> > Thoughts?\n> \n> Jonathan Nieder and I discussed this a little bit offline, and he\n> suggested another thought:\n> \n> [hook \"unique-name\"]\n>   pre-commit = ~/path-to-hook.sh args-for-precommit\n>   pre-push = ~/path-to-hook.sh\n>   order = 001\n\nYeah, giving each block a unique name lets you give them each an order.\nIt seems kind of weird to me that you'd define multiple hook types for a\ngiven name. And it doesn't leave a lot of room for defining\nper-hook-type options; you have to make new keys like pre-push-order\n(though that does work because the hook names are a finite set that\nconforms to our config key names).\n\nWhat if we added a layer of indirection: have a section for each type of\nhook, defining keys for that type. And then for each hook command we\ndefine there, it can have its own section, too. Maybe better explained\nwith an example:\n\n  [hook \"pre-receive\"]\n  # put any pre-receive related options here; e.g., a rule for what to\n  # do with hook exit codes (e.g., stop running, run all but return exit\n  # code, ignore failures, etc)\n  fail = stop\n\n  # And we can define actual hook commands. This one refers to the\n  # hookcmd block below.\n  command = foo\n\n  # But if there's no such hookcmd block, we could just do something\n  # sensible, like defaulting hookcmd.X.command to \"X\"\n  command = /path/to/some-hook.sh\n\n  [hookcmd \"foo\"]\n  # the actual hook command to run\n  command = /path/to/another-hook\n  # other hook options, like order priority\n  order = 123\n\nI think both this schema and the one you wrote above can express the\nsame set of things. But you don't _have_ to pick a unique name if you\ndon't want to. Just doing:\n\n  [hook \"pre-receive\"]\n  command = /some/script\n\nwould be valid and useful (and that's as far as 99% of use cases would\nneed to go).\n\n-Peff\n"},{"id":"395369","messageId":"20200413191515.GA5478@google.com","threadId":"52980","inReplyTo":"20200410213146.GA2075494@coredump.intra.peff.net","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-13T19:15:15Z","receivedAt":"2020-04-13T19:15:24Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Apr 10, 2020 at 05:31:46PM -0400, Jeff King wrote:\n> On Tue, Apr 07, 2020 at 04:51:16PM -0700, Emily Shaffer wrote:\n> \n> > On Tue, Apr 07, 2020 at 04:01:32PM -0700, Emily Shaffer wrote:\n> > > Thoughts?\n> > \n> > Jonathan Nieder and I discussed this a little bit offline, and he\n> > suggested another thought:\n> > \n> > [hook \"unique-name\"]\n> >   pre-commit = ~/path-to-hook.sh args-for-precommit\n> >   pre-push = ~/path-to-hook.sh\n> >   order = 001\n> \n> Yeah, giving each block a unique name lets you give them each an order.\n> It seems kind of weird to me that you'd define multiple hook types for a\n> given name.\n\nNot so odd - git-secrets configures itself for pre-commit,\nprepare-commit-msg, and commit-msg-hook. The invocation is slightly\ndifferent ('git-secrets pre-commit', 'git-secrets prepare-commit-msg',\netc) but to me it still makes some sense to treat it as a single logical\nunit.\n\n> And it doesn't leave a lot of room for defining\n> per-hook-type options; you have to make new keys like pre-push-order\n> (though that does work because the hook names are a finite set that\n> conforms to our config key names).\n\nOh, interesting. I think you're saying \"what if option 'frotz' only\nmakes sense for prepare-commit-msg; then there's no reason to allow\n'frotz' and 'prepare-commit-msg-frotz' and 'post-commit-frotz' and so\non? I think I didn't do a great job explaining myself in that mail, but\nmy idea was to let an unqualified option name in a hook block set the\ndefault, and then allow it to be overridden by qualifying it with the\nname of the hook in question:\n\n[hook \"unique-name\"]\n  option = \"some default\"\n  post-commit-option = \"post-commit specific version\"\n  pre-push = ~/foo.sh pre-push\n  post-commit = ~/foo.sh post-commit\n\nThen when post-commit is invoked, option = \"post-commit specific\nversion\"; when pre-push is invoked, option = \"some default\". My\nintention was to generate the hook-specific option key on the fly during\nsetup.\n\n> \n> What if we added a layer of indirection: have a section for each type of\n> hook, defining keys for that type. And then for each hook command we\n> define there, it can have its own section, too. Maybe better explained\n> with an example:\n> \n>   [hook \"pre-receive\"]\n>   # put any pre-receive related options here; e.g., a rule for what to\n>   # do with hook exit codes (e.g., stop running, run all but return exit\n>   # code, ignore failures, etc)\n>   fail = stop\n\nInteresting - so this is a default for all pre-receive hooks, that I can\nset at whichever scope I wish.\n\n> \n>   # And we can define actual hook commands. This one refers to the\n>   # hookcmd block below.\n>   command = foo\n> \n>   # But if there's no such hookcmd block, we could just do something\n>   # sensible, like defaulting hookcmd.X.command to \"X\"\n>   command = /path/to/some-hook.sh\n\nI like this idea a lot!\n\n> \n>   [hookcmd \"foo\"]\n>   # the actual hook command to run\n>   command = /path/to/another-hook\n>   # other hook options, like order priority\n>   order = 123\n\nLooks familiar enough. Now I worry - what if I specify 'fail' here too?\n\nIt seems like I may be saying \"let's set a default per hookcmd\" and you\nmay be saying \"let's set a default per hook\". Maybe you're saying \"some\noptions are hook-specific and some options are command-specific.\" You\nmight be saying \"we shouldn't need to set multiple option values for a\nsingle command,\" and I think I disagree with that based on the\ngit-secrets value alone; if I'm getting ready to commit, I want\ngit-secrets to run last so it can look at changes other hooks made to my\ncommit, but if I'm getting ready to push, I want git-secrets to run\nfirst so I don't wait around for a test suite just to find that my\ncommit is invalid anyways. Although, I guess with your schema the former\nwould be in [hookcmd \"git-secrets-committing\"] and the latter\nwould be in [hookcmd \"git-secrets-pushing\"], so I can set the ordering\nhow I wish.\n\n(This might be an OK problem to punt on. I don't think there are any\noptions we have in mind just yet - even \"order\", we aren't sure whether\nto prefer config order or an explicit number. I think if we make no\ndecision on how to treat per-hook options today, it doesn't stop us from\ndeciding on some schema tomorrow. Once we do decide, then we put it in\ndocumentation and need to stick to it, but for now I think it's OK to\nleave it undefined. We might not even need it.)\n\nThis schema also means it's easy to reorder or remove hooks later on,\nwhich I like. A single line in my worktree config is clear:\n\n  hookcmd.git-secrets-committing.skip = true\n  hookcmd.git-secrets-pushing.order = 001\n> \n> I think both this schema and the one you wrote above can express the\n> same set of things. But you don't _have_ to pick a unique name if you\n> don't want to. Just doing:\n> \n>   [hook \"pre-receive\"]\n>   command = /some/script\n> \n> would be valid and useful (and that's as far as 99% of use cases would\n> need to go).\n\nYeah, I see what you mean, and again I really like that. That lets us\nrun multiples in config order easily:\n\n[hook \"pre-receive\"]\n  command = /some/script\n  command = /some/other-script\n  command = some-hookcmd-header\n\nIf we add a little repeated-name detection then we can also reorder\neasily this way if that's the direction we want for ordering:\n\n{global}\nhook.pre-receive.command = a.sh\nhook.pre-receive.command = b.sh\n\n{local}\nhook.pre-receive.command = c.sh\nhook.pre-receive.command = a.sh\n\nfor a final order of {b.sh, c.sh, a.sh}.\n\nVery nice, IMO.\n\nI wonder - I think even something like this would work:\n\n{global}\n[hook \"pre-receive\"]\n  command = no-hookcmd-entry.sh\n\n{local for repo \"zork\"}\n[hookcmd \"no-hookcmd-entry.sh\"]\n  skip = true\n\nFor most repos, now I simply invoke no-hookcmd-entry.sh on pre-receive,\nbut when I'm parsing the config in \"zork\", now I see a populated hookcmd\nentry, and when I look it up with the key I found in the global config,\nI see that it's supposed to be skipped.\n\nAlthough I might need to do something hacky if I have multiple hooks\npointing to the same simple invocation:\n\n{global}\n[hook \"pre-receive\"]\n  command = no-hookcmd-entry.sh\n\n[hook \"post-commit\"]\n  command = no-hookcmd-entry.sh\n\n{local}\nhookcmd.no-hookcmd-entry.sh.skip = true\n\n[hook \"pre-receive\"]\n  command = modified-no-hookcmd.sh\n\n[hookcmd \"modified-no-hookcmd-entry\"]\n  command = no-hookcmd-entry.sh\n\nThat is, I think this makes it kind of tricky to shut off one invocation\nfor only one hook.  Maybe it makes sense to honor something like:\n\nhookcmd.foo.skip-pre-receive = true\n\n?\n\nI wonder if I'm getting buried in the weeds of stuff we won't ever have\nto worry about ;)\n\n - Emily\n"},{"id":"395373","messageId":"20200413215256.GA18990@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200413191515.GA5478@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-13T21:52:56Z","receivedAt":"2020-04-13T21:53:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Apr 13, 2020 at 12:15:15PM -0700, Emily Shaffer wrote:\n\n> > Yeah, giving each block a unique name lets you give them each an order.\n> > It seems kind of weird to me that you'd define multiple hook types for a\n> > given name.\n> \n> Not so odd - git-secrets configures itself for pre-commit,\n> prepare-commit-msg, and commit-msg-hook. The invocation is slightly\n> different ('git-secrets pre-commit', 'git-secrets prepare-commit-msg',\n> etc) but to me it still makes some sense to treat it as a single logical\n> unit.\n\nYeah, I do see how that use case makes sense. I wonder how common it is\nversus having separate one-off hooks. And whether setting the order\npriority for all hooks at once is that useful (e.g., I can easily\nimagine a case where the pre-commit hook for program A must go before B,\nbut it's the other way around for another hook).\n\nI'm just speculating, but my instinct is that it's worth trying to make\nthe simple things as simple as possible, while still allowing the more\ncomplex things.\n\n> > And it doesn't leave a lot of room for defining\n> > per-hook-type options; you have to make new keys like pre-push-order\n> > (though that does work because the hook names are a finite set that\n> > conforms to our config key names).\n> \n> Oh, interesting. I think you're saying \"what if option 'frotz' only\n> makes sense for prepare-commit-msg; then there's no reason to allow\n> 'frotz' and 'prepare-commit-msg-frotz' and 'post-commit-frotz' and so\n> on?\n\nNo, what I meant was just that if we had a hook \"foo/bar\", then the\nnatural option to control its hook-specific order would be:\n\n  [hook \"whatever\"]\n  order-foo/bar = 123\n\nwhich isn't allowed (\"/\" is not valid in a key name). But that should be\nOK since we control the names of hooks and can decide not to make one\nwith an invalid character in it. We might also support hooks for\nthird-party programs (e.g., if a porcelain wrapper wanted to have its\nown \"pre-switch-branches\" hook or something), but it's not too much of\nan imposition to say that the hook name should be a valid config key.\n\n> I think I didn't do a great job explaining myself in that mail, but\n> my idea was to let an unqualified option name in a hook block set the\n> default, and then allow it to be overridden by qualifying it with the\n> name of the hook in question:\n> \n> [hook \"unique-name\"]\n>   option = \"some default\"\n>   post-commit-option = \"post-commit specific version\"\n>   pre-push = ~/foo.sh pre-push\n>   post-commit = ~/foo.sh post-commit\n\nYeah, that overriding system makes sense to me. Any other option \"frotz\"\nwould have the same constraint, though I don't think that's too big a\ndeal.\n\n> Then when post-commit is invoked, option = \"post-commit specific\n> version\"; when pre-push is invoked, option = \"some default\". My\n> intention was to generate the hook-specific option key on the fly during\n> setup.\n\nRight that makes sense to me.\n\n> >   [hook \"pre-receive\"]\n> >   # put any pre-receive related options here; e.g., a rule for what to\n> >   # do with hook exit codes (e.g., stop running, run all but return exit\n> >   # code, ignore failures, etc)\n> >   fail = stop\n> \n> Interesting - so this is a default for all pre-receive hooks, that I can\n> set at whichever scope I wish.\n\nYes. Though I had imagined \"fail\" as semantics for operating on the\nwhole list of \"pre-receive\" hooks, you could define it in a per-command\nway, too. I was thinking of it as \"this is the strategy when a command\nfails\". But you could also think of it as \"what to do when this\nparticular command fails\".\n\n> >   [hookcmd \"foo\"]\n> >   # the actual hook command to run\n> >   command = /path/to/another-hook\n> >   # other hook options, like order priority\n> >   order = 123\n> \n> Looks familiar enough. Now I worry - what if I specify 'fail' here too?\n\nIf there's a per-command version of \"fail\", then presumably it would\noverride any per-hook. I.e., I'd expect code to resolve this at\nrun-time, like:\n\n  struct hook *hook = get_hook(\"pre-receive\");\n  for (i = 0; i < hook->nr; i++) {\n          struct hookcmd *cmd = hook->cmds[i];\n\n          if (run_hook(cmd->prog) != 0) {\n                  enum failure_strategy f = cmd->failure_strategy;\n                  if (f == FAILURE_STRATEGY_UNSET)\n                          f = hook->failure_strategy;\n                  switch (f) {\n                  ...do whatever...\n                  }\n          }\n  }\n\n> It seems like I may be saying \"let's set a default per hookcmd\" and you\n> may be saying \"let's set a default per hook\". Maybe you're saying \"some\n> options are hook-specific and some options are command-specific.\"\n\nYeah, the latter. Or it might even be that an option is sometimes\nhook-specific and sometimes command-specific.\n\n> You\n> might be saying \"we shouldn't need to set multiple option values for a\n> single command,\" and I think I disagree with that based on the\n> git-secrets value alone; if I'm getting ready to commit, I want\n> git-secrets to run last so it can look at changes other hooks made to my\n> commit, but if I'm getting ready to push, I want git-secrets to run\n> first so I don't wait around for a test suite just to find that my\n> commit is invalid anyways. Although, I guess with your schema the former\n> would be in [hookcmd \"git-secrets-committing\"] and the latter\n> would be in [hookcmd \"git-secrets-pushing\"], so I can set the ordering\n> how I wish.\n\nI think all of this is _possible_ in either scheme. We're encoding\npotentially tabular data into a hierarchical config structure. In either\ncase I can set hook->cmd->option or cmd->hook->option. The question is\njust which arrangement makes it simplest to do the most common things.\n\n> Yeah, I see what you mean, and again I really like that. That lets us\n> run multiples in config order easily:\n> \n> [hook \"pre-receive\"]\n>   command = /some/script\n>   command = /some/other-script\n>   command = some-hookcmd-header\n\nYep, config order makes sense as a default (though I think you could\nmake an argument for lexical order by command-name, which allows naming\nthings \"000foo\" if the user really wants to).\n\n> If we add a little repeated-name detection then we can also reorder\n> easily this way if that's the direction we want for ordering:\n> \n> {global}\n> hook.pre-receive.command = a.sh\n> hook.pre-receive.command = b.sh\n> \n> {local}\n> hook.pre-receive.command = c.sh\n> hook.pre-receive.command = a.sh\n> \n> for a final order of {b.sh, c.sh, a.sh}.\n\nI'm not sure what I'd expect a repeated mention of \"a.sh\" to do, but as\nlong as it's well-defined I don't really care. :)\n\n> I wonder - I think even something like this would work:\n> \n> {global}\n> [hook \"pre-receive\"]\n>   command = no-hookcmd-entry.sh\n> \n> {local for repo \"zork\"}\n> [hookcmd \"no-hookcmd-entry.sh\"]\n>   skip = true\n>\n> For most repos, now I simply invoke no-hookcmd-entry.sh on pre-receive,\n> but when I'm parsing the config in \"zork\", now I see a populated hookcmd\n> entry, and when I look it up with the key I found in the global config,\n> I see that it's supposed to be skipped.\n\nYes, exactly. The config parsing procedure is really just filling in a\n\"struct hookcmd\" as it goes, so we don't care that they're in two\nseparate files.\n\n> Although I might need to do something hacky if I have multiple hooks\n> pointing to the same simple invocation:\n> \n> {global}\n> [hook \"pre-receive\"]\n>   command = no-hookcmd-entry.sh\n> \n> [hook \"post-commit\"]\n>   command = no-hookcmd-entry.sh\n\nI think your commands would be:\n\n  command = \"no-hookcmd-entry.sh pre-receive\"\n\netc in that case, so they'd have different hookcmd blocks. You\n_wouldn't_ be able to just turn off all of them with one config command,\nthough.\n\n> I wonder if I'm getting buried in the weeds of stuff we won't ever have\n> to worry about ;)\n\nYeah. I don't mind a little over-engineering as long as the easy things\nremain simple, and the hard things remain possible. But that also means\nwe might be able to grow the hard things later (or never) as long as we\nhave a reasonable plan for them.\n\n-Peff\n"},{"id":"395383","messageId":"20200414005457.3505-1-emilyshaffer@google.com","threadId":"52980","inReplyTo":"20200413215256.GA18990@coredump.intra.peff.net","subject":"[RFC PATCH v2 0/2] configuration-based hook management (was: [TOPIC 2/17] Hooks in the future)","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-14T00:54:55Z","receivedAt":"2020-04-14T00:55:11Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Not much to look at compared to the original RFC I sent some months ago.\nThis implements Peff's suggestion of using the \"hookcmd\" section as a\nlayer of indirection. The scope is a little smaller than the original\nRFC as it doesn't have a way to remove hooks from downstream (yet), and\nordering numbers are dropped (for now).\n\nOne thing that's missing, as evidenced by the TODO, is a way to handle\narbitrary options given within a \"hookcmd\" unit. I think this can be\nachieved with a callback, since it seems plausible that \"pre-receive\"\nmight want a different set of options than \"post-commit\" or so on. To\nme, it sounds achievable with a callback; I imagine a follow-on teaching\ngit-hook how to remove a hook with something like \"hookcmd.foo.skip =\ntrue\" will give an OK indication of how that might look.\n\nOverall though, I think this is simpler than the first version of the\nRFC because I was reminded by wiser folks than I to \"keep it simple,\nstupid.\" ;)\n\nI think it's feasible that with these couple patches applied, someone\nwho wanted to jump in early could replace their\n.git/hook/whatever-hookname with some boilerplate like\n\n  xargs -n 1 'sh -c' <<<\"$(git hook --list whatever-hookname)\"\n\nand give it a shot. Untested snippet. :)\n\nCI run: https://github.com/gitgitgadget/git/pull/611/checks\n\n - Emily\n\nEmily Shaffer (2):\n  hook: scaffolding for git-hook subcommand\n  hook: add --list mode\n\n .gitignore                    |  1 +\n Documentation/git-hook.txt    | 53 ++++++++++++++++++++\n Makefile                      |  2 +\n builtin.h                     |  1 +\n builtin/hook.c                | 77 +++++++++++++++++++++++++++++\n git.c                         |  1 +\n hook.c                        | 92 +++++++++++++++++++++++++++++++++++\n hook.h                        | 13 +++++\n t/t1360-config-based-hooks.sh | 58 ++++++++++++++++++++++\n 9 files changed, 298 insertions(+)\n create mode 100644 Documentation/git-hook.txt\n create mode 100644 builtin/hook.c\n create mode 100644 hook.c\n create mode 100644 hook.h\n create mode 100755 t/t1360-config-based-hooks.sh\n\n-- \n2.26.0.110.g2183baf09c-goog\n\n"},{"id":"395384","messageId":"20200414005457.3505-2-emilyshaffer@google.com","threadId":"52980","inReplyTo":"20200414005457.3505-1-emilyshaffer@google.com","subject":"[RFC PATCH v2 1/2] hook: scaffolding for git-hook subcommand","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-14T00:54:56Z","receivedAt":"2020-04-14T00:55:14Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Introduce infrastructure for a new subcommand, git-hook, which will be\nused to ease config-based hook management. This command will handle\nparsing configs to compose a list of hooks to run for a given event, as\nwell as adding or modifying hook configs in an interactive fashion.\n\nSigned-off-by: Emily Shaffer <emilyshaffer@google.com>\n---\n .gitignore                    |  1 +\n Documentation/git-hook.txt    | 19 +++++++++++++++++++\n Makefile                      |  1 +\n builtin.h                     |  1 +\n builtin/hook.c                | 21 +++++++++++++++++++++\n git.c                         |  1 +\n t/t1360-config-based-hooks.sh | 11 +++++++++++\n 7 files changed, 55 insertions(+)\n create mode 100644 Documentation/git-hook.txt\n create mode 100644 builtin/hook.c\n create mode 100755 t/t1360-config-based-hooks.sh\n\ndiff --git a/.gitignore b/.gitignore\nindex 188bd1c3de..0f8b74f651 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -74,6 +74,7 @@\n /git-grep\n /git-hash-object\n /git-help\n+/git-hook\n /git-http-backend\n /git-http-fetch\n /git-http-push\ndiff --git a/Documentation/git-hook.txt b/Documentation/git-hook.txt\nnew file mode 100644\nindex 0000000000..2d50c414cc\n--- /dev/null\n+++ b/Documentation/git-hook.txt\n@@ -0,0 +1,19 @@\n+git-hook(1)\n+===========\n+\n+NAME\n+----\n+git-hook - Manage configured hooks\n+\n+SYNOPSIS\n+--------\n+[verse]\n+'git hook'\n+\n+DESCRIPTION\n+-----------\n+You can list, add, and modify hooks with this command.\n+\n+GIT\n+---\n+Part of the linkgit:git[1] suite\ndiff --git a/Makefile b/Makefile\nindex ef1ff2228f..7b9670c205 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1079,6 +1079,7 @@ BUILTIN_OBJS += builtin/get-tar-commit-id.o\n BUILTIN_OBJS += builtin/grep.o\n BUILTIN_OBJS += builtin/hash-object.o\n BUILTIN_OBJS += builtin/help.o\n+BUILTIN_OBJS += builtin/hook.o\n BUILTIN_OBJS += builtin/index-pack.o\n BUILTIN_OBJS += builtin/init-db.o\n BUILTIN_OBJS += builtin/interpret-trailers.o\ndiff --git a/builtin.h b/builtin.h\nindex 2b25a80cde..c4cd252f61 100644\n--- a/builtin.h\n+++ b/builtin.h\n@@ -173,6 +173,7 @@ int cmd_get_tar_commit_id(int argc, const char **argv, const char *prefix);\n int cmd_grep(int argc, const char **argv, const char *prefix);\n int cmd_hash_object(int argc, const char **argv, const char *prefix);\n int cmd_help(int argc, const char **argv, const char *prefix);\n+int cmd_hook(int argc, const char **argv, const char *prefix);\n int cmd_index_pack(int argc, const char **argv, const char *prefix);\n int cmd_init_db(int argc, const char **argv, const char *prefix);\n int cmd_interpret_trailers(int argc, const char **argv, const char *prefix);\ndiff --git a/builtin/hook.c b/builtin/hook.c\nnew file mode 100644\nindex 0000000000..b2bbc84d4d\n--- /dev/null\n+++ b/builtin/hook.c\n@@ -0,0 +1,21 @@\n+#include \"cache.h\"\n+\n+#include \"builtin.h\"\n+#include \"parse-options.h\"\n+\n+static const char * const builtin_hook_usage[] = {\n+\tN_(\"git hook\"),\n+\tNULL\n+};\n+\n+int cmd_hook(int argc, const char **argv, const char *prefix)\n+{\n+\tstruct option builtin_hook_options[] = {\n+\t\tOPT_END(),\n+\t};\n+\n+\targc = parse_options(argc, argv, prefix, builtin_hook_options,\n+\t\t\t     builtin_hook_usage, 0);\n+\n+\treturn 0;\n+}\ndiff --git a/git.c b/git.c\nindex b07198fe03..c79a9192d6 100644\n--- a/git.c\n+++ b/git.c\n@@ -513,6 +513,7 @@ static struct cmd_struct commands[] = {\n \t{ \"grep\", cmd_grep, RUN_SETUP_GENTLY },\n \t{ \"hash-object\", cmd_hash_object },\n \t{ \"help\", cmd_help },\n+\t{ \"hook\", cmd_hook, RUN_SETUP },\n \t{ \"index-pack\", cmd_index_pack, RUN_SETUP_GENTLY | NO_PARSEOPT },\n \t{ \"init\", cmd_init_db },\n \t{ \"init-db\", cmd_init_db },\ndiff --git a/t/t1360-config-based-hooks.sh b/t/t1360-config-based-hooks.sh\nnew file mode 100755\nindex 0000000000..34b0df5216\n--- /dev/null\n+++ b/t/t1360-config-based-hooks.sh\n@@ -0,0 +1,11 @@\n+#!/bin/bash\n+\n+test_description='config-managed multihooks, including git-hook command'\n+\n+. ./test-lib.sh\n+\n+test_expect_success 'git hook command does not crash' '\n+\tgit hook\n+'\n+\n+test_done\n-- \n2.26.0.110.g2183baf09c-goog\n\n"},{"id":"395385","messageId":"20200414005457.3505-3-emilyshaffer@google.com","threadId":"52980","inReplyTo":"20200414005457.3505-1-emilyshaffer@google.com","subject":"[RFC PATCH v2 2/2] hook: add --list mode","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-14T00:54:57Z","receivedAt":"2020-04-14T00:55:16Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Teach 'git hook --list <hookname>', which checks the known configs in\norder to create an ordered list of hooks to run on a given hook event.\n\nMultiple commands can be specified for a given hook by providing\nmultiple \"hook.<hookname>.command = <path-to-hook>\" lines. Hooks will be\nrun in config order. If more properties need to be set on a given hook\nin the future, commands can also be specified by providing\n\"hook.<hookname>.command = <hookcmd-name>\", as well as a \"[hookcmd\n<hookcmd-name>]\" subsection; at minimum, this subsection must contain a\n\"hookcmd.<hookcmd-name>.command = <path-to-hook>\" line.\n\nFor example:\n\n  $ git config --list | grep ^hook\n  hook.pre-commit.command=baz\n  hook.pre-commit.command=~/bar.sh\n  hookcmd.baz.command=~/baz/from/hookcmd.sh\n\n  $ git hook --list pre-commit\n  ~/baz/from/hookcmd.sh\n  ~/bar.sh\n\nSigned-off-by: Emily Shaffer <emilyshaffer@google.com>\n---\n Documentation/git-hook.txt    | 36 +++++++++++++-\n Makefile                      |  1 +\n builtin/hook.c                | 58 +++++++++++++++++++++-\n hook.c                        | 92 +++++++++++++++++++++++++++++++++++\n hook.h                        | 15 ++++++\n t/t1360-config-based-hooks.sh | 51 ++++++++++++++++++-\n 6 files changed, 249 insertions(+), 4 deletions(-)\n create mode 100644 hook.c\n create mode 100644 hook.h\n\ndiff --git a/Documentation/git-hook.txt b/Documentation/git-hook.txt\nindex 2d50c414cc..aafc762ea2 100644\n--- a/Documentation/git-hook.txt\n+++ b/Documentation/git-hook.txt\n@@ -8,12 +8,46 @@ git-hook - Manage configured hooks\n SYNOPSIS\n --------\n [verse]\n-'git hook'\n+'git hook' -l | --list <hook-name>\n \n DESCRIPTION\n -----------\n You can list, add, and modify hooks with this command.\n \n+This command parses the default configuration files for sections \"hook\" and\n+\"hookcmd\". \"hook\" is used to describe the commands which will be run during a\n+particular hook event; commands are run in config order. \"hookcmd\" is used to\n+describe attributes of a specific command. If additional attributes don't need\n+to be specified, a command to run can be specified directly in the \"hook\"\n+section; if a \"hookcmd\" by that name isn't found, Git will attempt to run the\n+provided value directly. For example:\n+\n+Global config\n+----\n+  [hook \"post-commit\"]\n+    command = \"linter\"\n+    command = \"~/typocheck.sh\"\n+\n+  [hookcmd \"linter\"]\n+    command = \"/bin/linter --c\"\n+----\n+\n+Local config\n+----\n+  [hook \"prepare-commit-msg\"]\n+    command = \"linter\"\n+  [hook \"post-commit\"]\n+    command = \"python ~/run-test-suite.py\"\n+----\n+\n+OPTIONS\n+-------\n+\n+-l::\n+--list::\n+\tList the hooks which have been configured for <hook-name>. Hooks appear\n+\tin the order they should be run.\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\ndiff --git a/Makefile b/Makefile\nindex 7b9670c205..5f170f885b 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -896,6 +896,7 @@ LIB_OBJS += hashmap.o\n LIB_OBJS += linear-assignment.o\n LIB_OBJS += help.o\n LIB_OBJS += hex.o\n+LIB_OBJS += hook.o\n LIB_OBJS += ident.o\n LIB_OBJS += interdiff.o\n LIB_OBJS += json-writer.o\ndiff --git a/builtin/hook.c b/builtin/hook.c\nindex b2bbc84d4d..60617578fb 100644\n--- a/builtin/hook.c\n+++ b/builtin/hook.c\n@@ -1,21 +1,77 @@\n #include \"cache.h\"\n \n #include \"builtin.h\"\n+#include \"config.h\"\n+#include \"hook.h\"\n #include \"parse-options.h\"\n+#include \"strbuf.h\"\n \n static const char * const builtin_hook_usage[] = {\n-\tN_(\"git hook\"),\n+\tN_(\"git hook --list <hookname>\"),\n \tNULL\n };\n \n+enum hook_command {\n+\tHOOK_NO_COMMAND = 0,\n+\tHOOK_LIST,\n+};\n+\n+static int print_hook_list(const struct strbuf *hookname)\n+{\n+\tstruct list_head *head, *pos;\n+\tstruct hook *item;\n+\n+\thead = hook_list(hookname);\n+\n+\tif (!head) {\n+\t\tprintf(_(\"no commands configured for hook '%s'\\n\"),\n+\t\t       hookname->buf);\n+\t\treturn 0;\n+\t}\n+\n+\tlist_for_each(pos, head) {\n+\t\titem = list_entry(pos, struct hook, list);\n+\t\tif (item)\n+\t\t\tprintf(\"%s\\n\",\n+\t\t\t       item->command.buf);\n+\t}\n+\n+\treturn 0;\n+}\n+\n int cmd_hook(int argc, const char **argv, const char *prefix)\n {\n+\tenum hook_command command = 0;\n+\tstruct strbuf hookname = STRBUF_INIT;\n+\n \tstruct option builtin_hook_options[] = {\n+\t\tOPT_CMDMODE('l', \"list\", &command,\n+\t\t\t    N_(\"list scripts which will be run for <hookname>\"),\n+\t\t\t    HOOK_LIST),\n \t\tOPT_END(),\n \t};\n \n \targc = parse_options(argc, argv, prefix, builtin_hook_options,\n \t\t\t     builtin_hook_usage, 0);\n \n+\tif (argc < 1) {\n+\t\tusage_msg_opt(\"a hookname must be provided to operate on.\",\n+\t\t\t      builtin_hook_usage, builtin_hook_options);\n+\t}\n+\n+\tstrbuf_addstr(&hookname, argv[0]);\n+\n+\tswitch(command) {\n+\t\tcase HOOK_LIST:\n+\t\t\treturn print_hook_list(&hookname);\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\tusage_msg_opt(\"no command given.\", builtin_hook_usage,\n+\t\t\t\t      builtin_hook_options);\n+\t}\n+\n+\tclear_hook_list();\n+\tstrbuf_release(&hookname);\n+\n \treturn 0;\n }\ndiff --git a/hook.c b/hook.c\nnew file mode 100644\nindex 0000000000..a31943a25e\n--- /dev/null\n+++ b/hook.c\n@@ -0,0 +1,92 @@\n+#include \"cache.h\"\n+\n+#include \"hook.h\"\n+#include \"config.h\"\n+\n+static LIST_HEAD(hook_head);\n+\n+void free_hook(struct hook *ptr)\n+{\n+\tif (ptr) {\n+\t\tstrbuf_release(&ptr->command);\n+\t\tfree(ptr);\n+\t}\n+}\n+\n+static void emplace_hook(struct list_head *pos, const char *command)\n+{\n+\tstruct hook *to_add = malloc(sizeof(struct hook));\n+\tto_add->origin = current_config_scope();\n+\tstrbuf_init(&to_add->command, 0);\n+\tstrbuf_addstr(&to_add->command, command);\n+\n+\tlist_add_tail(&to_add->list, pos);\n+}\n+\n+static void remove_hook(struct list_head *to_remove)\n+{\n+\tstruct hook *hook_to_remove = list_entry(to_remove, struct hook, list);\n+\tlist_del(to_remove);\n+\tfree_hook(hook_to_remove);\n+}\n+\n+void clear_hook_list(void)\n+{\n+\tstruct list_head *pos, *tmp;\n+\tlist_for_each_safe(pos, tmp, &hook_head)\n+\t\tremove_hook(pos);\n+}\n+\n+struct list_head* hook_list(const struct strbuf* hookname)\n+{\n+\tstruct strbuf hook_key = STRBUF_INIT;\n+\tconst struct string_list *commands = NULL;\n+\tstruct string_list_item *it = NULL;\n+\tstruct list_head *pos = NULL, *tmp = NULL;\n+\tstruct strbuf hookcmd_name = STRBUF_INIT;\n+\tstruct hook *hook = NULL;\n+\n+\tif (!hookname)\n+\t\treturn NULL;\n+\n+\tstrbuf_addf(&hook_key, \"hook.%s.command\", hookname->buf);\n+\n+\tcommands = git_config_get_value_multi(hook_key.buf);\n+\n+\tif (!commands)\n+\t\treturn NULL;\n+\n+\tfor_each_string_list_item(it, commands) {\n+\t\tconst char *command = it->string;\n+\n+\t\tstrbuf_reset(&hookcmd_name);\n+\t\tstrbuf_addf(&hookcmd_name, \"hookcmd.%s.command\", command);\n+\n+\t\t/* If no hookcmd with that name exists, &command is untouched */\n+\t\tgit_config_get_value(hookcmd_name.buf, &command);\n+\n+\t\tif (!command)\n+\t\t\treturn NULL;\n+\n+\t\t/*\n+\t\t * TODO: implement an option-getting callback, e.g.\n+\t\t *   get configs by pattern hookcmd.$value.*\n+\t\t *   for each key+value, do_callback(key, value, cb_data)\n+\t\t */\n+\n+\t\tlist_for_each_safe(pos, tmp, &hook_head) {\n+\t\t\thook = list_entry(pos, struct hook, list);\n+\t\t\t/*\n+\t\t\t * The list of hooks to run can be reordered by being redeclared\n+\t\t\t * in the config. Options about hook ordering should be checked\n+\t\t\t * here.\n+\t\t\t */\n+\t\t\tif (0 == strcmp(hook->command.buf, command))\n+\t\t\t\tremove_hook(pos);\n+\t\t}\n+\t\templace_hook(pos, command);\n+\n+\t}\n+\n+\treturn &hook_head;\n+}\ndiff --git a/hook.h b/hook.h\nnew file mode 100644\nindex 0000000000..aaf6511cff\n--- /dev/null\n+++ b/hook.h\n@@ -0,0 +1,15 @@\n+#include \"config.h\"\n+#include \"list.h\"\n+#include \"strbuf.h\"\n+\n+struct hook\n+{\n+\tstruct list_head list;\n+\tenum config_scope origin;\n+\tstruct strbuf command;\n+};\n+\n+struct list_head* hook_list(const struct strbuf *hookname);\n+\n+void free_hook(struct hook *ptr);\n+void clear_hook_list(void);\ndiff --git a/t/t1360-config-based-hooks.sh b/t/t1360-config-based-hooks.sh\nindex 34b0df5216..2e6a5e09d3 100755\n--- a/t/t1360-config-based-hooks.sh\n+++ b/t/t1360-config-based-hooks.sh\n@@ -4,8 +4,55 @@ test_description='config-managed multihooks, including git-hook command'\n \n . ./test-lib.sh\n \n-test_expect_success 'git hook command does not crash' '\n-\tgit hook\n+test_expect_success 'git hook rejects commands without a mode' '\n+\ttest_must_fail git hook pre-commit\n+'\n+\n+\n+test_expect_success 'git hook rejects commands without a hookname' '\n+\ttest_must_fail git hook --list\n+'\n+\n+test_expect_success 'setup hooks in global, and local' '\n+\tgit config --add --local hook.pre-commit.command \"/path/ghi\" &&\n+\tgit config --add --global hook.pre-commit.command \"/path/def\"\n+'\n+\n+test_expect_success 'git hook --list orders by config order' '\n+\tcat >expected <<-\\EOF &&\n+\t/path/def\n+\t/path/ghi\n+\tEOF\n+\n+\tgit hook --list pre-commit >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'git hook --list dereferences a hookcmd' '\n+\tgit config --add --local hook.pre-commit.command \"abc\" &&\n+\tgit config --add --global hookcmd.abc.command \"/path/abc\" &&\n+\n+\tcat >expected <<-\\EOF &&\n+\t/path/def\n+\t/path/ghi\n+\t/path/abc\n+\tEOF\n+\n+\tgit hook --list pre-commit >actual &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'git hook --list reorders on duplicate commands' '\n+\tgit config --add --local hook.pre-commit.command \"/path/def\" &&\n+\n+\tcat >expected <<-\\EOF &&\n+\t/path/ghi\n+\t/path/abc\n+\t/path/def\n+\tEOF\n+\n+\tgit hook --list pre-commit >actual &&\n+\ttest_cmp expected actual\n '\n \n test_done\n-- \n2.26.0.110.g2183baf09c-goog\n\n"},{"id":"395415","messageId":"efad3927-1d8f-5545-48e9-9a58c2308273@gmail.com","threadId":"52980","inReplyTo":"20200414005457.3505-1-emilyshaffer@google.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2020-04-14T15:15:11Z","receivedAt":"2020-04-14T15:15:42Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Emily\n\nThanks for working on this, having a way to manage multiple commands per \nhook without using an external framework would be really useful\n\nOn 14/04/2020 01:54, Emily Shaffer wrote:\n> Not much to look at compared to the original RFC I sent some months ago.\n> This implements Peff's suggestion of using the \"hookcmd\" section as a\n> layer of indirection.\n\nI'm not really clear what the advantage of this indirection is. It seems \nunlikely to me that different hooks will share exactly the same command \nline or other options. In the 'git secrets' example earlier in this \nthread each hook needs to use a different command line. In general a \ncommand cannot tell which hook it is being invoked as without a flag of \nsome kind. (In some cases it can use the number of arguments if that is \ndifferent for each hook that it handles but that is not true in general)\n\nWithout the redirection one could have\n   hook.pre-commit.linter.command = my-command\n   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n\nand other keys can be added for ordering etc. e.g.\n   hook.pre-commit.linter.before = check-whitespace\n\nWith the indirection one needs to set\n   hook.pre-commit.command = linter\n   hook.pre-commit.check-whitespace = 'git diff --check --cached'\n   hookcmd.linter.command = my-command\n   hookcmd.linter.pre-commit-before = check-whitespace\n\nwhich involves setting an extra key and checking it each time the hook \nis invoked without any benefit that I can see. I suspect which one seems \nmore logical depends on how one thinks of setting hooks - I tend to \nthink \"I want to set a pre-commit hook\" not \"I want to set a git-secrets \nhook\". If you've got an example where this indirection is helpful or \nnecessary that would be really useful to see.\n\nBest Wishes\n\nPhillip\n\n\n> The scope is a little smaller than the original\n> RFC as it doesn't have a way to remove hooks from downstream (yet), and\n> ordering numbers are dropped (for now).\n> \n> One thing that's missing, as evidenced by the TODO, is a way to handle\n> arbitrary options given within a \"hookcmd\" unit. I think this can be\n> achieved with a callback, since it seems plausible that \"pre-receive\"\n> might want a different set of options than \"post-commit\" or so on. To\n> me, it sounds achievable with a callback; I imagine a follow-on teaching\n> git-hook how to remove a hook with something like \"hookcmd.foo.skip =\n> true\" will give an OK indication of how that might look.\n> \n> Overall though, I think this is simpler than the first version of the\n> RFC because I was reminded by wiser folks than I to \"keep it simple,\n> stupid.\" ;)\n> \n> I think it's feasible that with these couple patches applied, someone\n> who wanted to jump in early could replace their\n> .git/hook/whatever-hookname with some boilerplate like\n> \n>    xargs -n 1 'sh -c' <<<\"$(git hook --list whatever-hookname)\"\n> \n> and give it a shot. Untested snippet. :)\n> CI run: https://github.com/gitgitgadget/git/pull/611/checks\n> \n>   - Emily\n> \n> Emily Shaffer (2):\n>    hook: scaffolding for git-hook subcommand\n>    hook: add --list mode\n> \n>   .gitignore                    |  1 +\n>   Documentation/git-hook.txt    | 53 ++++++++++++++++++++\n>   Makefile                      |  2 +\n>   builtin.h                     |  1 +\n>   builtin/hook.c                | 77 +++++++++++++++++++++++++++++\n>   git.c                         |  1 +\n>   hook.c                        | 92 +++++++++++++++++++++++++++++++++++\n>   hook.h                        | 13 +++++\n>   t/t1360-config-based-hooks.sh | 58 ++++++++++++++++++++++\n>   9 files changed, 298 insertions(+)\n>   create mode 100644 Documentation/git-hook.txt\n>   create mode 100644 builtin/hook.c\n>   create mode 100644 hook.c\n>   create mode 100644 hook.h\n>   create mode 100755 t/t1360-config-based-hooks.sh\n> \n"},{"id":"395433","messageId":"20200414192418.GB5478@google.com","threadId":"52980","inReplyTo":"efad3927-1d8f-5545-48e9-9a58c2308273@gmail.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-14T19:24:18Z","receivedAt":"2020-04-14T19:31:27Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Tue, Apr 14, 2020 at 04:15:11PM +0100, Phillip Wood wrote:\n> Hi Emily\n> \n> Thanks for working on this, having a way to manage multiple commands per\n> hook without using an external framework would be really useful\n> \n> On 14/04/2020 01:54, Emily Shaffer wrote:\n> > Not much to look at compared to the original RFC I sent some months ago.\n> > This implements Peff's suggestion of using the \"hookcmd\" section as a\n> > layer of indirection.\n> \n> I'm not really clear what the advantage of this indirection is. It seems\n> unlikely to me that different hooks will share exactly the same command line\n> or other options. In the 'git secrets' example earlier in this thread each\n> hook needs to use a different command line. In general a command cannot tell\n> which hook it is being invoked as without a flag of some kind. (In some\n> cases it can use the number of arguments if that is different for each hook\n> that it handles but that is not true in general)\n> \n> Without the redirection one could have\n>   hook.pre-commit.linter.command = my-command\n>   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n\nI think this isn't supported by the config semantics. Have a look at\nconfig.h:parse_config_key:\n\n  /*\n   * Match and parse a config key of the form:\n   *\n   *   section.(subsection.)?key\n   *\n   * (i.e., what gets handed to a config_fn_t). The caller provides the section;\n   * we return -1 if it does not match, 0 otherwise. The subsection and key\n   * out-parameters are filled by the function (and *subsection is NULL if it is\n   * missing).\n   *\n   * If the subsection pointer-to-pointer passed in is NULL, returns 0 only if\n   * there is no subsection at all.\n   */\n  int parse_config_key(const char *var,\n                       const char *section,\n                       const char **subsection, int *subsection_len,\n                       const char **key);\n\nWe'd need to fudge one of these fields to include the extra section, I\nthink. Unfortunate, because I find your example very tidy, but in\npractice maybe not very neat. The closest thing I can find to a nice way\nof writing it might be:\n\n  [hook.pre-commit \"linter\"]\n    command = my-command\n    before = check-whitespace\n  [hook.pre-commit \"check-whitespace\"]\n    command = 'git diff --check --cached'\n\nBut this is kind of a lie; the sections aren't \"hook\", \"pre-commit\", and\n\"linter\" as you'd expect. Whether it's OK to lie like this, though, I\ndon't know - I suspect it might make it awkward for others trying to\nparse the config. (my Vim syntax highlighter had kind of a hard time.)\n\n> \n> and other keys can be added for ordering etc. e.g.\n>   hook.pre-commit.linter.before = check-whitespace\n> \n> With the indirection one needs to set\n>   hook.pre-commit.command = linter\n>   hook.pre-commit.check-whitespace = 'git diff --check --cached'\n>   hookcmd.linter.command = my-command\n>   hookcmd.linter.pre-commit-before = check-whitespace\n> \n> which involves setting an extra key and checking it each time the hook is\n> invoked without any benefit that I can see. I suspect which one seems more\n> logical depends on how one thinks of setting hooks - I tend to think \"I want\n> to set a pre-commit hook\" not \"I want to set a git-secrets hook\". If you've\n> got an example where this indirection is helpful or necessary that would be\n> really useful to see.\n\nThanks for sharing your workflow; as always, it's hard to understand the\nways others work differently from yourself, so I'm glad to hear from\nyou. Let me think some more on it and reply back again.\n\n - Emily\n"},{"id":"395435","messageId":"20200414200347.GD12694@google.com","threadId":"52980","inReplyTo":"efad3927-1d8f-5545-48e9-9a58c2308273@gmail.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Josh Steadmon","fromEmail":"steadmon@google.com","sentAt":"2020-04-14T20:03:47Z","receivedAt":"2020-04-14T20:03:58Z","isPatch":true,"sender":{"key":"steadmon@google.com","avatar":"https://avatars.githubusercontent.com/u/2654920?v=4"},"body":"On 2020.04.14 16:15, Phillip Wood wrote:\n> Hi Emily\n> \n> Thanks for working on this, having a way to manage multiple commands per\n> hook without using an external framework would be really useful\n> \n> On 14/04/2020 01:54, Emily Shaffer wrote:\n> > Not much to look at compared to the original RFC I sent some months ago.\n> > This implements Peff's suggestion of using the \"hookcmd\" section as a\n> > layer of indirection.\n> \n> I'm not really clear what the advantage of this indirection is. It seems\n> unlikely to me that different hooks will share exactly the same command line\n> or other options. In the 'git secrets' example earlier in this thread each\n> hook needs to use a different command line. In general a command cannot tell\n> which hook it is being invoked as without a flag of some kind. (In some\n> cases it can use the number of arguments if that is different for each hook\n> that it handles but that is not true in general)\n> \n> Without the redirection one could have\n>   hook.pre-commit.linter.command = my-command\n>   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n> \n> and other keys can be added for ordering etc. e.g.\n>   hook.pre-commit.linter.before = check-whitespace\n> \n> With the indirection one needs to set\n>   hook.pre-commit.command = linter\n>   hook.pre-commit.check-whitespace = 'git diff --check --cached'\n>   hookcmd.linter.command = my-command\n>   hookcmd.linter.pre-commit-before = check-whitespace\n> \n> which involves setting an extra key and checking it each time the hook is\n> invoked without any benefit that I can see. I suspect which one seems more\n> logical depends on how one thinks of setting hooks - I tend to think \"I want\n> to set a pre-commit hook\" not \"I want to set a git-secrets hook\". If you've\n> got an example where this indirection is helpful or necessary that would be\n> really useful to see.\n> \n> Best Wishes\n> \n> Phillip\n\nIndexing repo content (see [1] for a detailed discussion) is one use\ncase where you have a single command that runs identically from\npost-commit, post-merge, and post-checkout.\n\nAlso, I suspect that many users don't have a firm enough grasp on the\nvarious git hooks options to know ahead of time which ones they want to\nset to accomplish a given task (without diving into the docs first). I'm\nnot trying to say that your workflow is incorrect, but my gut feeling is\nthat most Git users would work in the opposite direction. Every time I\nhave needed to automate something, I generally had a rough script in\nplace first, and then looked up which hook(s) would be appropriate\ntriggers for the script.\n\n\n[1]: https://tbaggery.com/2011/08/08/effortless-ctags-with-git.html\n"},{"id":"395441","messageId":"20200414202738.GD1879688@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200414192418.GB5478@google.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-14T20:27:38Z","receivedAt":"2020-04-14T20:27:52Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 14, 2020 at 12:24:18PM -0700, Emily Shaffer wrote:\n\n> > Without the redirection one could have\n> >   hook.pre-commit.linter.command = my-command\n> >   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n> [...]\n> We'd need to fudge one of these fields to include the extra section, I\n> think. Unfortunate, because I find your example very tidy, but in\n> practice maybe not very neat. The closest thing I can find to a nice way\n> of writing it might be:\n> \n>   [hook.pre-commit \"linter\"]\n>     command = my-command\n>     before = check-whitespace\n>   [hook.pre-commit \"check-whitespace\"]\n>     command = 'git diff --check --cached'\n\nSyntactically the whole section between the outer dots is the\nsubsection. So it's:\n\n  [hook \"pre-commit.check-whitespace\"]\n  command = ...\n\nAnd I don't think we want to change the config syntax at this point.\nEven in the neater dotted notation, we must keep that whole thing as a\nsubsection, because existing subsections may contain dots, too.\n\n> But this is kind of a lie; the sections aren't \"hook\", \"pre-commit\", and\n> \"linter\" as you'd expect. Whether it's OK to lie like this, though, I\n> don't know - I suspect it might make it awkward for others trying to\n> parse the config. (my Vim syntax highlighter had kind of a hard time.)\n\nI think we should avoid it if possible. There are some subtleties there,\nlike the fact that subsections are case-sensitive, but sections and keys\nare not.\n\n-Peff\n"},{"id":"395445","messageId":"20200414203247.GE1879688@coredump.intra.peff.net","threadId":"52980","inReplyTo":"efad3927-1d8f-5545-48e9-9a58c2308273@gmail.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-14T20:32:47Z","receivedAt":"2020-04-14T20:32:52Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 14, 2020 at 04:15:11PM +0100, Phillip Wood wrote:\n\n> On 14/04/2020 01:54, Emily Shaffer wrote:\n> > Not much to look at compared to the original RFC I sent some months ago.\n> > This implements Peff's suggestion of using the \"hookcmd\" section as a\n> > layer of indirection.\n> \n> I'm not really clear what the advantage of this indirection is. It seems\n> unlikely to me that different hooks will share exactly the same command line\n> or other options. In the 'git secrets' example earlier in this thread each\n> hook needs to use a different command line. In general a command cannot tell\n> which hook it is being invoked as without a flag of some kind. (In some\n> cases it can use the number of arguments if that is different for each hook\n> that it handles but that is not true in general)\n> \n> Without the redirection one could have\n>   hook.pre-commit.linter.command = my-command\n>   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n> \n> and other keys can be added for ordering etc. e.g.\n>   hook.pre-commit.linter.before = check-whitespace\n> \n> With the indirection one needs to set\n>   hook.pre-commit.command = linter\n>   hook.pre-commit.check-whitespace = 'git diff --check --cached'\n>   hookcmd.linter.command = my-command\n>   hookcmd.linter.pre-commit-before = check-whitespace\n\nIn the proposal I gave, you could do:\n\n  hook.pre-commit.command = my-command\n  hook.pre-commit.command = git diff --check --cached\n\nIf you want to refer to commands in ordering options (like your\n\"before\"), then you'd have to refer to their names. For \"my-command\"\nthat's not too bad. For the longer one, it's a bit awkward. You _could_\ndo:\n\n  hookcmd.my-command.before = git diff --check --cached\n\nwhich is the same number of lines as yours. But I'd probably give it a\nname, like:\n\n  hookcmd.check-whitespace.command = git diff --check --cached\n  hookcmd.my-command.before = check-whitespace\n\nThat's one more line than yours, but I think it separates the concerns\nmore clearly. And it extends naturally to more options specific to\ncheck-whitespace.\n\n-Peff\n"},{"id":"395470","messageId":"20200415034550.GB36683@google.com","threadId":"52980","inReplyTo":"20200413215256.GA18990@coredump.intra.peff.net","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2020-04-15T03:45:50Z","receivedAt":"2020-04-15T03:45:59Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJeff King wrote:\n> On Mon, Apr 13, 2020 at 12:15:15PM -0700, Emily Shaffer wrote:\n>> Jeff King wrote:\n\n>>> Yeah, giving each block a unique name lets you give them each an order.\n>>> It seems kind of weird to me that you'd define multiple hook types for a\n>>> given name.\n>>\n>> Not so odd - git-secrets configures itself for pre-commit,\n>> prepare-commit-msg, and commit-msg-hook.\n[...]\n> Yeah, I do see how that use case makes sense. I wonder how common it is\n> versus having separate one-off hooks.\n\nI think separately from the frequency question, we should look at the\n\"what model do we want to present to the user\" question.\n\nIt's not too unusual for a project with their source code in a Git\nrepository to have conventions they want to nudge users toward.  I'd\nexpect them to use a combination of hooks for this:\n\n\tprepare-commit-msg\n\tcommit-msg\n\tpre-push\n\nGit LFS installs multiple hooks:\n\n\tpre-push\n\tpost-checkout\n\tpost-commit\n\tpost-merge\n\ngit-secrets installs multiple hooks, as already mentioned.\n\nWe've also had some instances over time of one hook replacing another,\nto improve the interface.  A program wanting to install hooks would\nthen be likely to migrate from the older interface to the better one.\n\nWhat I mean to get at is that I think thinking of them in terms of\nindividual hooks, the user model assumed by these programs is to think\nof them as plugins hooking into Git.  The individual hooks are events\nthat the plugin listens on.  If I am trying to disable a plugin, I\ndon't want to have to learn which events it cared about.\n\n>                                       And whether setting the order\n> priority for all hooks at once is that useful (e.g., I can easily\n> imagine a case where the pre-commit hook for program A must go before B,\n> but it's the other way around for another hook).\n\nThis I agree about.  Actually I'm skeptical about ordering\ndependencies being something that is meaningful for users to work with\nin general, except in the case of closely cooperating hook authors.\n\nThat doesn't mean we shouldn't try to futureproof for that, but I\ndon't think we need to overfit on it.\n\n[...]\n>>> And it doesn't leave a lot of room for defining\n>>> per-hook-type options; you have to make new keys like pre-push-order\n>>> (though that does work because the hook names are a finite set that\n>>> conforms to our config key names).\n\nExactly: field names like prePushOrder should work okay, even if\nthey're a bit noisy.\n\n[...]\n>>>   [hook \"pre-receive\"]\n>>>   # put any pre-receive related options here; e.g., a rule for what to\n>>>   # do with hook exit codes (e.g., stop running, run all but return exit\n>>>   # code, ignore failures, etc)\n>>>   fail = stop\n>>\n>> Interesting - so this is a default for all pre-receive hooks, that I can\n>> set at whichever scope I wish.\n\nIf I have the mental model of \"these are plugins, and particular hooks\nare events they listen to\", then it seems hard to make use of this\nbroader setting.\n\nBut scoped to a particular (plugin, event) pair it sounds very handy.\n\nMy two cents,\nJonathan\n"},{"id":"395483","messageId":"b4013e48-cd70-ddf1-8800-b239de9b56c3@gmail.com","threadId":"52980","inReplyTo":"20200414202738.GD1879688@coredump.intra.peff.net","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2020-04-15T10:01:00Z","receivedAt":"2020-04-15T10:01:17Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 14/04/2020 21:27, Jeff King wrote:\n> On Tue, Apr 14, 2020 at 12:24:18PM -0700, Emily Shaffer wrote:\n> \n>>> Without the redirection one could have\n>>>   hook.pre-commit.linter.command = my-command\n>>>   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n>> [...]\n>> We'd need to fudge one of these fields to include the extra section, I\n>> think. Unfortunate, because I find your example very tidy, but in\n>> practice maybe not very neat. The closest thing I can find to a nice way\n>> of writing it might be:\n>>\n>>   [hook.pre-commit \"linter\"]\n>>     command = my-command\n>>     before = check-whitespace\n>>   [hook.pre-commit \"check-whitespace\"]\n>>     command = 'git diff --check --cached'\n> \n> Syntactically the whole section between the outer dots is the\n> subsection. So it's:\n> \n>   [hook \"pre-commit.check-whitespace\"]\n>   command = ...\n> \n> And I don't think we want to change the config syntax at this point.\n> Even in the neater dotted notation, we must keep that whole thing as a\n> subsection, because existing subsections may contain dots, too.\n\nThanks for clarifying that, I agree we don't want to change the config\nsyntax and break existing subsections\n\nBest Wishes\n\nPhillip\n\n>> But this is kind of a lie; the sections aren't \"hook\", \"pre-commit\", and\n>> \"linter\" as you'd expect. Whether it's OK to lie like this, though, I\n>> don't know - I suspect it might make it awkward for others trying to\n>> parse the config. (my Vim syntax highlighter had kind of a hard time.)\n> \n> I think we should avoid it if possible. There are some subtleties there,\n> like the fact that subsections are case-sensitive, but sections and keys\n> are not.> -Peff\n> \n\n"},{"id":"395484","messageId":"0f661f31-ee75-15fb-0272-48d459176f29@gmail.com","threadId":"52980","inReplyTo":"20200414203247.GE1879688@coredump.intra.peff.net","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2020-04-15T10:01:15Z","receivedAt":"2020-04-15T10:01:25Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 14/04/2020 21:32, Jeff King wrote:\n> On Tue, Apr 14, 2020 at 04:15:11PM +0100, Phillip Wood wrote:\n> \n>> On 14/04/2020 01:54, Emily Shaffer wrote:\n>>> Not much to look at compared to the original RFC I sent some months ago.\n>>> This implements Peff's suggestion of using the \"hookcmd\" section as a\n>>> layer of indirection.\n>>\n>> I'm not really clear what the advantage of this indirection is. It seems\n>> unlikely to me that different hooks will share exactly the same command line\n>> or other options. In the 'git secrets' example earlier in this thread each\n>> hook needs to use a different command line. In general a command cannot tell\n>> which hook it is being invoked as without a flag of some kind. (In some\n>> cases it can use the number of arguments if that is different for each hook\n>> that it handles but that is not true in general)\n>>\n>> Without the redirection one could have\n>>   hook.pre-commit.linter.command = my-command\n>>   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n>>\n>> and other keys can be added for ordering etc. e.g.\n>>   hook.pre-commit.linter.before = check-whitespace\n>>\n>> With the indirection one needs to set\n>>   hook.pre-commit.command = linter\n>>   hook.pre-commit.check-whitespace = 'git diff --check --cached'\n>>   hookcmd.linter.command = my-command\n>>   hookcmd.linter.pre-commit-before = check-whitespace\n> \n> In the proposal I gave, you could do:\n> \n>   hook.pre-commit.command = my-command\n>   hook.pre-commit.command = git diff --check --cached\n> \n> If you want to refer to commands in ordering options (like your\n> \"before\"), then you'd have to refer to their names. For \"my-command\"\n> that's not too bad. For the longer one, it's a bit awkward. You _could_\n> do:\n> \n>   hookcmd.my-command.before = git diff --check --cached\n> \n> which is the same number of lines as yours. But I'd probably give it a\n> name, like:\n> \n>   hookcmd.check-whitespace.command = git diff --check --cached\n>   hookcmd.my-command.before = check-whitespace\n> \n> That's one more line than yours, but I think it separates the concerns\n> more clearly. And it extends naturally to more options specific to\n> check-whitespace.\n\nI agree that using a name rather than the command line makes things\nclearer here\n\nBest Wishes\n\nPhillip\n\n> -Peff\n> \n\n"},{"id":"395485","messageId":"ec6efbd4-4821-eefe-b16f-ab1ed8bc2058@gmail.com","threadId":"52980","inReplyTo":"20200414200347.GD12694@google.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2020-04-15T10:08:20Z","receivedAt":"2020-04-15T10:08:34Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 14/04/2020 21:03, Josh Steadmon wrote:\n> On 2020.04.14 16:15, Phillip Wood wrote:\n>> Hi Emily\n>>\n>> Thanks for working on this, having a way to manage multiple commands per\n>> hook without using an external framework would be really useful\n>>\n>> On 14/04/2020 01:54, Emily Shaffer wrote:\n>>> Not much to look at compared to the original RFC I sent some months ago.\n>>> This implements Peff's suggestion of using the \"hookcmd\" section as a\n>>> layer of indirection.\n>>\n>> I'm not really clear what the advantage of this indirection is. It seems\n>> unlikely to me that different hooks will share exactly the same command line\n>> or other options. In the 'git secrets' example earlier in this thread each\n>> hook needs to use a different command line. In general a command cannot tell\n>> which hook it is being invoked as without a flag of some kind. (In some\n>> cases it can use the number of arguments if that is different for each hook\n>> that it handles but that is not true in general)\n>>\n>> Without the redirection one could have\n>>   hook.pre-commit.linter.command = my-command\n>>   hook.pre-commit.check-whitespace.command = 'git diff --check --cached'\n>>\n>> and other keys can be added for ordering etc. e.g.\n>>   hook.pre-commit.linter.before = check-whitespace\n>>\n>> With the indirection one needs to set\n>>   hook.pre-commit.command = linter\n>>   hook.pre-commit.check-whitespace = 'git diff --check --cached'\n>>   hookcmd.linter.command = my-command\n>>   hookcmd.linter.pre-commit-before = check-whitespace\n>>\n>> which involves setting an extra key and checking it each time the hook is\n>> invoked without any benefit that I can see. I suspect which one seems more\n>> logical depends on how one thinks of setting hooks - I tend to think \"I want\n>> to set a pre-commit hook\" not \"I want to set a git-secrets hook\". If you've\n>> got an example where this indirection is helpful or necessary that would be\n>> really useful to see.\n>>\n>> Best Wishes\n>>\n>> Phillip\n> \n> Indexing repo content (see [1] for a detailed discussion) is one use\n> case where you have a single command that runs identically from\n> post-commit, post-merge, and post-checkout.\n\nThanks for sharing that, it is a useful reference point\n\n> Also, I suspect that many users don't have a firm enough grasp on the\n> various git hooks options to know ahead of time which ones they want to\n> set to accomplish a given task (without diving into the docs first). \n\nI agree with this, especially as setting up a hook is probably an\ninfrequent task for most people\n\n> I'm\n> not trying to say that your workflow is incorrect, but my gut feeling is\n> that most Git users would work in the opposite direction. Every time I\n> have needed to automate something, I generally had a rough script in\n> place first, and then looked up which hook(s) would be appropriate\n> triggers for the script.\n\nAs you say once they have a script they still have to look up which\nhooks they want to hook it up to, the indirection does not avoid that,\nit just means they have to lookup how to set up a hookcmd as well as\nwhich hooks they want to use.\n\nBest Wishes\n\nPhillip\n\n> [1]: https://tbaggery.com/2011/08/08/effortless-ctags-with-git.html\n> \n\n"},{"id":"395497","messageId":"xmqqd088950d.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"0f661f31-ee75-15fb-0272-48d459176f29@gmail.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-04-15T14:51:14Z","receivedAt":"2020-04-15T14:51:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n>> If you want to refer to commands in ordering options (like your\n>> \"before\"), then you'd have to refer to their names. For \"my-command\"\n>> that's not too bad. For the longer one, it's a bit awkward. You _could_\n>> do:\n>> \n>>   hookcmd.my-command.before = git diff --check --cached\n>> \n>> which is the same number of lines as yours. But I'd probably give it a\n>> name, like:\n>> \n>>   hookcmd.check-whitespace.command = git diff --check --cached\n>>   hookcmd.my-command.before = check-whitespace\n>> \n>> That's one more line than yours, but I think it separates the concerns\n>> more clearly. And it extends naturally to more options specific to\n>> check-whitespace.\n>\n> I agree that using a name rather than the command line makes things\n> clearer here\n\nTrue.   \n\nThese ways call for a different attitude to deal with errors\ncompared to the approach to order them with numbers, though.  \n\nIf your approach is to order by number attached to each hook, only\npossible errors you'd need to worry about are (1) what to do when\nthe user forgets to give a number to a hook and (2) what to do when\nthe user gives the same number by accident to multiple hooks, and\nboth can even be made non-errors by declaring that an unnumbered\nhook has a default number, and that two hooks with the same number\nexecute in an unspecified and unstable order.\n\nOn the other hand, the approach to specify relative ordering among\nhooks can break more easily.  E.g. when a hook that used to be\nbefore \"my-command\" got removed.  It is harder to find a \"sensible\"\ndefault behaviour for such situations.\n\nI am perfectly fine with having more possible error cases than\nallowing misconfigured system to silently do a wrong thing, so...\n\n\n\n"},{"id":"395530","messageId":"20200415203029.GA24777@google.com","threadId":"52980","inReplyTo":"xmqqd088950d.fsf@gitster.c.googlers.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-15T20:30:29Z","receivedAt":"2020-04-15T20:31:06Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, Apr 15, 2020 at 07:51:14AM -0700, Junio C Hamano wrote:\n> \n> Phillip Wood <phillip.wood123@gmail.com> writes:\n> \n> >> If you want to refer to commands in ordering options (like your\n> >> \"before\"), then you'd have to refer to their names. For \"my-command\"\n> >> that's not too bad. For the longer one, it's a bit awkward. You _could_\n> >> do:\n> >> \n> >>   hookcmd.my-command.before = git diff --check --cached\n> >> \n> >> which is the same number of lines as yours. But I'd probably give it a\n> >> name, like:\n> >> \n> >>   hookcmd.check-whitespace.command = git diff --check --cached\n> >>   hookcmd.my-command.before = check-whitespace\n> >> \n> >> That's one more line than yours, but I think it separates the concerns\n> >> more clearly. And it extends naturally to more options specific to\n> >> check-whitespace.\n> >\n> > I agree that using a name rather than the command line makes things\n> > clearer here\n> \n> True.   \n> \n> These ways call for a different attitude to deal with errors\n> compared to the approach to order them with numbers, though.  \n> \n> If your approach is to order by number attached to each hook, only\n> possible errors you'd need to worry about are (1) what to do when\n> the user forgets to give a number to a hook and (2) what to do when\n> the user gives the same number by accident to multiple hooks, and\n> both can even be made non-errors by declaring that an unnumbered\n> hook has a default number, and that two hooks with the same number\n> execute in an unspecified and unstable order.\n> \n> On the other hand, the approach to specify relative ordering among\n> hooks can break more easily.  E.g. when a hook that used to be\n> before \"my-command\" got removed.  It is harder to find a \"sensible\"\n> default behaviour for such situations.\n\nTo be clear, the examples listed (both numbered order and relational\norder) were more for illustration purposes. At the contributor summit, I\nthink Peff's suggestion was to stick with config ordering until we\ndiscover something more robust is needed, which is fine by me. At that\ntime, I don't see a problem with doing something like:\n\n[hook]\n  ordering = numerical\n\n[hookcmd \"my-command\"]\n  command = ~/my-command.sh\n  order = 001\n\n(which means others can still rely on config ordering if they want.)\n\nOr, to put it another way, I don't think we need to solve the config\nordering problem today - as long as we don't make it impossible for us\nto change tomorrow :)\n\n - Emily\n"},{"id":"395537","messageId":"20200415205941.GB24777@google.com","threadId":"52980","inReplyTo":"20200415034550.GB36683@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-15T20:59:41Z","receivedAt":"2020-04-15T20:59:50Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Tue, Apr 14, 2020 at 08:45:50PM -0700, Jonathan Nieder wrote:\n> \n> Hi,\n> \n> Jeff King wrote:\n> > On Mon, Apr 13, 2020 at 12:15:15PM -0700, Emily Shaffer wrote:\n> >> Jeff King wrote:\n> \n> >>> Yeah, giving each block a unique name lets you give them each an order.\n> >>> It seems kind of weird to me that you'd define multiple hook types for a\n> >>> given name.\n> >>\n> >> Not so odd - git-secrets configures itself for pre-commit,\n> >> prepare-commit-msg, and commit-msg-hook.\n> [...]\n> > Yeah, I do see how that use case makes sense. I wonder how common it is\n> > versus having separate one-off hooks.\n> \n> I think separately from the frequency question, we should look at the\n> \"what model do we want to present to the user\" question.\n> \n> It's not too unusual for a project with their source code in a Git\n> repository to have conventions they want to nudge users toward.  I'd\n> expect them to use a combination of hooks for this:\n> \n> \tprepare-commit-msg\n> \tcommit-msg\n> \tpre-push\n> \n> Git LFS installs multiple hooks:\n> \n> \tpre-push\n> \tpost-checkout\n> \tpost-commit\n> \tpost-merge\n> \n> git-secrets installs multiple hooks, as already mentioned.\n> \n> We've also had some instances over time of one hook replacing another,\n> to improve the interface.  A program wanting to install hooks would\n> then be likely to migrate from the older interface to the better one.\n\nI find this argument particularly compelling :)\n\n> \n> What I mean to get at is that I think thinking of them in terms of\n> individual hooks, the user model assumed by these programs is to think\n> of them as plugins hooking into Git.  The individual hooks are events\n> that the plugin listens on.  If I am trying to disable a plugin, I\n> don't want to have to learn which events it cared about.\n> \n> >                                       And whether setting the order\n> > priority for all hooks at once is that useful (e.g., I can easily\n> > imagine a case where the pre-commit hook for program A must go before B,\n> > but it's the other way around for another hook).\n> \n> This I agree about.  Actually I'm skeptical about ordering\n> dependencies being something that is meaningful for users to work with\n> in general, except in the case of closely cooperating hook authors.\n> \n> That doesn't mean we shouldn't try to futureproof for that, but I\n> don't think we need to overfit on it.\n> \n> [...]\n> >>> And it doesn't leave a lot of room for defining\n> >>> per-hook-type options; you have to make new keys like pre-push-order\n> >>> (though that does work because the hook names are a finite set that\n> >>> conforms to our config key names).\n> \n> Exactly: field names like prePushOrder should work okay, even if\n> they're a bit noisy.\n> \n> [...]\n> >>>   [hook \"pre-receive\"]\n> >>>   # put any pre-receive related options here; e.g., a rule for what to\n> >>>   # do with hook exit codes (e.g., stop running, run all but return exit\n> >>>   # code, ignore failures, etc)\n> >>>   fail = stop\n> >>\n> >> Interesting - so this is a default for all pre-receive hooks, that I can\n> >> set at whichever scope I wish.\n> \n> If I have the mental model of \"these are plugins, and particular hooks\n> are events they listen to\", then it seems hard to make use of this\n> broader setting.\n> \n> But scoped to a particular (plugin, event) pair it sounds very handy.\n\nStriking out on finding another place to fit into the thread, I wonder\nif the reason some of us are thinking \"I'm going to write a pre-receive\nhook\" rather than \"I'm going to write a linter hook\" may be because of\nthe prior single-script-per-hook limitation. As a result, when you want\nto add another function to your hook, you think, \"I'll modify my\npre-receive hook\". I think part of this RFC is a subtle paradigm shift\naway from hooks-as-units-of-work and towards hooks-as-events.\n\nThat observation doesn't really provide much guidance though, except\nmaybe to point out we should think about what the glossary entries would\nsay for terms like \"hook\" and \"hook command\" now... and I think figuring\nout those definitions might help us settle on what is most logical in\nthe config.\n\n(That makes me think I had better write a design doc next, before I get\ntoo much further with RFC patches. I made one pass at one a while ago,\nbut it was more focused on history and choosing between alternatives;\nsince we seem to have agreed on an approach, I'll make another attempt\nfocusing on design and definition instead. I'll try to have something to\nthe list by next week.)\n\n - Emily\n"},{"id":"395544","messageId":"xmqqv9m04cjo.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200415203029.GA24777@google.com","subject":"Re: [RFC PATCH v2 0/2] configuration-based hook management","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-04-15T22:19:39Z","receivedAt":"2020-04-15T22:19:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n> Or, to put it another way, I don't think we need to solve the config\n> ordering problem today - as long as we don't make it impossible for us\n> to change tomorrow :)\n\nOK.\n"},{"id":"395546","messageId":"20200415224244.GB3595509@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200415034550.GB36683@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-15T22:42:44Z","receivedAt":"2020-04-15T22:42:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 14, 2020 at 08:45:50PM -0700, Jonathan Nieder wrote:\n\n> > Yeah, I do see how that use case makes sense. I wonder how common it is\n> > versus having separate one-off hooks.\n> \n> I think separately from the frequency question, we should look at the\n> \"what model do we want to present to the user\" question.\n\nI sort of agree. The mental model is important, but we should avoid\npresenting a model that is overly complex to a user who only wants to do\nsimple things. So how common that simple thing is impacts the answer to\nyour question.\n\n> [...]\n> What I mean to get at is that I think thinking of them in terms of\n> individual hooks, the user model assumed by these programs is to think\n> of them as plugins hooking into Git.  The individual hooks are events\n> that the plugin listens on.  If I am trying to disable a plugin, I\n> don't want to have to learn which events it cared about.\n\nSure, I agree that's a perfectly reasonable mental model. But for\nsomebody who just wants to do a one-off hook, they're now saddled with a\nthing they don't care about: defining a plugin group for their hook.\n\nThe examples you gave are all reasonable, but personally I've never used\nanything other than one-off hooks.\n\nOn the other hand, I've very rarely used hooks at all myself.\n\nTo be clear, I don't _really_ care all that much, and this isn't a hill\nI particularly care to die on. I was mostly just clarifying my earlier\nsuggestion. (I _am_ somewhat amazed that the simple concept of \"I would\nlike to run this shell command instead of $GIT_DIR/hooks/foo\" has\ngenerated so much discussion. So really I am in favor of whatever lets\nme stop thinking about this as soon as possible).\n\n> >                                       And whether setting the order\n> > priority for all hooks at once is that useful (e.g., I can easily\n> > imagine a case where the pre-commit hook for program A must go before B,\n> > but it's the other way around for another hook).\n> \n> This I agree about.  Actually I'm skeptical about ordering\n> dependencies being something that is meaningful for users to work with\n> in general, except in the case of closely cooperating hook authors.\n>\n> That doesn't mean we shouldn't try to futureproof for that, but I\n> don't think we need to overfit on it.\n\nI share that skepticism (and also agree that avoiding painting ourselves\ninto a corner is the main thing).\n\n> >>> And it doesn't leave a lot of room for defining\n> >>> per-hook-type options; you have to make new keys like pre-push-order\n> >>> (though that does work because the hook names are a finite set that\n> >>> conforms to our config key names).\n> \n> Exactly: field names like prePushOrder should work okay, even if\n> they're a bit noisy.\n\nA side note:\n\nHere you've done a custom munging of pre-push into prePush. I'm fine\nwith that, but would we ever want to allow third-party scripts to define\ntheir own hooks using this mechanism? E.g., if there's a git-hooks\ncommand could I run \"git hooks run foo\" to run the foo hook? If so, then\nit might be simpler to just use the name as-is rather than defining the\nexact munging rules.\n\n-Peff\n"},{"id":"395547","messageId":"20200415224852.GC24777@google.com","threadId":"52980","inReplyTo":"20200415224244.GB3595509@coredump.intra.peff.net","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-15T22:48:52Z","receivedAt":"2020-04-15T22:49:02Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, Apr 15, 2020 at 06:42:44PM -0400, Jeff King wrote:\n> \n> On Tue, Apr 14, 2020 at 08:45:50PM -0700, Jonathan Nieder wrote:\n> > >>> And it doesn't leave a lot of room for defining\n> > >>> per-hook-type options; you have to make new keys like pre-push-order\n> > >>> (though that does work because the hook names are a finite set that\n> > >>> conforms to our config key names).\n> > \n> > Exactly: field names like prePushOrder should work okay, even if\n> > they're a bit noisy.\n> \n> A side note:\n> \n> Here you've done a custom munging of pre-push into prePush. I'm fine\n> with that, but would we ever want to allow third-party scripts to define\n> their own hooks using this mechanism? E.g., if there's a git-hooks\n> command could I run \"git hooks run foo\" to run the foo hook? If so, then\n> it might be simpler to just use the name as-is rather than defining the\n> exact munging rules.\n\nI did envision that kind of thing, or at very least something like\n`git hook --list --porcelain foo | xargs -n 1 sh -c`. When I saw\nJonathan's suggestion I wondered if using the hookname as is (pre-push)\nwas not idiomatic to the config, and maybe I should change it. But I\nwould rather leave it identical to the hookname, personally.\n\n - Emily\n"},{"id":"395549","messageId":"20200415225726.GA3600473@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200415224852.GC24777@google.com","subject":"Re: [TOPIC 2/17] Hooks in the future","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-15T22:57:26Z","receivedAt":"2020-04-15T22:57:38Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 15, 2020 at 03:48:52PM -0700, Emily Shaffer wrote:\n\n> > Here you've done a custom munging of pre-push into prePush. I'm fine\n> > with that, but would we ever want to allow third-party scripts to define\n> > their own hooks using this mechanism? E.g., if there's a git-hooks\n> > command could I run \"git hooks run foo\" to run the foo hook? If so, then\n> > it might be simpler to just use the name as-is rather than defining the\n> > exact munging rules.\n> \n> I did envision that kind of thing, or at very least something like\n> `git hook --list --porcelain foo | xargs -n 1 sh -c`. When I saw\n> Jonathan's suggestion I wondered if using the hookname as is (pre-push)\n> was not idiomatic to the config, and maybe I should change it. But I\n> would rather leave it identical to the hookname, personally.\n\nYou do still have to communicate to users of git-hook that their hook\nnames are limited to the characters used in config keys. But that seems\nsimpler to me than describing any special dash-and-capitalization\nconversion.\n\n-Peff\n"},{"id":"395650","messageId":"CAP8UFD3v_J3zGqHKa94d71QB82hTsX0MZasERB-jOnY3Ya-uJw@mail.gmail.com","threadId":"52980","inReplyTo":"20200318101825.GB1227946@coredump.intra.peff.net","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2020-04-17T09:41:48Z","receivedAt":"2020-04-17T09:42:03Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi Taylor and Peff,\n\nOn Wed, Mar 18, 2020 at 11:18 AM Jeff King <peff@peff.net> wrote:\n>\n> On Tue, Mar 17, 2020 at 02:39:05PM -0600, Taylor Blau wrote:\n>\n> > Of course, I would be happy to send along our patches. They are included\n> > in the series below, and correspond roughly to what we are running at\n> > GitHub. (For us, there have been a few more clean-ups and additional\n> > patches, but I squashed them into 2/2 below).\n\nThanks for the patches, and sorry for the delay in responding!\n\n> > The approach is roughly that we have:\n> >\n> >   - 'uploadpack.filter.allow' -> specifying the default for unspecified\n> >     filter choices, itself defaulting to true in order to maintain\n> >     backwards compatibility, and\n> >\n> >   - 'uploadpack.filter.<filter>.allow' -> specifying whether or not each\n> >     filter kind is allowed or not. (Originally this was given as 'git\n> >     config uploadpack.filter=blob:none.allow true', but this '=' is\n> >     ambiguous to configuration given over '-c', which itself uses an '='\n> >     to separate keys from values.)\n>\n> One thing that's a little ugly here is the embedded dot in the\n> subsection (i.e., \"filter.<filter>\"). It makes it look like a four-level\n> key, but really there is no such thing in Git.  But everything else we\n> tried was even uglier.\n>\n> I think we want to declare a real subsection for each filter and not\n> just \"uploadpack.filter.<filter>\". That gives us room to expand to other\n> config options besides \"allow\" later on if we need to.\n>\n> We don't want to claim \"uploadpack.allow\" and \"uploadpack.<filter>.allow\";\n> that's too generic.\n>\n> Likewise \"filter.allow\" is too generic.\n>\n> We could do \"uploadpackfilter.allow\" and \"uploadpackfilter.<filter>.allow\",\n> but that's both ugly _and_ separates these options from the rest of\n> uploadpack.*.\n\nWhat do you think about something like:\n\n[promisorFilter \"noBlobs\"]\n        type = blob:none\n        uploadpack = true # maybe \"allow\" could also mean \"true\" here\n        ...\n?\n\n> > I noted in the second patch that there is the unfortunate possibility of\n> > encountering a SIGPIPE when trying to write the ERR sideband back to a\n> > client who requested a non-supported filter. Peff and I have had some\n> > discussion off-list about resurrecting SZEDZER's work which makes room\n> > in the buffer by reading one packet back from the client when the server\n> > encounters a SIGPIPE. It is for this reason that I am marking the series\n> > as 'RFC'.\n>\n> For reference, the patch I was thinking of was this:\n>\n>   https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n\nAre you using the patches in this series with or without something\nlike the above patch? I am ok to resend this patch series including\nthe above patch (crediting Szeder) if you use something like it.\n\nThanks,\nChristian.\n"},{"id":"395663","messageId":"20200417174030.GB2103@syl.local","threadId":"52980","inReplyTo":"CAP8UFD3v_J3zGqHKa94d71QB82hTsX0MZasERB-jOnY3Ya-uJw@mail.gmail.com","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-04-17T17:40:30Z","receivedAt":"2020-04-17T17:40:36Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Apr 17, 2020 at 11:41:48AM +0200, Christian Couder wrote:\n> Hi Taylor and Peff,\n>\n> On Wed, Mar 18, 2020 at 11:18 AM Jeff King <peff@peff.net> wrote:\n> >\n> > On Tue, Mar 17, 2020 at 02:39:05PM -0600, Taylor Blau wrote:\n> >\n> > > Of course, I would be happy to send along our patches. They are included\n> > > in the series below, and correspond roughly to what we are running at\n> > > GitHub. (For us, there have been a few more clean-ups and additional\n> > > patches, but I squashed them into 2/2 below).\n>\n> Thanks for the patches, and sorry for the delay in responding!\n\nNo need to apologize. Clearly these had slipped my mind, too :).\n\n> > > The approach is roughly that we have:\n> > >\n> > >   - 'uploadpack.filter.allow' -> specifying the default for unspecified\n> > >     filter choices, itself defaulting to true in order to maintain\n> > >     backwards compatibility, and\n> > >\n> > >   - 'uploadpack.filter.<filter>.allow' -> specifying whether or not each\n> > >     filter kind is allowed or not. (Originally this was given as 'git\n> > >     config uploadpack.filter=blob:none.allow true', but this '=' is\n> > >     ambiguous to configuration given over '-c', which itself uses an '='\n> > >     to separate keys from values.)\n> >\n> > One thing that's a little ugly here is the embedded dot in the\n> > subsection (i.e., \"filter.<filter>\"). It makes it look like a four-level\n> > key, but really there is no such thing in Git.  But everything else we\n> > tried was even uglier.\n> >\n> > I think we want to declare a real subsection for each filter and not\n> > just \"uploadpack.filter.<filter>\". That gives us room to expand to other\n> > config options besides \"allow\" later on if we need to.\n> >\n> > We don't want to claim \"uploadpack.allow\" and \"uploadpack.<filter>.allow\";\n> > that's too generic.\n> >\n> > Likewise \"filter.allow\" is too generic.\n> >\n> > We could do \"uploadpackfilter.allow\" and \"uploadpackfilter.<filter>.allow\",\n> > but that's both ugly _and_ separates these options from the rest of\n> > uploadpack.*.\n>\n> What do you think about something like:\n>\n> [promisorFilter \"noBlobs\"]\n>         type = blob:none\n>         uploadpack = true # maybe \"allow\" could also mean \"true\" here\n>         ...\n> ?\n\nI'm not sure about introducing a layer of indirection here with\n\"noBlobs\". It's nice that it could perhaps be enabled/disabled for\ndifferent builtins (e.g., by adding 'revList = false', say), but I'm not\nconvinced that this is improving all of those cases, either.\n\nFor example, what happens if I have something like:\n\n  [uploadpack \"filter.tree\"]\n    maxDepth = 1\n    allow = true\n\nbut I want to use a different value of maxDepth for, say, rev-list? I'd\nrather have two sections (each for the 'tree' filter, but scoped to\n'upload-pack' and 'rev-list' separately) than write something like:\n\n  [promisorFilter \"treeDepth\"]\n          type = tree\n          uploadpack = true\n          uploadpackMaxDepth = 1\n          revList = true\n          revListMaxDepth = 0\n          ...\n\nSo, yeah, the current system is not great because it has the '.' in the\nsecond component. I am definitely eager to hear other suggestions about\nnaming it differently, but I think that the general structure is on\ntrack.\n\nOne thing that I can think of (other than replacing the '.' with another\ndelimiting character other than '=') is renaming the key from\n'uploadPack' to 'uploadPackFilter'. I believe that this was suggested by\nJuino (?) earlier in the thread. I think that it's a fine resolution to\nthis, but I'm also not opposed to what is currently written in too above patches.\n\n> > > I noted in the second patch that there is the unfortunate possibility of\n> > > encountering a SIGPIPE when trying to write the ERR sideband back to a\n> > > client who requested a non-supported filter. Peff and I have had some\n> > > discussion off-list about resurrecting SZEDZER's work which makes room\n> > > in the buffer by reading one packet back from the client when the server\n> > > encounters a SIGPIPE. It is for this reason that I am marking the series\n> > > as 'RFC'.\n> >\n> > For reference, the patch I was thinking of was this:\n> >\n> >   https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n>\n> Are you using the patches in this series with or without something\n> like the above patch? I am ok to resend this patch series including\n> the above patch (crediting Szeder) if you use something like it.\n\nWe're not using them, but without them we suffer from a problem that if\nwe can get a SIGPIPE when writing the \"sorry, I don't support that\nfilter\" message back to the client, then they won't receive it.\n\nSzeder's patches help address that issue by catching the SIGPIPE and\npopping off enough from the client buffer so that we can write the\nmessage out before dying.\n\nI appreciate your offer to resubmit the series on my behalf, but I was\nalready planning on doing this myself and wouldn't want to burden you\nwith another to-do. I'll be happy to take it on myself, probably within\na week or so.\n\n> Thanks,\n> Christian.\n\nThanks,\nTaylor\n"},{"id":"395664","messageId":"20200417180645.GJ1739940@coredump.intra.peff.net","threadId":"52980","inReplyTo":"20200417174030.GB2103@syl.local","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-04-17T18:06:45Z","receivedAt":"2020-04-17T18:06:47Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Apr 17, 2020 at 11:40:30AM -0600, Taylor Blau wrote:\n\n> > What do you think about something like:\n> >\n> > [promisorFilter \"noBlobs\"]\n> >         type = blob:none\n> >         uploadpack = true # maybe \"allow\" could also mean \"true\" here\n> >         ...\n> > ?\n> \n> I'm not sure about introducing a layer of indirection here with\n> \"noBlobs\". It's nice that it could perhaps be enabled/disabled for\n> different builtins (e.g., by adding 'revList = false', say), but I'm not\n> convinced that this is improving all of those cases, either.\n\nYeah, I don't like forcing the user to invent a subsection name. My\nfirst thought was to suggest:\n\n  [promisorFilter \"blob:none\"]\n  uploadpack = true\n\nbut your tree example shows why that gets awkward: there are more keys\nthan just \"allow this\".\n\n> One thing that I can think of (other than replacing the '.' with another\n> delimiting character other than '=') is renaming the key from\n> 'uploadPack' to 'uploadPackFilter'. I believe that this was suggested by\n\nYeah, that proposal isn't bad. To me the two viable options seem like:\n\n - uploadpack.filter.<filter>.*: this has the ugly fake multilevel\n   subsection, but stays under uploadpack.*\n\n - uploadpackfilter.<filter>.*: more natural subsection, but not grouped\n   syntactically with other uploadpack stuff\n\nI am actually leaning towards the second. It should make the parsing\ncode less confusing, and it's not like there aren't already other config\nsections that impact uploadpack.\n\n> > > For reference, the patch I was thinking of was this:\n> > >\n> > >   https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n> >\n> > Are you using the patches in this series with or without something\n> > like the above patch? I am ok to resend this patch series including\n> > the above patch (crediting Szeder) if you use something like it.\n> \n> We're not using them, but without them we suffer from a problem that if\n> we can get a SIGPIPE when writing the \"sorry, I don't support that\n> filter\" message back to the client, then they won't receive it.\n> \n> Szeder's patches help address that issue by catching the SIGPIPE and\n> popping off enough from the client buffer so that we can write the\n> message out before dying.\n\nI definitely think we should pursue that patch, but it really can be\ndone orthogonally. It's an existing bug that affects other instances\nwhere upload-pack returns an error. The tests can work around it with\n\"test_must_fail ok=sigpipe\" in the meantime.\n\n-Peff\n"},{"id":"395788","messageId":"20200420235310.94493-1-emilyshaffer@google.com","threadId":"52980","inReplyTo":"20200415205941.GB24777@google.com","subject":"[PATCH] doc: propose hooks managed by the config","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-20T23:53:10Z","receivedAt":"2020-04-20T23:53:23Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Begin a design document for config-based hooks, managed via git-hook.\nFocus on an overview of the implementation and motivation for design\ndecisions. Briefly discuss the alternatives considered before this\npoint. Also, attempt to redefine terms to fit into a multihook world.\n\nSigned-off-by: Emily Shaffer <emilyshaffer@google.com>\n---\nHi all,\n\nI wasn't sure whether it made more sense to leave the design doc in the\nconversation or not, but I figured it fit well into the context. I tried\nto also add relevant IDs to the \"References\" headers to this mail.\n\nHopefully this is complete enough that we can discuss it directly until\nwe feel comfortable getting ready for implementation. I'm planning to\nsend a reply today with some comments, too.\n\n - Emily\n\n Documentation/Makefile                        |   1 +\n .../technical/config-based-hooks.txt          | 317 ++++++++++++++++++\n 2 files changed, 318 insertions(+)\n create mode 100644 Documentation/technical/config-based-hooks.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 8fe829cc1b..301111f236 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -79,6 +79,7 @@ SP_ARTICLES += $(API_DOCS)\n TECH_DOCS += MyFirstContribution\n TECH_DOCS += MyFirstObjectWalk\n TECH_DOCS += SubmittingPatches\n+TECH_DOCS += technical/config-based-hooks\n TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\ndiff --git a/Documentation/technical/config-based-hooks.txt b/Documentation/technical/config-based-hooks.txt\nnew file mode 100644\nindex 0000000000..38893423be\n--- /dev/null\n+++ b/Documentation/technical/config-based-hooks.txt\n@@ -0,0 +1,317 @@\n+Configuration-based hook management\n+===================================\n+\n+== Motivation\n+\n+Treat hooks as a first-class citizen by replacing the .git/hook/hookname path as\n+the only source of hooks to execute, in a way which is friendly to users with\n+multiple repos which have similar needs.\n+\n+Redefine \"hook\" as an event rather than a single script, allowing users to\n+perform unrelated actions on a single event.\n+\n+Take a step closer to safety when copying zipped Git repositories from untrusted\n+users.\n+\n+Make it easier for users to discover Git's hook feature and automate their\n+workflows.\n+\n+== User interfaces\n+\n+=== Config schema\n+\n+Hooks can be introduced by editing the configuration manually. There are two new\n+sections added, `hook` and `hookcmd`.\n+\n+==== `hook`\n+\n+Primarily contains subsections for each hook event. These subsections define\n+hook command execution order; hook commands can be specified by passing the\n+command directly if no additional configuration is needed, or by passing the\n+name of a `hookcmd`. If Git does not find a `hookcmd` whose subsection matches\n+the value of the given command string, Git will try to execute the string\n+directly. Hook event subsections can also contain per-hook-event settings.\n+\n+Also contains top-level hook execution settings, for example,\n+`hook.warnHookDir`, `hook.runHookDir`, or `hook.disableAll`.\n+\n+----\n+[hook \"pre-commit\"]\n+  command = perl-linter\n+  command = /usr/bin/git-secrets --pre-commit\n+\n+[hook \"pre-applypatch\"]\n+  command = perl-linter\n+  error = ignore\n+\n+[hook]\n+  warnHookDir = true\n+  runHookDir = prompt\n+----\n+\n+==== `hookcmd`\n+\n+Defines a hook command and its attributes, which will be used when a hook event\n+occurs. Unqualified attributes are assumed to apply to this hook during all hook\n+events, but event-specific attributes can also be supplied. The example runs\n+`/usr/bin/lint-it --language=perl <args passed by Git>`, but for repos which\n+include this config, the hook command will be skipped for all events to which\n+it's normally subscribed _except_ `pre-commit`.\n+\n+----\n+[hookcmd \"perl-linter\"]\n+  command = /usr/bin/lint-it --language=perl\n+  skip = true\n+  pre-commit-skip = false\n+----\n+\n+=== Command-line API\n+\n+Users should be able to view, reorder, and create hook commands via the command\n+line. External tools should be able to view a list of hooks in the correct order\n+to run.\n+\n+*`git hook list <hook-event>`*\n+\n+*`git hook list (--system|--global|--local|--worktree)`*\n+\n+*`git hook edit <hook-event>`*\n+\n+*`git hook add <hook-command> <hook-event> <options...>`*\n+\n+=== Hook editor\n+\n+The tool which is presented by `git hook edit <hook-command>`. Ideally, this\n+tool should be easier to use than manually editing the config, and then produce\n+a concise config afterwards. It may take a form similar to `git rebase\n+--interactive`.\n+\n+== Implementation\n+\n+=== Library\n+\n+`hook.c` and `hook.h` are responsible for interacting with the config files. In\n+the case when the code generating a hook event doesn't have special concerns\n+about how to run the hooks, the hook library will provide a basic API to call\n+all hooks in config order with an `argv_array` provided by the code which\n+generates the hook event:\n+\n+*`int run_hooks(const char *hookname, struct argv_array *args)`*\n+\n+This call includes the hook command provided by `run-command.h:find_hook()`;\n+eventually, this legacy hook will be gated by a config `hook.runHookDir`. The\n+config is checked against a number of cases:\n+\n+- \"no\": the legacy hook will not be run\n+- \"interactive\": Git will prompt the user before running the legacy hook\n+- \"warn\": Git will print a warning to stderr before running the legacy hook\n+- \"yes\" (default): Git will silently run the legacy hook\n+\n+If `hook.runHookDir` is provided more than once, Git will use the most\n+restrictive setting provided, for security reasons.\n+\n+If the caller wants to do something more complicated, the hook library can also\n+provide a callback API:\n+\n+*`int for_each_hookcmd(const char *hookname, hookcmd_function *cb)`*\n+\n+Finally, to facilitate the builtin, the library will also provide the following\n+APIs to interact with the config:\n+\n+----\n+int set_hook_commands(const char *hookname, struct string_list *commands,\n+\tenum config_scope scope);\n+int set_hookcmd(const char *hookcmd, struct hookcmd options);\n+\n+int list_hook_commands(const char *hookname, struct string_list *commands);\n+int list_hooks_in_scope(enum config_scope scope, struct string_list *commands);\n+----\n+\n+`struct hookcmd` is expected to grow in size over time as more functionality is\n+added to hooks; so that other parts of the code don't need to understand the\n+config schema, `struct hookcmd` should contain logical values instead of string\n+pairs.\n+\n+----\n+struct hookcmd {\n+  const char *name;\n+  const char *command;\n+\n+  /* for illustration only; not planned at present */\n+  int parallelizable;\n+  const char *hookcmd_before;\n+  const char *hookcmd_after;\n+  enum recovery_action on_fail;\n+}\n+----\n+\n+=== Builtin\n+\n+`builtin/hook.c` is responsible for providing the frontend. It's responsible for\n+formatting user-provided data and then calling the library API to set the\n+configs as appropriate. The builtin frontend is not responsible for calling the\n+config directly, so that other areas of Git can rely on the hook library to\n+understand the most recent config schema for hooks.\n+\n+=== Migration path\n+\n+==== Stage 0\n+\n+Hooks are called by running `run-command.h:find_hook()` with the hookname and\n+executing the result. The hook library and builtin do not exist. Hooks only\n+exist as specially named scripts within `.git/hooks/`.\n+\n+==== Stage 1\n+\n+`git hook list --porcelain <hook-event>` is implemented. Users can replace their\n+`.git/hooks/<hook-event>` scripts with a trampoline based on `git hook list`'s\n+output. Modifier commands like `git hook add` and `git hook edit` can be\n+implemented around this time as well.\n+\n+==== Stage 2\n+\n+`hook.h:run_hooks()` is taught to include `run-command.h:find_hook()` at the\n+end; calls to `find_hook()` are replaced with calls to `run_hooks()`. Users can\n+opt-in to config-based hooks simply by creating some in their config; otherwise\n+users should remain unaffected by the change.\n+\n+==== Stage 3\n+\n+The call to `find_hook()` inside of `run_hooks()` learns to check for a config,\n+`hook.runHookDir`. Users can opt into managing their hooks completely via the\n+config this way.\n+\n+==== Stage 4\n+\n+`.git/hooks` is removed from the template and the hook directory is considered\n+deprecated. To avoid breaking older repos, the default of `hook.runHookDir` is\n+not changed, and `find_hook()` is not removed.\n+\n+== Caveats\n+\n+=== Security and repo config\n+\n+Part of the motivation behind this refactor is to mitigate hooks as an attack\n+vector;footnote:[https://lore.kernel.org/git/20171002234517.GV19555@aiede.mtv.corp.google.com/]\n+however, as the design stands, users can still provide hooks in the repo-level\n+config, which is included when a repo is zipped and sent elsewhere.  The\n+security of the repo-level config is still under discussion; this design\n+generally assumes the repo-level config is secure, which is not true yet. The\n+goal is to avoid an overcomplicated design to work around a problem which has\n+ceased to exist.\n+\n+=== Ease of use\n+\n+The config schema is nontrivial; that's why it's important for the `git hook`\n+modifier commands to be usable. Contributors with UX expertise are encouraged to\n+share their suggestions.\n+\n+== Alternative approaches\n+\n+A previous summary of alternatives exists in the\n+archives.footnote:[https://lore.kernel.org/git/20191116011125.GG22855@google.com]\n+\n+=== Status quo\n+\n+Today users can implement multihooks themselves by using a \"trampoline script\"\n+as their hook, and pointing that script to a directory or list of other scripts\n+they wish to run.\n+\n+=== Hook directories\n+\n+Other contributors have suggested Git learn about the existence of a directory\n+such as `.git/hooks/<hookname>.d` and execute those hooks in alphabetical order.\n+\n+=== Comparison table\n+\n+.Comparison of alternatives\n+|===\n+|Feature |Config-based hooks |Hook directories |Status quo\n+\n+|Supports multiple hooks\n+|Natively\n+|Natively\n+|With user effort\n+\n+|Safer for zipped repos\n+|A little\n+|No\n+|No\n+\n+|Previous hooks just work\n+|If configured\n+|Yes\n+|Yes\n+\n+|Can install one hook to many repos\n+|Yes\n+|No\n+|No\n+\n+|Discoverability\n+|Better (in `git help git`)\n+|Same as before\n+|Same as before\n+\n+|Hard to run unexpected hook\n+|If configured\n+|No\n+|No\n+|===\n+\n+== Future work\n+\n+=== Execution ordering\n+\n+We may find that config order is insufficient for some users; for example,\n+config order makes it difficult to add a new hook to the system or global config\n+which runs at the end of the hook list. A new ordering schema should be:\n+\n+1) Specified by a `hook.order` config, so that users will not unexpectedly see\n+their order change;\n+\n+2) Either dependency or numerically based.\n+\n+Dependency-based ordering is prone to classic linked-list problems, like a\n+cycles and handling of missing dependencies. But, it paves the way for enabling\n+parallelization if some tasks truly depend on others.\n+\n+Numerical ordering makes it tricky for Git to generate suggested ordering\n+numbers for each command, but is easy to determine a definitive order.\n+\n+=== Parallelization\n+\n+Users with many hooks might want to run them simultaneously, if the hooks don't\n+modify state; if one hook depends on another's output, then users will want to\n+specify those dependencies. If we decide to solve this problem, we may want to\n+look to modern build systems for inspiration on how to manage dependencies and\n+parallel tasks.\n+\n+=== Securing hookdir hooks\n+\n+With the design as written in this doc, it's still possible for a malicious user\n+to modify `.git/config` to include `hook.pre-receive.command = rm -rf /`, then\n+zip their repo and send it to another user. It may be necessary to teach Git to\n+only allow one-line hooks like this if they were configured outside of the local\n+scope; or another approach, like a list of safe projects, might be useful. It\n+may also be sufficient (or at least useful) to teach a `hook.disableAll` config\n+or similar flag to the Git executable.\n+\n+=== Submodule inheritance\n+\n+It's possible some submodules may want to run the identical set of hooks that\n+their superrepo runs. While a globally-configured hook set is helpful, it's not\n+a great solution for users who have multiple repos-with-submodules under the\n+same user. It would be useful for submodules to learn how to run hooks from\n+their superrepo's config, or inherit that hook setting.\n+\n+== Glossary\n+\n+*hook event*\n+\n+A point during Git's execution where user scripts may be run, for example,\n+_prepare-commit-msg_ or _pre-push_.\n+\n+*hook command*\n+\n+A user script or executable which will be run on one or more hook events.\n-- \n2.26.1.301.g55bc3eb7cb9-goog\n\n"},{"id":"395798","messageId":"20200421002248.GC236872@google.com","threadId":"52980","inReplyTo":"20200420235310.94493-1-emilyshaffer@google.com","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-21T00:22:48Z","receivedAt":"2020-04-21T00:22:56Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Mon, Apr 20, 2020 at 04:53:10PM -0700, Emily Shaffer wrote:\n> \n> Begin a design document for config-based hooks, managed via git-hook.\n> Focus on an overview of the implementation and motivation for design\n> decisions. Briefly discuss the alternatives considered before this\n> point. Also, attempt to redefine terms to fit into a multihook world.\n> \n> Signed-off-by: Emily Shaffer <emilyshaffer@google.com>\n> ---\n> Hi all,\n> \n> I wasn't sure whether it made more sense to leave the design doc in the\n> conversation or not, but I figured it fit well into the context. I tried\n> to also add relevant IDs to the \"References\" headers to this mail.\n> \n> Hopefully this is complete enough that we can discuss it directly until\n> we feel comfortable getting ready for implementation. I'm planning to\n> send a reply today with some comments, too.\n> \n>  - Emily\n> \n>  Documentation/Makefile                        |   1 +\n>  .../technical/config-based-hooks.txt          | 317 ++++++++++++++++++\n>  2 files changed, 318 insertions(+)\n>  create mode 100644 Documentation/technical/config-based-hooks.txt\n> \n> diff --git a/Documentation/Makefile b/Documentation/Makefile\n> index 8fe829cc1b..301111f236 100644\n> --- a/Documentation/Makefile\n> +++ b/Documentation/Makefile\n> @@ -79,6 +79,7 @@ SP_ARTICLES += $(API_DOCS)\n>  TECH_DOCS += MyFirstContribution\n>  TECH_DOCS += MyFirstObjectWalk\n>  TECH_DOCS += SubmittingPatches\n> +TECH_DOCS += technical/config-based-hooks\n>  TECH_DOCS += technical/hash-function-transition\n>  TECH_DOCS += technical/http-protocol\n>  TECH_DOCS += technical/index-format\n> diff --git a/Documentation/technical/config-based-hooks.txt b/Documentation/technical/config-based-hooks.txt\n> new file mode 100644\n> index 0000000000..38893423be\n> --- /dev/null\n> +++ b/Documentation/technical/config-based-hooks.txt\n> @@ -0,0 +1,317 @@\n> +Configuration-based hook management\n> +===================================\n> +\n> +== Motivation\n> +\n> +Treat hooks as a first-class citizen by replacing the .git/hook/hookname path as\n> +the only source of hooks to execute, in a way which is friendly to users with\n> +multiple repos which have similar needs.\n> +\n> +Redefine \"hook\" as an event rather than a single script, allowing users to\n> +perform unrelated actions on a single event.\n> +\n> +Take a step closer to safety when copying zipped Git repositories from untrusted\n> +users.\n> +\n> +Make it easier for users to discover Git's hook feature and automate their\n> +workflows.\n> +\n> +== User interfaces\n> +\n> +=== Config schema\n> +\n> +Hooks can be introduced by editing the configuration manually. There are two new\n> +sections added, `hook` and `hookcmd`.\n> +\n> +==== `hook`\n> +\n> +Primarily contains subsections for each hook event. These subsections define\n> +hook command execution order; hook commands can be specified by passing the\n> +command directly if no additional configuration is needed, or by passing the\n> +name of a `hookcmd`. If Git does not find a `hookcmd` whose subsection matches\n> +the value of the given command string, Git will try to execute the string\n> +directly. Hook event subsections can also contain per-hook-event settings.\n> +\n> +Also contains top-level hook execution settings, for example,\n> +`hook.warnHookDir`, `hook.runHookDir`, or `hook.disableAll`.\n> +\n> +----\n> +[hook \"pre-commit\"]\n> +  command = perl-linter\n> +  command = /usr/bin/git-secrets --pre-commit\n> +\n> +[hook \"pre-applypatch\"]\n> +  command = perl-linter\n> +  error = ignore\n> +\n> +[hook]\n> +  warnHookDir = true\n> +  runHookDir = prompt\n\nwhoops, just realized this doesn't match the proposal below. Wrote these\non different days :)\n\n> +----\n> +\n> +==== `hookcmd`\n> +\n> +Defines a hook command and its attributes, which will be used when a hook event\n> +occurs. Unqualified attributes are assumed to apply to this hook during all hook\n> +events, but event-specific attributes can also be supplied. The example runs\n> +`/usr/bin/lint-it --language=perl <args passed by Git>`, but for repos which\n> +include this config, the hook command will be skipped for all events to which\n> +it's normally subscribed _except_ `pre-commit`.\n> +\n> +----\n> +[hookcmd \"perl-linter\"]\n> +  command = /usr/bin/lint-it --language=perl\n> +  skip = true\n> +  pre-commit-skip = false\n> +----\n> +\n> +=== Command-line API\n> +\n> +Users should be able to view, reorder, and create hook commands via the command\n> +line. External tools should be able to view a list of hooks in the correct order\n> +to run.\n> +\n> +*`git hook list <hook-event>`*\n> +\n> +*`git hook list (--system|--global|--local|--worktree)`*\n> +\n> +*`git hook edit <hook-event>`*\n> +\n> +*`git hook add <hook-command> <hook-event> <options...>`*\n> +\n> +=== Hook editor\n> +\n> +The tool which is presented by `git hook edit <hook-command>`. Ideally, this\n> +tool should be easier to use than manually editing the config, and then produce\n> +a concise config afterwards. It may take a form similar to `git rebase\n> +--interactive`.\n\nThis section is a little thin because I'm hoping to meet with some UX\nfolks on our end and get a better suggestion. Suggestions welcome from\nupstream too - I'm having trouble coming up with anything that's better\nthan modifying the config directly.\n\n> +\n> +== Implementation\n> +\n> +=== Library\n> +\n> +`hook.c` and `hook.h` are responsible for interacting with the config files. In\n> +the case when the code generating a hook event doesn't have special concerns\n> +about how to run the hooks, the hook library will provide a basic API to call\n> +all hooks in config order with an `argv_array` provided by the code which\n> +generates the hook event:\n> +\n> +*`int run_hooks(const char *hookname, struct argv_array *args)`*\n> +\n> +This call includes the hook command provided by `run-command.h:find_hook()`;\n> +eventually, this legacy hook will be gated by a config `hook.runHookDir`. The\n> +config is checked against a number of cases:\n> +\n> +- \"no\": the legacy hook will not be run\n> +- \"interactive\": Git will prompt the user before running the legacy hook\n> +- \"warn\": Git will print a warning to stderr before running the legacy hook\n> +- \"yes\" (default): Git will silently run the legacy hook\n> +\n> +If `hook.runHookDir` is provided more than once, Git will use the most\n> +restrictive setting provided, for security reasons.\n> +\n> +If the caller wants to do something more complicated, the hook library can also\n> +provide a callback API:\n> +\n> +*`int for_each_hookcmd(const char *hookname, hookcmd_function *cb)`*\n\nAnother alternative is to do this by providing a linked-list of\nstructs or even just an ordered string_list; that means the caller\nbecomes responsible for config syntax and parallelization, which I\ndidn't want. I'm open to hearing more argument. (on the rest of the doc\ntoo... but also here. :) )\n\n> +\n> +Finally, to facilitate the builtin, the library will also provide the following\n> +APIs to interact with the config:\n> +\n> +----\n> +int set_hook_commands(const char *hookname, struct string_list *commands,\n> +\tenum config_scope scope);\n> +int set_hookcmd(const char *hookcmd, struct hookcmd options);\n> +\n> +int list_hook_commands(const char *hookname, struct string_list *commands);\n> +int list_hooks_in_scope(enum config_scope scope, struct string_list *commands);\n> +----\n> +\n> +`struct hookcmd` is expected to grow in size over time as more functionality is\n> +added to hooks; so that other parts of the code don't need to understand the\n> +config schema, `struct hookcmd` should contain logical values instead of string\n> +pairs.\n> +\n> +----\n> +struct hookcmd {\n> +  const char *name;\n> +  const char *command;\n> +\n> +  /* for illustration only; not planned at present */\n> +  int parallelizable;\n> +  const char *hookcmd_before;\n> +  const char *hookcmd_after;\n> +  enum recovery_action on_fail;\n> +}\n> +----\n> +\n> +=== Builtin\n> +\n> +`builtin/hook.c` is responsible for providing the frontend. It's responsible for\n> +formatting user-provided data and then calling the library API to set the\n> +configs as appropriate. The builtin frontend is not responsible for calling the\n> +config directly, so that other areas of Git can rely on the hook library to\n> +understand the most recent config schema for hooks.\n> +\n> +=== Migration path\n> +\n> +==== Stage 0\n> +\n> +Hooks are called by running `run-command.h:find_hook()` with the hookname and\n> +executing the result. The hook library and builtin do not exist. Hooks only\n> +exist as specially named scripts within `.git/hooks/`.\n> +\n> +==== Stage 1\n> +\n> +`git hook list --porcelain <hook-event>` is implemented. Users can replace their\n> +`.git/hooks/<hook-event>` scripts with a trampoline based on `git hook list`'s\n> +output. Modifier commands like `git hook add` and `git hook edit` can be\n> +implemented around this time as well.\n> +\n> +==== Stage 2\n> +\n> +`hook.h:run_hooks()` is taught to include `run-command.h:find_hook()` at the\n> +end; calls to `find_hook()` are replaced with calls to `run_hooks()`. Users can\n> +opt-in to config-based hooks simply by creating some in their config; otherwise\n> +users should remain unaffected by the change.\n> +\n> +==== Stage 3\n> +\n> +The call to `find_hook()` inside of `run_hooks()` learns to check for a config,\n> +`hook.runHookDir`. Users can opt into managing their hooks completely via the\n> +config this way.\n> +\n> +==== Stage 4\n> +\n> +`.git/hooks` is removed from the template and the hook directory is considered\n> +deprecated. To avoid breaking older repos, the default of `hook.runHookDir` is\n> +not changed, and `find_hook()` is not removed.\n> +\n> +== Caveats\n> +\n> +=== Security and repo config\n> +\n> +Part of the motivation behind this refactor is to mitigate hooks as an attack\n> +vector;footnote:[https://lore.kernel.org/git/20171002234517.GV19555@aiede.mtv.corp.google.com/]\n> +however, as the design stands, users can still provide hooks in the repo-level\n> +config, which is included when a repo is zipped and sent elsewhere.  The\n> +security of the repo-level config is still under discussion; this design\n> +generally assumes the repo-level config is secure, which is not true yet. The\n> +goal is to avoid an overcomplicated design to work around a problem which has\n> +ceased to exist.\n> +\n> +=== Ease of use\n> +\n> +The config schema is nontrivial; that's why it's important for the `git hook`\n> +modifier commands to be usable. Contributors with UX expertise are encouraged to\n> +share their suggestions.\n> +\n> +== Alternative approaches\n> +\n> +A previous summary of alternatives exists in the\n> +archives.footnote:[https://lore.kernel.org/git/20191116011125.GG22855@google.com]\n> +\n> +=== Status quo\n> +\n> +Today users can implement multihooks themselves by using a \"trampoline script\"\n> +as their hook, and pointing that script to a directory or list of other scripts\n> +they wish to run.\n> +\n> +=== Hook directories\n> +\n> +Other contributors have suggested Git learn about the existence of a directory\n> +such as `.git/hooks/<hookname>.d` and execute those hooks in alphabetical order.\n> +\n> +=== Comparison table\n> +\n> +.Comparison of alternatives\n> +|===\n> +|Feature |Config-based hooks |Hook directories |Status quo\n> +\n> +|Supports multiple hooks\n> +|Natively\n> +|Natively\n> +|With user effort\n> +\n> +|Safer for zipped repos\n> +|A little\n> +|No\n> +|No\n> +\n> +|Previous hooks just work\n> +|If configured\n> +|Yes\n> +|Yes\n> +\n> +|Can install one hook to many repos\n> +|Yes\n> +|No\n> +|No\n> +\n> +|Discoverability\n> +|Better (in `git help git`)\n> +|Same as before\n> +|Same as before\n> +\n> +|Hard to run unexpected hook\n> +|If configured\n> +|No\n> +|No\n> +|===\n\nPlease share more features that come to your mind; I took most of this\nlist from the RFC I sent last fall:\nhttps://lore.kernel.org/git/20191116011125.GG22855@google.com\n\n> +\n> +== Future work\n> +\n> +=== Execution ordering\n> +\n> +We may find that config order is insufficient for some users; for example,\n> +config order makes it difficult to add a new hook to the system or global config\n> +which runs at the end of the hook list. A new ordering schema should be:\n> +\n> +1) Specified by a `hook.order` config, so that users will not unexpectedly see\n> +their order change;\n> +\n> +2) Either dependency or numerically based.\n> +\n> +Dependency-based ordering is prone to classic linked-list problems, like a\n> +cycles and handling of missing dependencies. But, it paves the way for enabling\n> +parallelization if some tasks truly depend on others.\n> +\n> +Numerical ordering makes it tricky for Git to generate suggested ordering\n> +numbers for each command, but is easy to determine a definitive order.\n> +\n> +=== Parallelization\n> +\n> +Users with many hooks might want to run them simultaneously, if the hooks don't\n> +modify state; if one hook depends on another's output, then users will want to\n> +specify those dependencies. If we decide to solve this problem, we may want to\n> +look to modern build systems for inspiration on how to manage dependencies and\n> +parallel tasks.\n> +\n> +=== Securing hookdir hooks\n> +\n> +With the design as written in this doc, it's still possible for a malicious user\n> +to modify `.git/config` to include `hook.pre-receive.command = rm -rf /`, then\n> +zip their repo and send it to another user. It may be necessary to teach Git to\n> +only allow one-line hooks like this if they were configured outside of the local\n> +scope; or another approach, like a list of safe projects, might be useful. It\n> +may also be sufficient (or at least useful) to teach a `hook.disableAll` config\n> +or similar flag to the Git executable.\n> +\n> +=== Submodule inheritance\n> +\n> +It's possible some submodules may want to run the identical set of hooks that\n> +their superrepo runs. While a globally-configured hook set is helpful, it's not\n> +a great solution for users who have multiple repos-with-submodules under the\n> +same user. It would be useful for submodules to learn how to run hooks from\n> +their superrepo's config, or inherit that hook setting.\n> +\n> +== Glossary\n> +\n> +*hook event*\n> +\n> +A point during Git's execution where user scripts may be run, for example,\n> +_prepare-commit-msg_ or _pre-push_.\n> +\n> +*hook command*\n> +\n> +A user script or executable which will be run on one or more hook events.\n\nIf other terms in the design doc are surprising to you, let me know and\nI'll define them here too.\n\n> -- \n> 2.26.1.301.g55bc3eb7cb9-goog\n> \n"},{"id":"395802","messageId":"xmqqh7xdprcv.fsf@gitster.c.googlers.com","threadId":"52980","inReplyTo":"20200421002248.GC236872@google.com","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-04-21T01:20:00Z","receivedAt":"2020-04-21T01:20:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n> Whoops, just realized this doesn't match the proposal below. Wrote these\n> on different days :)\n\nIt often is a good idea to attempt writing anything in one sitting\nfor coherency, and proofread the result on a separate day before\nsending it out ;-)\n"},{"id":"395820","messageId":"CAP8UFD1Yb+8Ox=dDkMfdBqBqW20FLfSnOY9hWhvx=8dx8OfXrw@mail.gmail.com","threadId":"52980","inReplyTo":"20200417174030.GB2103@syl.local","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2020-04-21T12:17:14Z","receivedAt":"2020-04-21T12:17:29Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Apr 17, 2020 at 7:40 PM Taylor Blau <me@ttaylorr.com> wrote:\n>\n> On Fri, Apr 17, 2020 at 11:41:48AM +0200, Christian Couder wrote:\n\n> > What do you think about something like:\n> >\n> > [promisorFilter \"noBlobs\"]\n> >         type = blob:none\n> >         uploadpack = true # maybe \"allow\" could also mean \"true\" here\n> >         ...\n> > ?\n>\n> I'm not sure about introducing a layer of indirection here with\n> \"noBlobs\". It's nice that it could perhaps be enabled/disabled for\n> different builtins (e.g., by adding 'revList = false', say), but I'm not\n> convinced that this is improving all of those cases, either.\n>\n> For example, what happens if I have something like:\n>\n>   [uploadpack \"filter.tree\"]\n>     maxDepth = 1\n>     allow = true\n>\n> but I want to use a different value of maxDepth for, say, rev-list? I'd\n> rather have two sections (each for the 'tree' filter, but scoped to\n> 'upload-pack' and 'rev-list' separately) than write something like:\n>\n>   [promisorFilter \"treeDepth\"]\n>           type = tree\n>           uploadpack = true\n>           uploadpackMaxDepth = 1\n>           revList = true\n>           revListMaxDepth = 0\n>           ...\n\nYou can have two sections using:\n\n[promisorFilter \"treeDepth1\"]\n          type = tree\n          uploadpack = true\n          maxDepth = 1\n\n[promisorFilter \"treeDepth0\"]\n          type = tree\n          revList = true\n          maxDepth = 0\n\n(Of course \"treeDepth1\" for example could be also spelled\n\"treeDepthOneLevel\" or however the user prefers.)\n\n> So, yeah, the current system is not great because it has the '.' in the\n> second component. I am definitely eager to hear other suggestions about\n> naming it differently, but I think that the general structure is on\n> track.\n>\n> One thing that I can think of (other than replacing the '.' with another\n> delimiting character other than '=') is renaming the key from\n> 'uploadPack' to 'uploadPackFilter'.\n\nI don't like either of those very much. I think an upload-pack filter\nis not very different than a rev-list filter. They are all promisor\n(or partial clone) filter, so there is no real reason to differentiate\nat the top level of the key name hierarchy.\n\nI also think that users are likely to want to use the same filters for\nboth upload-pack filters and rev-list filters, so using 'uploadPack'\nor 'uploadPackFilter' might necessitate duplicating entries with other\nkeys for rev-list filters or other filters.\n\n> > > For reference, the patch I was thinking of was this:\n> > >\n> > >   https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n> >\n> > Are you using the patches in this series with or without something\n> > like the above patch? I am ok to resend this patch series including\n> > the above patch (crediting Szeder) if you use something like it.\n>\n> We're not using them, but without them we suffer from a problem that if\n> we can get a SIGPIPE when writing the \"sorry, I don't support that\n> filter\" message back to the client, then they won't receive it.\n>\n> Szeder's patches help address that issue by catching the SIGPIPE and\n> popping off enough from the client buffer so that we can write the\n> message out before dying.\n>\n> I appreciate your offer to resubmit the series on my behalf, but I was\n> already planning on doing this myself and wouldn't want to burden you\n> with another to-do. I'll be happy to take it on myself, probably within\n> a week or so.\n\nOk, I am happy that you will resubmit then.\n\nThanks,\nChristian.\n"},{"id":"395821","messageId":"CAP8UFD0kqSQAnpfUxqDn_qwigQZhq7zyxY_CZhd1nJzqHT1cqw@mail.gmail.com","threadId":"52980","inReplyTo":"20200417180645.GJ1739940@coredump.intra.peff.net","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2020-04-21T12:34:18Z","receivedAt":"2020-04-21T12:34:34Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Apr 17, 2020 at 8:06 PM Jeff King <peff@peff.net> wrote:\n>\n> On Fri, Apr 17, 2020 at 11:40:30AM -0600, Taylor Blau wrote:\n>\n> > > What do you think about something like:\n> > >\n> > > [promisorFilter \"noBlobs\"]\n> > >         type = blob:none\n> > >         uploadpack = true # maybe \"allow\" could also mean \"true\" here\n> > >         ...\n> > > ?\n> >\n> > I'm not sure about introducing a layer of indirection here with\n> > \"noBlobs\". It's nice that it could perhaps be enabled/disabled for\n> > different builtins (e.g., by adding 'revList = false', say), but I'm not\n> > convinced that this is improving all of those cases, either.\n>\n> Yeah, I don't like forcing the user to invent a subsection name. My\n> first thought was to suggest:\n>\n>   [promisorFilter \"blob:none\"]\n>   uploadpack = true\n>\n> but your tree example shows why that gets awkward: there are more keys\n> than just \"allow this\".\n\nI like your first thought better than something that starts with\n\"uploadPack\". And I think if we let people find a subsection name (as\nwhat I suggest) they might indeed end up with something like:\n\n[promisorFilter \"blob:none\"]\n     type = blob:none\n     uploadpack = true\n\nas they might lack inspiration. As filters are becoming more and more\ncomplex though, people might find it much simpler to use the\nsubsection name in commands if we let them do that. For example we\nalready allow:\n\ngit rev-list --filter=combine:<filter1>+<filter2>+...<filterN> ...\n\nwhich could be simplified to:\n\ngit rev-list --filter=combinedFilter ...\n\n(where \"combinedFilter\" is defined in the config with\n\"type=combine:<filter1>+<filter2>+...<filterN>\".)\n\n[...]\n\n> > We're not using them, but without them we suffer from a problem that if\n> > we can get a SIGPIPE when writing the \"sorry, I don't support that\n> > filter\" message back to the client, then they won't receive it.\n> >\n> > Szeder's patches help address that issue by catching the SIGPIPE and\n> > popping off enough from the client buffer so that we can write the\n> > message out before dying.\n>\n> I definitely think we should pursue that patch, but it really can be\n> done orthogonally. It's an existing bug that affects other instances\n> where upload-pack returns an error. The tests can work around it with\n> \"test_must_fail ok=sigpipe\" in the meantime.\n\nOk, maybe I will take a look a this one then.\n\nThanks,\nChristian.\n"},{"id":"395941","messageId":"20200422204114.GA4850@syl.local","threadId":"52980","inReplyTo":"CAP8UFD0kqSQAnpfUxqDn_qwigQZhq7zyxY_CZhd1nJzqHT1cqw@mail.gmail.com","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-04-22T20:41:14Z","receivedAt":"2020-04-22T20:41:21Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Tue, Apr 21, 2020 at 02:34:18PM +0200, Christian Couder wrote:\n> On Fri, Apr 17, 2020 at 8:06 PM Jeff King <peff@peff.net> wrote:\n> >\n> > On Fri, Apr 17, 2020 at 11:40:30AM -0600, Taylor Blau wrote:\n> >\n> > > > What do you think about something like:\n> > > >\n> > > > [promisorFilter \"noBlobs\"]\n> > > >         type = blob:none\n> > > >         uploadpack = true # maybe \"allow\" could also mean \"true\" here\n> > > >         ...\n> > > > ?\n> > >\n> > > I'm not sure about introducing a layer of indirection here with\n> > > \"noBlobs\". It's nice that it could perhaps be enabled/disabled for\n> > > different builtins (e.g., by adding 'revList = false', say), but I'm not\n> > > convinced that this is improving all of those cases, either.\n> >\n> > Yeah, I don't like forcing the user to invent a subsection name. My\n> > first thought was to suggest:\n> >\n> >   [promisorFilter \"blob:none\"]\n> >   uploadpack = true\n> >\n> > but your tree example shows why that gets awkward: there are more keys\n> > than just \"allow this\".\n>\n> I like your first thought better than something that starts with\n> \"uploadPack\". And I think if we let people find a subsection name (as\n> what I suggest) they might indeed end up with something like:\n>\n> [promisorFilter \"blob:none\"]\n>      type = blob:none\n>      uploadpack = true\n>\n> as they might lack inspiration. As filters are becoming more and more\n> complex though, people might find it much simpler to use the\n> subsection name in commands if we let them do that. For example we\n> already allow:\n>\n> git rev-list --filter=combine:<filter1>+<filter2>+...<filterN> ...\n>\n> which could be simplified to:\n>\n> git rev-list --filter=combinedFilter ...\n>\n> (where \"combinedFilter\" is defined in the config with\n> \"type=combine:<filter1>+<filter2>+...<filterN>\".)\n>\n> [...]\n\nI really think that we're getting ahead of ourselves here. For now, I\ndon't think that we have powerful enough filters that it makes sense to\nput them together with combine and give them meaningful names. At least,\nno one has asked about such a thing on the list, which I take to mean\nthat people don't have a use for it.\n\nI'm also skeptical about relying on named filters when working with a\nserver. If the server defines the filter names (as we at GitHub would\ndo under this proposal), then what use are they to the client? For the\nserver, I'm not at all convinced that this is beneficial: the extra\nlayer of indirection through the configuration makes this brittle and\nhard-to-follow.\n\nNot to mention that the server could just as easily spell out the whole\nfilter.\n\nI'm not trying to give you a too-simple proposal, but I think yours\nintroduces additional complexity that is trying to enable use-cases that\nwe don't have in practice.\n\n> > > We're not using them, but without them we suffer from a problem that if\n> > > we can get a SIGPIPE when writing the \"sorry, I don't support that\n> > > filter\" message back to the client, then they won't receive it.\n> > >\n> > > Szeder's patches help address that issue by catching the SIGPIPE and\n> > > popping off enough from the client buffer so that we can write the\n> > > message out before dying.\n> >\n> > I definitely think we should pursue that patch, but it really can be\n> > done orthogonally. It's an existing bug that affects other instances\n> > where upload-pack returns an error. The tests can work around it with\n> > \"test_must_fail ok=sigpipe\" in the meantime.\n>\n> Ok, maybe I will take a look a this one then.\n\nThanks.\n\n> Thanks,\n> Christian.\n\nThanks,\nTaylor\n"},{"id":"395942","messageId":"20200422204252.GB4850@syl.local","threadId":"52980","inReplyTo":"20200417180645.GJ1739940@coredump.intra.peff.net","subject":"Re: [RFC PATCH 0/2] upload-pack.c: limit allowed filter choices","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-04-22T20:42:52Z","receivedAt":"2020-04-22T20:42:56Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Apr 17, 2020 at 02:06:45PM -0400, Jeff King wrote:\n> On Fri, Apr 17, 2020 at 11:40:30AM -0600, Taylor Blau wrote:\n>\n> > > What do you think about something like:\n> > >\n> > > [promisorFilter \"noBlobs\"]\n> > >         type = blob:none\n> > >         uploadpack = true # maybe \"allow\" could also mean \"true\" here\n> > >         ...\n> > > ?\n> >\n> > I'm not sure about introducing a layer of indirection here with\n> > \"noBlobs\". It's nice that it could perhaps be enabled/disabled for\n> > different builtins (e.g., by adding 'revList = false', say), but I'm not\n> > convinced that this is improving all of those cases, either.\n>\n> Yeah, I don't like forcing the user to invent a subsection name. My\n> first thought was to suggest:\n>\n>   [promisorFilter \"blob:none\"]\n>   uploadpack = true\n>\n> but your tree example shows why that gets awkward: there are more keys\n> than just \"allow this\".\n>\n> > One thing that I can think of (other than replacing the '.' with another\n> > delimiting character other than '=') is renaming the key from\n> > 'uploadPack' to 'uploadPackFilter'. I believe that this was suggested by\n>\n> Yeah, that proposal isn't bad. To me the two viable options seem like:\n>\n>  - uploadpack.filter.<filter>.*: this has the ugly fake multilevel\n>    subsection, but stays under uploadpack.*\n>\n>  - uploadpackfilter.<filter>.*: more natural subsection, but not grouped\n>    syntactically with other uploadpack stuff\n>\n> I am actually leaning towards the second. It should make the parsing\n> code less confusing, and it's not like there aren't already other config\n> sections that impact uploadpack.\n\nMe too.\n\n> > > > For reference, the patch I was thinking of was this:\n> > > >\n> > > >   https://lore.kernel.org/git/20190830121005.GI8571@szeder.dev/\n> > >\n> > > Are you using the patches in this series with or without something\n> > > like the above patch? I am ok to resend this patch series including\n> > > the above patch (crediting Szeder) if you use something like it.\n> >\n> > We're not using them, but without them we suffer from a problem that if\n> > we can get a SIGPIPE when writing the \"sorry, I don't support that\n> > filter\" message back to the client, then they won't receive it.\n> >\n> > Szeder's patches help address that issue by catching the SIGPIPE and\n> > popping off enough from the client buffer so that we can write the\n> > message out before dying.\n>\n> I definitely think we should pursue that patch, but it really can be\n> done orthogonally. It's an existing bug that affects other instances\n> where upload-pack returns an error. The tests can work around it with\n> \"test_must_fail ok=sigpipe\" in the meantime.\n\nYes, I agree. My main hesitation is that it would be uncouth of me to\nsend a patch that includes 'test_must_fail ok=sigpipe' to the list, but\nif you (and others) feel that this is an OK intermediate step (given\nthat we can easily remove it once SZEDER's patch lands), then I am OK\nwith it, too.\n\nAnd I see that Christian already posted such a patch to the list.\n\n> -Peff\n\nThanks,\nTaylor\n"},{"id":"396205","messageId":"20200424231403.GD236872@google.com","threadId":"52980","inReplyTo":"xmqqh7xdprcv.fsf@gitster.c.googlers.com","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-04-24T23:14:03Z","receivedAt":"2020-04-24T23:14:13Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Mon, Apr 20, 2020 at 06:20:00PM -0700, Junio C Hamano wrote:\n> \n> Emily Shaffer <emilyshaffer@google.com> writes:\n> \n> > Whoops, just realized this doesn't match the proposal below. Wrote these\n> > on different days :)\n> \n> It often is a good idea to attempt writing anything in one sitting\n> for coherency, and proofread the result on a separate day before\n> sending it out ;-)\n\nAgreed for next time :)\n\nI didn't make it very clear in my initial comment that the only problem\nhere is the code snippets and the difference is very minor - I don't\nthink it's worth a reroll on its own without hearing feedback about the\nrest. Or, to put it another way, if any interested reader said \"I'll\nwait to review\" - don't ;)\n\n - Emily\n"},{"id":"396269","messageId":"20200425205727.GB6421@camp.crustytoothpaste.net","threadId":"52980","inReplyTo":"20200420235310.94493-1-emilyshaffer@google.com","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2020-04-25T20:57:27Z","receivedAt":"2020-04-25T20:57:37Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2020-04-20 at 23:53:10, Emily Shaffer wrote:\n> +=== Config schema\n> +\n> +Hooks can be introduced by editing the configuration manually. There are two new\n> +sections added, `hook` and `hookcmd`.\n> +\n> +==== `hook`\n> +\n> +Primarily contains subsections for each hook event. These subsections define\n> +hook command execution order; hook commands can be specified by passing the\n> +command directly if no additional configuration is needed, or by passing the\n> +name of a `hookcmd`. If Git does not find a `hookcmd` whose subsection matches\n> +the value of the given command string, Git will try to execute the string\n> +directly. Hook event subsections can also contain per-hook-event settings.\n\nCan we say explicitly that the commands are invoked by the shell?  Or is\nthe plan to try to parse them without passing to the shell?\n\n> +Also contains top-level hook execution settings, for example,\n> +`hook.warnHookDir`, `hook.runHookDir`, or `hook.disableAll`.\n> +\n> +----\n> +[hook \"pre-commit\"]\n> +  command = perl-linter\n> +  command = /usr/bin/git-secrets --pre-commit\n> +\n> +[hook \"pre-applypatch\"]\n> +  command = perl-linter\n> +  error = ignore\n> +\n> +[hook]\n> +  warnHookDir = true\n> +  runHookDir = prompt\n> +----\n> +\n> +==== `hookcmd`\n> +\n> +Defines a hook command and its attributes, which will be used when a hook event\n> +occurs. Unqualified attributes are assumed to apply to this hook during all hook\n> +events, but event-specific attributes can also be supplied. The example runs\n> +`/usr/bin/lint-it --language=perl <args passed by Git>`, but for repos which\n> +include this config, the hook command will be skipped for all events to which\n> +it's normally subscribed _except_ `pre-commit`.\n> +\n> +----\n> +[hookcmd \"perl-linter\"]\n> +  command = /usr/bin/lint-it --language=perl\n> +  skip = true\n> +  pre-commit-skip = false\n> +----\n\nThis seems fine to me.  I like this design and it seems sane.\n\n> +== Implementation\n> +\n> +=== Library\n> +\n> +`hook.c` and `hook.h` are responsible for interacting with the config files. In\n> +the case when the code generating a hook event doesn't have special concerns\n> +about how to run the hooks, the hook library will provide a basic API to call\n> +all hooks in config order with an `argv_array` provided by the code which\n> +generates the hook event:\n> +\n> +*`int run_hooks(const char *hookname, struct argv_array *args)`*\n> +\n> +This call includes the hook command provided by `run-command.h:find_hook()`;\n> +eventually, this legacy hook will be gated by a config `hook.runHookDir`. The\n> +config is checked against a number of cases:\n> +\n> +- \"no\": the legacy hook will not be run\n> +- \"interactive\": Git will prompt the user before running the legacy hook\n> +- \"warn\": Git will print a warning to stderr before running the legacy hook\n> +- \"yes\" (default): Git will silently run the legacy hook\n> +\n> +If `hook.runHookDir` is provided more than once, Git will use the most\n> +restrictive setting provided, for security reasons.\n\nI don't think this is consistent with the way the rest of our options\nwork.  What if someone generally wants to disable legacy hooks but then\nworks with a program in a repository that requires them?\n\n> +== Caveats\n> +\n> +=== Security and repo config\n> +\n> +Part of the motivation behind this refactor is to mitigate hooks as an attack\n> +vector;footnote:[https://lore.kernel.org/git/20171002234517.GV19555@aiede.mtv.corp.google.com/]\n> +however, as the design stands, users can still provide hooks in the repo-level\n> +config, which is included when a repo is zipped and sent elsewhere.  The\n> +security of the repo-level config is still under discussion; this design\n> +generally assumes the repo-level config is secure, which is not true yet. The\n> +goal is to avoid an overcomplicated design to work around a problem which has\n> +ceased to exist.\n\nI want to be clear that I'm very much opposed to trying to \"secure\" the\nconfig as a whole.  I believe that it's going to ultimately lead to a\nvariety of new and interesting attack vectors and will lead to Git\nbecoming a CVE factory.  Vim has this problem with modelines, for\nexample.\n\nI think we should maintain the status quo that the only safe things you\ncan do with an untrusted repository are clone and fetch because it sets\na clear security boundary.\n\nHaving said that, I'm otherwise pretty happy with this design and I'm\nlooking forward to seeing it implemented.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"397239","messageId":"20200506213354.GG77802@google.com","threadId":"52980","inReplyTo":"20200425205727.GB6421@camp.crustytoothpaste.net","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-05-06T21:33:54Z","receivedAt":"2020-05-06T21:34:02Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Sat, Apr 25, 2020 at 08:57:27PM +0000, brian m. carlson wrote:\n> \n> On 2020-04-20 at 23:53:10, Emily Shaffer wrote:\n> > +=== Config schema\n> > +\n> > +Hooks can be introduced by editing the configuration manually. There are two new\n> > +sections added, `hook` and `hookcmd`.\n> > +\n> > +==== `hook`\n> > +\n> > +Primarily contains subsections for each hook event. These subsections define\n> > +hook command execution order; hook commands can be specified by passing the\n> > +command directly if no additional configuration is needed, or by passing the\n> > +name of a `hookcmd`. If Git does not find a `hookcmd` whose subsection matches\n> > +the value of the given command string, Git will try to execute the string\n> > +directly. Hook event subsections can also contain per-hook-event settings.\n> \n> Can we say explicitly that the commands are invoked by the shell?  Or is\n> the plan to try to parse them without passing to the shell?\n\nSure. If I didn't make it clear it was by mistake, not by intent.\n\n> \n> > +Also contains top-level hook execution settings, for example,\n> > +`hook.warnHookDir`, `hook.runHookDir`, or `hook.disableAll`.\n> > +\n> > +----\n> > +[hook \"pre-commit\"]\n> > +  command = perl-linter\n> > +  command = /usr/bin/git-secrets --pre-commit\n> > +\n> > +[hook \"pre-applypatch\"]\n> > +  command = perl-linter\n> > +  error = ignore\n> > +\n> > +[hook]\n> > +  warnHookDir = true\n> > +  runHookDir = prompt\n> > +----\n> > +\n> > +==== `hookcmd`\n> > +\n> > +Defines a hook command and its attributes, which will be used when a hook event\n> > +occurs. Unqualified attributes are assumed to apply to this hook during all hook\n> > +events, but event-specific attributes can also be supplied. The example runs\n> > +`/usr/bin/lint-it --language=perl <args passed by Git>`, but for repos which\n> > +include this config, the hook command will be skipped for all events to which\n> > +it's normally subscribed _except_ `pre-commit`.\n> > +\n> > +----\n> > +[hookcmd \"perl-linter\"]\n> > +  command = /usr/bin/lint-it --language=perl\n> > +  skip = true\n> > +  pre-commit-skip = false\n> > +----\n> \n> This seems fine to me.  I like this design and it seems sane.\n> \n> > +== Implementation\n> > +\n> > +=== Library\n> > +\n> > +`hook.c` and `hook.h` are responsible for interacting with the config files. In\n> > +the case when the code generating a hook event doesn't have special concerns\n> > +about how to run the hooks, the hook library will provide a basic API to call\n> > +all hooks in config order with an `argv_array` provided by the code which\n> > +generates the hook event:\n> > +\n> > +*`int run_hooks(const char *hookname, struct argv_array *args)`*\n> > +\n> > +This call includes the hook command provided by `run-command.h:find_hook()`;\n> > +eventually, this legacy hook will be gated by a config `hook.runHookDir`. The\n> > +config is checked against a number of cases:\n> > +\n> > +- \"no\": the legacy hook will not be run\n> > +- \"interactive\": Git will prompt the user before running the legacy hook\n> > +- \"warn\": Git will print a warning to stderr before running the legacy hook\n> > +- \"yes\" (default): Git will silently run the legacy hook\n> > +\n> > +If `hook.runHookDir` is provided more than once, Git will use the most\n> > +restrictive setting provided, for security reasons.\n> \n> I don't think this is consistent with the way the rest of our options\n> work.  What if someone generally wants to disable legacy hooks but then\n> works with a program in a repository that requires them?\n\nUnfortunately this is something I think my end will want to hold firm\non. In general we disagree with your statement later about not wanting\nto make the .git/config secure. I see your use case, and I anticipate\ntwo possible workarounds I'd present:\n\n1) If working in that repo for the short term, run `git -c\nhook.runHookDir=yes <command> <arg...>` (and therefore allow the config\nfrom command line scope, which I'm happy with in general). Maybe\nsomeone would want to use an alias, hookgit or hg? Just kidding.. ;P\n\n2) If you're stuck with that repo for the long term, add\n`hook.<hookname>.command = /path/.git/hooks/<hookname>` lines to the local\nconfig.\n\nYes, those are both somewhat user-unfriendly, and I think we can do\nbetter... I'll have to think more and see what I can come up with.\nSuggestions welcome.\n\n> \n> > +== Caveats\n> > +\n> > +=== Security and repo config\n> > +\n> > +Part of the motivation behind this refactor is to mitigate hooks as an attack\n> > +vector;footnote:[https://lore.kernel.org/git/20171002234517.GV19555@aiede.mtv.corp.google.com/]\n> > +however, as the design stands, users can still provide hooks in the repo-level\n> > +config, which is included when a repo is zipped and sent elsewhere.  The\n> > +security of the repo-level config is still under discussion; this design\n> > +generally assumes the repo-level config is secure, which is not true yet. The\n> > +goal is to avoid an overcomplicated design to work around a problem which has\n> > +ceased to exist.\n> \n> I want to be clear that I'm very much opposed to trying to \"secure\" the\n> config as a whole.  I believe that it's going to ultimately lead to a\n> variety of new and interesting attack vectors and will lead to Git\n> becoming a CVE factory.  Vim has this problem with modelines, for\n> example.\n\nI'm really interested to hear more - it seems like security and config\nefforts will end up on my plate before the end of the year, so I'd like\nto know what is on your mind.\n\n> \n> I think we should maintain the status quo that the only safe things you\n> can do with an untrusted repository are clone and fetch because it sets\n> a clear security boundary.\n\nI wish there was a way to make that more apparent. The trouble is that\nwhile you and I and the sysadmin know the dangers, the high schooler\nmaking a website might not. Talking about how to warn users is\ndefinitely out-of-scope for this conversation, but it's on my mind.\n\n> \n> Having said that, I'm otherwise pretty happy with this design and I'm\n> looking forward to seeing it implemented.\n\nThanks very much for the feedback and for reading it through! :)\n\n - Emily\n"},{"id":"397249","messageId":"20200506231316.GA7234@camp.crustytoothpaste.net","threadId":"52980","inReplyTo":"20200506213354.GG77802@google.com","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2020-05-06T23:13:16Z","receivedAt":"2020-05-06T23:14:02Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2020-05-06 at 21:33:54, Emily Shaffer wrote:\n> On Sat, Apr 25, 2020 at 08:57:27PM +0000, brian m. carlson wrote:\n> > \n> > On 2020-04-20 at 23:53:10, Emily Shaffer wrote:\n> > > +== Caveats\n> > > +\n> > > +=== Security and repo config\n> > > +\n> > > +Part of the motivation behind this refactor is to mitigate hooks as an attack\n> > > +vector;footnote:[https://lore.kernel.org/git/20171002234517.GV19555@aiede.mtv.corp.google.com/]\n> > > +however, as the design stands, users can still provide hooks in the repo-level\n> > > +config, which is included when a repo is zipped and sent elsewhere.  The\n> > > +security of the repo-level config is still under discussion; this design\n> > > +generally assumes the repo-level config is secure, which is not true yet. The\n> > > +goal is to avoid an overcomplicated design to work around a problem which has\n> > > +ceased to exist.\n> > \n> > I want to be clear that I'm very much opposed to trying to \"secure\" the\n> > config as a whole.  I believe that it's going to ultimately lead to a\n> > variety of new and interesting attack vectors and will lead to Git\n> > becoming a CVE factory.  Vim has this problem with modelines, for\n> > example.\n> \n> I'm really interested to hear more - it seems like security and config\n> efforts will end up on my plate before the end of the year, so I'd like\n> to know what is on your mind.\n\nIn general, having untrusted configuration is enormously difficult and\nis typically only possible as a designed-in feature with extremely\nlimited options.  We have not designed that feature in from the\nbeginning and our config parsing is far too ad-hoc to support any\nreasonable security posture.  We've also written a program entirely in\nC, which has all of the fun memory safety problems.\n\nIf we try to secure the config and allow people to use untrusted\nrepositories securely, we've changed the security posture of the project\nvery significantly.  The number of keys we can safely trust come down to\nprobably core.repositoryformatversion and extensions.objectformat, and\nI'm not even sure that the latter can be trusted because there are all\nsorts of fun behaviors one can produce by setting the wrong hash\nalgorithm.\n\nThat's just one example of a potential source of security problems, but\nI anticipate people can use other options as well.  Setting the rename\nlimit can be a DoS.  Changing the colors of diff or log output could be\nused to hide malicious code from inspection.  We obviously can't trust\nanything containing a URL, since an attacker could try to make \"git pull\norigin\" point to their server instead, which means having remotes is out\nof the question.  Most of our recent security issues have involved the\n.gitmodules file, which, despite being extremely limited, is indeed an\nuntrusted config file.\n\nThe scope of potential vulnerabilities explodes as you allow users to\nhave untrusted config.  I don't think there's any reasonable set of\nuseful configuration we can have on a per-repo basis that doesn't open\nus up to a whole set of security vulnerabilities.  It seems to me that\nwe're setting ourselves up to either have a feature so limited nobody\nuses it or a massive, never-ending set of CVEs as everybody finds new\nways to attack things.  I just don't think promising that feature to\nusers is honest because I don't think we can practically achieve it in\nGit.  Most projects don't even try it as an option.\n\nOn the other hand, what we promise now, which is to restrict untrusted\nrepositories to cloning and fetching, while surprising to users,\ndramatically reduces the scope because it's basically what we promise\nover the network.  The interface is highly restricted, well known, and\nreasonably secure.  We've also limited attack surface to a much smaller\nnumber of binaries.\n\nSo while I think the intention is good and the idea, if implementable,\nwould be beneficial to users, I think it's practically going to be\nunachievable.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"397471","messageId":"49bf5848-b8b0-13b0-4038-352f702d11ee@gmail.com","threadId":"52980","inReplyTo":"9CE46D29-4BCD-4E95-B2DA-939EA10D7934@jramsay.com.au","subject":"Re: [TOPIC 9/17] Obsolescence markers and evolve","fromName":"Noam Soloveichik","fromEmail":"inoamsol@gmail.com","sentAt":"2020-05-09T21:31:37Z","receivedAt":"2020-05-09T21:31:47Z","isPatch":false,"sender":{"key":"inoamsol@gmail.com","avatar":null},"body":"\nOn 12/03/2020 6:04, James Ramsay wrote:\n> 1. Brandon: I thought it would be interesting to have a similar feature\n> as Mercurial has. Mercurial evolve will help you do a big rebase\n> commit by commit. Giving you more insights how commits change over time.\n>\n> 2. Peff: This has been discussed a lot of time on the list already.\nSince I'm very interested in this topic, can you link me to some key\ndiscussions you remember? Most of what I've found is Stefan Xenos having\na take on implementing it.\n\nAlso, I read some discussions on tools trying to record rebase history\nsuch as\ngit-series, available on GitHub.\n> 3. Jonathan N: It will help with Googlers productivity, but it’s\n> smaller compared to other performance fixes.\n>\n> 4. Brian: It’s a great feature and I would like to have it, but I’m\n> not sure it gives enough value to someone to sit down and implement it.\nI personally am very interested in this and consider contributing to it\nmyself,\nalthough it sounds very complex and intricate to perfect.\n> 5. Emily: Is it a good candidate for GSoC?\n>\n> 6. Brian: If we have a good design.\n\nThree's a design proposal from last year:\nhttps://public-inbox.org/git/20190215043105.163688-1-sxenos@google.com/\n\nDid you get a chance to have a look at it?\n\n> 7. Stolee: It should be easier to use than interactive rebase.\n>\n> 8. Stolee: It would be nice to have instead of fixup commits I would\n> send to you new commits which mark your original commits are obsolete.\n\nHi all,\n\nSeeing this post really encourages me to try my best and tackle this issue.\nThe main reason I'm interested in this is that such a feature encourages\npeople to craft very high quality commit history.\n\nIt's a great relief to see people discussing the topic, since I've looked\nover the web and did not find much talk about it so far.\n\nI want to contribute but since it's a big feature proposal with project-wide\nconsequences and not some simple bug fix, I'm not sure where to start.\n\nDo you guys have any tips?\n\nAnyway, I think part of the deal of making this feature materialize is\nto raise\nawareness, so if anybody reading this is onboard, please share\n\nThanks!\nNoam\n\n\n"},{"id":"397945","messageId":"20200515222651.GB116674@coredump.intra.peff.net","threadId":"52980","inReplyTo":"49bf5848-b8b0-13b0-4038-352f702d11ee@gmail.com","subject":"Re: [TOPIC 9/17] Obsolescence markers and evolve","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-05-15T22:26:51Z","receivedAt":"2020-05-15T22:26:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, May 10, 2020 at 12:31:37AM +0300, Noam Soloveichik wrote:\n\n> On 12/03/2020 6:04, James Ramsay wrote:\n> > 1. Brandon: I thought it would be interesting to have a similar feature\n> > as Mercurial has. Mercurial evolve will help you do a big rebase\n> > commit by commit. Giving you more insights how commits change over time.\n> >\n> > 2. Peff: This has been discussed a lot of time on the list already.\n> Since I'm very interested in this topic, can you link me to some key\n> discussions you remember? Most of what I've found is Stefan Xenos having\n> a take on implementing it.\n\nSorry, I don't much to offer. The discussion from Stefan is the only one\nI remember talking about evolve itself.\n\nI think my comment may have been specifically about the extra graph\npointers that would be needed to represent rebases, etc. The general\nconcept of a parent pointer that doesn't imply reachability has come up\nover the years. I don't have any links handy, though, and searching for\n\"parent\" in the list archive is not likely to be all that helpful.\n\nHmm, searching for \"parent\" and \"reachable\" also turns up a lot, but\nthis one is probably relevant:\n\n  https://lore.kernel.org/git/20060425035421.18382.51677.stgit@localhost.localdomain/\n\nSomething that old is as likely to hurt as help, though. ;)\n\n-Peff\n"},{"id":"397952","messageId":"20200516022148.20863-1-nbelakovski@gmail.com","threadId":"52980","inReplyTo":"5cab1530-f8b6-cef3-7b93-48fad410a160@iee.email","subject":"Re: [TOPIC 3/17] Obliterate","fromName":"","fromEmail":"nbelakovski@gmail.com","sentAt":"2020-05-16T02:21:48Z","receivedAt":"2020-05-16T02:21:56Z","isPatch":false,"sender":{"key":"nbelakovski@gmail.com","avatar":"https://avatars.githubusercontent.com/u/864630?v=4"},"body":"From: Nickolai Belakovski <nbelakovski@gmail.com>\n\nHi guys,\n\nSorry I missed you at the contributor summit, but this is an idea I've been\nthinking of on my own for some time now, mostly in the context of dealing with\nlarge files as opposed to security issues. I've come to a lot of the same\nconclusions that this group has already come up with, namely\n\n* Using git replace functionality is a very obvious apth forward here\n\n* A list of 'revoked' (could I propose the word 'obliterated', in order to make\n  the names consistent?) hashes needs to be maintained so that any\n  functionality expecting the original object can figure out that it's not\n  available.\n\n* This needs support from GitHub, Gitlab, etc. in order for it to work. I'm\n  thinking that git prune gets updated to remove 'oblierated' objects, and\n  when a git hosting service receives an updated list of obliterated objects,\n  it just runs prune. Of course, there would need to be support for replace\n  refs as well\n\nI've started working on a prototype/proof of concept. In the v1 it will do the\nfollowing upon receiving a hash (i.e. git obliterate abc123):\n\n* Add it to the list of obliterated objects (I'm thinking just .git/obliterate,\n  any issues with that?)\n\n* Create a new blob containing the content \"This file was obliterated by\n  $git.user on $today\" and create a replace ref from the provided has to the\n  hash of this new blob (so instead of an empty file, there's some info as to\n  why the file is missing)\n\n* Run git prune (which will be modified to delete obliterated objects)\n\nIt should just take a couple of days. If anyone is interested in joining, I'm\nlivestreaming my work on twitch at https://www.twitch.tv/actinium226 from 1pm\nto 5pm Pacific time on weekdays. This version will still have some issues. As\nDamien pointed out, index doesn't handle replace, so the file will look\nmodified, but I hope that having an initial prototype will help further\ndiscussion and get this feature closer to a state of being completed.\n"},{"id":"398217","messageId":"20200519201034.GB196295@google.com","threadId":"52980","inReplyTo":"20200506213354.GG77802@google.com","subject":"Re: [PATCH] doc: propose hooks managed by the config","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-05-19T20:10:34Z","receivedAt":"2020-05-19T20:10:44Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, May 06, 2020 at 02:33:54PM -0700, Emily Shaffer wrote:\n> \n> On Sat, Apr 25, 2020 at 08:57:27PM +0000, brian m. carlson wrote:\n> > \n> > On 2020-04-20 at 23:53:10, Emily Shaffer wrote:\n> > > +=== Config schema\n> > > +\n> > > +Hooks can be introduced by editing the configuration manually. There are two new\n> > > +sections added, `hook` and `hookcmd`.\n> > > +\n> > > +==== `hook`\n> > > +\n> > > +Primarily contains subsections for each hook event. These subsections define\n> > > +hook command execution order; hook commands can be specified by passing the\n> > > +command directly if no additional configuration is needed, or by passing the\n> > > +name of a `hookcmd`. If Git does not find a `hookcmd` whose subsection matches\n> > > +the value of the given command string, Git will try to execute the string\n> > > +directly. Hook event subsections can also contain per-hook-event settings.\n> > \n> > Can we say explicitly that the commands are invoked by the shell?  Or is\n> > the plan to try to parse them without passing to the shell?\n> \n> Sure. If I didn't make it clear it was by mistake, not by intent.\n> \n> > \n> > > +Also contains top-level hook execution settings, for example,\n> > > +`hook.warnHookDir`, `hook.runHookDir`, or `hook.disableAll`.\n> > > +\n> > > +----\n> > > +[hook \"pre-commit\"]\n> > > +  command = perl-linter\n> > > +  command = /usr/bin/git-secrets --pre-commit\n> > > +\n> > > +[hook \"pre-applypatch\"]\n> > > +  command = perl-linter\n> > > +  error = ignore\n> > > +\n> > > +[hook]\n> > > +  warnHookDir = true\n> > > +  runHookDir = prompt\n> > > +----\n> > > +\n> > > +==== `hookcmd`\n> > > +\n> > > +Defines a hook command and its attributes, which will be used when a hook event\n> > > +occurs. Unqualified attributes are assumed to apply to this hook during all hook\n> > > +events, but event-specific attributes can also be supplied. The example runs\n> > > +`/usr/bin/lint-it --language=perl <args passed by Git>`, but for repos which\n> > > +include this config, the hook command will be skipped for all events to which\n> > > +it's normally subscribed _except_ `pre-commit`.\n> > > +\n> > > +----\n> > > +[hookcmd \"perl-linter\"]\n> > > +  command = /usr/bin/lint-it --language=perl\n> > > +  skip = true\n> > > +  pre-commit-skip = false\n> > > +----\n> > \n> > This seems fine to me.  I like this design and it seems sane.\n> > \n> > > +== Implementation\n> > > +\n> > > +=== Library\n> > > +\n> > > +`hook.c` and `hook.h` are responsible for interacting with the config files. In\n> > > +the case when the code generating a hook event doesn't have special concerns\n> > > +about how to run the hooks, the hook library will provide a basic API to call\n> > > +all hooks in config order with an `argv_array` provided by the code which\n> > > +generates the hook event:\n> > > +\n> > > +*`int run_hooks(const char *hookname, struct argv_array *args)`*\n> > > +\n> > > +This call includes the hook command provided by `run-command.h:find_hook()`;\n> > > +eventually, this legacy hook will be gated by a config `hook.runHookDir`. The\n> > > +config is checked against a number of cases:\n> > > +\n> > > +- \"no\": the legacy hook will not be run\n> > > +- \"interactive\": Git will prompt the user before running the legacy hook\n> > > +- \"warn\": Git will print a warning to stderr before running the legacy hook\n> > > +- \"yes\" (default): Git will silently run the legacy hook\n> > > +\n> > > +If `hook.runHookDir` is provided more than once, Git will use the most\n> > > +restrictive setting provided, for security reasons.\n> > \n> > I don't think this is consistent with the way the rest of our options\n> > work.  What if someone generally wants to disable legacy hooks but then\n> > works with a program in a repository that requires them?\n> \n> Unfortunately this is something I think my end will want to hold firm\n> on. In general we disagree with your statement later about not wanting\n> to make the .git/config secure. I see your use case, and I anticipate\n> two possible workarounds I'd present:\n> \n> 1) If working in that repo for the short term, run `git -c\n> hook.runHookDir=yes <command> <arg...>` (and therefore allow the config\n> from command line scope, which I'm happy with in general). Maybe\n> someone would want to use an alias, hookgit or hg? Just kidding.. ;P\n> \n> 2) If you're stuck with that repo for the long term, add\n> `hook.<hookname>.command = /path/.git/hooks/<hookname>` lines to the local\n> config.\n> \n> Yes, those are both somewhat user-unfriendly, and I think we can do\n> better... I'll have to think more and see what I can come up with.\n> Suggestions welcome.\n\nI thought more about this and today I'm revisiting this work (and\nstarting on patches!) so I figured I'd close the loop, since it'll be\nburied in the next round of the design doc.\n\nRefusing to trust the local config is actually contrary to one of the\ntenets I was trying to use when designing this - that we should assume\nthe .git/config is safe, so that we don't end up with bloat later if\n.git/config does become safe. The suggestion I made here to disallow\noverrides doesn't fit, so I'll drop it. The implementation will allow a\nmore local config to turn hookdir hooks back on.\n\nThanks, all. By way of status update, I think I'll be able to start\nworking on this more actively starting this week.\n\n - Emily\n"}]}