{"thread":{"id":"59261","subject":"Proposal/Discussion: Turning parts of Git into libraries","startedAt":"2023-02-17T21:12:42Z","lastAt":"2023-03-24T22:30:05Z","messageCount":37,"participants":["Emily Shaffer","brian m. carlson","rsbecker@nexbridge.com","Junio C Hamano","demerphq","Elijah Newren","Phillip Wood","Taylor Blau","Victoria Dye","Derrick Stolee","Jeff King","Jonathan Tan","Felipe Contreras"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"472268","messageId":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","threadId":"59261","inReplyTo":null,"subject":"Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-17T21:12:23Z","receivedAt":"2023-02-17T21:12:42Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Hi folks,\n\nAs I mentioned in standup this week[1], my colleagues and I at Google\nhave become very interested in converting parts of Git into libraries\nusable by external programs. In other words, for some modules which\nalready have clear boundaries inside of Git - like config.[ch],\nstrbuf.[ch], etc. - we want to remove some implicit dependencies, like\nreferences to globals, and make explicit other dependencies, like\nreferences to other modules within Git. Eventually, we'd like both for\nan external program to use Git libraries within its own process, and\nfor Git to be given an alternative implementation of a library it uses\ninternally (like a plugin at runtime).\n\nThis turned out pretty long-winded, so a quick summary before I dive in:\n\n- We want to compile parts of Git as independent libraries\n- We want to do it by making incremental code quality improvements to Git\n- Let's avoid promising stability of the interfaces of those libraries\n- We think it'll let Git do cool stuff like unit tests and allowing\npurpose-built plugins\n- Hopefully by example we can convince the rest of the project to join\nin the effort\n\nMy team has spent the past year or so trying to make improvements to\nGit's behavior with submodules, and we found that the current\nstructure of Git is quite tricky to work with: because Git doesn't\nexecute on second repositories in the same process well, recursing\ninto submodules typically involves spawning child processes, and\npiping new arguments through the helpers around those child processes,\nand then through Git's typical codepaths, is very tricky. After\nspending more than a year trying to make improvements, we have very\nlittle to show for it, largely as a result of the difficulty of\npassing information between superprojects and submodules.\n\nIt seems like being able to invoke parts of Git as a library, or Git\nbeing able to invoke custom libraries, does a lot of good for the Git\nproject:\n\n- Having clear, modular libraries makes it easy to find the code\nresponsible for some behavior, or determine where to add something\nnew.\n- Operations recursing into submodules can be run in-process, and the\ncallsite makes it very clear which repository context is being used;\nif we're not subprocessing to recurse, then we can lose the awkward\ngit-submodule.sh -> builtin/submodule__helper.c -> top-level Git\nprocess codepath.\n- Being able to test libraries in isolation via unit tests or mocks\nspeeds up determining the root cause of bugs.\n- If we can swap out an entire library, or just a single function of\none, it's easy to experiment with the entire codebase without sweeping\nchanges.\n\nThe ability to use Git as a library also makes it easier for other\ntooling to interact with a Git repository *correctly*. As an example,\n`repo` has a long history of abusing Git by directly manipulating the\ngitdir[2], but if it were written in a world where Git exists as\neasy-to-use libraries, it probably wouldn't have needed to, as it\ncould have invoked Git directly or replaced the relevant modules with\nits own implementation. Both `repo`[3] and `git-gui[4]` have\nreimplemented logic from git.git. Other interfaces that cooperate with\nGit's filesystem storage layer, like `scalar` or `jj`[5], would be\nable to interop with a Git repository without having to reimplement\ncustom logic or keep up with new Git changes.\n\nOf course, there's a reason Google wants it, too. We've talked\npreviously about wanting better integration between Git and something\nlike a VFS; as we've experimented with it internally, we've found a\ncouple of tricky areas:\n\n- The VFS relies on running Git commands to determine the state of the\nrepository. However, these Git commands interact with the gitdir or\nworktree, which is populated by the VFS. For example, parsing a\n.gitattributes or .gitmodules which is already stored in the VFS\nrequires the VFS to provide a POSIX file handle, spawn a Git\nsubprocess, populate other files needed by that subprocess (like\n.git/config), and finally collect the output stream of the subprocess.\nAs you can imagine, this interaction of VFS -> Git -> VFS [-> Git]\ncreates all sort of complications. The alternative is for the VFS to\nwrite its own parser (or use a library like libgit2; more on that\nlater). But having a Git library means that a subset of Git\nfunctionality can happen in-process, and that filesystem access could\nbe replaced by the VFS directly providing high-level objects or plain\nbytestreams.\n\n- A user running `git status` in a directory controlled by the VFS\nwill require the VFS to populate the entire (large) worktree - even\nthough the VFS is sure that only one file has been modified. The\nclosest we've come with an alternative is an orchestrated use of\nsparse-checkout - but modifying the sparse-checkout configs\nautomatically in response to the user's filesystem operations takes us\nright back to the previous point. If Git could take a plugin and\nreplacement for the object store that directly interacts with the VFS\ndaemon, a layer of complexity would disappear and performance would\nimprove.\n\nWe discussed using an existing library like libgit2 or JGit, but it's\nnot a very exciting proposal: these libraries are already lagging\nbehind git.git in features, and trying to use them in combination with\nbrand-new improvements to Git (like new partial clone filters) means\nthat we'll always get to implement those improvements twice to bring\nlibgit2 up to speed. It also means that people using the `git` client\ndirectly won't get performance benefits derived from having Git\ninternal libraries replaced by purpose-built ones in certain contexts.\nThat said, if libgit2 already provides functionality and performance\nequivalent to git.git's in an appropriate wrapper, I'd be excited to\npursue integrating that library into git.git's codebase directly.\n\nThe good news is that for practical near-term purposes, \"libification\"\nmostly means cleanups to the Git codebase, and continuing code health\nwork that the project has already cared about doing:\n\n- Removing references to global variables and instead piping them\nthrough arguments\n- Finding and fixing memory leaks, especially in widely-used low-level code\n- Reducing or removing `die()` invocations in low-level code, and\ninstead reporting errors back to callers in a consistent way\n- Clarifying the scope and layering of existing modules, for example\nby moving single-use helpers from the shared module's scope into the\nsingle user's scope\n- Making module interfaces more consistent and easier to understand,\nincluding moving \"private\" functions out of headers and into source\nfiles and improving in-header documentation\n\nBasically, if this effort turns out not to be fruitful as a whole, I'd\nlike for us to still have left a positive impact on the codebase.\n\nIn the longer term, if Git has libraries with easily-replaced\ndependencies, we get a few more benefits:\n\n- Unit tests. We already have some in t/helper/, but if we can replace\nall the dependencies of a certain library with simple stubs, it's\neasier for us to write comprehensive unit tests, in addition to the\nwork we already do introducing edge cases in bash integration tests.\n- If our users can use plugins to improve performance in specific\nscenarios (like a VFS-aware object store in the VFS case I cited\nabove), then Git works better for them without having to adopt a\ndifferent workflow, such as using an alternative tool or wrapper.\n- An easy-to-understand modular codebase makes it easy for new\ncontributors to start hacking and understand the consequences of their\npatch.\n\nOf course, we haven't maintained any guarantee about the consistency\nof our implementation between releases. I don't anticipate that we'll\nwrite the perfect library interface on our first try. So I hope that\nwe can be very explicit about refusing to provide any compatibility\nguarantee whatsoever between versions for quite a long time. On\nGoogle's end, that's well-understood and accepted. As I understand,\nsome other projects already use Git's codebase as a \"library\" by\nincluding it as a submodule and using the code directly[6]; even a\nbreakable API seems like an improvement over that, too.\n\nSo what's next? Naturally, I'm looking forward to a spirited\ndiscussion about this topic - I'd like to know which concerns haven't\nbeen addressed and figure out whether we can find a way around them,\nand generally build awareness of this effort with the community.\n\nI'm also planning to send a proposal for a document full of \"best\npractices\" for turning Git code into libraries (and have quite a lot\nof discussion around that document, too). My hope is that we can use\nthat document to help us during implementation as well as during\nreview, and refine it over time as we learn more about what works and\nwhat doesn't. Having this kind of documentation will make it easy for\nothers to join us in moving Git's codebase towards a clean set of\nlibraries. I hope that, as a project, we can settle on some tenets\nthat we all agree would make Git nicer.\n\nFrom the rest of my own team, we're planning on working first on some\nlimited scope, low-level libraries so that we can all see how the\nprocess works. We're starting with strbuf.[ch] (as it's used\neverywhere with few or no dependencies and helps us guarantee string\nsafety at API boundaries), config.[ch] (as many external tools are\nprobably interested in parsing Git config formatted files directly),\nand a subset of operations related to the object store. These starting\npoints are intended to have a small impact on the codebase and teach\nus a lot about logistics and best practices while doing these kinds of\nconversions.\n\nAfter that, we're still hoping to target low-level libraries first - I\ncertainly don't think it will make sense to ship a high-level `git\ncommit` library in the near future, if ever - in the order that\nthey're required from the VFS project we're working closely with. As\nfar as I can tell right now, that's likely to cover object store and\nworktree access, as well as commit creation and pushing, but we'll see\nhow planning shakes out over the next month or so. But Google's\nschedule should have no bearing on what others in the Git project feel\nis important to clean up and libify, and if there is interest in the\nrest of the project in converting other existing modules into\nlibraries, my team and I are excited to participate in the review.\n\nMuch, much later on, I'm expecting us to form a plan around allowing\n\"plugins\" - that is, replacing library functionality we use today with\nan alternative library, such as an object store relying on a\ndistributed file store like S3. Making that work well will also likely\ninvolve us coming up with a solution for dependency injection, and to\nbegin using vtables for some libraries. I'm hoping that we can figure\nout a way to do that that won't make the Git source ugly. Around this\ntime, I think it will make sense to buy into unit tests even more and\nstart using an approach like mocking to test various edge cases. And\nat some point, it's likely that we'll want to make the interfaces to\nvarious Git libraries consistent with each other, which would involve\nsome large-scale but hopefully-mechanical refactors.\n\nI'm looking forward to the discussion!\n\n - Emily\n\n1: https://colabti.org/irclogger/irclogger_log/git-devel?date=2023-02-13#l29\n2: https://gerrit.googlesource.com/git-repo/+/refs/heads/main/docs/internal-fs-layout.md\n3: https://gerrit.googlesource.com/git-repo/+/refs/heads/main/git_config.py\n4: https://github.com/git/git/blob/master/git-gui/git-gui.sh#L305\n5: https://github.com/martinvonz/jj\n6: https://github.com/glandium/git-cinnabar\n"},{"id":"472269","messageId":"Y+/v15TyCbSYzlVg@tapette.crustytoothpaste.net","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-17T21:21:40Z","receivedAt":"2023-02-17T21:21:49Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-17 at 21:12:23, Emily Shaffer wrote:\n> Hi folks,\n\nHey,\n\n> I'm looking forward to the discussion!\n\nWhile I'm not personally interested in the VFS work, I think it's a\ngreat idea to turn more of the code into libraries (or at least make it\nmore library-like), and so I'm fully in support of this approach.  When\nI send patches in the future, I'll try to make sure that they're\nfriendly to this goal.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"472272","messageId":"CAJoAoZmMQ-ROdCp0=4oaFa836-PqxwYntnRSBSzzJc5chp16eQ@mail.gmail.com","threadId":"59261","inReplyTo":"Y+/v15TyCbSYzlVg@tapette.crustytoothpaste.net","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-17T21:38:34Z","receivedAt":"2023-02-17T21:38:52Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Feb 17, 2023 at 1:21 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> On 2023-02-17 at 21:12:23, Emily Shaffer wrote:\n> > Hi folks,\n>\n> Hey,\n>\n> > I'm looking forward to the discussion!\n>\n> While I'm not personally interested in the VFS work, I think it's a\n> great idea to turn more of the code into libraries (or at least make it\n> more library-like), and so I'm fully in support of this approach.\n\nYeah, I expect this sentiment is true for most contributors. And I'm\nreally hoping there are other things which aren't on my radar, but\nother contributors are enthusiastic about, that would also be served\nby this kind of library approach.\n\nFor example, I seem to remember you saying during the SHA-256 series\nthat the next hashing algorithm would also be painful to implement;\nwould that still be true if the hashing algorithm is encapsulated well\nby a library interface? Or is it for a different reason?\n\n> When\n> I send patches in the future, I'll try to make sure that they're\n> friendly to this goal.\n\nThat's awesome to hear. Thanks, brian.\n\n - Emily\n"},{"id":"472277","messageId":"00b401d9431b$cc3c4460$64b4cd20$@nexbridge.com","threadId":"59261","inReplyTo":"Y+/v15TyCbSYzlVg@tapette.crustytoothpaste.net","subject":"RE: Proposal/Discussion: Turning parts of Git into libraries","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-02-17T22:04:19Z","receivedAt":"2023-02-17T22:04:36Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Friday, February 17, 2023 4:22 PM, brian m. carlson wrote:\n>On 2023-02-17 at 21:12:23, Emily Shaffer wrote:\n>> Hi folks,\n>\n>Hey,\n>\n>> I'm looking forward to the discussion!\n>\n>While I'm not personally interested in the VFS work, I think it's a great idea to turn\n>more of the code into libraries (or at least make it more library-like), and so I'm fully\n>in support of this approach.  When I send patches in the future, I'll try to make sure\n>that they're friendly to this goal.\n\nI am uncertain about this, from a licensing standpoint. Typically, when one links in a library from one project, the license from that project may inherit into your own project. AFAIK, GPLv3 has this implied provision - I do not think it is explicit, but the implication seems to be there. Making git libraries has the potential to cause git's license rights to be incorporated into other products. I am suggesting that we would need to tread carefully in this area. Using someone else's DLL is not so bad, as the code is not bound together, but may also cause ambiguities depending on whether the licenses are conflicting or not. I am not suggesting that this is a bad idea, just one that should be handled carefully.\n\n--Randall\n\n"},{"id":"472279","messageId":"Y/ACqlhtLMjfgJFQ@tapette.crustytoothpaste.net","threadId":"59261","inReplyTo":"CAJoAoZmMQ-ROdCp0=4oaFa836-PqxwYntnRSBSzzJc5chp16eQ@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-17T22:41:46Z","receivedAt":"2023-02-17T22:41:52Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-17 at 21:38:34, Emily Shaffer wrote:\n> For example, I seem to remember you saying during the SHA-256 series\n> that the next hashing algorithm would also be painful to implement;\n> would that still be true if the hashing algorithm is encapsulated well\n> by a library interface? Or is it for a different reason?\n\nRight now, most of the code for a future hash algorithm wouldn't be too\ndifficult to implement, I'd think, because we can already support two of\nthem.  If we decide, say, to implement SHA-3-512, we basically just add\nthat algorithm, update all the entries in the tests (which is kind of a\npain since there's a lot of them, but not really difficult), and then\nmove on with our lives.\n\nThe difficulty is dealing with interop work, which is basically\nswitching from dealing with just one algorithm to rewriting things\nbetween the two on the fly.  I think _that_ work would be made easier by\nlibrary work because sometimes it involves working with submodules, such\nas when updating the submodule commit, and being able to deal with both\nobject stores more easily at the same time would be very helpful in that\nregard.\n\nI can imagine there are other things that would be easier as well, and I\ncan also imagine that we'll have better control over memory allocations\nand leak less, which would be nice.  If we can get leaks low enough, we\ncould even add CI jobs to catch them and fail, which I think would be\nsuper valuable, especially since I find even after over two decades of C\nthat I'm still not very good about catching all the leaks (which is one\nof the reasons I've mostly switched to Rust).  We might also be able to\nmake nicer steps on multithreading our code as well.\n\nPersonally, I'd like to see some sort of standard error type (whether\nintegral or not) that would let us do more bubbling up of errors and\nless die().  I don't know if that's in the cards, but I thought I'd\nsuggest it in case other folks are interested.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"472280","messageId":"Y/AEWHTJtzBArGNv@tapette.crustytoothpaste.net","threadId":"59261","inReplyTo":"00b401d9431b$cc3c4460$64b4cd20$@nexbridge.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-02-17T22:48:56Z","receivedAt":"2023-02-17T22:49:02Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-02-17 at 22:04:19, rsbecker@nexbridge.com wrote:\n> I am uncertain about this, from a licensing standpoint. Typically,\n> when one links in a library from one project, the license from that\n> project may inherit into your own project. AFAIK, GPLv3 has this\n> implied provision - I do not think it is explicit, but the implication\n> seems to be there. Making git libraries has the potential to cause\n> git's license rights to be incorporated into other products. I am\n> suggesting that we would need to tread carefully in this area. Using\n> someone else's DLL is not so bad, as the code is not bound together,\n> but may also cause ambiguities depending on whether the licenses are\n> conflicting or not. I am not suggesting that this is a bad idea, just\n> one that should be handled carefully.\n\nI think it's pretty clear that if software used Git's libraries, that\nthe result would be GPLv2.  That might be fine for some projects, and\nfor others, libgit2 would be more appealing.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"472281","messageId":"CAJoAoZkMR9Acy7thVs-_e=Fz8wwjoDGDKb46wmwn8yxk0ODGow@mail.gmail.com","threadId":"59261","inReplyTo":"Y/ACqlhtLMjfgJFQ@tapette.crustytoothpaste.net","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-17T22:49:51Z","receivedAt":"2023-02-17T22:50:10Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Feb 17, 2023 at 2:41 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> On 2023-02-17 at 21:38:34, Emily Shaffer wrote:\n> > For example, I seem to remember you saying during the SHA-256 series\n> > that the next hashing algorithm would also be painful to implement;\n> > would that still be true if the hashing algorithm is encapsulated well\n> > by a library interface? Or is it for a different reason?\n>\n> Right now, most of the code for a future hash algorithm wouldn't be too\n> difficult to implement, I'd think, because we can already support two of\n> them.  If we decide, say, to implement SHA-3-512, we basically just add\n> that algorithm, update all the entries in the tests (which is kind of a\n> pain since there's a lot of them, but not really difficult), and then\n> move on with our lives.\n>\n> The difficulty is dealing with interop work, which is basically\n> switching from dealing with just one algorithm to rewriting things\n> between the two on the fly.  I think _that_ work would be made easier by\n> library work because sometimes it involves working with submodules, such\n> as when updating the submodule commit, and being able to deal with both\n> object stores more easily at the same time would be very helpful in that\n> regard.\n\nOoh, I see what you mean. Thanks for the extra context here.\n\n> I can imagine there are other things that would be easier as well, and I\n> can also imagine that we'll have better control over memory allocations\n> and leak less, which would be nice.  If we can get leaks low enough, we\n> could even add CI jobs to catch them and fail, which I think would be\n> super valuable, especially since I find even after over two decades of C\n> that I'm still not very good about catching all the leaks (which is one\n> of the reasons I've mostly switched to Rust).  We might also be able to\n> make nicer steps on multithreading our code as well.\n>\n> Personally, I'd like to see some sort of standard error type (whether\n> integral or not) that would let us do more bubbling up of errors and\n> less die().  I don't know if that's in the cards, but I thought I'd\n> suggest it in case other folks are interested.\n\nYes!!! We have talked about this a lot internally - but this is one\nthing that will be difficult to introduce into Git without making\nparts of the codebase a little uglier. Since obviously C doesn't have\nan intrinsic to do this, we'll have to roll our own, which means that\nmanipulating it consistently at function exits might end up pretty\nugly. So hearing that there's interest outside of my team to come up\nwith such a type makes me optimistic that we can figure out a\nneat-enough solution.\n\n> --\n> brian m. carlson (he/him or they/them)\n> Toronto, Ontario, CA\n"},{"id":"472282","messageId":"xmqq3573lx2d.fsf@gitster.g","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-17T22:57:30Z","receivedAt":"2023-02-17T22:58:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <nasamuffin@google.com> writes:\n\n> Basically, if this effort turns out not to be fruitful as a whole, I'd\n> like for us to still have left a positive impact on the codebase.\n> ...\n> So what's next? Naturally, I'm looking forward to a spirited\n> discussion about this topic - I'd like to know which concerns haven't\n> been addressed and figure out whether we can find a way around them,\n> and generally build awareness of this effort with the community.\n\nOn of the gravest concerns is that the devil is in the details.\n\nFor example, \"die() is inconvenient to callers, let's propagate\nerrors up the callchain\" is an easy thing to say, but it would take\nmuch more than \"let's propagate errors up\" to libify something like\ncheck_connected() to do the same thing without spawning a separate\nprocess that is expected to exit with failure.\n\nIt is not clear if we can start small, work on a subset of the\nthings and still reap the benefit of libification.  Is there an\nexisting example that we have successfully modularlized the API into\none subsystem?  Offhand, I suspect that the refs API with its two\nimplementations may be reasonably close, but is the inteface into\nthat subsystem the granularity of the library interface you guys\nhave in mind?\n\n"},{"id":"472287","messageId":"CANgJU+XoT42u91WP7-p4V41w7q-UVhutL2LUfNkp3_BRCOn-FQ@mail.gmail.com","threadId":"59261","inReplyTo":"xmqq3573lx2d.fsf@gitster.g","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2023-02-18T01:59:51Z","receivedAt":"2023-02-18T02:00:06Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Emily Shaffer <nasamuffin@google.com> writes:\n>\n> > Basically, if this effort turns out not to be fruitful as a whole, I'd\n> > like for us to still have left a positive impact on the codebase.\n> > ...\n> > So what's next? Naturally, I'm looking forward to a spirited\n> > discussion about this topic - I'd like to know which concerns haven't\n> > been addressed and figure out whether we can find a way around them,\n> > and generally build awareness of this effort with the community.\n>\n> On of the gravest concerns is that the devil is in the details.\n>\n> For example, \"die() is inconvenient to callers, let's propagate\n> errors up the callchain\" is an easy thing to say, but it would take\n> much more than \"let's propagate errors up\" to libify something like\n> check_connected() to do the same thing without spawning a separate\n> process that is expected to exit with failure.\n\n\nWhat does \"propagate errors up the callchain\" mean?  One\ninterpretation I can think of seems quite horrible, but another seems\nquite doable and reasonable and likely not even very invasive of the\nexisting code:\n\nYou can use setjmp/longjmp to implement a form of \"try\", so that\nerrors dont have to be *explicitly* returned *in* the call chain. And\nyou could probably do so without changing very much of the existing\ncode at all, and maintain a high level of conceptual alignment with\nthe current code strategy.\n\nTo do this you need to set up a globally available linked list of\njmp_env data (see `man setjmp` for jmp_env), and a global error\nobject, and make the existing \"die\" functions populate the global\nerror object, and then pop the most recent jmp_env data and longjmp to\nit.\n\nAt the top of any git invocation you would set up the topmost jmp_env\n\"frame\". Any code that wants to \"try\" existing logic pushes a new\njmp_env (using a wrapper around setjmp), and prepares to be longjmp'ed\nto. If the code does not die then it pops the jmp_env it just pushed\nand returns as normal, if it is longjmp'ed to you can detect this and\ndo some other behavior to handle the exception (by reading the global\nerror object). If the code that died *really* wants to exit, then it\nreturns the appropriate code as part of the longjmp, and the try\nhandler longjmps again propagating up the chain. Eventually you either\nhave an error that \"propagates to the top\" which results in an exit\nwith an appropriate error message, or you have an error that is\ntrapped and the code does something else, and then eventually returns\nnormally.\n\nFWIW, this is essentially a loose description of how Perl handles the\nexecution part of \"eval\" and supports exception handling internally.\nMost of the perl internals do not know anything about exceptions, they\njust call functions similar to gits die functions if they need to,\nwhich then call into Perl_die_unwind(). which then calls the\nJUMPENV_JUMP() macro which does the \"pop and longjmp\" dance.\n\nSeems to me that it wouldn't be very difficult nor particularly\ninvasive to implement this in git. Much of the logic in the perl\nproject to do this is at the top of cop.h,  see the macros\nJMPENV_PUSH(), JMPENV_POP(), JMPENV_JUMP(). Obviously this code\ncontains a bunch of perl specific logic, but the general gist of it\nshould be easily understood and easily converted to a more git like\ncontext:\n\nstruct jmpenv: https://github.com/Perl/perl5/blob/blead/cop.h#L32\nJMPENV_BOOTSTRAP: https://github.com/Perl/perl5/blob/blead/cop.h#L66\nJMPENV_PUSH: https://github.com/Perl/perl5/blob/blead/cop.h#L113\nJMPENV_POP: https://github.com/Perl/perl5/blob/blead/cop.h#L147\nJMPENV_JUMP: https://github.com/Perl/perl5/blob/blead/cop.h#L159\n\nPerl_die_unwind: https://github.com/Perl/perl5/blob/blead/pp_ctl.c#L1741\nWhere Perl_die_unwind() calls JMPENV_JUMP:\nhttps://github.com/Perl/perl5/blob/blead/pp_ctl.c#L1865\n\nYou can also grep for functions of the form S_try_ in the perl code\nbase to find examples where the C code explicitly sets up an \"eval\nframe\" to interoperate with the functionality above.\n\ngit grep -nP '^S_try_'\npp_ctl.c:3548:S_try_yyparse(pTHX_ int gramtype, OP *caller_op)\npp_ctl.c:3604:S_try_run_unitcheck(pTHX_ OP* caller_op)\npp_sys.c:3120:S_try_amagic_ftest(pTHX_ char chr) {\n\nSeems to me that this gives enough prior art to convert git to use the\nsame strategy, and that doing so would not actually be that big a\nchange to the existing code.  Both environments are fairly similar if\nyou look at them from the right perspective. Both are C, and both have\na lot of global state, and both have lots of functions which you\nreally dont want to have to change to understand about exception\nobjects..\n\nHere is an example of how a C function might be written to use this\nkind of infrastructure to \"try\" functionality that might call die. In\nthis case there is no need for the code to inspect the global error\nobject, but the basic pattern is consistent. The \"default\" case below\nhandles the situation where the \"tried\" function is signalling an\n\"untrappable error\" that needs to be rethrown to ultimately unwind the\nentire try/catch chain and exit the program. It is derived and\nsimplified from S_try_yyparse mentioned above. This function handles\nthe \"compile the code\" part of an `eval EXPR`, and traps exceptions\nfrom the parser so that they can be handled properly and distinctly\nfrom errors trapped during execution of the compiled code. [ I am\nassuming that given the historical relationship between git and perl\nthese concepts are not alien to everybody on this list. ]\n\n/* S_try_yyparse():\n *\n * Run yyparse() in a setjmp wrapper. Returns:\n *   0: yyparse() successful\n *   1: yyparse() failed\n *   3: yyparse() died\n *\n * ...\n */\nSTATIC int\nS_try_yyparse(pTHX_ int gramtype, ...)\n{\n    dJMPENV;\n\n    JMPENV_PUSH(ret);\n    switch (ret) {\n    case 0:\n        ret = yyparse(gramtype) ? 1 : 0;\n        break;\n    case 3:\n        /* yyparse() died and we trapped the error. */\n        ....\n        break;\n    default:\n        JMPENV_POP;          /* remove our own setjmp data */\n        JMPENV_JUMP(ret); /* RETHROW */\n    }\n    JMPENV_POP;\n    return ret;\n}\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"472293","messageId":"CABPp-BE6EA+vXLXJtn8CHO9pHJgLH_uh7_t7AYBRN2gAAA5C+Q@mail.gmail.com","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2023-02-18T04:05:00Z","receivedAt":"2023-02-18T04:15:56Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Feb 17, 2023 at 1:45 PM Emily Shaffer <nasamuffin@google.com> wrote:\n>\n[...]\n> This turned out pretty long-winded, so a quick summary before I dive in:\n>\n> - We want to compile parts of Git as independent libraries\n> - We want to do it by making incremental code quality improvements to Git\n> - Let's avoid promising stability of the interfaces of those libraries\n> - We think it'll let Git do cool stuff like unit tests and allowing\n> purpose-built plugins\n> - Hopefully by example we can convince the rest of the project to join\n> in the effort\n\nSeems like quite reasonable high-level goals.\n\n[...]\n> The good news is that for practical near-term purposes, \"libification\"\n> mostly means cleanups to the Git codebase, and continuing code health\n> work that the project has already cared about doing:\n>\n> - Removing references to global variables and instead piping them\n> through arguments\n> - Finding and fixing memory leaks, especially in widely-used low-level code\n\nDoes removing memory leaks also mean converting UNLEAK to free()?\nThinking of things in a library context probably pushes us in that\ndirection (though, alternatively, it might just highlight the question\nof what is considered \"low-level\" instead).\n\n> - Reducing or removing `die()` invocations in low-level code, and\n> instead reporting errors back to callers in a consistent way\n\nWhat delinates \"low-level\" code?  (A \"we don't know yet but we'll\nstart with obvious places and plan to have good discussions on the\nappropriate boundary in the future as we submit patches\" is a fine\nanswer, I'm just curious if you already have a rough idea of where you\nintend that boundary to lie.)\n\n> - Clarifying the scope and layering of existing modules, for example\n> by moving single-use helpers from the shared module's scope into the\n> single user's scope\n> - Making module interfaces more consistent and easier to understand,\n> including moving \"private\" functions out of headers and into source\n> files and improving in-header documentation\n\nI think these are very positive directions.  I like the fact that your\ninitial plan benefits all of us, whether or not libification is\nultimately achieved.\n\n[...]\n> So what's next? Naturally, I'm looking forward to a spirited\n> discussion about this topic - I'd like to know which concerns haven't\n> been addressed and figure out whether we can find a way around them,\n> and generally build awareness of this effort with the community.\n\nI'm curious whether clarifying scope/layering and cleaning up\ninterfaces might mean you'd be interested in things like:\n  * https://github.com/newren/git/commits/header-cleanups (which was\nstill WIP; I paused working on it because I figured people would see\nit as big \"cleanup\" patches with no practical benefit)\n  * https://github.com/gitgitgadget/git/pull/1149 (which has been\nready to submit for a _long_ time, but I just haven't yet)\nor if these two things are orthogonal to what you have in mind.\n\n> I'm also planning to send a proposal for a document full of \"best\n> practices\" for turning Git code into libraries (and have quite a lot\n> of discussion around that document, too). My hope is that we can use\n> that document to help us during implementation as well as during\n> review, and refine it over time as we learn more about what works and\n> what doesn't. Having this kind of documentation will make it easy for\n> others to join us in moving Git's codebase towards a clean set of\n> libraries. I hope that, as a project, we can settle on some tenets\n> that we all agree would make Git nicer.\n\nI like the sound of this.\n\n> After that, we're still hoping to target low-level libraries first - I\n> certainly don't think it will make sense to ship a high-level `git\n> commit` library in the near future, if ever - in the order that\n> they're required from the VFS project we're working closely with. As\n> far as I can tell right now, that's likely to cover object store and\n> worktree access, as well as commit creation and pushing, but we'll see\n> how planning shakes out over the next month or so. But Google's\n> schedule should have no bearing on what others in the Git project feel\n> is important to clean up and libify, and if there is interest in the\n> rest of the project in converting other existing modules into\n> libraries, my team and I are excited to participate in the review.\n\nIf we can't libify something like commit, does that prevent libifying\nhigher level things like merge?\n\nI spent some time thinking about this a while back.  I tried to\ncarefully design merge-ort to improve the odds it could be used\nelsewhere, maybe even libgit2.  (I hope it shows in the many comments\nin merge-ort.h, and I think the \"priv\" field in particular allowing me\nto hide the first ~300 lines of merge-ort.c declaring data structures\nfrom users was really nice.)  However, I still had to accept data in\nsome known format.  So input parameters are things like trees and\ncommits.  But tree.h and commit.h both include object.h first, which\nincludes cache.h, which is basically all of Git.  And the functions I\ncall to interoperate with the system are similarly entangled.  So, the\nodds of merge-ort being reused by libgit2 or otherwise used in a\nlibrary seems essentially nil, at least without some broader\nlibification effort.\n\nI'd like to make that story better, time permitting (which is much\nmore of a challenge these days than it was a couple years ago), but\nI'm curious if you or others have thoughts on something like that.\n\n> Much, much later on, I'm expecting us to form a plan around allowing\n> \"plugins\" - that is, replacing library functionality we use today with\n> an alternative library, such as an object store relying on a\n> distributed file store like S3. Making that work well will also likely\n> involve us coming up with a solution for dependency injection, and to\n> begin using vtables for some libraries. I'm hoping that we can figure\n> out a way to do that that won't make the Git source ugly. Around this\n> time, I think it will make sense to buy into unit tests even more and\n> start using an approach like mocking to test various edge cases. And\n> at some point, it's likely that we'll want to make the interfaces to\n> various Git libraries consistent with each other, which would involve\n> some large-scale but hopefully-mechanical refactors.\n\nWould these plugins resemble the pluggable merge backends that was\nadded to builtin/merge.c?  Would it replace that mechanism with a\ndifferent one?  Would it be more like the refs backends?\n\nWould this plugin scheme allow us to, for example, use gitoxide[1] as\na clone replacement to make clones 2x as fast (and with half the\nmemory -- although I suspect they cheated and used sha1 instead of\nsha1dc, so maybe it wouldn't really be 2x)?\n\nOh, and it's totally okay if you don't know the answers to any or all\nof my questions right now.  I'm just curious, because I've long\nthought these kinds of directions would be good.  Since I've spent\ntime thinking about it, I have questions that I don't know the answers\nto, but I figured it couldn't hurt to bounce them off others who are\nthinking about this area.\n\nAnyway, it's a large pile of work that you're undertaking, and as\nJunio comments elsewhere in this thread it's unclear if libification\ncan be achieved for a big enough component (and you seem to admit as\nmuch in your email as well), but I applaud the general direction and\nyour initial plans.\n\n\n[1] https://github.com/Byron/gitoxide/discussions/579\n"},{"id":"472300","messageId":"4222af90-bd6b-d970-2829-1ddfaeb770bf@dunelm.org.uk","threadId":"59261","inReplyTo":"CANgJU+XoT42u91WP7-p4V41w7q-UVhutL2LUfNkp3_BRCOn-FQ@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2023-02-18T10:36:52Z","receivedAt":"2023-02-18T10:36:58Z","isPatch":false,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 18/02/2023 01:59, demerphq wrote:\n> On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Emily Shaffer <nasamuffin@google.com> writes:\n>>\n>>> Basically, if this effort turns out not to be fruitful as a whole, I'd\n>>> like for us to still have left a positive impact on the codebase.\n>>> ...\n>>> So what's next? Naturally, I'm looking forward to a spirited\n>>> discussion about this topic - I'd like to know which concerns haven't\n>>> been addressed and figure out whether we can find a way around them,\n>>> and generally build awareness of this effort with the community.\n>>\n>> On of the gravest concerns is that the devil is in the details.\n>>\n>> For example, \"die() is inconvenient to callers, let's propagate\n>> errors up the callchain\" is an easy thing to say, but it would take\n>> much more than \"let's propagate errors up\" to libify something like\n>> check_connected() to do the same thing without spawning a separate\n>> process that is expected to exit with failure.\n> \n> \n> What does \"propagate errors up the callchain\" mean?  One\n> interpretation I can think of seems quite horrible, but another seems\n> quite doable and reasonable and likely not even very invasive of the\n> existing code:\n> \n> You can use setjmp/longjmp to implement a form of \"try\", so that\n> errors dont have to be *explicitly* returned *in* the call chain. And\n> you could probably do so without changing very much of the existing\n> code at all, and maintain a high level of conceptual alignment with\n> the current code strategy.\n\nUsing setjmp/longjmp is an interesting suggestion, I think lua does \nsomething similar to what you describe for perl. However I think both of \nthose use a allocator with garbage collection. I worry that using \nlongjmp in git would be more invasive (or result in more memory leaks) \nas we'd need to to guard each allocation with some code to clean it up \nand then propagate the error. That means we're back to manually \npropagating errors up the call chain in many cases.\n\nBest Wishes\n\nPhillip\n\n> To do this you need to set up a globally available linked list of\n> jmp_env data (see `man setjmp` for jmp_env), and a global error\n> object, and make the existing \"die\" functions populate the global\n> error object, and then pop the most recent jmp_env data and longjmp to\n> it.\n> \n> At the top of any git invocation you would set up the topmost jmp_env\n> \"frame\". Any code that wants to \"try\" existing logic pushes a new\n> jmp_env (using a wrapper around setjmp), and prepares to be longjmp'ed\n> to. If the code does not die then it pops the jmp_env it just pushed\n> and returns as normal, if it is longjmp'ed to you can detect this and\n> do some other behavior to handle the exception (by reading the global\n> error object). If the code that died *really* wants to exit, then it\n> returns the appropriate code as part of the longjmp, and the try\n> handler longjmps again propagating up the chain. Eventually you either\n> have an error that \"propagates to the top\" which results in an exit\n> with an appropriate error message, or you have an error that is\n> trapped and the code does something else, and then eventually returns\n> normally.\n> \n> FWIW, this is essentially a loose description of how Perl handles the\n> execution part of \"eval\" and supports exception handling internally.\n> Most of the perl internals do not know anything about exceptions, they\n> just call functions similar to gits die functions if they need to,\n> which then call into Perl_die_unwind(). which then calls the\n> JUMPENV_JUMP() macro which does the \"pop and longjmp\" dance.\n> \n> Seems to me that it wouldn't be very difficult nor particularly\n> invasive to implement this in git. Much of the logic in the perl\n> project to do this is at the top of cop.h,  see the macros\n> JMPENV_PUSH(), JMPENV_POP(), JMPENV_JUMP(). Obviously this code\n> contains a bunch of perl specific logic, but the general gist of it\n> should be easily understood and easily converted to a more git like\n> context:\n> \n> struct jmpenv: https://github.com/Perl/perl5/blob/blead/cop.h#L32\n> JMPENV_BOOTSTRAP: https://github.com/Perl/perl5/blob/blead/cop.h#L66\n> JMPENV_PUSH: https://github.com/Perl/perl5/blob/blead/cop.h#L113\n> JMPENV_POP: https://github.com/Perl/perl5/blob/blead/cop.h#L147\n> JMPENV_JUMP: https://github.com/Perl/perl5/blob/blead/cop.h#L159\n> \n> Perl_die_unwind: https://github.com/Perl/perl5/blob/blead/pp_ctl.c#L1741\n> Where Perl_die_unwind() calls JMPENV_JUMP:\n> https://github.com/Perl/perl5/blob/blead/pp_ctl.c#L1865\n> \n> You can also grep for functions of the form S_try_ in the perl code\n> base to find examples where the C code explicitly sets up an \"eval\n> frame\" to interoperate with the functionality above.\n> \n> git grep -nP '^S_try_'\n> pp_ctl.c:3548:S_try_yyparse(pTHX_ int gramtype, OP *caller_op)\n> pp_ctl.c:3604:S_try_run_unitcheck(pTHX_ OP* caller_op)\n> pp_sys.c:3120:S_try_amagic_ftest(pTHX_ char chr) {\n> \n> Seems to me that this gives enough prior art to convert git to use the\n> same strategy, and that doing so would not actually be that big a\n> change to the existing code.  Both environments are fairly similar if\n> you look at them from the right perspective. Both are C, and both have\n> a lot of global state, and both have lots of functions which you\n> really dont want to have to change to understand about exception\n> objects..\n> \n> Here is an example of how a C function might be written to use this\n> kind of infrastructure to \"try\" functionality that might call die. In\n> this case there is no need for the code to inspect the global error\n> object, but the basic pattern is consistent. The \"default\" case below\n> handles the situation where the \"tried\" function is signalling an\n> \"untrappable error\" that needs to be rethrown to ultimately unwind the\n> entire try/catch chain and exit the program. It is derived and\n> simplified from S_try_yyparse mentioned above. This function handles\n> the \"compile the code\" part of an `eval EXPR`, and traps exceptions\n> from the parser so that they can be handled properly and distinctly\n> from errors trapped during execution of the compiled code. [ I am\n> assuming that given the historical relationship between git and perl\n> these concepts are not alien to everybody on this list. ]\n> \n> /* S_try_yyparse():\n>   *\n>   * Run yyparse() in a setjmp wrapper. Returns:\n>   *   0: yyparse() successful\n>   *   1: yyparse() failed\n>   *   3: yyparse() died\n>   *\n>   * ...\n>   */\n> STATIC int\n> S_try_yyparse(pTHX_ int gramtype, ...)\n> {\n>      dJMPENV;\n> \n>      JMPENV_PUSH(ret);\n>      switch (ret) {\n>      case 0:\n>          ret = yyparse(gramtype) ? 1 : 0;\n>          break;\n>      case 3:\n>          /* yyparse() died and we trapped the error. */\n>          ....\n>          break;\n>      default:\n>          JMPENV_POP;          /* remove our own setjmp data */\n>          JMPENV_JUMP(ret); /* RETHROW */\n>      }\n>      JMPENV_POP;\n>      return ret;\n> }\n> \n"},{"id":"472392","messageId":"Y/UXBw3Y9YnXUBIN@nand.local","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-02-21T19:09:59Z","receivedAt":"2023-02-21T19:10:27Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Feb 17, 2023 at 01:12:23PM -0800, Emily Shaffer wrote:\n> This turned out pretty long-winded, so a quick summary before I dive in:\n>\n> - We want to compile parts of Git as independent libraries\n> - We want to do it by making incremental code quality improvements to Git\n> - Let's avoid promising stability of the interfaces of those libraries\n> - We think it'll let Git do cool stuff like unit tests and allowing\n>   purpose-built plugins\n> - Hopefully by example we can convince the rest of the project to join\n>   in the effort\n\nLike others, I am less interested in the VFS-specific components you\nmention here, but I suspect that is just one particular instance of\nsomething that would be benefited by making git internals exposed via a\nlinkable library.\n\nI don't have objections to things like reducing our usage of globals,\nmaking fewer internal functions die() when they encounter an error, and\nso on. But like Junio, I suspect that this is definitely an instance of\na \"devil's in the details\" kind of problem.\n\nThat's definitely my main concern: that this turns out to be much more\ncomplicated than imagined and that we leave the codebase in a worse\nstate without much to show. A lesser version of that outcome would be\nthat we cause a lot of churn in the tree with not much to show either.\n\nSo I think we'd want to see some more concrete examples with clear\nbenefits to gauge whether this is a worthwhile direction. I think that\nstrbuf.h is too trivial an example to demonstrate anything useful. Being\nable to extract config.h into its own library so that another non-Git\nprogram could link against it and implement 'git config'-like\nfunctionality would be much more interesting.\n\nThanks,\nTaylor\n"},{"id":"472407","messageId":"CAJoAoZn7Nt37Eh17dpLDK+YX2BaEaAaii2rJPXO3L0BmQQkcgQ@mail.gmail.com","threadId":"59261","inReplyTo":"xmqq3573lx2d.fsf@gitster.g","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-21T21:42:31Z","receivedAt":"2023-02-21T21:42:49Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Feb 17, 2023 at 2:57 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Emily Shaffer <nasamuffin@google.com> writes:\n>\n> > Basically, if this effort turns out not to be fruitful as a whole, I'd\n> > like for us to still have left a positive impact on the codebase.\n> > ...\n> > So what's next? Naturally, I'm looking forward to a spirited\n> > discussion about this topic - I'd like to know which concerns haven't\n> > been addressed and figure out whether we can find a way around them,\n> > and generally build awareness of this effort with the community.\n>\n> On of the gravest concerns is that the devil is in the details.\n>\n> For example, \"die() is inconvenient to callers, let's propagate\n> errors up the callchain\" is an easy thing to say, but it would take\n> much more than \"let's propagate errors up\" to libify something like\n> check_connected() to do the same thing without spawning a separate\n> process that is expected to exit with failure.\n\nBecause the error propagation path is complicated, you mean? Or\nbecause the cleanup is painful?\n\nI wonder about this idea of spawning a worker thread that can\nterminate itself, though. Is it a bad idea? Is it a hacky way of\npretending that we have exceptions? I guess if we have a thread then\nwe still have the same concerns about memory management (which we\ndon't have if we use a child process). (I'll reply to demerphq's mail\nin detail, but it seems like the hardest part of this is memory\ncleanup, no?)\n\nIn other cases, we might want to perform some work that can be sped up\nby using more threads; how do we want to expose that functionality to\nthe caller? Do we want to manage our own threads, or do we want to\npass off orchestrating those worker threads to the caller (who\ntheoretically might have a faster way to manage them, like GPU\nexecution or distributed execution or something, or who might be using\ntheir own thread pool manager)?\n\n>\n> It is not clear if we can start small, work on a subset of the\n> things and still reap the benefit of libification.  Is there an\n> existing example that we have successfully modularlized the API into\n> one subsystem?  Offhand, I suspect that the refs API with its two\n> implementations may be reasonably close, but is the inteface into\n> that subsystem the granularity of the library interface you guys\n> have in mind?\n\nI think many of our internal APIs, especially the lower level ones,\nare actually quite well modularized, or close enough to it that you\ncan't really tell they aren't. run-command.h and config.h come to\nmind. The ones that aren't, I tend to think are frustrating to work\nwith anyways - is it reasonable to consider, for example, further\ncleanup of cache.h as part of this effort? Is it reasonable to rework\nan ugly circular dependency between two headers as a prerequisite to\ndoing library work around one of them?\n\nI had a look at the refs API documentation but it seems that we don't\nactually have a way for the code to use reftable. Is that what you\nmeant by the two implementations of refs API, or am I missing\nsomething else? Anyway, abstracting at the \"which backend do I want to\nuse\" layer seems absolutely appropriate to me, if we're discussing\nplaces where Git can use an alternative implementation. (For example,\nthis means it's also easy for Git to use some random NoSQL table as a\nref store, if that's what the caller wants.) For the most part refs.h\nseems like it has things I would want to expose to external callers\n(or that I would want to reimplement as a library author).\n\n - Emily\n"},{"id":"472413","messageId":"CAJoAoZm+TkCL0Jpg_qFgKottxbtiG2QOiY0qGrz3-uQy+=waPg@mail.gmail.com","threadId":"59261","inReplyTo":"CABPp-BE6EA+vXLXJtn8CHO9pHJgLH_uh7_t7AYBRN2gAAA5C+Q@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-21T22:06:55Z","receivedAt":"2023-02-21T22:07:13Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Feb 17, 2023 at 8:05 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Fri, Feb 17, 2023 at 1:45 PM Emily Shaffer <nasamuffin@google.com> wrote:\n> [...]\n> > The good news is that for practical near-term purposes, \"libification\"\n> > mostly means cleanups to the Git codebase, and continuing code health\n> > work that the project has already cared about doing:\n> >\n> > - Removing references to global variables and instead piping them\n> > through arguments\n> > - Finding and fixing memory leaks, especially in widely-used low-level code\n>\n> Does removing memory leaks also mean converting UNLEAK to free()?\n\nI suspect so - as I understand it, UNLEAK is a macro that resolves to\n\"don't complain to me, compiler, I meant not to free it.\"\n\n> Thinking of things in a library context probably pushes us in that\n> direction (though, alternatively, it might just highlight the question\n> of what is considered \"low-level\" instead).\n\nI'm not sure whether use of UNLEAK has so much to do with \"low-level\"\nor not. In cases when Git is being called as an ephemeral single-run\nprocess, UNLEAK makes a lot of sense. In cases when Git is being\ncalled in a long-lived process, UNLEAK is just a sign that says\n\"there's a leak here\".  So I think the distinction is not low-level or\nhigh-level, but more simply, within a library or not.\n\nI do anticipate that we'll still have \"non-libified\" code for the\nbuiltins, and that those builtins will invoke libraries at whatever\nlayer. So UNLEAKing memory allocated by the builtin - seems fine to\nme, even if that builtin is a \"low-level\" plumbing command.\n\n>\n> > - Reducing or removing `die()` invocations in low-level code, and\n> > instead reporting errors back to callers in a consistent way\n>\n> What delinates \"low-level\" code?  (A \"we don't know yet but we'll\n> start with obvious places and plan to have good discussions on the\n> appropriate boundary in the future as we submit patches\" is a fine\n> answer, I'm just curious if you already have a rough idea of where you\n> intend that boundary to lie.)\n\nThe biggest one is our \"standard library\" - stuff like strbuf,\nstring-list, strvec, etc. etc. I'd like to expose those to callers so\nthat we don't end up having library interfaces passing around\nunterminated buffers, which means that they'll be used in almost any\nother library.\n\nThat sort of hints at the next criteria - stuff that's used by lots of\noperations, or lots of other parts of Git code. So that means things\nlike run-command and config.\n\nPast that, we're determining libification order based on internal\npriorities. A request like \"our VFS helper needs to do `git commit`\nwith this specific set of constraints, please give us library calls to\ndo it\" would probably result in us working on library interfaces to\nhook execution, index parsing, and ref manipulation, and anything\nthat's a dependency of those three. It's very unlikely that it would\nresult in something like `git_do_commit(struct git_commit_flags)`.\n(That's what I meant about avoiding high-level libraries to begin\nwith.)\n\n> > - Clarifying the scope and layering of existing modules, for example\n> > by moving single-use helpers from the shared module's scope into the\n> > single user's scope\n> > - Making module interfaces more consistent and easier to understand,\n> > including moving \"private\" functions out of headers and into source\n> > files and improving in-header documentation\n>\n> I think these are very positive directions.  I like the fact that your\n> initial plan benefits all of us, whether or not libification is\n> ultimately achieved.\n>\n> [...]\n> > So what's next? Naturally, I'm looking forward to a spirited\n> > discussion about this topic - I'd like to know which concerns haven't\n> > been addressed and figure out whether we can find a way around them,\n> > and generally build awareness of this effort with the community.\n>\n> I'm curious whether clarifying scope/layering and cleaning up\n> interfaces might mean you'd be interested in things like:\n>   * https://github.com/newren/git/commits/header-cleanups (which was\n> still WIP; I paused working on it because I figured people would see\n> it as big \"cleanup\" patches with no practical benefit)\n>   * https://github.com/gitgitgadget/git/pull/1149 (which has been\n> ready to submit for a _long_ time, but I just haven't yet)\n> or if these two things are orthogonal to what you have in mind.\n\nExtremely yes. :) Even \"small\" stuff like the need for header cleanups\nhave already come up for Glen and Calvin working on config and strbuf.\n\n>\n> > I'm also planning to send a proposal for a document full of \"best\n> > practices\" for turning Git code into libraries (and have quite a lot\n> > of discussion around that document, too). My hope is that we can use\n> > that document to help us during implementation as well as during\n> > review, and refine it over time as we learn more about what works and\n> > what doesn't. Having this kind of documentation will make it easy for\n> > others to join us in moving Git's codebase towards a clean set of\n> > libraries. I hope that, as a project, we can settle on some tenets\n> > that we all agree would make Git nicer.\n>\n> I like the sound of this.\n>\n> > After that, we're still hoping to target low-level libraries first - I\n> > certainly don't think it will make sense to ship a high-level `git\n> > commit` library in the near future, if ever - in the order that\n> > they're required from the VFS project we're working closely with. As\n> > far as I can tell right now, that's likely to cover object store and\n> > worktree access, as well as commit creation and pushing, but we'll see\n> > how planning shakes out over the next month or so. But Google's\n> > schedule should have no bearing on what others in the Git project feel\n> > is important to clean up and libify, and if there is interest in the\n> > rest of the project in converting other existing modules into\n> > libraries, my team and I are excited to participate in the review.\n>\n> If we can't libify something like commit, does that prevent libifying\n> higher level things like merge?\n>\n> I spent some time thinking about this a while back.  I tried to\n> carefully design merge-ort to improve the odds it could be used\n> elsewhere, maybe even libgit2.  (I hope it shows in the many comments\n> in merge-ort.h, and I think the \"priv\" field in particular allowing me\n> to hide the first ~300 lines of merge-ort.c declaring data structures\n> from users was really nice.)  However, I still had to accept data in\n> some known format.  So input parameters are things like trees and\n> commits.  But tree.h and commit.h both include object.h first, which\n> includes cache.h, which is basically all of Git.  And the functions I\n> call to interoperate with the system are similarly entangled.  So, the\n> odds of merge-ort being reused by libgit2 or otherwise used in a\n> library seems essentially nil, at least without some broader\n> libification effort.\n\nYeah, it depends a lot on the usage. For merge, it would be tricky to\nget the scope just right - should the merge library be responsible for\nlocating the merge-base? Should it just perform the conflict\nresolution? Something else?\n\nAs for \"I tried to include this thing which eventually included\ncache.h\" - yeah, I think we will be pulling stuff out of cache.h quite\nheavily. But IMO this falls under \"cleanups we want to do in Git\nanyway\" - I think it's widely understood that cache.h is not\nwell-scoped and could use improvement.\n\n>\n> I'd like to make that story better, time permitting (which is much\n> more of a challenge these days than it was a couple years ago), but\n> I'm curious if you or others have thoughts on something like that.\n>\n> > Much, much later on, I'm expecting us to form a plan around allowing\n> > \"plugins\" - that is, replacing library functionality we use today with\n> > an alternative library, such as an object store relying on a\n> > distributed file store like S3. Making that work well will also likely\n> > involve us coming up with a solution for dependency injection, and to\n> > begin using vtables for some libraries. I'm hoping that we can figure\n> > out a way to do that that won't make the Git source ugly. Around this\n> > time, I think it will make sense to buy into unit tests even more and\n> > start using an approach like mocking to test various edge cases. And\n> > at some point, it's likely that we'll want to make the interfaces to\n> > various Git libraries consistent with each other, which would involve\n> > some large-scale but hopefully-mechanical refactors.\n>\n> Would these plugins resemble the pluggable merge backends that was\n> added to builtin/merge.c?  Would it replace that mechanism with a\n> different one?  Would it be more like the refs backends?\n\nI suspect that it's likely to be most similar to the refs backend\nreplacement, although I investigated it only a little bit just now.\n\nThe pluggable merge backends are an interesting thought - right now\nall those alternatives are built in and we decide based on config,\nright? But if we were able to easily decide which library to link\nbased on config during runtime, then I could see that being an\nappealing use of plugins too. I wonder whether \"custom\" merge backends\nmake the story for people doing compiled asset storage in Git (like\ngame assets or hardware layouts, both of which famously merge\nhorribly) any easier.\n\n>\n> Would this plugin scheme allow us to, for example, use gitoxide[1] as\n> a clone replacement to make clones 2x as fast (and with half the\n> memory -- although I suspect they cheated and used sha1 instead of\n> sha1dc, so maybe it wouldn't really be 2x)?\n\nInteresting. It would depend on whether we can match the interface to\ngitoxide, or write a translation layer. I could see it! I'm also a\nlittle curious how much of that speedup is because of corner-cutting\n(since you mentioned skipping the collision detection) vs. how much is\ndue to Rust magic. In theory, building the Git CLI out of a handful of\nlibraries means that we could write some of those libraries in\nsomething besides C; in practice, I understand there's a\nmaintainability issue around introducing new languages into the pile\nof stuff the community is expected to understand and maintain. (For\nexample, I think many people don't like to touch git-gui, probably\nprimarily because it's in Tcl.)\n\n>\n> Oh, and it's totally okay if you don't know the answers to any or all\n> of my questions right now.  I'm just curious, because I've long\n> thought these kinds of directions would be good.  Since I've spent\n> time thinking about it, I have questions that I don't know the answers\n> to, but I figured it couldn't hurt to bounce them off others who are\n> thinking about this area.\n>\n> Anyway, it's a large pile of work that you're undertaking, and as\n> Junio comments elsewhere in this thread it's unclear if libification\n> can be achieved for a big enough component (and you seem to admit as\n> much in your email as well), but I applaud the general direction and\n> your initial plans.\n\nThanks for your thoughtful reply.\n\n - Emily\n\n>\n>\n> [1] https://github.com/Byron/gitoxide/discussions/579\n"},{"id":"472415","messageId":"CAJoAoZm+bS3pT_DOaQfafW6dyV=m3ZUs=oxNZ_sKdfFO7uxM9A@mail.gmail.com","threadId":"59261","inReplyTo":"Y/UXBw3Y9YnXUBIN@nand.local","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-21T22:27:47Z","receivedAt":"2023-02-21T22:28:04Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Tue, Feb 21, 2023 at 11:10 AM Taylor Blau <me@ttaylorr.com> wrote:\n>\n> On Fri, Feb 17, 2023 at 01:12:23PM -0800, Emily Shaffer wrote:\n> > This turned out pretty long-winded, so a quick summary before I dive in:\n> >\n> > - We want to compile parts of Git as independent libraries\n> > - We want to do it by making incremental code quality improvements to Git\n> > - Let's avoid promising stability of the interfaces of those libraries\n> > - We think it'll let Git do cool stuff like unit tests and allowing\n> >   purpose-built plugins\n> > - Hopefully by example we can convince the rest of the project to join\n> >   in the effort\n>\n> Like others, I am less interested in the VFS-specific components you\n> mention here, but I suspect that is just one particular instance of\n> something that would be benefited by making git internals exposed via a\n> linkable library.\n>\n> I don't have objections to things like reducing our usage of globals,\n> making fewer internal functions die() when they encounter an error, and\n> so on. But like Junio, I suspect that this is definitely an instance of\n> a \"devil's in the details\" kind of problem.\n>\n> That's definitely my main concern: that this turns out to be much more\n> complicated than imagined and that we leave the codebase in a worse\n> state without much to show.\n\nYeah, I'm really hoping we don't end up with ugly half-changes too.\nSome examples of \"partial credit\" that I'd be happy with:\n\n- Fewer internal libraries relying on globals like\nthe_repository/the_index/etc (we've already started this effort,\nlibification or no)\n- An \"ugly\" library interface becoming clearer and easier to use (and\ninternal callers updated)\n- Figuring out an \"error reporting type\" that works well for us\n\nThere are some things that *are* ugly, for example, calling a library\nvia a vtable. But I do feel comfortable waiting to introduce that kind\nof thing until we really need it, at which point I suspect we'll have\nalready made some successful strides with libification in general.\n\nIt's not so great to just trust me to say \"I promise not to make ugly\nchanges\" - I'd appreciate the community's help pushing back if we\npropose doing something in an untidy way without clear justification.\n\n> A lesser version of that outcome would be\n> that we cause a lot of churn in the tree with not much to show either.\n\nI'm actually not so concerned about this! The \"churn\", as I see it,\ncomes in the form of code cleanup that already makes Git more\nunderstandable for Git hackers. We do spend some time on that now, as\na project, but I wouldn't be unhappy if we spent even more :)\n\n>\n> So I think we'd want to see some more concrete examples with clear\n> benefits to gauge whether this is a worthwhile direction. I think that\n> strbuf.h is too trivial an example to demonstrate anything useful. Being\n> able to extract config.h into its own library so that another non-Git\n> program could link against it and implement 'git config'-like\n> functionality would be much more interesting.\n\nSure - I'm also looking forward to seeing it.\n\nThanks for your thoughtful reply.\n - Emily\n"},{"id":"472420","messageId":"xmqqk00aczwj.fsf@gitster.g","threadId":"59261","inReplyTo":"CAJoAoZn7Nt37Eh17dpLDK+YX2BaEaAaii2rJPXO3L0BmQQkcgQ@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-22T00:22:20Z","receivedAt":"2023-02-22T00:22:26Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <nasamuffin@google.com> writes:\n\n>> For example, \"die() is inconvenient to callers, let's propagate\n>> errors up the callchain\" is an easy thing to say, but it would take\n>> much more than \"let's propagate errors up\" to libify something like\n>> check_connected() to do the same thing without spawning a separate\n>> process that is expected to exit with failure.\n>\n> Because the error propagation path is complicated, you mean? Or\n> because the cleanup is painful?\n\nBoth.\n\nThe amount of data the caller may want to learn about an error may\nnot be uneven, depending on the caller even for a single function.\nAnd yes, cleaning up of shared resources like object flag bits after\na traversal, especially a failed one, would be very painful unless\nthe processing is designed from day one to allow it (and the\nrevision traversal codepath is not).\n\n> ... is it reasonable to consider, for example, further\n> cleanup of cache.h as part of this effort? Is it reasonable to rework\n> an ugly circular dependency between two headers as a prerequisite to\n> doing library work around one of them?\n\nI am not sure about which two headers you are talking about, but if\nthere is circular dependency that can be untangled, it would be a\nreasonable preliminary clean-up work.  I am not sure if that is\n\"prerequisite\"---it is up to folks who want to design how the\n\"libification\" work goes.\n"},{"id":"472426","messageId":"86e7f589-7576-5ed3-3dc9-5dec9ca346eb@github.com","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Victoria Dye","fromEmail":"vdye@github.com","sentAt":"2023-02-22T01:44:35Z","receivedAt":"2023-02-22T01:44:44Z","isPatch":false,"sender":{"key":"vdye@github.com","avatar":"https://avatars.githubusercontent.com/u/3619353?v=4"},"body":"Emily Shaffer wrote:\n> Hi folks,\n> \n> As I mentioned in standup this week[1], my colleagues and I at Google\n> have become very interested in converting parts of Git into libraries\n> usable by external programs. In other words, for some modules which\n> already have clear boundaries inside of Git - like config.[ch],\n> strbuf.[ch], etc. - we want to remove some implicit dependencies, like\n> references to globals, and make explicit other dependencies, like\n> references to other modules within Git. Eventually, we'd like both for\n> an external program to use Git libraries within its own process, and\n> for Git to be given an alternative implementation of a library it uses\n> internally (like a plugin at runtime).\n\nI see why you've linked these two objectives (there could be some shared\ninfrastructure between them), but they're ultimately separate concerns, with\nseparate needs & tradeoffs. By trying to design with a unified solution for\nboth, though, this proposal's solution may be broader than it needs to be\nand comes with some serious developer experience ramifications.\n\nFor example, looking at the  \"call libgit.so from other applications\"\nobjective: one way I could imagine implementing this is by creating a stable\nAPI wrapper around various hand-picked internal functions. If those internal\nfunctions change (e.g., adding a function argument to handle a new flag in a\nbuiltin), the API can remain stable (by setting a sane default) without\nreally limiting the contributor adding a new feature. \n\nHowever, my reading of this proposal (although correct me if I'm wrong) is\nthat you want to use that same external-facing API as a plugin interface for\nGit. In that case, the contributor can't just add a new argument and\ntrivially modify the wrapper; they either break the API contract, or create\na new version of the API and continue to support the old one. While that's\ntechnically doable, it's a pretty sizeable increase to the amount of \"stuff\"\ncontributors need to keep track of and limits the project's flexibility. \n\n> \n> This turned out pretty long-winded, so a quick summary before I dive in:\n> \n> - We want to compile parts of Git as independent libraries\n> - We want to do it by making incremental code quality improvements to Git\n> - Let's avoid promising stability of the interfaces of those libraries\n> - We think it'll let Git do cool stuff like unit tests and allowing\n> purpose-built plugins\n> - Hopefully by example we can convince the rest of the project to join\n> in the effort\n> \n> My team has spent the past year or so trying to make improvements to\n> Git's behavior with submodules, and we found that the current\n> structure of Git is quite tricky to work with: because Git doesn't\n> execute on second repositories in the same process well, recursing\n> into submodules typically involves spawning child processes, and\n> piping new arguments through the helpers around those child processes,\n> and then through Git's typical codepaths, is very tricky. After\n> spending more than a year trying to make improvements, we have very\n> little to show for it, largely as a result of the difficulty of\n> passing information between superprojects and submodules.\n> \n> It seems like being able to invoke parts of Git as a library, or Git\n> being able to invoke custom libraries, does a lot of good for the Git\n> project:\n> \n> - Having clear, modular libraries makes it easy to find the code\n> responsible for some behavior, or determine where to add something\n> new.\n\nThis is more an issue of code & repository organization, and isn't really an\ninherent outcome of lib-ifying Git. That said, like the cleanup you mention\nlater, I would not be opposed to a dedicated code reorganization proposal.\n\n> - Operations recursing into submodules can be run in-process, and the\n> callsite makes it very clear which repository context is being used;\n> if we're not subprocessing to recurse, then we can lose the awkward\n> git-submodule.sh -> builtin/submodule__helper.c -> top-level Git\n> process codepath.\n\nInteresting - I wasn't aware this was how submodules worked internally. Is\nthere a specific reason this can't be refactored to perform the recursion\nin-process?\n\n> - Being able to test libraries in isolation via unit tests or mocks\n> speeds up determining the root cause of bugs.\n\nThe addition of unit tests is a project of its own, and one I don't believe\nis intrinsically tied to this one. More on this later.\n\n> - If we can swap out an entire library, or just a single function of\n> one, it's easy to experiment with the entire codebase without sweeping\n> changes.\n\nWhat sort of \"experiment\" are you thinking of here? Anyone can make internal\nchanges to Git in their own clone of the repository, then build and test it.\nShipping a custom fork of Git doesn't seem any less complex than building a\nplugin library, shipping that, and injecting it into an existing Git\ninstall.\n\n> \n> The ability to use Git as a library also makes it easier for other\n> tooling to interact with a Git repository *correctly*. As an example,\n> `repo` has a long history of abusing Git by directly manipulating the\n> gitdir[2], but if it were written in a world where Git exists as\n> easy-to-use libraries, it probably wouldn't have needed to, as it\n> could have invoked Git directly or replaced the relevant modules with\n> its own implementation. Both `repo`[3] and `git-gui[4]` have\n> reimplemented logic from git.git. Other interfaces that cooperate with\n> Git's filesystem storage layer, like `scalar` or `jj`[5], would be\n> able to interop with a Git repository without having to reimplement\n> custom logic or keep up with new Git changes.\n\nI'd alternatively suggest that these use cases could be addressed with\nbetter plumbing command support and better (public) documentation of our\non-disk data structures. In many cases, we're treating on-disk data\nstructures as external-facing APIs anyway - the index [1], packfiles [2],\netc. are versioned. We could definitely add documentation that's more \ndirectly targeted at external integrators.\n\n[1] https://git-scm.com/docs/index-format\n[2] https://git-scm.com/docs/gitformat-pack\n\n> \n> Of course, there's a reason Google wants it, too. We've talked\n> previously about wanting better integration between Git and something\n> like a VFS; as we've experimented with it internally, we've found a\n> couple of tricky areas:\n> \n> - The VFS relies on running Git commands to determine the state of the\n> repository. However, these Git commands interact with the gitdir or\n> worktree, which is populated by the VFS. For example, parsing a\n> .gitattributes or .gitmodules which is already stored in the VFS\n> requires the VFS to provide a POSIX file handle, spawn a Git\n> subprocess, populate other files needed by that subprocess (like\n> .git/config), and finally collect the output stream of the subprocess.\n> As you can imagine, this interaction of VFS -> Git -> VFS [-> Git]\n> creates all sort of complications. The alternative is for the VFS to\n> write its own parser (or use a library like libgit2; more on that\n> later). But having a Git library means that a subset of Git\n> functionality can happen in-process, and that filesystem access could\n> be replaced by the VFS directly providing high-level objects or plain\n> bytestreams.\n> \n> - A user running `git status` in a directory controlled by the VFS\n> will require the VFS to populate the entire (large) worktree - even\n> though the VFS is sure that only one file has been modified. The\n> closest we've come with an alternative is an orchestrated use of\n> sparse-checkout - but modifying the sparse-checkout configs\n> automatically in response to the user's filesystem operations takes us\n> right back to the previous point. If Git could take a plugin and\n> replacement for the object store that directly interacts with the VFS\n> daemon, a layer of complexity would disappear and performance would\n> improve.\n\nBoth of these points sound more like implementation details/quirks of the\nVFS prototype you're building than hard blockers coming from Git. For\nexample, you mention 'git status' needing to populate a full worktree -\ncould it instead use a mechanism similar to FSMonitor to avoid scanning for\nfiles you know are virtualized? While there'd likely be some Git changes\ninvolved (some of which might be a bit plugin-y), they'd be much more\ntightly scoped than full libification.\n\n> The good news is that for practical near-term purposes, \"libification\"\n> mostly means cleanups to the Git codebase, and continuing code health\n> work that the project has already cared about doing:\n> \n> - Removing references to global variables and instead piping them\n> through arguments\n> - Finding and fixing memory leaks, especially in widely-used low-level code\n> - Reducing or removing `die()` invocations in low-level code, and\n> instead reporting errors back to callers in a consistent way\n> - Clarifying the scope and layering of existing modules, for example\n> by moving single-use helpers from the shared module's scope into the\n> single user's scope\n> - Making module interfaces more consistent and easier to understand,\n> including moving \"private\" functions out of headers and into source\n> files and improving in-header documentation\n> \n> Basically, if this effort turns out not to be fruitful as a whole, I'd\n> like for us to still have left a positive impact on the codebase.\n\nThese cleanups/best practices are a great idea regardless of any potential\nlib-ification. I'll be sure to keep them in mind in any future changes I\nmake or review.\n\n> \n> In the longer term, if Git has libraries with easily-replaced\n> dependencies, we get a few more benefits:\n> \n> - Unit tests. We already have some in t/helper/, but if we can replace\n> all the dependencies of a certain library with simple stubs, it's\n> easier for us to write comprehensive unit tests, in addition to the\n> work we already do introducing edge cases in bash integration tests.\n\nThis doesn't really follow as a direct consequence of making libgit\nexternally facing. AFAIK, there's nothing explicitly stopping us from\nwriting or integrating a unit test framework now.\n\nLike the \"make libgit public\" vs. \"allow for custom backends in Git\", I\nthink this is a separate project with its own tradeoffs to evaluate.\n\n> - If our users can use plugins to improve performance in specific\n> scenarios (like a VFS-aware object store in the VFS case I cited\n> above), then Git works better for them without having to adopt a\n> different workflow, such as using an alternative tool or wrapper.\n\nI'm curious as to what you want the user experience for this to look like.\nHow would Git know to dynamically load a plugin? How does a user configure\n(or disable) plugins? Will a plugin need to replace all of the functions in\na given library, or would Git be able to \"fall back\" on its own internal\nimplementation?\n\n> - An easy-to-understand modular codebase makes it easy for new\n> contributors to start hacking and understand the consequences of their\n> patch.\n\nAs noted earlier, I think improving the developer experience in this way is\nindependent of the development of external APIs.\n\n> \n> Of course, we haven't maintained any guarantee about the consistency\n> of our implementation between releases. I don't anticipate that we'll\n> write the perfect library interface on our first try. So I hope that\n> we can be very explicit about refusing to provide any compatibility\n> guarantee whatsoever between versions for quite a long time. On\n> Google's end, that's well-understood and accepted. As I understand,\n> some other projects already use Git's codebase as a \"library\" by\n> including it as a submodule and using the code directly[6]; even a\n> breakable API seems like an improvement over that, too.\n\nThis hints at, but sidesteps, a really important aspect of the long-term\ngoals of this project - are you planning on having us start guaranteeing\nconsistency once there's an external-facing API available? \n\nIf not, that sounds like an unpleasant user experience (and one prone to\nHyrum's Law [3] at that). \n\nIf so, Git contributors will either be much more constrained in the\nintroduction of new features, or we'll end up with a mess of\nbackward-compatibility APIs.\n\n[3] https://www.hyrumslaw.com/\n\n> \n> So what's next? Naturally, I'm looking forward to a spirited\n> discussion about this topic - I'd like to know which concerns haven't\n> been addressed and figure out whether we can find a way around them,\n> and generally build awareness of this effort with the community.\n> \n> I'm also planning to send a proposal for a document full of \"best\n> practices\" for turning Git code into libraries (and have quite a lot\n> of discussion around that document, too). My hope is that we can use\n> that document to help us during implementation as well as during\n> review, and refine it over time as we learn more about what works and\n> what doesn't. Having this kind of documentation will make it easy for\n> others to join us in moving Git's codebase towards a clean set of\n> libraries. I hope that, as a project, we can settle on some tenets\n> that we all agree would make Git nicer.\n\nI think the use of multiple libraries is part of the potentially suboptimal\nbalance between the \"load Git's internals as a library\" and \"let Git use\nplugins in place of its own implementations\" goals. The former could leave a\nuser in dependency hell if there are lots of small libraries to load, while\nthe latter may work better with smaller-scoped libraries (to avoid needing\nto implement more than is useful). And, the functionality that's useful to a\nprogram invoking Git via library may not be the same as what someone would\nwant to replace via plugin. \n\n> \n> From the rest of my own team, we're planning on working first on some\n> limited scope, low-level libraries so that we can all see how the\n> process works. We're starting with strbuf.[ch] (as it's used\n> everywhere with few or no dependencies and helps us guarantee string\n> safety at API boundaries), config.[ch] (as many external tools are\n> probably interested in parsing Git config formatted files directly),\n> and a subset of operations related to the object store. These starting\n> points are intended to have a small impact on the codebase and teach\n> us a lot about logistics and best practices while doing these kinds of\n> conversions.\n> \n> After that, we're still hoping to target low-level libraries first - I\n> certainly don't think it will make sense to ship a high-level `git\n> commit` library in the near future, if ever - in the order that\n> they're required from the VFS project we're working closely with. As\n> far as I can tell right now, that's likely to cover object store and\n> worktree access, as well as commit creation and pushing, but we'll see\n> how planning shakes out over the next month or so. But Google's\n> schedule should have no bearing on what others in the Git project feel\n> is important to clean up and libify, and if there is interest in the\n> rest of the project in converting other existing modules into\n> libraries, my team and I are excited to participate in the review.\n\nAs a general note, this libification project - its justification,\nmilestones, organization, etc. - seems to be primarily driven by the needs\nof a VFS that is entirely opaque to the Git community you're proposing to.\nAs it stands, reviewers are put in a position of needing to accept, without\nmuch evidence, that this is the best (or only) possible approach for meeting\nthose needs. \n\nWhile I'm normally content with \"scratch your own itch\"-type changes, what\nyou're proposing is a major paradigm shift in how Git is written,\nmaintained, and used. I'm not comfortable accepting that level of impact\nwithout at least being able to evaluate whether the problem can be solved\nsome other way.\n\n> \n> Much, much later on, I'm expecting us to form a plan around allowing\n> \"plugins\" - that is, replacing library functionality we use today with\n> an alternative library, such as an object store relying on a\n> distributed file store like S3. Making that work well will also likely\n> involve us coming up with a solution for dependency injection, and to\n> begin using vtables for some libraries. I'm hoping that we can figure\n> out a way to do that that won't make the Git source ugly. Around this\n> time, I think it will make sense to buy into unit tests even more and\n> start using an approach like mocking to test various edge cases. And\n> at some point, it's likely that we'll want to make the interfaces to\n> various Git libraries consistent with each other, which would involve\n> some large-scale but hopefully-mechanical refactors.\n\nImplementing dependency injection (via vtable or otherwise) and unit test\nmocking is a massive undertaking on its own, let alone everything else\ndescribed so far. You've outlined some clear, fairly unobtrusive first steps\nto the overarching proposal, but there's a lot of detail missing from the\nplans for later steps. While there's always risk of a project getting\nderailed or blocked later on due to some unforeseen issue, that risk seems\nparticularly high here given the size & scope of all of these components. \n\nIntermediate milestones/goals, each of which is valuable on its own,\noutlined in the proposal would help allay those fears (for me, at least).\n\n> \n> I'm looking forward to the discussion!\n> \n\nSo, to summarize my thoughts into some (hopefully) actionable feedback:\n\n- It's really important that the proposal presents a clear, concrete\n  long-term vision for this project. \n  - What is the desired end state, in terms of what is built and installed\n    with Git?\n  - What will be expected of other Git contributors to support this design?\n    What happens if someone wants to \"break\" an API?\n  - What is the impact to users (including security implications)?\n- The scope for what you've proposed is pretty huge (which comes with a lot\n  of risk), but I think it could be broken into smaller, *independent*\n  pieces. \n- It's hard to judge or suggest adjustments to this proposal without knowing\n  the specific challenges you're facing with Git as it is.\n\nFinally, thanks for sending this email & starting a discussion. It's an\ninteresting topic and I'm looking forward to seeing everyone's perspectives\non the matter.\n\n>  - Emily\n> \n> 1: https://colabti.org/irclogger/irclogger_log/git-devel?date=2023-02-13#l29\n> 2: https://gerrit.googlesource.com/git-repo/+/refs/heads/main/docs/internal-fs-layout.md\n> 3: https://gerrit.googlesource.com/git-repo/+/refs/heads/main/git_config.py\n> 4: https://github.com/git/git/blob/master/git-gui/git-gui.sh#L305\n> 5: https://github.com/martinvonz/jj\n> 6: https://github.com/glandium/git-cinnabar\n\n"},{"id":"472438","messageId":"CABPp-BFohpW6y1piVDYPbQx_vL1jbD0PZjkiAHcj_p-EiDoJXg@mail.gmail.com","threadId":"59261","inReplyTo":"CAJoAoZm+TkCL0Jpg_qFgKottxbtiG2QOiY0qGrz3-uQy+=waPg@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2023-02-22T08:23:46Z","receivedAt":"2023-02-22T08:24:04Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Tue, Feb 21, 2023 at 2:07 PM Emily Shaffer <nasamuffin@google.com> wrote:\n>\n> On Fri, Feb 17, 2023 at 8:05 PM Elijah Newren <newren@gmail.com> wrote:\n> >\n> > On Fri, Feb 17, 2023 at 1:45 PM Emily Shaffer <nasamuffin@google.com> wrote:\n[...]\n> > > So what's next? Naturally, I'm looking forward to a spirited\n> > > discussion about this topic - I'd like to know which concerns haven't\n> > > been addressed and figure out whether we can find a way around them,\n> > > and generally build awareness of this effort with the community.\n> >\n> > I'm curious whether clarifying scope/layering and cleaning up\n> > interfaces might mean you'd be interested in things like:\n> >   * https://github.com/newren/git/commits/header-cleanups (which was\n> > still WIP; I paused working on it because I figured people would see\n> > it as big \"cleanup\" patches with no practical benefit)\n> >   * https://github.com/gitgitgadget/git/pull/1149 (which has been\n> > ready to submit for a _long_ time, but I just haven't yet)\n> > or if these two things are orthogonal to what you have in mind.\n>\n> Extremely yes. :) Even \"small\" stuff like the need for header cleanups\n> have already come up for Glen and Calvin working on config and strbuf.\n\nOk, I'll clean up what I've got and submit.\n\n[...]\n> > > Much, much later on, I'm expecting us to form a plan around allowing\n> > > \"plugins\" - that is, replacing library functionality we use today with\n> > > an alternative library, such as an object store relying on a\n> > > distributed file store like S3. Making that work well will also likely\n> > > involve us coming up with a solution for dependency injection, and to\n> > > begin using vtables for some libraries. I'm hoping that we can figure\n> > > out a way to do that that won't make the Git source ugly. Around this\n> > > time, I think it will make sense to buy into unit tests even more and\n> > > start using an approach like mocking to test various edge cases. And\n> > > at some point, it's likely that we'll want to make the interfaces to\n> > > various Git libraries consistent with each other, which would involve\n> > > some large-scale but hopefully-mechanical refactors.\n> >\n> > Would these plugins resemble the pluggable merge backends that was\n> > added to builtin/merge.c?  Would it replace that mechanism with a\n> > different one?  Would it be more like the refs backends?\n>\n> I suspect that it's likely to be most similar to the refs backend\n> replacement, although I investigated it only a little bit just now.\n>\n> The pluggable merge backends are an interesting thought - right now\n> all those alternatives are built in and we decide based on config,\n> right?\n\nWhile we have several built in merge backends (recursive, resolve,\nort, octopus, ours, subtree), it is not limited to built-in\nalternatives.  If someone creates a \"git-merge-$STRATEGY\" executable\nand runs `git merge -s $STRATEGY` then builtin/merge.c will fork their\nsubcommand to try to resolve the merge.  (Users can even specify\nsomething like `-s $STRATEGY1 -s $STRATEGY2` to have git find the best\nstrategy to use.)  The subcommand is then expected to update the\nworking tree and index and return an exit status signalling whether\nthe merge was clean.  It's very much built around assuming you\ncurrently have a branch checked out and you want to merge into that\nbranch.\n\nWe do not know how widely used this feature is, but it's kept us from\nfixing some API mistakes.  For example, a user strategy can return an\nexit status of 2 signalling that it cannot even consider the merge in\nquestion (e.g. most strategies cannot handle octopus merges).\nHowever, the merge strategies are allowed to muck with the working\ndirectory and index and leave them in a totally messed up state prior\nto returning the \"2\" signalling that the given merge strategy is\ninappropriate and another should be selected instead.  That means\nbuiltin/merge.c is required after calling any given strategy and if it\ndoes not succeed, go and \"clean everything up\".\n"},{"id":"472444","messageId":"13e0737a-7d66-7122-9dab-f7659948cda3@github.com","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Derrick Stolee","fromEmail":"derrickstolee@github.com","sentAt":"2023-02-22T14:55:39Z","receivedAt":"2023-02-22T14:55:49Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 2/17/2023 4:12 PM, Emily Shaffer wrote:\n\n> This turned out pretty long-winded, so a quick summary before I dive in:\n>\n> - We want to compile parts of Git as independent libraries\n> - We want to do it by making incremental code quality improvements to Git\n> - Let's avoid promising stability of the interfaces of those libraries\n> - We think it'll let Git do cool stuff like unit tests and allowing\n> purpose-built plugins\n> - Hopefully by example we can convince the rest of the project to join\n> in the effort\n\nAs I was thinking about this proposal, I realized that my biggest concern\nis that we are jumping to the \"what\" and \"how\" without fully exploring the\n\"why\"s, so I've reorganized some of your points along those lines, so I\ncan respond to each in turn.\n\n## Why: Modular code\n\n> - Having clear, modular libraries makes it easy to find the code\n> responsible for some behavior, or determine where to add something\n> new.\n\n> - An easy-to-understand modular codebase makes it easy for new\n> contributors to start hacking and understand the consequences of their\n> patch.\n\nGenerally, having a modular codebase does not guarantee that things are\neasy to find or that consequences are understood. The correct abstractions\nare key in this, as well as developing boundaries that do not leak into\neach other.\n\nUnless done correctly, we can have significant issues with how things are\ncommunicated through layers of abstraction. We already have that problem\nwhen we want to add a new option to a builtin and need to modify three or\nfour method prototypes in order to communicate a change of behavior to the\nright layer.\n\nI've also found that where we have abstractions, such as the refs code or\nthe transport code, that the indirection created by the vtables is more\nconfusing to discover by reading the code. I've needed the use of a\ndebugger to even discover the call stack for certain operations.\n\n> - If we can swap out an entire library, or just a single function of\n> one, it's easy to experiment with the entire codebase without sweeping\n> changes.\n\nI don't understand this one too much. We already have this ability through\nforking the repository and mutating the code to do different things. Why\ndo we need modules (with vtables and all the rest of the pain that goes\ninto it) to make such swaps?\n\nThis only really matters in the fullness of the libification effort: if\neverything is mocked, then subsystems can be replaced wholesale. That's a\nlot of work just to avoid experimenting with a fork.\n\n...\n\nYou said this about how to achieve modular code:\n\n> - Removing references to global variables and instead piping them\n> through arguments\n> - Finding and fixing memory leaks, especially in widely-used low-level code\n> - Reducing or removing `die()` invocations in low-level code, and\n> instead reporting errors back to callers in a consistent way\n...\n> - Making module interfaces more consistent and easier to understand,\n> including moving \"private\" functions out of headers and into source\n> files and improving in-header documentation\n\nI think these mechanisms should be universally welcomed. One thing that is\ntricky is how we can guarantee that an API can be marked as \"done\" and\nmaintained that way in perpetuity. (Unit tests may help with that, so\nlet's come back to this then.)\n\n> - Clarifying the scope and layering of existing modules, for example\n> by moving single-use helpers from the shared module's scope into the\n> single user's scope\n\nThis one is a bit murky to me. It sounds like that if there is only one\ncaller to a strbuf_*() method that it should be made 'static' inside the\ncaller instead of global in strbuf.h for use by future callers. Making\nthis a goal in itself is probably more likely to see methods moving around\nmore frequently due to the frequency of their use, rather than natural\ngroupings based on which data structures are being mutated.\n\nWhat measurement would we optimize for? How could we maintain such\nstandards as we go? The Git codebase doesn't really have a firm\n\"architecture\" other than \"builtin code calls libgit.a code, which calls\nother libgit.a code as necessary\". There aren't really strong layers\nbetween things. Everything assumes it can look up an object or load\nconfig. Are there organizational things that we can do in this space that\nwould help decoupling things before jumping to libification?\n\n\n## Why: Submodules\n\n> - Operations recursing into submodules can be run in-process, and the\n> callsite makes it very clear which repository context is being used;\n> if we're not subprocessing to recurse, then we can lose the awkward\n> git-submodule.sh -> builtin/submodule__helper.c -> top-level Git\n> process codepath.\n\nPreviously, there was an effort to replace dependencies on the_repository\nin favor of repository structs being passed through method parameters.\nThis seems to have fallen off of the priority list, and prevous APIs that\nwere converted have regressed in their dependencies.\n\nShould we consider restarting that effort as an early goal? Should a\nrepository struct be the \"god object\" that also includes the default\nimplemenations of these modules until the modules can be teased apart and\nno longer care about the entire repository?\n\n## Why: Unit testing\n\n> - Unit tests. We already have some in t/helper/, but if we can replace\n> all the dependencies of a certain library with simple stubs, it's\n> easier for us to write comprehensive unit tests, in addition to the\n> work we already do introducing edge cases in bash integration tests.\n\nUnit tests are very valuable in keeping the codebase stable and reducing\nthe chance that code was mutated by accident. This is achieved by very\nrigid test infrastructure, requiring _replacing_ methods with mocks in\ncareful ways in order to test that behavior matches expectation. The\n\"simple stubs\" are actually carefully crafted to verify their inputs and\nprovide a carefully selected return value.\n\nThe difficulty of writing unit tests (or mutating the code after writing\nunit tests) is balanced by the ability to more explicitly create error\nconditions that are unlikely to occur in real test cases. Cases such as\nconnection errors, file I/O problems, or other unexpected responses from\nthe mocked methods, are all relatively difficult to check via end-to-end\ntests.\n\nI've personally found that the hardest unit test to write is the _first_\none, and after that the rest are simpler. The difficulty arises in making\nthe code testable in the first place, as well as creating infrastructure\nfor the most-common unit test scenarios. My expectation is that it will\ntake significant changes to the Git codebase to make any non-trivial unit\ntests outside of these very isolated cases that are presented here: strbuf\nmanipulation is easy to unit-test and config methods operating on\nconfig_set structs should be similar. However, config methods that\ninteract with a repository and its config file(s) will be much harder to\ntest unless we start mocking filesystem interactions. We could create\ntest-tool helpers that load a full repository and check the config files\npresent on the filesystem, but now we're talking about integration tests\ninstead of unit tests.\n\n> - Being able to test libraries in isolation via unit tests or mocks\n> speeds up determining the root cause of bugs.\n\nI'm not sure I completely agree with this statement. Unit tests can help\nprevent introducing a new fault when mutating code to adjust for a bug,\nbut unit tests only guarantee expected behavior on the expected\npreconditions. None of this necessarily helps finding a root cause,\nespecially in the likely event that the root cause is due to combining\nunits and thus not covered by unit tests.\n\n## Why: Correct use\n\n> The ability to use Git as a library also makes it easier for other\n> tooling to interact with a Git repository *correctly*.\n\nIs there is a correct way to interact with a Git repository? We definitely\nprefer that the repository is updated by executing Git processes, leaving\nall internals up to Git. Libification thus has a purpose for scenarios\nthat do not have an appropriate Git builtin or where the process startup\ntime is too expensive for that use.\n\nHowever, would it not be preferrable to update Git to include these use\ncases in the form of new builtin operations? Even in the case where\nprocess startup is expensive, operations can be batched (as in 'git\ncat-file --batch').\n\n> Of course, we haven't maintained any guarantee about the consistency\n> of our implementation between releases. I don't anticipate that we'll\n> write the perfect library interface on our first try. So I hope that\n> we can be very explicit about refusing to provide any compatibility\n> guarantee whatsoever between versions for quite a long time. On\n> Google's end, that's well-understood and accepted. As I understand,\n> some other projects already use Git's codebase as a \"library\" by\n> including it as a submodule and using the code directly[6]; even a\n> breakable API seems like an improvement over that, too.\n\nOne thing we _do_ prioritize is the CLI as being backwards-compatible as\npossible. We already have that interface as a stable one that can be\ndepended upon, even when the Git executable is built from a fork with\ndifferent implementations of subsystems (or operating differently in the\npresence of custom config).\n\n## Why: Virtual Filesystem Support\n\n> Of course, there's a reason Google wants it, too. We've talked\n> previously about wanting better integration between Git and something\n> like a VFS; as we've experimented with it internally, we've found a\n> couple of tricky areas:\n>\n> - The VFS relies on running Git commands to determine the state of the\n> repository. However, these Git commands interact with the gitdir or\n> worktree, which is populated by the VFS.\n\nThe way VFS for Git solved this was to not virtualize the gitdir and to\npass immediately to the filesystem if a worktree file was already\n\"hydrated\". Of course, an early version was virtualizing the gitdir,\nincluding faking that every possible loose object was present and could be\nfound via a network call, but the same issues you are bringing up now were\nblockers for that approach.\n\n> - A user running `git status` in a directory controlled by the VFS\n> will require the VFS to populate the entire (large) worktree - even\n> though the VFS is sure that only one file has been modified. The\n> closest we've come with an alternative is an orchestrated use of\n> sparse-checkout - but modifying the sparse-checkout configs\n> automatically in response to the user's filesystem operations takes us\n> right back to the previous point. If Git could take a plugin and\n> replacement for the object store that directly interacts with the VFS\n> daemon, a layer of complexity would disappear and performance would\n> improve.\n\nThis is another case where the issue is that Git isn't aware that it is\noperating in a virtual environment and doesn't speak to the virtualization\nsystem directly. Git could talk to the virtualization layer as if it was a\nfilesystem monitor, and that would prevent a significant amount of these\nchanges. Preventing 'git checkout' from writing files to the worktree also\nrequires some coordination. The virtualization layer needs a signal that\nit will need to update its projection of the worktree (the\npost-index-change hook can do this)  and the Git process needs a way to\nmark the index with skip-worktree bits for the changed files (while\nkeeping the bits off for files that were previously hydrated and not\nchanged by the index update).\n\nIn this sense, we already have _some_ of the pluggability (through hooks)\nand could extend that ability more either via more hooks or by making Git\nitself aware that it's in a virtual filesystem. This pluggability could be\nextended by using pipe-based communication like the builtin FS Monitor,\nexcept that the communication is a new protocol that can speak to an\narbitrary implementation on the other side.\n\nI've said before that the goal of using git.git with a virtual filesystem\n(as-is, no custom bits) is unlikely to succeed _unless_ there are changes\ncontributed to git.git to make it aware of a \"filesystem virtualizer\". I\nalso don't think that inserting plugins is the right way to solve for\nthis. Users will want to use the Git CLI on their PATH for both virtual\nand non-virtual repositories, so the distinction between them needs to\nhappen at runtime, likely via Git config, hooks, or protocols.\n\n## How to achieve these goals\n\n> I'm also planning to send a proposal for a document full of \"best\n> practices\" for turning Git code into libraries (and have quite a lot\n> of discussion around that document, too). My hope is that we can use\n> that document to help us during implementation as well as during\n> review, and refine it over time as we learn more about what works and\n> what doesn't. Having this kind of documentation will make it easy for\n> others to join us in moving Git's codebase towards a clean set of\n> libraries. I hope that, as a project, we can settle on some tenets\n> that we all agree would make Git nicer.\n\nI like the idea of a \"best practices\" document, but I would hesitate to\nfocus on the libification and instead aim for the high-value items such\nas lack of globals and reduced memory leaks. How do we write such code?\nHow do we write (unit?) tests that guarantee those properties will be\nmaintained into the future?\n\n> From the rest of my own team, we're planning on working first on some\n> limited scope, low-level libraries so that we can all see how the\n> process works. We're starting with strbuf.[ch] (as it's used\n> everywhere with few or no dependencies and helps us guarantee string\n> safety at API boundaries), config.[ch] (as many external tools are\n> probably interested in parsing Git config formatted files directly),\n> and a subset of operations related to the object store. These starting\n> points are intended to have a small impact on the codebase and teach\n> us a lot about logistics and best practices while doing these kinds of\n> conversions.\n\nI can see that making a library version of config.h could be helpful to\nthird parties wanting to read Git config. In particular, I know that Git\nCredential Manager runs multiple 'git config' calls to get multiple values\nout. Perhaps time would be better spent creating a 'git config --batch'\nmode so these consumers could call one process, pass a list of config keys\nas input, and get a list of key-value pairs out? (Use a '-z' mode to allow\nnewlines in the values.)\n\nHowever, I'm not seeing value to anyone else to have strbuf.h available\nexternally. I'm also not seeing any value to the Git project by having\neither of these available as a library. I can only see increased developer\nfriction due to the new restrictions on making changes in these areas.\n\nIf this first effort is instead swapped to say \"we plan on introducing\nunit tests for these libraries\" then I think that is an easier thing to\nsupport.\n\nIt happens that these efforts will make it easier to make the jump to\nproviding APIs as a library. Not only are unit tests _necessary_ to\nsupport libification, they also make a smaller lift. Unit tests will\nalready create more friction when changing the API, since either new tests\nmust be written or old tests must be modified.\n\n\n### Summary\n\nGenerally, making architectural change is difficult. In addition to the\namount of code that needs to be moved, adjusted, and tested in new ways,\nthere is a cultural element that needs to be adjusted. Expecting Git's\ncode to be used as a library is a fundamentally different use case than we\nhave supported in the past, and most of our developer guidelines reflect\nthat. We don't hesitate to modify APIs. We prefer end-to-end tests over\nunit tests whenever possible.\n\nThere are focused bits of your proposal that I think will be universally\nwelcomed, and I think the best course of action in the short term is to\ndemonstrate improvements in those areas:\n\n * Reduced use of globals.\n * Reduced memory leaks.\n * Use error responses instead of die()s within low-level code.\n * Update header files to have appropriate visibility and documentation.\n\nThe things that we need to really focus on are how we can measure progress\nin these areas as well as how can we prevent regression in the future.\nUnit tests are one way to do this (especially with leak detection), but\nalso documentation and coding guidelines need to be updated based on the\nnew patterns discovered in this process.\n\nI mentioned the_repository earlier as a partially-complete architectural\nchange. I think the_index compatibility macros are in a similar boat of\nbeing half-completed (though, the last time I tried to remove these macros\nI was told that it wasn't a goal to remove the macros from builtins). The\nsparse-index's command_requires_full_index checks are also in this boat,\nso I understand the difficulty of getting a big change 100% complete (and\nin this case, the plan is to do at least one more GSoC project in the area\nto see what progress can be made that way).\n\nWe should be careful in committing to a very long-term direction without\nfirst delivering value on a shorter timeline.\n\nThanks,\n-Stolee\n"},{"id":"472452","messageId":"Y/ZsFuTKyfR+AQy5@coredump.intra.peff.net","threadId":"59261","inReplyTo":"CAJoAoZm+TkCL0Jpg_qFgKottxbtiG2QOiY0qGrz3-uQy+=waPg@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-02-22T19:25:10Z","receivedAt":"2023-02-22T19:26:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 21, 2023 at 02:06:55PM -0800, Emily Shaffer wrote:\n\n> > Does removing memory leaks also mean converting UNLEAK to free()?\n> \n> I suspect so - as I understand it, UNLEAK is a macro that resolves to\n> \"don't complain to me, compiler, I meant not to free it.\"\n\nCorrect. It is supposed to be used sparingly at the outermost level to\nsay \"I'm about to exit, so yes, we are leaking this, but no, it does not\nmatter\".\n\nSo...\n\n> > Thinking of things in a library context probably pushes us in that\n> > direction (though, alternatively, it might just highlight the question\n> > of what is considered \"low-level\" instead).\n> \n> I'm not sure whether use of UNLEAK has so much to do with \"low-level\"\n> or not. In cases when Git is being called as an ephemeral single-run\n> process, UNLEAK makes a lot of sense. In cases when Git is being\n> called in a long-lived process, UNLEAK is just a sign that says\n> \"there's a leak here\".  So I think the distinction is not low-level or\n> high-level, but more simply, within a library or not.\n\nI'd take \"low-level\" here to mean \"far down in the call stack\". That is,\ncode which is called potentially from a lot of places, and can't know\nwhat is going to happen afterwards.\n\nIn that case, such code calling UNLEAK() is already doing the wrong\nthing. And such code is a likely candidate for being called in a\nlib-ified long-running process, which means that ignoring the leaks is\nlikely to be more noticeable. :)\n\nThere are probably cases where code that is currently high-level becomes\nmore low-level, and will need to be adapted. For example, if cmd_diff()\nhas a static-local helper function for \"diff these two blobs\", and it\nknows it will run it exactly once, it is OK to UNLEAK() from there now.\nBut that may be a reasonable API to expose more widely, at which point\nit needs to stop UNLEAK()-ing and really free.\n\nJust my two cents as the originator of UNLEAK(). :)\n\n-Peff\n"},{"id":"472453","messageId":"Y/ZuR9zs3peUfO0g@coredump.intra.peff.net","threadId":"59261","inReplyTo":"CAJoAoZkMR9Acy7thVs-_e=Fz8wwjoDGDKb46wmwn8yxk0ODGow@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-02-22T19:34:31Z","receivedAt":"2023-02-22T19:34:35Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 17, 2023 at 02:49:51PM -0800, Emily Shaffer wrote:\n\n> > Personally, I'd like to see some sort of standard error type (whether\n> > integral or not) that would let us do more bubbling up of errors and\n> > less die().  I don't know if that's in the cards, but I thought I'd\n> > suggest it in case other folks are interested.\n> \n> Yes!!! We have talked about this a lot internally - but this is one\n> thing that will be difficult to introduce into Git without making\n> parts of the codebase a little uglier. Since obviously C doesn't have\n> an intrinsic to do this, we'll have to roll our own, which means that\n> manipulating it consistently at function exits might end up pretty\n> ugly. So hearing that there's interest outside of my team to come up\n> with such a type makes me optimistic that we can figure out a\n> neat-enough solution.\n\nHere are some past discussions on what I thought would be a good\napproach to error handling. The basic idea is to replace the \"pass a\nstrbuf that people shove error messages into\" approach with an error\ncontext struct that has a callback. And that callback can then stuff\nthem into a strbuf, or report them directly, or even die.\n\nThis thread sketches out the idea, though sadly I no longer have the\nmore fleshed-out patches I mentioned there:\n\n  https://lore.kernel.org/git/20160927191955.mympqgylrxhkp24n@sigill.intra.peff.net/\n\nAnd then the sub-thread starting here discusses a similar approach:\n\n  https://lore.kernel.org/git/20171103191309.sth4zjokgcupvk2e@sigill.intra.peff.net/\n\nIt does mean passing a \"struct error_context\" just about everywhere.\nThough since the context doesn't change very much and most calls are\njust forwarding it along, it would probably also be reasonable to have a\nthread-local global context, and push/pop from it (sort of a poor man's\ndynamic scoping).\n\nOne thing that strategy doesn't help with, though, that your\nlibification might want: it's not very machine-readable. The error\nreporters would still fundamentally be working with strings. So a\nlibified process can know \"OK, writing this ref failed, and I have some\nerror messages in a buffer\". But the calling code can't know specifics\nlike \"it failed because we tried to open file 'foo' and it got EPERM\".\nWe _could_ design an error context that stores individual errno values\nor codes in a list, but having each caller report those specific errors\nis a much bigger job (and ongoing maintenance burden as we catalogue and\ngive an identifier to each error).\n\n-Peff\n"},{"id":"472675","messageId":"CAJoAoZknYizS4peYgR4Zy5KUMEpFUbj5eREZoC_K5vUDXnAhng@mail.gmail.com","threadId":"59261","inReplyTo":"Y/ZuR9zs3peUfO0g@coredump.intra.peff.net","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-24T20:31:16Z","receivedAt":"2023-02-24T20:31:33Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, Feb 22, 2023 at 11:34 AM Jeff King <peff@peff.net> wrote:\n>\n> On Fri, Feb 17, 2023 at 02:49:51PM -0800, Emily Shaffer wrote:\n>\n> > > Personally, I'd like to see some sort of standard error type (whether\n> > > integral or not) that would let us do more bubbling up of errors and\n> > > less die().  I don't know if that's in the cards, but I thought I'd\n> > > suggest it in case other folks are interested.\n> >\n> > Yes!!! We have talked about this a lot internally - but this is one\n> > thing that will be difficult to introduce into Git without making\n> > parts of the codebase a little uglier. Since obviously C doesn't have\n> > an intrinsic to do this, we'll have to roll our own, which means that\n> > manipulating it consistently at function exits might end up pretty\n> > ugly. So hearing that there's interest outside of my team to come up\n> > with such a type makes me optimistic that we can figure out a\n> > neat-enough solution.\n>\n> Here are some past discussions on what I thought would be a good\n> approach to error handling. The basic idea is to replace the \"pass a\n> strbuf that people shove error messages into\" approach with an error\n> context struct that has a callback. And that callback can then stuff\n> them into a strbuf, or report them directly, or even die.\n\nThanks! I'll give these a read in detail soon, I appreciate you digging them up.\n\n>\n> This thread sketches out the idea, though sadly I no longer have the\n> more fleshed-out patches I mentioned there:\n>\n>   https://lore.kernel.org/git/20160927191955.mympqgylrxhkp24n@sigill.intra.peff.net/\n>\n> And then the sub-thread starting here discusses a similar approach:\n>\n>   https://lore.kernel.org/git/20171103191309.sth4zjokgcupvk2e@sigill.intra.peff.net/\n>\n> It does mean passing a \"struct error_context\" just about everywhere.\n> Though since the context doesn't change very much and most calls are\n> just forwarding it along, it would probably also be reasonable to have a\n> thread-local global context, and push/pop from it (sort of a poor man's\n> dynamic scoping).\n>\n> One thing that strategy doesn't help with, though, that your\n> libification might want: it's not very machine-readable. The error\n> reporters would still fundamentally be working with strings. So a\n> libified process can know \"OK, writing this ref failed, and I have some\n> error messages in a buffer\". But the calling code can't know specifics\n> like \"it failed because we tried to open file 'foo' and it got EPERM\".\n> We _could_ design an error context that stores individual errno values\n> or codes in a list, but having each caller report those specific errors\n> is a much bigger job (and ongoing maintenance burden as we catalogue and\n> give an identifier to each error).\n\nIs there a reason not to use this kind of struct and provide\nlibrary-specific error code enums, though, I wonder? You're right that\nparsing the error string is really bad for the caller, for anything\nbesides just logging it. But it seems somewhat reasonable to expect\nthat any call from config library returning an integer error code is\nreferring to enum config_errors...\n\n>\n> -Peff\n"},{"id":"472676","messageId":"CAJoAoZ=aoqWWfjUXg2VOeTkiVPsccqpWAm-yw2QzURGh-m-bNQ@mail.gmail.com","threadId":"59261","inReplyTo":"13e0737a-7d66-7122-9dab-f7659948cda3@github.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2023-02-24T21:06:05Z","receivedAt":"2023-02-24T21:06:24Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, Feb 22, 2023 at 6:55 AM Derrick Stolee <derrickstolee@github.com> wrote:\n>\n> On 2/17/2023 4:12 PM, Emily Shaffer wrote:\n>\n> > This turned out pretty long-winded, so a quick summary before I dive in:\n> >\n> > - We want to compile parts of Git as independent libraries\n> > - We want to do it by making incremental code quality improvements to Git\n> > - Let's avoid promising stability of the interfaces of those libraries\n> > - We think it'll let Git do cool stuff like unit tests and allowing\n> > purpose-built plugins\n> > - Hopefully by example we can convince the rest of the project to join\n> > in the effort\n>\n> As I was thinking about this proposal, I realized that my biggest concern\n> is that we are jumping to the \"what\" and \"how\" without fully exploring the\n> \"why\"s, so I've reorganized some of your points along those lines, so I\n> can respond to each in turn.\n\nThanks - and thanks for your patience in my replying, because there's\na lot here ;)\n\n>\n> ## Why: Modular code\n>\n> > - Having clear, modular libraries makes it easy to find the code\n> > responsible for some behavior, or determine where to add something\n> > new.\n>\n> > - An easy-to-understand modular codebase makes it easy for new\n> > contributors to start hacking and understand the consequences of their\n> > patch.\n>\n> Generally, having a modular codebase does not guarantee that things are\n> easy to find or that consequences are understood. The correct abstractions\n> are key in this, as well as developing boundaries that do not leak into\n> each other.\n\nAbsolutely agreed. I do think we will want to be very careful about\nhow to abstract things - and very open to reworking the interface when\nsomething isn't working. (See also, not guaranteeing ABI stability)\n\n>\n> Unless done correctly, we can have significant issues with how things are\n> communicated through layers of abstraction. We already have that problem\n> when we want to add a new option to a builtin and need to modify three or\n> four method prototypes in order to communicate a change of behavior to the\n> right layer.\n\nYes, that's true - and usually plumbing some new argument through and\ncoping with it appropriately in those 3-4 functions is because there's\nnot a way to silently provide the needed context to the bottom-most\nfunction, if I'm thinking of the same scenario you are.\n\nOne thing we discussed internally was the need to pass around some\ncontext struct to libraries, mostly to avoid reliance on globals. But\nI wonder if that same schema means that we can reduce the amount of\nargument plumbing needed when adding something to a low-level\ndependency?\n\n>\n> I've also found that where we have abstractions, such as the refs code or\n> the transport code, that the indirection created by the vtables is more\n> confusing to discover by reading the code. I've needed the use of a\n> debugger to even discover the call stack for certain operations.\n\nThis is a really good point, and part of why I'd like to hold off on\nintroducing vtables until we really painfully need them. I wonder when\nthat point will be - I guess if/when we start using runtime library\nlinking/selection it will be necessary.\n\n>\n> > - If we can swap out an entire library, or just a single function of\n> > one, it's easy to experiment with the entire codebase without sweeping\n> > changes.\n>\n> I don't understand this one too much. We already have this ability through\n> forking the repository and mutating the code to do different things. Why\n> do we need modules (with vtables and all the rest of the pain that goes\n> into it) to make such swaps?\n>\n> This only really matters in the fullness of the libification effort: if\n> everything is mocked, then subsystems can be replaced wholesale. That's a\n> lot of work just to avoid experimenting with a fork.\n\nYeah, I see your point. Git has also experimented with two backends\nselectable via config lots of times, and that's basically the same\nfunctional role as experimenting by replacing a library. But I sure\nwould like it if our project's default suggestion for experimenting\nwith different implementations wasn't just \"fork\" - forks are\ncumbersome to maintain and painful to reconvene with the mainline.\n\n>\n> ...\n>\n> You said this about how to achieve modular code:\n>\n> > - Removing references to global variables and instead piping them\n> > through arguments\n> > - Finding and fixing memory leaks, especially in widely-used low-level code\n> > - Reducing or removing `die()` invocations in low-level code, and\n> > instead reporting errors back to callers in a consistent way\n> ...\n> > - Making module interfaces more consistent and easier to understand,\n> > including moving \"private\" functions out of headers and into source\n> > files and improving in-header documentation\n>\n> I think these mechanisms should be universally welcomed. One thing that is\n> tricky is how we can guarantee that an API can be marked as \"done\" and\n> maintained that way in perpetuity. (Unit tests may help with that, so\n> let's come back to this then.)\n\nI'd really like to avoid that guarantee, actually! Even in the long\nterm, I think it's very likely that we'll want to continue extending\nsome APIs forever - for example, config learns new kinds of things\nwhich might change the API quite frequently. Let's not reach for\n\"done, never touch the interface again\" as a goal.\n\n>\n> > - Clarifying the scope and layering of existing modules, for example\n> > by moving single-use helpers from the shared module's scope into the\n> > single user's scope\n>\n> This one is a bit murky to me. It sounds like that if there is only one\n> caller to a strbuf_*() method that it should be made 'static' inside the\n> caller instead of global in strbuf.h for use by future callers. Making\n> this a goal in itself is probably more likely to see methods moving around\n> more frequently due to the frequency of their use, rather than natural\n> groupings based on which data structures are being mutated.\n\nLet me be a little more explicit about the strbuf example. strbuf\nincludes the purpose-built helper `strbuf_add_unique_abbrev()`, to\ntruncate the oid as much as possible to remain unique.\n(https://github.com/git/git/blob/master/strbuf.h#L633) I'm proposing\nthat this helper doesn't belong in strbuf.a, but instead in oid.a or\nsimilar. Still public, but since it relies on also having linked\noid.a, and it's in fact useless _without_ oid.a, it should live in\noid.a. (Reading my original phrasing back, I see how it didn't imply\nthis at all, so thanks for asking.)\n\n>\n> What measurement would we optimize for? How could we maintain such\n> standards as we go? The Git codebase doesn't really have a firm\n> \"architecture\" other than \"builtin code calls libgit.a code, which calls\n> other libgit.a code as necessary\". There aren't really strong layers\n> between things. Everything assumes it can look up an object or load\n> config. Are there organizational things that we can do in this space that\n> would help decoupling things before jumping to libification?\n\nThe example I cited above is a symptom of this weak layering you\ndescribe, yeah. In practice so far, these symptoms show up in the\nprocess of trying to isolate a single library, and require deciding\nhow to reorganize before we can build it as a library. But I think\npart of the reason we don't have strong layering now is precisely\nbecause we need it; I'm not sure that a soft reorganization will stick\nwithout some kind of enforcement (for example, in the form of unit\ntests operating on isolated components).\n\n>\n> ## Why: Submodules\n>\n> > - Operations recursing into submodules can be run in-process, and the\n> > callsite makes it very clear which repository context is being used;\n> > if we're not subprocessing to recurse, then we can lose the awkward\n> > git-submodule.sh -> builtin/submodule__helper.c -> top-level Git\n> > process codepath.\n>\n> Previously, there was an effort to replace dependencies on the_repository\n> in favor of repository structs being passed through method parameters.\n> This seems to have fallen off of the priority list, and prevous APIs that\n> were converted have regressed in their dependencies.\n>\n> Should we consider restarting that effort as an early goal?\n\nYes, I think that effort is a good early goal, but I think the reason\nit fell off is because there's not a lot of direct reward in moving\nbits and pieces off of `the_repository`. It's hard (impossible?) to\nguarantee with an integration test, and taken alone, it's too big of a\nchange for one person to complete in a single series. So I'm less\ninterested in saying \"start by getting rid of the_repository in 100%\nof the code!\" and a lot more interested in saying \"start by getting\nrid of 100% of the globals in config.[ch]\", if that makes sense.\n\n> Should a\n> repository struct be the \"god object\" that also includes the default\n> implemenations of these modules until the modules can be teased apart and\n> no longer care about the entire repository?\n\nHm, I'm not sure I follow this and the first impression I get is\npretty scary.  Could you explain a little more?\n\n> ## Why: Unit testing\n>\n> > - Unit tests. We already have some in t/helper/, but if we can replace\n> > all the dependencies of a certain library with simple stubs, it's\n> > easier for us to write comprehensive unit tests, in addition to the\n> > work we already do introducing edge cases in bash integration tests.\n>\n> Unit tests are very valuable in keeping the codebase stable and reducing\n> the chance that code was mutated by accident. This is achieved by very\n> rigid test infrastructure, requiring _replacing_ methods with mocks in\n> careful ways in order to test that behavior matches expectation. The\n> \"simple stubs\" are actually carefully crafted to verify their inputs and\n> provide a carefully selected return value.\n>\n> The difficulty of writing unit tests (or mutating the code after writing\n> unit tests) is balanced by the ability to more explicitly create error\n> conditions that are unlikely to occur in real test cases. Cases such as\n> connection errors, file I/O problems, or other unexpected responses from\n> the mocked methods, are all relatively difficult to check via end-to-end\n> tests.\n>\n> I've personally found that the hardest unit test to write is the _first_\n> one, and after that the rest are simpler. The difficulty arises in making\n> the code testable in the first place, as well as creating infrastructure\n> for the most-common unit test scenarios. My expectation is that it will\n> take significant changes to the Git codebase to make any non-trivial unit\n> tests outside of these very isolated cases that are presented here: strbuf\n> manipulation is easy to unit-test and config methods operating on\n> config_set structs should be similar. However, config methods that\n> interact with a repository and its config file(s) will be much harder to\n> test unless we start mocking filesystem interactions. We could create\n> test-tool helpers that load a full repository and check the config files\n> present on the filesystem, but now we're talking about integration tests\n> instead of unit tests.\n\nYeah, I see what you're saying. I think some of it depends on how we\nbuild the interfaces; it's difficult to mock an entire filesystem for\nan API that takes a path, but it's easy to give fake data to an API\nthat takes a string or bytestream. I suspect we'll want to be very\ncareful looking at how much work it is to build unit tests for a\ncertain library - overly complicated mocking setup seems like a code\nsmell to me.\n\n>\n> > - Being able to test libraries in isolation via unit tests or mocks\n> > speeds up determining the root cause of bugs.\n>\n> I'm not sure I completely agree with this statement. Unit tests can help\n> prevent introducing a new fault when mutating code to adjust for a bug,\n> but unit tests only guarantee expected behavior on the expected\n> preconditions. None of this necessarily helps finding a root cause,\n> especially in the likely event that the root cause is due to combining\n> units and thus not covered by unit tests.\n\nYeah, that's fair.\n\n>\n> ## Why: Correct use\n>\n> > The ability to use Git as a library also makes it easier for other\n> > tooling to interact with a Git repository *correctly*.\n>\n> Is there is a correct way to interact with a Git repository? We definitely\n> prefer that the repository is updated by executing Git processes, leaving\n> all internals up to Git. Libification thus has a purpose for scenarios\n> that do not have an appropriate Git builtin or where the process startup\n> time is too expensive for that use.\n\nThere are certainly some incorrect ways! We get a lot of firsthand\nexperience with `repo` tool's poor interaction with Git repositories\n:)\n\n> However, would it not be preferrable to update Git to include these use\n> cases in the form of new builtin operations? Even in the case where\n> process startup is expensive, operations can be batched (as in 'git\n> cat-file --batch').\n\nIn addition to expensive startup time, in some cases we've gotten\npushback about starting a subprocess at all, for security reasons.\nEspecially because Git loves to then start subprocesses itself.\n\n>\n> ## Why: Virtual Filesystem Support\n\nIn the initial email I CC'd a couple of the people working on our VFS,\nso I'll leave this section alone and leave it to them to reply to you.\n\n>\n> ## How to achieve these goals\n>\n> > I'm also planning to send a proposal for a document full of \"best\n> > practices\" for turning Git code into libraries (and have quite a lot\n> > of discussion around that document, too). My hope is that we can use\n> > that document to help us during implementation as well as during\n> > review, and refine it over time as we learn more about what works and\n> > what doesn't. Having this kind of documentation will make it easy for\n> > others to join us in moving Git's codebase towards a clean set of\n> > libraries. I hope that, as a project, we can settle on some tenets\n> > that we all agree would make Git nicer.\n>\n> I like the idea of a \"best practices\" document, but I would hesitate to\n> focus on the libification and instead aim for the high-value items such\n> as lack of globals and reduced memory leaks. How do we write such code?\n> How do we write (unit?) tests that guarantee those properties will be\n> maintained into the future?\n>\n> > From the rest of my own team, we're planning on working first on some\n> > limited scope, low-level libraries so that we can all see how the\n> > process works. We're starting with strbuf.[ch] (as it's used\n> > everywhere with few or no dependencies and helps us guarantee string\n> > safety at API boundaries), config.[ch] (as many external tools are\n> > probably interested in parsing Git config formatted files directly),\n> > and a subset of operations related to the object store. These starting\n> > points are intended to have a small impact on the codebase and teach\n> > us a lot about logistics and best practices while doing these kinds of\n> > conversions.\n>\n> I can see that making a library version of config.h could be helpful to\n> third parties wanting to read Git config. In particular, I know that Git\n> Credential Manager runs multiple 'git config' calls to get multiple values\n> out. Perhaps time would be better spent creating a 'git config --batch'\n> mode so these consumers could call one process, pass a list of config keys\n> as input, and get a list of key-value pairs out? (Use a '-z' mode to allow\n> newlines in the values.)\n>\n> However, I'm not seeing value to anyone else to have strbuf.h available\n> externally. I'm also not seeing any value to the Git project by having\n> either of these available as a library. I can only see increased developer\n> friction due to the new restrictions on making changes in these areas.\n\nThe value in strbuf is simply this: a Git config library should\nprobably not be returning char*+int to its callers, or even worse, an\nunwrapped linked list or map. It's too easy to make memory mistakes\ndealing with raw C types, and I'd like for us to avoid that by using\nthe types we already have.\n\n>\n> If this first effort is instead swapped to say \"we plan on introducing\n> unit tests for these libraries\" then I think that is an easier thing to\n> support.\n>\n> It happens that these efforts will make it easier to make the jump to\n> providing APIs as a library. Not only are unit tests _necessary_ to\n> support libification, they also make a smaller lift. Unit tests will\n> already create more friction when changing the API, since either new tests\n> must be written or old tests must be modified.\n>\n>\n> ### Summary\n>\n> Generally, making architectural change is difficult. In addition to the\n> amount of code that needs to be moved, adjusted, and tested in new ways,\n> there is a cultural element that needs to be adjusted. Expecting Git's\n> code to be used as a library is a fundamentally different use case than we\n> have supported in the past, and most of our developer guidelines reflect\n> that. We don't hesitate to modify APIs. We prefer end-to-end tests over\n> unit tests whenever possible.\n\nInherent in your summary I'm hearing a lot of concern about providing\nAPI stability. But I agree with you there - I don't think that we\nshould provide API stability, especially at the beginning while we're\nlearning how to write good APIs. I hope we can wait as long as\npossible to start promising backwards compatibility - maybe even never\n- and I hope we can stand our ground if we do receive bug reports\nabout breaking API stability on an API still labeled as unstable. I'm\nnot sure how I can make this statement more strongly :)\n\nIt seems that the Linux kernel has also managed to get away with this\nkind of purposeful instability\n(https://www.kernel.org/doc/Documentation/process/stable-api-nonsense.rst)\nand I'd be happy to try and follow their example. Another standard\nlibrary provided by Google (https://abseil.io/about/compatibility)\ndoes the same. So I do think it's possible for us to guarantee\ninstability, as it were. (Thanks, Jonathan Nieder, for the links\nhere.)\n\n>\n> There are focused bits of your proposal that I think will be universally\n> welcomed, and I think the best course of action in the short term is to\n> demonstrate improvements in those areas:\n>\n>  * Reduced use of globals.\n>  * Reduced memory leaks.\n>  * Use error responses instead of die()s within low-level code.\n>  * Update header files to have appropriate visibility and documentation.\n>\n> The things that we need to really focus on are how we can measure progress\n> in these areas as well as how can we prevent regression in the future.\n> Unit tests are one way to do this (especially with leak detection), but\n> also documentation and coding guidelines need to be updated based on the\n> new patterns discovered in this process.\n>\n> I mentioned the_repository earlier as a partially-complete architectural\n> change. I think the_index compatibility macros are in a similar boat of\n> being half-completed (though, the last time I tried to remove these macros\n> I was told that it wasn't a goal to remove the macros from builtins). The\n> sparse-index's command_requires_full_index checks are also in this boat,\n> so I understand the difficulty of getting a big change 100% complete (and\n> in this case, the plan is to do at least one more GSoC project in the area\n> to see what progress can be made that way).\n>\n> We should be careful in committing to a very long-term direction without\n> first delivering value on a shorter timeline.\n>\n> Thanks,\n> -Stolee\n"},{"id":"472678","messageId":"Y/kvC9+VVo80npe3@coredump.intra.peff.net","threadId":"59261","inReplyTo":"CAJoAoZknYizS4peYgR4Zy5KUMEpFUbj5eREZoC_K5vUDXnAhng@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-02-24T21:41:31Z","receivedAt":"2023-02-24T21:41:35Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2023 at 12:31:16PM -0800, Emily Shaffer wrote:\n\n> > One thing that strategy doesn't help with, though, that your\n> > libification might want: it's not very machine-readable. The error\n> > reporters would still fundamentally be working with strings. So a\n> > libified process can know \"OK, writing this ref failed, and I have some\n> > error messages in a buffer\". But the calling code can't know specifics\n> > like \"it failed because we tried to open file 'foo' and it got EPERM\".\n> > We _could_ design an error context that stores individual errno values\n> > or codes in a list, but having each caller report those specific errors\n> > is a much bigger job (and ongoing maintenance burden as we catalogue and\n> > give an identifier to each error).\n> \n> Is there a reason not to use this kind of struct and provide\n> library-specific error code enums, though, I wonder? You're right that\n> parsing the error string is really bad for the caller, for anything\n> besides just logging it. But it seems somewhat reasonable to expect\n> that any call from config library returning an integer error code is\n> referring to enum config_errors...\n\nRight, you could definitely layer the two approaches by storing the\nenums in the struct. And that might be a good thing to do in the long\nrun. But I think it's a much harder change, as it implies assigning\nthose codes (and developing a taxonomy of errors). But that can also be\ndone incrementally, especially if it's done on top of human-readable\nstrings.\n\nI.e., I can imagine a world where low-level code reports an error like:\n\n   int read_foo(const char *path, struct error_context *err)\n   {\n           fd = open(path, ...);\n           if (fd < 0)\n                   return report_errno(&err, ERR_FOO_OPEN, \"unable to open foo file '%s'\", path);\n           ...read and parse...\n           if (some_unexpected_format)\n                   return report(&err, ERR_FOO_FORMAT, \"unable to parse foo file '%s'\", path);\n   }\n\nand then the caller has many options:\n\n  - pass in a context that just dumps the human readable errors to\n    stderr\n\n  - collect the error strings in a buffer to report by some other\n    mechanism\n\n  - check err.type to act on ERR_FOO_OPEN, etc, including errno (which\n    I'd imagine report_errno() to record)\n\nThe details above are just a sketch, of course. Rather than FOO_OPEN,\ncallers may actually want to think of things more like a stack of\nerrors that would mirror the callstack. You could imagine something\nlike:\n\n  type=ERR_REF_WRITE\n  type=ERR_GET_LOCK\n  type=ERR_OPEN, errno=EPERM\n\nI do think it's possible to over-engineer this to the point of\nabsurdity, and end up creating a lot of work. Which is why I'd probably\nstart with uniformly trying to use an error context struct with strings,\nafter which adding on fancier features gets easier (and again, is\npossibly something that can be done incrementally as various subsystems\nsupport it).\n\n-Peff\n"},{"id":"472683","messageId":"xmqqbkliwtyy.fsf@gitster.g","threadId":"59261","inReplyTo":"CAJoAoZknYizS4peYgR4Zy5KUMEpFUbj5eREZoC_K5vUDXnAhng@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-02-24T22:59:17Z","receivedAt":"2023-02-24T22:59:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <nasamuffin@google.com> writes:\n\n> Is there a reason not to use this kind of struct and provide\n> library-specific error code enums, though, I wonder? You're right that\n> parsing the error string is really bad for the caller, for anything\n> besides just logging it. But it seems somewhat reasonable to expect\n> that any call from config library returning an integer error code is\n> referring to enum config_errors...\n\nIn addition to what Peff already said, I think the harder part of it\nis to parametralize the errors in a machine readable way.  A part of\na library may say (with an enum) that it is returning \"Ref cannot be\nread\" error, with a parameter that says \"The ref that caused this\nerror was 'refs/heads/next'\" which makes \"Ref cannot be read\" error\nhas one parameter.  \"Ref cannot be renamed\" may have two (old and\nnew name).  Other errors from some library functions may not even be\nof type \"string\".\n\nComing up with the enums to cover the error conditions (which Peff\ncovered well) is already a lot of work.  Making sure each of them\ntake sufficient parameters to describe the error usefully adds more\non top.  And the code to pass these variable number of parameters of\nvariable types would be, eh, fun---it would be error prone without\ncompiler help.\n\n"},{"id":"472697","messageId":"20230225014802.2064597-1-jonathantanmy@google.com","threadId":"59261","inReplyTo":"86e7f589-7576-5ed3-3dc9-5dec9ca346eb@github.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2023-02-25T01:48:02Z","receivedAt":"2023-02-25T01:48:10Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"Victoria Dye <vdye@github.com> writes:\n> Emily Shaffer wrote:\n> > Hi folks,\n> > \n> > As I mentioned in standup this week[1], my colleagues and I at Google\n> > have become very interested in converting parts of Git into libraries\n> > usable by external programs. In other words, for some modules which\n> > already have clear boundaries inside of Git - like config.[ch],\n> > strbuf.[ch], etc. - we want to remove some implicit dependencies, like\n> > references to globals, and make explicit other dependencies, like\n> > references to other modules within Git. Eventually, we'd like both for\n> > an external program to use Git libraries within its own process, and\n> > for Git to be given an alternative implementation of a library it uses\n> > internally (like a plugin at runtime).\n> \n> I see why you've linked these two objectives (there could be some shared\n> infrastructure between them), but they're ultimately separate concerns, with\n> separate needs & tradeoffs. \n\nI think the main link is that we need (or think we need) both of them\n- for example, we do want external code to be able to use Git's config\nparsing mechanism, and we want to be able to swap out the object store.\nThey indeed are separate concerns.\n\n> By trying to design with a unified solution for\n> both, though, this proposal's solution may be broader than it needs to be\n> and comes with some serious developer experience ramifications.\n> \n> For example, looking at the  \"call libgit.so from other applications\"\n> objective: one way I could imagine implementing this is by creating a stable\n> API wrapper around various hand-picked internal functions. If those internal\n> functions change (e.g., adding a function argument to handle a new flag in a\n> builtin), the API can remain stable (by setting a sane default) without\n> really limiting the contributor adding a new feature. \n> \n> However, my reading of this proposal (although correct me if I'm wrong) is\n> that you want to use that same external-facing API as a plugin interface for\n> Git. In that case, the contributor can't just add a new argument and\n> trivially modify the wrapper; they either break the API contract, or create\n> a new version of the API and continue to support the old one. While that's\n> technically doable, it's a pretty sizeable increase to the amount of \"stuff\"\n> contributors need to keep track of and limits the project's flexibility. \n\nFor now, we at Google think that Git shouldn't guarantee API backwards\ncompatibility (unlike for plumbing commands), and that any users of\nthe API should be prepared for API backwards compatibility to not be\nsomething Git has (and thus be prepared to, say, pin a version of Git and/\nor be prepared to continually update their code, which they might want\nto do anyway because hopefully, the Git developer community would update\nan API because there is a user need or because it will make internal\ncode that uses that API better: hence they might want to meet the same\nuser need or have the same code improvement).\n\nBut I agree that the way an API change would be handled is different in\nthe two cases.\n\n> > - Having clear, modular libraries makes it easy to find the code\n> > responsible for some behavior, or determine where to add something\n> > new.\n> \n> This is more an issue of code & repository organization, and isn't really an\n> inherent outcome of lib-ifying Git. That said, like the cleanup you mention\n> later, I would not be opposed to a dedicated code reorganization proposal.\n\nThanks. I think that Emily didn't intend to present libification as\nthe only way to accomplish such an organization, but more of \"here's\nthe business need to explain why we're doing this, and here's how it\nbenefits the project\". I do agree that there are other ways of achieving\nsuch a reorganization.\n\n> > - Operations recursing into submodules can be run in-process, and the\n> > callsite makes it very clear which repository context is being used;\n> > if we're not subprocessing to recurse, then we can lose the awkward\n> > git-submodule.sh -> builtin/submodule__helper.c -> top-level Git\n> > process codepath.\n> \n> Interesting - I wasn't aware this was how submodules worked internally. Is\n> there a specific reason this can't be refactored to perform the recursion\n> in-process?\n\nSame answer as above.\n\n> > - Being able to test libraries in isolation via unit tests or mocks\n> > speeds up determining the root cause of bugs.\n> \n> The addition of unit tests is a project of its own, and one I don't believe\n> is intrinsically tied to this one. More on this later.\n\nSame answer as above.\n\n> > - If we can swap out an entire library, or just a single function of\n> > one, it's easy to experiment with the entire codebase without sweeping\n> > changes.\n> \n> What sort of \"experiment\" are you thinking of here? Anyone can make internal\n> changes to Git in their own clone of the repository, then build and test it.\n> Shipping a custom fork of Git doesn't seem any less complex than building a\n> plugin library, shipping that, and injecting it into an existing Git\n> install.\n\nHaving clearly separated out modules means that experimentation,\nespecially those that span several months (say), can be more easily\ndone, because what we would need to change can be confined to one or\nseveral files instead of possibly needing to be distributed across many\nfiles. This not only helps speed up development but also keeping up with\nupstream changes (there will hopefully be fewer merge conflicts this\nway). The way of shipping experimental builds could indeed be from a\nfork of the codebase.\n\n> > The ability to use Git as a library also makes it easier for other\n> > tooling to interact with a Git repository *correctly*. As an example,\n> > `repo` has a long history of abusing Git by directly manipulating the\n> > gitdir[2], but if it were written in a world where Git exists as\n> > easy-to-use libraries, it probably wouldn't have needed to, as it\n> > could have invoked Git directly or replaced the relevant modules with\n> > its own implementation. Both `repo`[3] and `git-gui[4]` have\n> > reimplemented logic from git.git. Other interfaces that cooperate with\n> > Git's filesystem storage layer, like `scalar` or `jj`[5], would be\n> > able to interop with a Git repository without having to reimplement\n> > custom logic or keep up with new Git changes.\n> \n> I'd alternatively suggest that these use cases could be addressed with\n> better plumbing command support and better (public) documentation of our\n> on-disk data structures. In many cases, we're treating on-disk data\n> structures as external-facing APIs anyway - the index [1], packfiles [2],\n> etc. are versioned. We could definitely add documentation that's more \n> directly targeted at external integrators.\n> \n> [1] https://git-scm.com/docs/index-format\n> [2] https://git-scm.com/docs/gitformat-pack\n\nPlumbing commands require spawning processes and serializing inputs\nand outputs, and we can't do things like pass callbacks or an mmap-ped\nbuffer we already have. (As an aside, we did consider that approach with\ngit-batch [1], but that didn't get much traction.)\n\nAs for interoperability on the basis of on-disk data, I think\nthe workload of keeping up with Git improvements shouldn't be\nunderestimated. For example, I don't think many (if any) implementations\nthat can read Git's files support partial clone (at least on the\nclient), and we are still improving things, e.g. hasconfig:remote.*.url\nincludeIf was recently added to Git config.\n\n[1] https://lore.kernel.org/git/20220207190320.2960362-1-jonathantanmy@google.com/\n\n> Both of these points sound more like implementation details/quirks of the\n> VFS prototype you're building than hard blockers coming from Git. For\n> example, you mention 'git status' needing to populate a full worktree -\n> could it instead use a mechanism similar to FSMonitor to avoid scanning for\n> files you know are virtualized? While there'd likely be some Git changes\n> involved (some of which might be a bit plugin-y), they'd be much more\n> tightly scoped than full libification.\n\nI think the answer to this is the same as the disadvantages of outside\ncode calling plumbing commands (above).\n\n> These cleanups/best practices are a great idea regardless of any potential\n> lib-ification. I'll be sure to keep them in mind in any future changes I\n> make or review.\n\nThanks.\n\n> > In the longer term, if Git has libraries with easily-replaced\n> > dependencies, we get a few more benefits:\n> > \n> > - Unit tests. We already have some in t/helper/, but if we can replace\n> > all the dependencies of a certain library with simple stubs, it's\n> > easier for us to write comprehensive unit tests, in addition to the\n> > work we already do introducing edge cases in bash integration tests.\n> \n> This doesn't really follow as a direct consequence of making libgit\n> externally facing. AFAIK, there's nothing explicitly stopping us from\n> writing or integrating a unit test framework now.\n> \n> Like the \"make libgit public\" vs. \"allow for custom backends in Git\", I\n> think this is a separate project with its own tradeoffs to evaluate.\n\nThe way we're making code in Git externally facing, at least right now,\nis not to make the entirety of the Git code (perhaps minus the builtins)\nconsumable from the outside, but just selected .c/.h files. Being able\nto say \"these .c/.h files can be used independently\" will make unit\ntests easier, I think. But yes, there is nothing stopping us from such\nan effort independent of libification.\n\n> > - If our users can use plugins to improve performance in specific\n> > scenarios (like a VFS-aware object store in the VFS case I cited\n> > above), then Git works better for them without having to adopt a\n> > different workflow, such as using an alternative tool or wrapper.\n> \n> I'm curious as to what you want the user experience for this to look like.\n> How would Git know to dynamically load a plugin? How does a user configure\n> (or disable) plugins? Will a plugin need to replace all of the functions in\n> a given library, or would Git be able to \"fall back\" on its own internal\n> implementation?\n\nRight now, it's probably going to be just recompiling the Git binary\nwith a few .c files swapped for others, which is why it's important\nto us to have defined boundaries between .c files (even if they move\nover time). I do think that converting code in Git to access other\ncode through a vtable is desirable (that way, for example, a single\nbinary can access 2 kinds of object stores, for example) but even if\nthe Git project decides not to do that, we can still do a lot (and just\nship, for example, some extra .c files that can do both the standard\nimplementation and a custom one).\n\n> > - An easy-to-understand modular codebase makes it easy for new\n> > contributors to start hacking and understand the consequences of their\n> > patch.\n> \n> As noted earlier, I think improving the developer experience in this way is\n> independent of the development of external APIs.\n\nNoted - same answer as earlier too.\n\n> > Of course, we haven't maintained any guarantee about the consistency\n> > of our implementation between releases. I don't anticipate that we'll\n> > write the perfect library interface on our first try. So I hope that\n> > we can be very explicit about refusing to provide any compatibility\n> > guarantee whatsoever between versions for quite a long time. On\n> > Google's end, that's well-understood and accepted. As I understand,\n> > some other projects already use Git's codebase as a \"library\" by\n> > including it as a submodule and using the code directly[6]; even a\n> > breakable API seems like an improvement over that, too.\n> \n> This hints at, but sidesteps, a really important aspect of the long-term\n> goals of this project - are you planning on having us start guaranteeing\n> consistency once there's an external-facing API available? \n> \n> If not, that sounds like an unpleasant user experience (and one prone to\n> Hyrum's Law [3] at that). \n> \n> If so, Git contributors will either be much more constrained in the\n> introduction of new features, or we'll end up with a mess of\n> backward-compatibility APIs.\n> \n> [3] https://www.hyrumslaw.com/\n\nI answered this above.\n\n> > So what's next? Naturally, I'm looking forward to a spirited\n> > discussion about this topic - I'd like to know which concerns haven't\n> > been addressed and figure out whether we can find a way around them,\n> > and generally build awareness of this effort with the community.\n> > \n> > I'm also planning to send a proposal for a document full of \"best\n> > practices\" for turning Git code into libraries (and have quite a lot\n> > of discussion around that document, too). My hope is that we can use\n> > that document to help us during implementation as well as during\n> > review, and refine it over time as we learn more about what works and\n> > what doesn't. Having this kind of documentation will make it easy for\n> > others to join us in moving Git's codebase towards a clean set of\n> > libraries. I hope that, as a project, we can settle on some tenets\n> > that we all agree would make Git nicer.\n> \n> I think the use of multiple libraries is part of the potentially suboptimal\n> balance between the \"load Git's internals as a library\" and \"let Git use\n> plugins in place of its own implementations\" goals. The former could leave a\n> user in dependency hell if there are lots of small libraries to load, while\n> the latter may work better with smaller-scoped libraries (to avoid needing\n> to implement more than is useful). And, the functionality that's useful to a\n> program invoking Git via library may not be the same as what someone would\n> want to replace via plugin. \n\nI don't see how we can avoid separating up Git's code with either goal\n(internals as library vs. plugins), but I think that such separation\nmakes the code clearer anyway even if one does not need to reuse\ninternals or use plugins.\n\n> As a general note, this libification project - its justification,\n> milestones, organization, etc. - seems to be primarily driven by the needs\n> of a VFS that is entirely opaque to the Git community you're proposing to.\n> As it stands, reviewers are put in a position of needing to accept, without\n> much evidence, that this is the best (or only) possible approach for meeting\n> those needs. \n> \n> While I'm normally content with \"scratch your own itch\"-type changes, what\n> you're proposing is a major paradigm shift in how Git is written,\n> maintained, and used. I'm not comfortable accepting that level of impact\n> without at least being able to evaluate whether the problem can be solved\n> some other way.\n\nI think that reviewers should review code based on current principles,\nnot just based on our (we at Google) claim that it is useful for our\nVFS (although perhaps our need might show that other people might have a\nsimilar need too).\n\n> Implementing dependency injection (via vtable or otherwise) and unit test\n> mocking is a massive undertaking on its own, let alone everything else\n> described so far. You've outlined some clear, fairly unobtrusive first steps\n> to the overarching proposal, but there's a lot of detail missing from the\n> plans for later steps. While there's always risk of a project getting\n> derailed or blocked later on due to some unforeseen issue, that risk seems\n> particularly high here given the size & scope of all of these components. \n\nTrue, but even the first steps will already enable us to accomplish\nquite a bit. The later steps may change based on what we learn after the\nfirst steps.\n\n> Intermediate milestones/goals, each of which is valuable on its own,\n> outlined in the proposal would help allay those fears (for me, at least).\n\nNoted.\n\n> So, to summarize my thoughts into some (hopefully) actionable feedback:\n> \n> - It's really important that the proposal presents a clear, concrete\n>   long-term vision for this project. \n>   - What is the desired end state, in terms of what is built and installed\n>     with Git?\n\nFor the intermediate step (what we want to accomplish now/soon):\n - Git will compile into the same binaries as it does now, with no\n   plugin functionality. There will be no additional libraries.\n - Git will have unit tests that compile a subset of Git's .c files,\n   and may possibly supply their own .c files that implement existing .h\n   signatures to serve as stubs.\n - Other projects that want to use Git's code will need to include Git's\n   .c/.h files in their own build systems, possibly using the lists of\n   files in the unit test Makefiles for reference. They will also need\n   to keep up with changes themselves.\n\nThe end step might include vtables, but we probably will revisit that\nlater once we learn more about the task of code reuse, aided by what\nwe've accomplished in the intermediate step.\n\n>   - What will be expected of other Git contributors to support this design?\n\nTo not break the unit tests by putting functionality in the wrong\nplace (but hopefully this is seen as a plus, in that we have something\nmechanical to check such things), and to design new code to fit with the\nprinciples of modularity.\n\n>     What happens if someone wants to \"break\" an API?\n\nI think I already answered this above.\n\n>   - What is the impact to users (including security implications)?\n\nAs of the intermediate step, we don't allow plugins in the standard\nbuild of Git, so I don't think there will be much impact. In certain\nenvironments that run patched versions of Git, .c files could be swapped\nout, but the impact is similar to that of what a patch can do.\n\n> - The scope for what you've proposed is pretty huge (which comes with a lot\n>   of risk), but I think it could be broken into smaller, *independent*\n>   pieces. \n\nI think each patch set towards this goal can be considered independently.\n\n> - It's hard to judge or suggest adjustments to this proposal without knowing\n>   the specific challenges you're facing with Git as it is.\n\nI think Emily has already outlined a few such challenges (and hopefully\nI have elaborated on them), but if there are any specific ones you'd\nlike more information on, feel free to let us know.\n\n> Finally, thanks for sending this email & starting a discussion. It's an\n> interesting topic and I'm looking forward to seeing everyone's perspectives\n> on the matter.\n\nThank you for your thorough review.\n"},{"id":"474060","messageId":"CAMP44s1Qqd2cYcf7OGxz1-PY-8TF2KG+9jPEWMrnCaCfPe_1sw@mail.gmail.com","threadId":"59261","inReplyTo":"4222af90-bd6b-d970-2829-1ddfaeb770bf@dunelm.org.uk","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-03-23T23:22:02Z","receivedAt":"2023-03-23T23:22:59Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> On 18/02/2023 01:59, demerphq wrote:\n> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n> >>\n> >> Emily Shaffer <nasamuffin@google.com> writes:\n> >>\n> >>> Basically, if this effort turns out not to be fruitful as a whole, I'd\n> >>> like for us to still have left a positive impact on the codebase.\n> >>> ...\n> >>> So what's next? Naturally, I'm looking forward to a spirited\n> >>> discussion about this topic - I'd like to know which concerns haven't\n> >>> been addressed and figure out whether we can find a way around them,\n> >>> and generally build awareness of this effort with the community.\n> >>\n> >> On of the gravest concerns is that the devil is in the details.\n> >>\n> >> For example, \"die() is inconvenient to callers, let's propagate\n> >> errors up the callchain\" is an easy thing to say, but it would take\n> >> much more than \"let's propagate errors up\" to libify something like\n> >> check_connected() to do the same thing without spawning a separate\n> >> process that is expected to exit with failure.\n> >\n> >\n> > What does \"propagate errors up the callchain\" mean?  One\n> > interpretation I can think of seems quite horrible, but another seems\n> > quite doable and reasonable and likely not even very invasive of the\n> > existing code:\n> >\n> > You can use setjmp/longjmp to implement a form of \"try\", so that\n> > errors dont have to be *explicitly* returned *in* the call chain. And\n> > you could probably do so without changing very much of the existing\n> > code at all, and maintain a high level of conceptual alignment with\n> > the current code strategy.\n>\n> Using setjmp/longjmp is an interesting suggestion, I think lua does\n> something similar to what you describe for perl. However I think both of\n> those use a allocator with garbage collection. I worry that using\n> longjmp in git would be more invasive (or result in more memory leaks)\n> as we'd need to to guard each allocation with some code to clean it up\n> and then propagate the error. That means we're back to manually\n> propagating errors up the call chain in many cases.\n\nWe could just use talloc [1].\n\nAll the downstream code can create objects that are descendents of the\nmaster object, when the \"exception\" occurs the code is sent back to\nwhere the master object was created, it's destroyed, and all the\nchildren are destroyed as well.\n\nI've played around with both talloc and stack contexts and I don't see\nhow this isn't doable.\n\nCheers.\n\n[1] https://talloc.samba.org/talloc/doc/html/index.html\n\n-- \nFelipe Contreras\n"},{"id":"474062","messageId":"008101d95ddf$7863d900$692b8b00$@nexbridge.com","threadId":"59261","inReplyTo":"CAMP44s1Qqd2cYcf7OGxz1-PY-8TF2KG+9jPEWMrnCaCfPe_1sw@mail.gmail.com","subject":"RE: Proposal/Discussion: Turning parts of Git into libraries","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-03-23T23:30:30Z","receivedAt":"2023-03-23T23:30:52Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n>On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>\n>> On 18/02/2023 01:59, demerphq wrote:\n>> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n>> >>\n>> >> Emily Shaffer <nasamuffin@google.com> writes:\n>> >>\n>> >>> Basically, if this effort turns out not to be fruitful as a whole,\n>> >>> I'd like for us to still have left a positive impact on the codebase.\n>> >>> ...\n>> >>> So what's next? Naturally, I'm looking forward to a spirited\n>> >>> discussion about this topic - I'd like to know which concerns\n>> >>> haven't been addressed and figure out whether we can find a way\n>> >>> around them, and generally build awareness of this effort with the community.\n>> >>\n>> >> On of the gravest concerns is that the devil is in the details.\n>> >>\n>> >> For example, \"die() is inconvenient to callers, let's propagate\n>> >> errors up the callchain\" is an easy thing to say, but it would take\n>> >> much more than \"let's propagate errors up\" to libify something like\n>> >> check_connected() to do the same thing without spawning a separate\n>> >> process that is expected to exit with failure.\n>> >\n>> >\n>> > What does \"propagate errors up the callchain\" mean?  One\n>> > interpretation I can think of seems quite horrible, but another\n>> > seems quite doable and reasonable and likely not even very invasive\n>> > of the existing code:\n>> >\n>> > You can use setjmp/longjmp to implement a form of \"try\", so that\n>> > errors dont have to be *explicitly* returned *in* the call chain.\n>> > And you could probably do so without changing very much of the\n>> > existing code at all, and maintain a high level of conceptual\n>> > alignment with the current code strategy.\n>>\n>> Using setjmp/longjmp is an interesting suggestion, I think lua does\n>> something similar to what you describe for perl. However I think both\n>> of those use a allocator with garbage collection. I worry that using\n>> longjmp in git would be more invasive (or result in more memory leaks)\n>> as we'd need to to guard each allocation with some code to clean it up\n>> and then propagate the error. That means we're back to manually\n>> propagating errors up the call chain in many cases.\n>\n>We could just use talloc [1].\n\ntalloc is not portable. This would break various platforms.\n\n--Randall\n\n"},{"id":"474063","messageId":"CAMP44s1X6LGpFfA_Zb_GakXehBJDeGrfFcehPgv+YM++xKHN3A@mail.gmail.com","threadId":"59261","inReplyTo":"008101d95ddf$7863d900$692b8b00$@nexbridge.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-03-23T23:34:56Z","receivedAt":"2023-03-23T23:35:11Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n>\n> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood <phillip.wood123@gmail.com> wrote:\n> >>\n> >> On 18/02/2023 01:59, demerphq wrote:\n> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n> >> >>\n> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n> >> >>\n> >> >>> Basically, if this effort turns out not to be fruitful as a whole,\n> >> >>> I'd like for us to still have left a positive impact on the codebase.\n> >> >>> ...\n> >> >>> So what's next? Naturally, I'm looking forward to a spirited\n> >> >>> discussion about this topic - I'd like to know which concerns\n> >> >>> haven't been addressed and figure out whether we can find a way\n> >> >>> around them, and generally build awareness of this effort with the community.\n> >> >>\n> >> >> On of the gravest concerns is that the devil is in the details.\n> >> >>\n> >> >> For example, \"die() is inconvenient to callers, let's propagate\n> >> >> errors up the callchain\" is an easy thing to say, but it would take\n> >> >> much more than \"let's propagate errors up\" to libify something like\n> >> >> check_connected() to do the same thing without spawning a separate\n> >> >> process that is expected to exit with failure.\n> >> >\n> >> >\n> >> > What does \"propagate errors up the callchain\" mean?  One\n> >> > interpretation I can think of seems quite horrible, but another\n> >> > seems quite doable and reasonable and likely not even very invasive\n> >> > of the existing code:\n> >> >\n> >> > You can use setjmp/longjmp to implement a form of \"try\", so that\n> >> > errors dont have to be *explicitly* returned *in* the call chain.\n> >> > And you could probably do so without changing very much of the\n> >> > existing code at all, and maintain a high level of conceptual\n> >> > alignment with the current code strategy.\n> >>\n> >> Using setjmp/longjmp is an interesting suggestion, I think lua does\n> >> something similar to what you describe for perl. However I think both\n> >> of those use a allocator with garbage collection. I worry that using\n> >> longjmp in git would be more invasive (or result in more memory leaks)\n> >> as we'd need to to guard each allocation with some code to clean it up\n> >> and then propagate the error. That means we're back to manually\n> >> propagating errors up the call chain in many cases.\n> >\n> >We could just use talloc [1].\n>\n> talloc is not portable.\n\nWhat makes you say that?\n\nEither way, there's multiple libraries that do the same thing, and of\ncourse one could be implemented within Git. It's not that complex.\n\n-- \nFelipe Contreras\n"},{"id":"474064","messageId":"CAMP44s2AMrXCN6f6v-W0sqb++TVfHf7Q1miJE7iZjZOVwFQa0Q@mail.gmail.com","threadId":"59261","inReplyTo":"CAJoAoZ=Cig_kLocxKGax31sU7Xe4==BGzC__Bg2_pr7krNq6MA@mail.gmail.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-03-23T23:37:21Z","receivedAt":"2023-03-23T23:37:44Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Fri, Feb 17, 2023 at 3:45 PM Emily Shaffer <nasamuffin@google.com> wrote:\n\n> As I mentioned in standup this week[1], my colleagues and I at Google\n> have become very interested in converting parts of Git into libraries\n> usable by external programs. In other words, for some modules which\n> already have clear boundaries inside of Git - like config.[ch],\n> strbuf.[ch], etc. - we want to remove some implicit dependencies, like\n> references to globals, and make explicit other dependencies, like\n> references to other modules within Git. Eventually, we'd like both for\n> an external program to use Git libraries within its own process, and\n> for Git to be given an alternative implementation of a library it uses\n> internally (like a plugin at runtime).\n\nThis is obviously the way it should have been done from the beginning,\nbut unfortunately at this point the Git project has too much inertia\nand too many vested interests from multi-billion dollar corporations\nto change.\n\nI wonder if a single person who isn't paid to work on Git commented on\nthis thread.\n\nI don't think these kinds of laudable efforts can be achieved within\nthe Git project.\n\nCheers.\n\n-- \nFelipe Contreras\n"},{"id":"474065","messageId":"008201d95de1$359285c0$a0b79140$@nexbridge.com","threadId":"59261","inReplyTo":"CAMP44s1X6LGpFfA_Zb_GakXehBJDeGrfFcehPgv+YM++xKHN3A@mail.gmail.com","subject":"RE: Proposal/Discussion: Turning parts of Git into libraries","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-03-23T23:42:57Z","receivedAt":"2023-03-23T23:43:15Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Thursday, March 23, 2023 7:35 PM, Felipe Contreras wrote:\n>On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n>>\n>> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n>> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood <phillip.wood123@gmail.com>\n>wrote:\n>> >>\n>> >> On 18/02/2023 01:59, demerphq wrote:\n>> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n>> >> >>\n>> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n>> >> >>\n>> >> >>> Basically, if this effort turns out not to be fruitful as a\n>> >> >>> whole, I'd like for us to still have left a positive impact on the codebase.\n>> >> >>> ...\n>> >> >>> So what's next? Naturally, I'm looking forward to a spirited\n>> >> >>> discussion about this topic - I'd like to know which concerns\n>> >> >>> haven't been addressed and figure out whether we can find a way\n>> >> >>> around them, and generally build awareness of this effort with the\n>community.\n>> >> >>\n>> >> >> On of the gravest concerns is that the devil is in the details.\n>> >> >>\n>> >> >> For example, \"die() is inconvenient to callers, let's propagate\n>> >> >> errors up the callchain\" is an easy thing to say, but it would\n>> >> >> take much more than \"let's propagate errors up\" to libify\n>> >> >> something like\n>> >> >> check_connected() to do the same thing without spawning a\n>> >> >> separate process that is expected to exit with failure.\n>> >> >\n>> >> >\n>> >> > What does \"propagate errors up the callchain\" mean?  One\n>> >> > interpretation I can think of seems quite horrible, but another\n>> >> > seems quite doable and reasonable and likely not even very\n>> >> > invasive of the existing code:\n>> >> >\n>> >> > You can use setjmp/longjmp to implement a form of \"try\", so that\n>> >> > errors dont have to be *explicitly* returned *in* the call chain.\n>> >> > And you could probably do so without changing very much of the\n>> >> > existing code at all, and maintain a high level of conceptual\n>> >> > alignment with the current code strategy.\n>> >>\n>> >> Using setjmp/longjmp is an interesting suggestion, I think lua does\n>> >> something similar to what you describe for perl. However I think\n>> >> both of those use a allocator with garbage collection. I worry that\n>> >> using longjmp in git would be more invasive (or result in more\n>> >> memory leaks) as we'd need to to guard each allocation with some\n>> >> code to clean it up and then propagate the error. That means we're\n>> >> back to manually propagating errors up the call chain in many cases.\n>> >\n>> >We could just use talloc [1].\n>>\n>> talloc is not portable.\n>\n>What makes you say that?\n\ntalloc is not part of a POSIX standard I could find. Aside from that:\n\n$ man talloc\n(on NonStop) No manual entry for talloc\n(on Cygwin) No manual entry for talloc\n\nJust reporting my findings.\n\n"},{"id":"474066","messageId":"000001d95de1$6de529a0$49af7ce0$@nexbridge.com","threadId":"59261","inReplyTo":"CAMP44s2AMrXCN6f6v-W0sqb++TVfHf7Q1miJE7iZjZOVwFQa0Q@mail.gmail.com","subject":"RE: Proposal/Discussion: Turning parts of Git into libraries","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-03-23T23:44:31Z","receivedAt":"2023-03-23T23:44:48Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Thursday, March 23, 2023 7:37 PM, Felipe Contreras wrote:\n>On Fri, Feb 17, 2023 at 3:45 PM Emily Shaffer <nasamuffin@google.com> wrote:\n>\n>> As I mentioned in standup this week[1], my colleagues and I at Google\n>> have become very interested in converting parts of Git into libraries\n>> usable by external programs. In other words, for some modules which\n>> already have clear boundaries inside of Git - like config.[ch],\n>> strbuf.[ch], etc. - we want to remove some implicit dependencies, like\n>> references to globals, and make explicit other dependencies, like\n>> references to other modules within Git. Eventually, we'd like both for\n>> an external program to use Git libraries within its own process, and\n>> for Git to be given an alternative implementation of a library it uses\n>> internally (like a plugin at runtime).\n>\n>This is obviously the way it should have been done from the beginning, but\n>unfortunately at this point the Git project has too much inertia and too many vested\n>interests from multi-billion dollar corporations to change.\n>\n>I wonder if a single person who isn't paid to work on Git commented on this thread.\n\nRaises hand.\n\n"},{"id":"474067","messageId":"CAMP44s3Gk67rPEPjoAxLHS4KrCQBb6VoPJ6Rqm-FTK+8PTaRRQ@mail.gmail.com","threadId":"59261","inReplyTo":"008201d95de1$359285c0$a0b79140$@nexbridge.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-03-23T23:55:29Z","receivedAt":"2023-03-23T23:55:45Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Thu, Mar 23, 2023 at 5:43 PM <rsbecker@nexbridge.com> wrote:\n>\n> On Thursday, March 23, 2023 7:35 PM, Felipe Contreras wrote:\n> >On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n> >>\n> >> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n> >> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood <phillip.wood123@gmail.com>\n> >wrote:\n> >> >>\n> >> >> On 18/02/2023 01:59, demerphq wrote:\n> >> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com> wrote:\n> >> >> >>\n> >> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n> >> >> >>\n> >> >> >>> Basically, if this effort turns out not to be fruitful as a\n> >> >> >>> whole, I'd like for us to still have left a positive impact on the codebase.\n> >> >> >>> ...\n> >> >> >>> So what's next? Naturally, I'm looking forward to a spirited\n> >> >> >>> discussion about this topic - I'd like to know which concerns\n> >> >> >>> haven't been addressed and figure out whether we can find a way\n> >> >> >>> around them, and generally build awareness of this effort with the\n> >community.\n> >> >> >>\n> >> >> >> On of the gravest concerns is that the devil is in the details.\n> >> >> >>\n> >> >> >> For example, \"die() is inconvenient to callers, let's propagate\n> >> >> >> errors up the callchain\" is an easy thing to say, but it would\n> >> >> >> take much more than \"let's propagate errors up\" to libify\n> >> >> >> something like\n> >> >> >> check_connected() to do the same thing without spawning a\n> >> >> >> separate process that is expected to exit with failure.\n> >> >> >\n> >> >> >\n> >> >> > What does \"propagate errors up the callchain\" mean?  One\n> >> >> > interpretation I can think of seems quite horrible, but another\n> >> >> > seems quite doable and reasonable and likely not even very\n> >> >> > invasive of the existing code:\n> >> >> >\n> >> >> > You can use setjmp/longjmp to implement a form of \"try\", so that\n> >> >> > errors dont have to be *explicitly* returned *in* the call chain.\n> >> >> > And you could probably do so without changing very much of the\n> >> >> > existing code at all, and maintain a high level of conceptual\n> >> >> > alignment with the current code strategy.\n> >> >>\n> >> >> Using setjmp/longjmp is an interesting suggestion, I think lua does\n> >> >> something similar to what you describe for perl. However I think\n> >> >> both of those use a allocator with garbage collection. I worry that\n> >> >> using longjmp in git would be more invasive (or result in more\n> >> >> memory leaks) as we'd need to to guard each allocation with some\n> >> >> code to clean it up and then propagate the error. That means we're\n> >> >> back to manually propagating errors up the call chain in many cases.\n> >> >\n> >> >We could just use talloc [1].\n> >>\n> >> talloc is not portable.\n> >\n> >What makes you say that?\n>\n> talloc is not part of a POSIX standard I could find.\n\nIt's a library, like: z, ssl, curl, pcre2-8, etc. Libraries can be\ncompiled on different platforms.\n\n-- \nFelipe Contreras\n"},{"id":"474116","messageId":"004d01d95e86$bd355d40$37a017c0$@nexbridge.com","threadId":"59261","inReplyTo":"CAMP44s3Gk67rPEPjoAxLHS4KrCQBb6VoPJ6Rqm-FTK+8PTaRRQ@mail.gmail.com","subject":"RE: Proposal/Discussion: Turning parts of Git into libraries","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-03-24T19:27:51Z","receivedAt":"2023-03-24T19:28:17Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Thursday, March 23, 2023 7:55 PM, Felipe Contreras wrote:\n>On Thu, Mar 23, 2023 at 5:43 PM <rsbecker@nexbridge.com> wrote:\n>>\n>> On Thursday, March 23, 2023 7:35 PM, Felipe Contreras wrote:\n>> >On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n>> >>\n>> >> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n>> >> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood\n>> >> ><phillip.wood123@gmail.com>\n>> >wrote:\n>> >> >>\n>> >> >> On 18/02/2023 01:59, demerphq wrote:\n>> >> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com>\n>wrote:\n>> >> >> >>\n>> >> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n>> >> >> >>\n>> >> >> >>> Basically, if this effort turns out not to be fruitful as a\n>> >> >> >>> whole, I'd like for us to still have left a positive impact on the codebase.\n>> >> >> >>> ...\n>> >> >> >>> So what's next? Naturally, I'm looking forward to a spirited\n>> >> >> >>> discussion about this topic - I'd like to know which\n>> >> >> >>> concerns haven't been addressed and figure out whether we\n>> >> >> >>> can find a way around them, and generally build awareness of\n>> >> >> >>> this effort with the\n>> >community.\n>> >> >> >>\n>> >> >> >> On of the gravest concerns is that the devil is in the details.\n>> >> >> >>\n>> >> >> >> For example, \"die() is inconvenient to callers, let's\n>> >> >> >> propagate errors up the callchain\" is an easy thing to say,\n>> >> >> >> but it would take much more than \"let's propagate errors up\"\n>> >> >> >> to libify something like\n>> >> >> >> check_connected() to do the same thing without spawning a\n>> >> >> >> separate process that is expected to exit with failure.\n>> >> >> >\n>> >> >> >\n>> >> >> > What does \"propagate errors up the callchain\" mean?  One\n>> >> >> > interpretation I can think of seems quite horrible, but\n>> >> >> > another seems quite doable and reasonable and likely not even\n>> >> >> > very invasive of the existing code:\n>> >> >> >\n>> >> >> > You can use setjmp/longjmp to implement a form of \"try\", so\n>> >> >> > that errors dont have to be *explicitly* returned *in* the call chain.\n>> >> >> > And you could probably do so without changing very much of the\n>> >> >> > existing code at all, and maintain a high level of conceptual\n>> >> >> > alignment with the current code strategy.\n>> >> >>\n>> >> >> Using setjmp/longjmp is an interesting suggestion, I think lua\n>> >> >> does something similar to what you describe for perl. However I\n>> >> >> think both of those use a allocator with garbage collection. I\n>> >> >> worry that using longjmp in git would be more invasive (or\n>> >> >> result in more memory leaks) as we'd need to to guard each\n>> >> >> allocation with some code to clean it up and then propagate the\n>> >> >> error. That means we're back to manually propagating errors up the call\n>chain in many cases.\n>> >> >\n>> >> >We could just use talloc [1].\n>> >>\n>> >> talloc is not portable.\n>> >\n>> >What makes you say that?\n>>\n>> talloc is not part of a POSIX standard I could find.\n>\n>It's a library, like: z, ssl, curl, pcre2-8, etc. Libraries can be compiled on different\n>platforms.\n\ntalloc adds additional *required* dependencies to git, including python3 - required to configure and build talloc - which is not available on the NonStop ia64 platform (required support through end of 2025). I must express my resistance to what would amount to losing support for git on this NonStop platform.\n\n--Randall\n\n"},{"id":"474131","messageId":"CAMP44s3fXYcOsq-XRNZX0y6D=W37=ONELUpTBtkzv4KLbym2iA@mail.gmail.com","threadId":"59261","inReplyTo":"004d01d95e86$bd355d40$37a017c0$@nexbridge.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-03-24T21:21:32Z","receivedAt":"2023-03-24T21:21:48Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Fri, Mar 24, 2023 at 1:28 PM <rsbecker@nexbridge.com> wrote:\n>\n> On Thursday, March 23, 2023 7:55 PM, Felipe Contreras wrote:\n> >On Thu, Mar 23, 2023 at 5:43 PM <rsbecker@nexbridge.com> wrote:\n> >>\n> >> On Thursday, March 23, 2023 7:35 PM, Felipe Contreras wrote:\n> >> >On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n> >> >>\n> >> >> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n> >> >> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood\n> >> >> ><phillip.wood123@gmail.com>\n> >> >wrote:\n> >> >> >>\n> >> >> >> On 18/02/2023 01:59, demerphq wrote:\n> >> >> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano <gitster@pobox.com>\n> >wrote:\n> >> >> >> >>\n> >> >> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n> >> >> >> >>\n> >> >> >> >>> Basically, if this effort turns out not to be fruitful as a\n> >> >> >> >>> whole, I'd like for us to still have left a positive impact on the codebase.\n> >> >> >> >>> ...\n> >> >> >> >>> So what's next? Naturally, I'm looking forward to a spirited\n> >> >> >> >>> discussion about this topic - I'd like to know which\n> >> >> >> >>> concerns haven't been addressed and figure out whether we\n> >> >> >> >>> can find a way around them, and generally build awareness of\n> >> >> >> >>> this effort with the\n> >> >community.\n> >> >> >> >>\n> >> >> >> >> On of the gravest concerns is that the devil is in the details.\n> >> >> >> >>\n> >> >> >> >> For example, \"die() is inconvenient to callers, let's\n> >> >> >> >> propagate errors up the callchain\" is an easy thing to say,\n> >> >> >> >> but it would take much more than \"let's propagate errors up\"\n> >> >> >> >> to libify something like\n> >> >> >> >> check_connected() to do the same thing without spawning a\n> >> >> >> >> separate process that is expected to exit with failure.\n> >> >> >> >\n> >> >> >> >\n> >> >> >> > What does \"propagate errors up the callchain\" mean?  One\n> >> >> >> > interpretation I can think of seems quite horrible, but\n> >> >> >> > another seems quite doable and reasonable and likely not even\n> >> >> >> > very invasive of the existing code:\n> >> >> >> >\n> >> >> >> > You can use setjmp/longjmp to implement a form of \"try\", so\n> >> >> >> > that errors dont have to be *explicitly* returned *in* the call chain.\n> >> >> >> > And you could probably do so without changing very much of the\n> >> >> >> > existing code at all, and maintain a high level of conceptual\n> >> >> >> > alignment with the current code strategy.\n> >> >> >>\n> >> >> >> Using setjmp/longjmp is an interesting suggestion, I think lua\n> >> >> >> does something similar to what you describe for perl. However I\n> >> >> >> think both of those use a allocator with garbage collection. I\n> >> >> >> worry that using longjmp in git would be more invasive (or\n> >> >> >> result in more memory leaks) as we'd need to to guard each\n> >> >> >> allocation with some code to clean it up and then propagate the\n> >> >> >> error. That means we're back to manually propagating errors up the call\n> >chain in many cases.\n> >> >> >\n> >> >> >We could just use talloc [1].\n> >> >>\n> >> >> talloc is not portable.\n> >> >\n> >> >What makes you say that?\n> >>\n> >> talloc is not part of a POSIX standard I could find.\n> >\n> >It's a library, like: z, ssl, curl, pcre2-8, etc. Libraries can be compiled on different\n> >platforms.\n>\n> talloc adds additional *required* dependencies to git, including python3 - required to configure and build talloc - which is not available on the NonStop ia64 platform (required support through end of 2025). I must express my resistance to what would amount to losing support for git on this NonStop platform.\n\nThat is not true. You don't need python3 for talloc, not even to build\nit, it's just a single simple c file, it's easy to compile.\n\nThe only reason python is used is to run waf, which is used to build\nSamba, which is much more complex, but you don't need to run it,\nespecially if you know the characteristics of your system.\n\nThis simple Makefile builds libtalloc.so just fine:\n\n  CC := gcc\n  CFLAGS := -fPIC -I./lib/replace\n  LDFLAGS := -Wl,--no-undefined\n\n  # For talloc.c\n  CFLAGS += -DTALLOC_BUILD_VERSION_MAJOR=2\n-DTALLOC_BUILD_VERSION_MINOR=4 -DTALLOC_BUILD_VERSION_RELEASE=0\n  CFLAGS += -DHAVE_CONSTRUCTOR_ATTRIBUTE -DHAVE_VA_COPY\n-DHAVE_VALGRIND_MEMCHECK_H -DHAVE_INTPTR_T\n\n  # For replace.h\n  CFLAGS += -DNO_CONFIG_H -D__STDC_WANT_LIB_EXT1__=1\n  CFLAGS += -DHAVE_STDBOOL_H -DHAVE_BOOL -DHAVE_STRING_H\n-DHAVE_LIMITS_H -DHAVE_STDINT_H\n  CFLAGS += -DHAVE_DLFCN_H -DHAVE_UINTPTR_T -DHAVE_C99_VSNPRINTF\n-DHAVE_MEMMOVE -DHAVE_STRNLEN -DHAVE_VSNPRINTF\n\n  libtalloc.so: talloc.o\n    $(CC) $(LDFLAGS) -shared -o $@ $^\n\nBut of course, most of those defines are not even needed with a simple\n\"replace.h\" that is less than 10 lines of code.\n\n-- \nFelipe Contreras\n"},{"id":"474134","messageId":"005601d95e9c$dc20b360$94621a20$@nexbridge.com","threadId":"59261","inReplyTo":"CAMP44s3fXYcOsq-XRNZX0y6D=W37=ONELUpTBtkzv4KLbym2iA@mail.gmail.com","subject":"RE: Proposal/Discussion: Turning parts of Git into libraries","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2023-03-24T22:06:12Z","receivedAt":"2023-03-24T22:06:32Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Friday, March 24, 2023 5:22 PM, Felipe Contreras wrote:\n>On Fri, Mar 24, 2023 at 1:28 PM <rsbecker@nexbridge.com> wrote:\n>>\n>> On Thursday, March 23, 2023 7:55 PM, Felipe Contreras wrote:\n>> >On Thu, Mar 23, 2023 at 5:43 PM <rsbecker@nexbridge.com> wrote:\n>> >>\n>> >> On Thursday, March 23, 2023 7:35 PM, Felipe Contreras wrote:\n>> >> >On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n>> >> >>\n>> >> >> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n>> >> >> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood\n>> >> >> ><phillip.wood123@gmail.com>\n>> >> >wrote:\n>> >> >> >>\n>> >> >> >> On 18/02/2023 01:59, demerphq wrote:\n>> >> >> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano\n>> >> >> >> > <gitster@pobox.com>\n>> >wrote:\n>> >> >> >> >>\n>> >> >> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n>> >> >> >> >>\n>> >> >> >> >>> Basically, if this effort turns out not to be fruitful as\n>> >> >> >> >>> a whole, I'd like for us to still have left a positive impact on the\n>codebase.\n>> >> >> >> >>> ...\n>> >> >> >> >>> So what's next? Naturally, I'm looking forward to a\n>> >> >> >> >>> spirited discussion about this topic - I'd like to know\n>> >> >> >> >>> which concerns haven't been addressed and figure out\n>> >> >> >> >>> whether we can find a way around them, and generally\n>> >> >> >> >>> build awareness of this effort with the\n>> >> >community.\n>> >> >> >> >>\n>> >> >> >> >> On of the gravest concerns is that the devil is in the details.\n>> >> >> >> >>\n>> >> >> >> >> For example, \"die() is inconvenient to callers, let's\n>> >> >> >> >> propagate errors up the callchain\" is an easy thing to\n>> >> >> >> >> say, but it would take much more than \"let's propagate errors up\"\n>> >> >> >> >> to libify something like\n>> >> >> >> >> check_connected() to do the same thing without spawning a\n>> >> >> >> >> separate process that is expected to exit with failure.\n>> >> >> >> >\n>> >> >> >> >\n>> >> >> >> > What does \"propagate errors up the callchain\" mean?  One\n>> >> >> >> > interpretation I can think of seems quite horrible, but\n>> >> >> >> > another seems quite doable and reasonable and likely not\n>> >> >> >> > even very invasive of the existing code:\n>> >> >> >> >\n>> >> >> >> > You can use setjmp/longjmp to implement a form of \"try\", so\n>> >> >> >> > that errors dont have to be *explicitly* returned *in* the call chain.\n>> >> >> >> > And you could probably do so without changing very much of\n>> >> >> >> > the existing code at all, and maintain a high level of\n>> >> >> >> > conceptual alignment with the current code strategy.\n>> >> >> >>\n>> >> >> >> Using setjmp/longjmp is an interesting suggestion, I think\n>> >> >> >> lua does something similar to what you describe for perl.\n>> >> >> >> However I think both of those use a allocator with garbage\n>> >> >> >> collection. I worry that using longjmp in git would be more\n>> >> >> >> invasive (or result in more memory leaks) as we'd need to to\n>> >> >> >> guard each allocation with some code to clean it up and then\n>> >> >> >> propagate the error. That means we're back to manually\n>> >> >> >> propagating errors up the call\n>> >chain in many cases.\n>> >> >> >\n>> >> >> >We could just use talloc [1].\n>> >> >>\n>> >> >> talloc is not portable.\n>> >> >\n>> >> >What makes you say that?\n>> >>\n>> >> talloc is not part of a POSIX standard I could find.\n>> >\n>> >It's a library, like: z, ssl, curl, pcre2-8, etc. Libraries can be\n>> >compiled on different platforms.\n>>\n>> talloc adds additional *required* dependencies to git, including python3 - required\n>to configure and build talloc - which is not available on the NonStop ia64 platform\n>(required support through end of 2025). I must express my resistance to what would\n>amount to losing support for git on this NonStop platform.\n>\n>That is not true. You don't need python3 for talloc, not even to build it, it's just a\n>single simple c file, it's easy to compile.\n>\n>The only reason python is used is to run waf, which is used to build Samba, which is\n>much more complex, but you don't need to run it, especially if you know the\n>characteristics of your system.\n>\n>This simple Makefile builds libtalloc.so just fine:\n>\n>  CC := gcc\n>  CFLAGS := -fPIC -I./lib/replace\n>  LDFLAGS := -Wl,--no-undefined\n>\n>  # For talloc.c\n>  CFLAGS += -DTALLOC_BUILD_VERSION_MAJOR=2\n>-DTALLOC_BUILD_VERSION_MINOR=4 -DTALLOC_BUILD_VERSION_RELEASE=0\n>  CFLAGS += -DHAVE_CONSTRUCTOR_ATTRIBUTE -DHAVE_VA_COPY -\n>DHAVE_VALGRIND_MEMCHECK_H -DHAVE_INTPTR_T\n>\n>  # For replace.h\n>  CFLAGS += -DNO_CONFIG_H -D__STDC_WANT_LIB_EXT1__=1\n>  CFLAGS += -DHAVE_STDBOOL_H -DHAVE_BOOL -DHAVE_STRING_H -\n>DHAVE_LIMITS_H -DHAVE_STDINT_H\n>  CFLAGS += -DHAVE_DLFCN_H -DHAVE_UINTPTR_T -DHAVE_C99_VSNPRINTF -\n>DHAVE_MEMMOVE -DHAVE_STRNLEN -DHAVE_VSNPRINTF\n>\n>  libtalloc.so: talloc.o\n>    $(CC) $(LDFLAGS) -shared -o $@ $^\n>\n>But of course, most of those defines are not even needed with a simple \"replace.h\"\n>that is less than 10 lines of code.\n\nThe 2.4.0 version is substantially more complex than one .c file. The version I just unpacked has 61 C and H files, and requires more than a trivial makefile. \n\nconfig.h requires that the ./configure script runs (which it does not because of python3). replace.h is over 1000 lines in this version.\n\nThis also gets into git supply chain issues where people who want to build git, outside of specific platforms, need to modify the build environment associated with talloc, which changes the delivered signature.\n\nSpeaking purely from a selfish standpoint, this is more than a trivial amount of work at least for me (the above Makefile does not satisfy 2.4.0) that is proposed in order to make talloc a seamless addition. Running the talloc test suite, which I consider a requirement to adding this as a dependency, is also problematic for the reasons I previously indicated.\n\nWe may just have to agree to disagree, but I stand by my concern that this suggestion will cause maintenance issues given the current state of the talloc code - and I do not have the cycles to become a community maintainer of that code also.\n\n--Randall\n\n"},{"id":"474139","messageId":"CAMP44s3kNSctpOhuLHkCE6tLV9-9-5LpVoORSyWt-Ov5FJ=cow@mail.gmail.com","threadId":"59261","inReplyTo":"005601d95e9c$dc20b360$94621a20$@nexbridge.com","subject":"Re: Proposal/Discussion: Turning parts of Git into libraries","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-03-24T22:29:44Z","receivedAt":"2023-03-24T22:30:05Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Fri, Mar 24, 2023 at 4:06 PM <rsbecker@nexbridge.com> wrote:\n>\n> On Friday, March 24, 2023 5:22 PM, Felipe Contreras wrote:\n> >On Fri, Mar 24, 2023 at 1:28 PM <rsbecker@nexbridge.com> wrote:\n> >>\n> >> On Thursday, March 23, 2023 7:55 PM, Felipe Contreras wrote:\n> >> >On Thu, Mar 23, 2023 at 5:43 PM <rsbecker@nexbridge.com> wrote:\n> >> >>\n> >> >> On Thursday, March 23, 2023 7:35 PM, Felipe Contreras wrote:\n> >> >> >On Thu, Mar 23, 2023 at 5:30 PM <rsbecker@nexbridge.com> wrote:\n> >> >> >>\n> >> >> >> On Thursday, March 23, 2023 7:22 PM, Felipe Contreras wrote:\n> >> >> >> >On Sat, Feb 18, 2023 at 5:12 AM Phillip Wood\n> >> >> >> ><phillip.wood123@gmail.com>\n> >> >> >wrote:\n> >> >> >> >>\n> >> >> >> >> On 18/02/2023 01:59, demerphq wrote:\n> >> >> >> >> > On Sat, 18 Feb 2023 at 00:24, Junio C Hamano\n> >> >> >> >> > <gitster@pobox.com>\n> >> >wrote:\n> >> >> >> >> >>\n> >> >> >> >> >> Emily Shaffer <nasamuffin@google.com> writes:\n> >> >> >> >> >>\n> >> >> >> >> >>> Basically, if this effort turns out not to be fruitful as\n> >> >> >> >> >>> a whole, I'd like for us to still have left a positive impact on the\n> >codebase.\n> >> >> >> >> >>> ...\n> >> >> >> >> >>> So what's next? Naturally, I'm looking forward to a\n> >> >> >> >> >>> spirited discussion about this topic - I'd like to know\n> >> >> >> >> >>> which concerns haven't been addressed and figure out\n> >> >> >> >> >>> whether we can find a way around them, and generally\n> >> >> >> >> >>> build awareness of this effort with the\n> >> >> >community.\n> >> >> >> >> >>\n> >> >> >> >> >> On of the gravest concerns is that the devil is in the details.\n> >> >> >> >> >>\n> >> >> >> >> >> For example, \"die() is inconvenient to callers, let's\n> >> >> >> >> >> propagate errors up the callchain\" is an easy thing to\n> >> >> >> >> >> say, but it would take much more than \"let's propagate errors up\"\n> >> >> >> >> >> to libify something like\n> >> >> >> >> >> check_connected() to do the same thing without spawning a\n> >> >> >> >> >> separate process that is expected to exit with failure.\n> >> >> >> >> >\n> >> >> >> >> >\n> >> >> >> >> > What does \"propagate errors up the callchain\" mean?  One\n> >> >> >> >> > interpretation I can think of seems quite horrible, but\n> >> >> >> >> > another seems quite doable and reasonable and likely not\n> >> >> >> >> > even very invasive of the existing code:\n> >> >> >> >> >\n> >> >> >> >> > You can use setjmp/longjmp to implement a form of \"try\", so\n> >> >> >> >> > that errors dont have to be *explicitly* returned *in* the call chain.\n> >> >> >> >> > And you could probably do so without changing very much of\n> >> >> >> >> > the existing code at all, and maintain a high level of\n> >> >> >> >> > conceptual alignment with the current code strategy.\n> >> >> >> >>\n> >> >> >> >> Using setjmp/longjmp is an interesting suggestion, I think\n> >> >> >> >> lua does something similar to what you describe for perl.\n> >> >> >> >> However I think both of those use a allocator with garbage\n> >> >> >> >> collection. I worry that using longjmp in git would be more\n> >> >> >> >> invasive (or result in more memory leaks) as we'd need to to\n> >> >> >> >> guard each allocation with some code to clean it up and then\n> >> >> >> >> propagate the error. That means we're back to manually\n> >> >> >> >> propagating errors up the call\n> >> >chain in many cases.\n> >> >> >> >\n> >> >> >> >We could just use talloc [1].\n> >> >> >>\n> >> >> >> talloc is not portable.\n> >> >> >\n> >> >> >What makes you say that?\n> >> >>\n> >> >> talloc is not part of a POSIX standard I could find.\n> >> >\n> >> >It's a library, like: z, ssl, curl, pcre2-8, etc. Libraries can be\n> >> >compiled on different platforms.\n> >>\n> >> talloc adds additional *required* dependencies to git, including python3 - required\n> >to configure and build talloc - which is not available on the NonStop ia64 platform\n> >(required support through end of 2025). I must express my resistance to what would\n> >amount to losing support for git on this NonStop platform.\n> >\n> >That is not true. You don't need python3 for talloc, not even to build it, it's just a\n> >single simple c file, it's easy to compile.\n> >\n> >The only reason python is used is to run waf, which is used to build Samba, which is\n> >much more complex, but you don't need to run it, especially if you know the\n> >characteristics of your system.\n> >\n> >This simple Makefile builds libtalloc.so just fine:\n> >\n> >  CC := gcc\n> >  CFLAGS := -fPIC -I./lib/replace\n> >  LDFLAGS := -Wl,--no-undefined\n> >\n> >  # For talloc.c\n> >  CFLAGS += -DTALLOC_BUILD_VERSION_MAJOR=2\n> >-DTALLOC_BUILD_VERSION_MINOR=4 -DTALLOC_BUILD_VERSION_RELEASE=0\n> >  CFLAGS += -DHAVE_CONSTRUCTOR_ATTRIBUTE -DHAVE_VA_COPY -\n> >DHAVE_VALGRIND_MEMCHECK_H -DHAVE_INTPTR_T\n> >\n> >  # For replace.h\n> >  CFLAGS += -DNO_CONFIG_H -D__STDC_WANT_LIB_EXT1__=1\n> >  CFLAGS += -DHAVE_STDBOOL_H -DHAVE_BOOL -DHAVE_STRING_H -\n> >DHAVE_LIMITS_H -DHAVE_STDINT_H\n> >  CFLAGS += -DHAVE_DLFCN_H -DHAVE_UINTPTR_T -DHAVE_C99_VSNPRINTF -\n> >DHAVE_MEMMOVE -DHAVE_STRNLEN -DHAVE_VSNPRINTF\n> >\n> >  libtalloc.so: talloc.o\n> >    $(CC) $(LDFLAGS) -shared -o $@ $^\n> >\n> >But of course, most of those defines are not even needed with a simple \"replace.h\"\n> >that is less than 10 lines of code.\n>\n> The 2.4.0 version is substantially more complex than one .c file. The version I just unpacked has 61 C and H files, and requires more than a trivial makefile.\n\nNot to build libtalloc.so, it doesn't.\n\n> config.h requires that the ./configure script runs (which it does not because of python3). replace.h is over 1000 lines in this version.\n\nNeither config.h nor replace.h are needed.\n\n> This also gets into git supply chain issues where people who want to build git, outside of specific platforms, need to modify the build environment associated with talloc, which changes the delivered signature.\n\nThe same happens with -lssl and many other libraries Git relies on.\n\nPeople who can't or won't use the ssl library can *alternatively* use\nblock-sha1/sha1.c, the same can be done with talloc, which the Git\nproject can easily embed.\n\n> Speaking purely from a selfish standpoint, this is more than a trivial amount of work at least for me (the above Makefile does not satisfy 2.4.0) that is proposed in order to make talloc a seamless addition.\n\nNobody is asking you to do anything.\n\n> Running the talloc test suite, which I consider a requirement to adding this as a dependency, is also problematic for the reasons I previously indicated.\n\nNobody is proposing to make this a dependency.\n\n> We may just have to agree to disagree, but I stand by my concern that this suggestion will cause maintenance issues given the current state of the talloc code - and I do not have the cycles to become a community maintainer of that code also.\n\nThat is simply not true. You have no idea what maintenance issues\ncould be caused by a hypothetical patch, because no such patch has\nbeen put forward, and likely never will.\n\nThe fact is that the suggestion is perfectly *doable*.\n\n-- \nFelipe Contreras\n"}]}