{"thread":{"id":"36109","subject":"[RFC/WIP] Pluggable reference backends","startedAt":"2014-03-10T11:00:32Z","lastAt":"2014-03-12T16:48:08Z","messageCount":17,"participants":["Michael Haggerty","Johan Herland","Shawn Pearce","Max Horn","Jeff King","David Kastrup","David Lang","Junio C Hamano","Karsten Blees","Andreas Krey"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"236370","messageId":"531D9B50.5030404@alum.mit.edu","threadId":"36109","inReplyTo":null,"subject":"[RFC/WIP] Pluggable reference backends","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2014-03-10T11:00:32Z","receivedAt":"2014-03-10T11:00:32Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"I have started working on pluggable ref backends.  In this email I\nwould like to share my plans and solicit feedback.\n\n(This morning I removed this project from the GSoC ideas page, because\nit is unfair to ask a student to shoot at a moving target.)\n\nWhy?\n====\n\nCurrently, the reference- and reflog-handling code in Git is too\ncoupled to the rest of the system.  There are too many places that\nknow, for example, the difference between loose and packed refs, or\nthat loose references are stored as files directly under\n$GIT_DIR/refs/heads/, or the locking protocols that have to be adhered\nto when managing references.  This tight coupling, in turn, makes it\nnearly impossible to experiment with alternate reference storage\nschemes.\n\nBut there is a lot of potential to use alternate reference storage\nschemes to fix some currently-unfixable problems, and to implement\nsome cool new features.\n\nUnfixable problems\n------------------\n\nThe on-disk format that we currently use to store references makes\nsome problems impossible to fix:\n\n* It is impossible to get a self-consistent snapshot of all references\n  at a given moment in time.  This makes it impossible, even in\n  principle, to do object pruning in a 100% race-free way.  (Our\n  current workaround of not deleting objects that are less than two\n  weeks works in most cases but, aside from being ugly, has holes.\n\n* There are awkward filesystem-imposed constraints on reference\n  naming, for example:\n\n  * D/F conflicts (I): it is not possible to have branches named\n    \"my-feature\" and \"my-feature/base\" at the same time.\n\n  * D/F conflicts (II): it is not possible to have reflogs for\n    branches named \"my-feature\" and \"my-feature/base\" at the same\n    time.  This leads to the problem that it is not, in general,\n    possible to retain reflogs for branches that have been deleted.\n\n  * There are additional constraints on reference names depending on\n    the filesystem used to store them.  For example, a Git repository\n    on a case-insensitive filesystem fails in confusing ways if there\n    are two loose references whose names differ only in case; however,\n    packed references differing in case might work for a while.  Also,\n    reference names that include Unicode characters can have their\n    normalization form changed if they are written on Mac OS.\n\n* The packed-refs file has to be rewritten whenever a packed reference\n  is deleted.  It might be nice to write 0{40} to a loose reference\n  file to indicate that the reference has been deleted, but that would\n  open the way for more D/F conflicts.)\n\nWild new ideas\n--------------\n\nSo, I would like to reorganize the Git code to allow pluggable\nreference backends.  If we had this, we could try out ideas like\n\n* Retain the idea of loose/packed references, but encode loose\n  reference names using a portable naming scheme before storing them\n  to the filesystem; maybe something like\n\n      refs/heads/Foo.42 -> refs.dir/heads.dir/%46oo%2e42\n      logs/refs/heads/Foo.42 -> refs.dir/heads.dir/%46oo%2e42.log\n\n  Yes, it looks uglier.  But users shouldn't be looking in these\n  directories anyway.  This single change would prevent D/F conflicts,\n  allow a reference to be deleted by writing 0{40} to its loose\n  reference file, allow reflogs to be kept for deleted refs, and\n  remove the problem of filesystem-dependent naming constraints.\n\n* Store references in a SQLite database, to get correct transaction\n  handling.\n\n* Store references directly in the Git object database.\n\n* Implement repository \"groups\" that share a common object database\n  and also a common reference store.  Each repository in a group would\n  get a sub-namespace in the shared database, and store its references\n  in names like \"refs/member/$MEMBERID/refs/heads/...\".  The member\n  repos would act like restricted views of the shared database.  This\n  would be like a combination between alternates (with lowered risk of\n  corruption) and gitnamespaces(7) (but usable for all git commands).\n\n* Reference transactions that can be used across multiple Git\n  commands.  Imagine,\n\n      export GIT_TRANSACTION=$(git transaction begin)\n      trap 'git transaction rollback' ERR\n      git foo ...\n      git bar ...\n      git baz ...\n      if ! git transaction commit\n      then\n          # Transaction failed; all references rolled back\n      else\n          # Transaction succeeded; all references updated atomically\n      fi\n      trap '' ERR\n      unset GIT_TRANSACTION\n\n  The \"GIT_TRANSACTION\" environment variable would tell git to read\n  from the usual references, overridden with any reference changes\n  that have occurred during the transaction, but write any changes\n  (including both old and new values) to the transaction.  The command\n  \"git transaction commit\" would verify that the old values listed in\n  the transaction still agree with the current values, and then make\n  all of the changes atomically.\n\n  Such transactions could also be broadcast to mirrors when they are\n  committed to keep multiple Git repositories in sync.\n\n* One alternate backend might even be a shim that delegates to libgit2\n  to do the actual reading/writing of references.  Then new backends\n  could be implemented in libgit2 to allow both git and libgit2 to\n  benefit.\n\n\nThe plan\n========\n\nIt is currently not possible to experiment with any of these things\nbecause of the tight coupling between the reference code and the rest\nof git. The goal of this project is first to choke the interactions\ndown to a coherent interface, and second to make the implementation\nselectable at runtime.  The implementation of specific alternate\nbackends will hopefully follow.\n\nquagga references\n-----------------\n\nThe overriding task is to isolate the reference-handling code; i.e.,\nmake sure that only code within refs.c touches git references, and\nthat the refs API provides all of the features that other code needs\nto do its work.\n\nSo as a whimsical first milestone, I want to make it possible to\nchoose a different directory name for storing references and reflogs\nby changing one #define statement in refs.c.  The goal is to get the\ntest suite to run correctly regardless of how this variable is set,\nwhich would be a pretty good check that all reference-handling code\npaths go though the refs API.  For no special reason I've been using\n\"quagga\" as the new place, so references go to \"$GIT_DIR/quagga/HEAD\",\n\"$GIT_DIR/quagga/refs/heads/master\", etc.  (Of course we wouldn't\nactually *change* this name; it is only for testing purposes.)  I've\nstarted working on this but there is a lot of code to change\n(including test code).\n\nReference transactions\n----------------------\n\nI want to orient the new reference API as much as possible around\ntransactions.  I think a transaction is a flexible abstraction that\nshould be implementable by any backend (albeit not always with 100%\nACID compliance) and will allow a couple of existing races to be\nfixed.\n\nSo as a first step, I will soon submit a patch series that starts\nfleshing out the concept of a ref_transaction, and rewrites \"git\nupdate-ref --stdin\" to use the new API.  For now, ref_transaction will\nonly be usable within a single git command invocation, but I want to\nleave the way open to the GIT_TRANSACTION idea mentioned above.\n\n\nTransition\n==========\n\nThe current project is only to isolate the reference-handling code and\nmake it, in principle, exchangeable with another implementation.  It\ndoesn't require any transition.\n\nMoreover, the changes will improve the modularity of the Git code, and\nwill be beneficial purely on those grounds.\n\nWhen/if alternate backends are implemented, then the transition will\nhave to be handled on a case-by-case basis.  How references are stored\nis mostly a decision internal to a single repository.  Any new\nrepository storage formats should be supported *in addition to* the\ntraditional storage scheme, to prevent the need for a flag day when\nall repositories have to be converted simultaneously.\n\nGit hosters [1] will be likely to take advantage of alternate\nreference backends pretty easily, because they know which tools touch\ntheir repositories and need only update those tools.  It is expected\nthat alternate reference backends will be useful for hosters even if\nthey don't become practical for end-users.\n\nFor end-users it is important that their repository be readable by all\nof the tools that they use.  So if we want to make a new format a\nviable option for normal Git users (let alone make it the new default\nformat), some coordination will be needed between all of the\ncommonly-used Git implementations (git-core, libgit2, JGit, and maybe\nDulwich, Grit, ...).  Whether or not this happens in real life depends\non how advantageous the hypothetical new format is to Git users and is\nbeyond the scope of this proposal.\n\nMichael\n\n[1] Full discloser: this includes my employer, GitHub.\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"236373","messageId":"CALKQrgfczpUU6Lh+hf867UOYUge2qsBTBTknwAk2pvRF6=xi4A@mail.gmail.com","threadId":"36109","inReplyTo":"531D9B50.5030404@alum.mit.edu","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2014-03-10T11:44:23Z","receivedAt":"2014-03-10T11:44:23Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Mon, Mar 10, 2014 at 12:00 PM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> I have started working on pluggable ref backends.  In this email I\n> would like to share my plans and solicit feedback.\n\nNo comments or useful feedback yet, except that I enthusiastically\napprove of the objective and the plan you have for how to get there.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"236406","messageId":"CAJo=hJtiPgByhk9M4ZKD98DARzgeU6z2mmw7fcLTEbBza-_h6A@mail.gmail.com","threadId":"36109","inReplyTo":"531D9B50.5030404@alum.mit.edu","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2014-03-10T14:30:45Z","receivedAt":"2014-03-10T14:30:45Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Mon, Mar 10, 2014 at 4:00 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> I have started working on pluggable ref backends.  In this email I\n> would like to share my plans and solicit feedback.\n\nYay!\n\nJGit already has pluggable ref backends, so it is good to see this\nstarting in git-core.\n\nFWIW the Gerrit Code Review community is interested in this project.\n\n> * Store references in a SQLite database, to get correct transaction\n>   handling.\n\nNo to SQLLite in git-core. Using it from JGit requires building\nSQLLite and a JNI wrapper, which makes JGit significantly less\nportable. I know SQLLite is pretty amazing, but implementing\ncompatibility with it from JGit will be a big nightmare for us.\n\n> * Reference transactions that can be used across multiple Git\n>   commands.  Imagine,\n>\n>       export GIT_TRANSACTION=$(git transaction begin)\n>       trap 'git transaction rollback' ERR\n>       git foo ...\n>       git bar ...\n>       git baz ...\n>       if ! git transaction commit\n>       then\n>           # Transaction failed; all references rolled back\n>       else\n>           # Transaction succeeded; all references updated atomically\n>       fi\n>       trap '' ERR\n>       unset GIT_TRANSACTION\n>\n>   The \"GIT_TRANSACTION\" environment variable would tell git to read\n>   from the usual references, overridden with any reference changes\n>   that have occurred during the transaction, but write any changes\n>   (including both old and new values) to the transaction.  The command\n>   \"git transaction commit\" would verify that the old values listed in\n>   the transaction still agree with the current values, and then make\n>   all of the changes atomically.\n\nYay!\n\nGerrit Code Review really wants to get transactions implemented. So I\nam very much in favor of trying to improve the situation in git-core.\n\nWe want not only a transaction over 2+ references in the same\nrepository, but we also want to perform transactions across\nrepositories. Consider a git submodule child and parent being updated\nat the same time. We really want to update refs/heads/master in both\nrepositories atomically at the central server.\n\n>   Such transactions could also be broadcast to mirrors when they are\n>   committed to keep multiple Git repositories in sync.\n\nOoh, this would be very interesting.\n\n> Git hosters [1] will be likely to take advantage of alternate\n> reference backends pretty easily, because they know which tools touch\n> their repositories and need only update those tools.  It is expected\n> that alternate reference backends will be useful for hosters even if\n> they don't become practical for end-users.\n\nAlternate reference backends are absolutely useful to large hosters.\nThe loose reference format isn't very scalable. The packed-refs helps,\nbut you can do better. IIRC our android.googlesource.com reference\nbackend uses only 79 bytes per reference on average, including both\nthe name string and the value. This super compact format is easy to\nhold in RAM for hundreds of busy repositories.\n\n> For end-users it is important that their repository be readable by all\n> of the tools that they use.  So if we want to make a new format a\n> viable option for normal Git users (let alone make it the new default\n> format), some coordination will be needed between all of the\n> commonly-used Git implementations (git-core, libgit2, JGit, and maybe\n> Dulwich, Grit, ...).  Whether or not this happens in real life depends\n> on how advantageous the hypothetical new format is to Git users and is\n> beyond the scope of this proposal.\n\nIt is sad we have this many implementations, but as one of the authors\n(JGit) I am happy to at least see you are worrying about compatibility\nwith them.\n"},{"id":"236413","messageId":"84EEAF2D-BCEA-4D02-95BE-31E9C518A0BC@quendi.de","threadId":"36109","inReplyTo":"CAJo=hJtiPgByhk9M4ZKD98DARzgeU6z2mmw7fcLTEbBza-_h6A@mail.gmail.com","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Max Horn","fromEmail":"max@quendi.de","sentAt":"2014-03-10T15:51:37Z","receivedAt":"2014-03-10T15:51:37Z","isPatch":false,"sender":{"key":"max@quendi.de","avatar":"https://avatars.githubusercontent.com/u/241512?v=4"},"body":"\nOn 10.03.2014, at 15:30, Shawn Pearce <spearce@spearce.org> wrote:\n\n> On Mon, Mar 10, 2014 at 4:00 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> I have started working on pluggable ref backends.  In this email I\n>> would like to share my plans and solicit feedback.\n> \n> Yay!\n\nYay, too!\n\n> JGit already has pluggable ref backends, so it is good to see this\n> starting in git-core.\n> \n> FWIW the Gerrit Code Review community is interested in this project.\n> \n>> * Store references in a SQLite database, to get correct transaction\n>>  handling.\n> \n> No to SQLLite in git-core. Using it from JGit requires building\n> SQLLite and a JNI wrapper, which makes JGit significantly less\n> portable. I know SQLLite is pretty amazing, but implementing\n> compatibility with it from JGit will be a big nightmare for us.\n\nI understood this as an example (indeed, it is listed under \"Wile new ideas\"), not a proposal to put this into the git core. It might be an interesting experiment in any case, and if the proposed modularity is truly achieved, it could (if there was any interest in it, that is) be implemented in an external 3rd party project.\n\n\nAnyway, I am quite excited about this project. Usually, I am quite skeptical about such large scope ideas (\"Yeah, cool idea, but who will pull it off, and with which resources?\"). But this one seems to have a good chance of being implemented gradually and inside the main repository, with the help of \"feature flags\". \n\nThus, I am looking forward to Michael's announced initial patch series. I feel that I don't know enough yet about git overall to be of much help on my own at this point. But perhaps over time some mini- or micro-projects pop up were others can help (e.g. \"adapt these 50 tests to work with the 'quagga' ref\"); if they are pointed out (assuming that doing so isn't more work than just addressing them yourself ;-), I am willing to help out.\n\n\nCheers,\nMax\n"},{"id":"236414","messageId":"20140310155230.GA29801@sigill.intra.peff.net","threadId":"36109","inReplyTo":"CAJo=hJtiPgByhk9M4ZKD98DARzgeU6z2mmw7fcLTEbBza-_h6A@mail.gmail.com","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2014-03-10T15:52:30Z","receivedAt":"2014-03-10T15:52:30Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 10, 2014 at 07:30:45AM -0700, Shawn Pearce wrote:\n\n> > * Store references in a SQLite database, to get correct transaction\n> >   handling.\n> \n> No to SQLLite in git-core. Using it from JGit requires building\n> SQLLite and a JNI wrapper, which makes JGit significantly less\n> portable. I know SQLLite is pretty amazing, but implementing\n> compatibility with it from JGit will be a big nightmare for us.\n\nThat seems like a poor reason not to implement a pluggable feature for\ngit-core. If we implement it, then a site using only git-core can take\nadvantage of it. Sites with JGit cannot, and would use a different\npluggable storage mechanism that's supported by both. But if we don't\nimplement, it hurts people using only git-core, and it does not help\nsites using JGit at all.\n\nThat's assuming that attention spent on implementing the feature does\nnot take away from implementing some other parallel scheme that does the\nsame thing but does not use SQLite. I don't know what that would be\noffhand; mapping the ref and reflog into a relational database is pretty\nsimple, and we get a lot of robustness and efficiency benefits for free.\nWe could perhaps have some kind of \"relational\" backend could use an\nODBC-like abstraction to point to a database. I have no idea if people\nwould want to ever store refs in a \"real\" server-backend RDBMS, but I\nsuspect Java has native support for such things.\n\nCertainly I think we should aim for compatibility where we can, but if\nthere's not a compatible way to do something, I don't think the\nlimitations of one platform should drag other ones down. And that goes\nboth ways; we had to reimplement disk-compatible EWAH from scratch in C\nfor git-core to have bitmaps, whereas JGit just got to use a ready-made\nlibrary. I don't think that was a bad thing.  People in\nmixed-implementation environments couldn't use it, but people with\nJGit-only environments were free to take advantage of it.\n\nAt any rate, the repository needs to advertise \"this is the ref storage\nmechanism I use\" in the config. We're going to need to bump\ncore.repositoryformatversion for such cases (because an old version of\ngit should not blindly lock and write to a refs/ directory that nobody\nelse is ever going to look at). And I'd suggest with that bump adding in\nsomething like core.refstorage, so that an implementation can say\n\"foobar ref storage? Never heard of it\" and barf. Whether it's because\nthat implementation doesn't support \"foobar\", because it's an old\nversion that doesn't understand \"foobar\" yet, or because it was simply\nbuilt without \"foobar\" support.\n\n-Peff\n"},{"id":"236415","messageId":"87k3c2820l.fsf@fencepost.gnu.org","threadId":"36109","inReplyTo":"20140310155230.GA29801@sigill.intra.peff.net","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2014-03-10T16:14:02Z","receivedAt":"2014-03-10T16:14:02Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Mon, Mar 10, 2014 at 07:30:45AM -0700, Shawn Pearce wrote:\n>\n>> > * Store references in a SQLite database, to get correct transaction\n>> >   handling.\n>> \n>> No to SQLLite in git-core. Using it from JGit requires building\n>> SQLLite and a JNI wrapper, which makes JGit significantly less\n>> portable. I know SQLLite is pretty amazing, but implementing\n>> compatibility with it from JGit will be a big nightmare for us.\n>\n> That seems like a poor reason not to implement a pluggable feature for\n> git-core. If we implement it, then a site using only git-core can take\n> advantage of it. Sites with JGit cannot, and would use a different\n> pluggable storage mechanism that's supported by both. But if we don't\n> implement, it hurts people using only git-core, and it does not help\n> sites using JGit at all.\n\nOf course, the basic premise for this feature is \"let's assume that our\nfile and/or operating system suck at providing file system functionality\nat file name granularity\".  There have been two historically approaches\nto that problem that are not independent: a) use Linux b) kick Linus.\n\nOption b) has been fairly successful over quite a bit of time, but at\nthe current point of time, it has become harder to aim that kick on a\nsingle person and/or where it counts.\n\nThe database approach is an alternative approach based on kicking an\nalternate set of people, namely database rather than operating system\nproviders, based on the assumption that the former have softer behinds\n(the backend-based approach) making them more sensitive to kicking.\n\nSo the database approach is most promising on the \"what are we going to\ndo if our operating system vendor won't bother with sensible file system\nperformance\" angle.  Which isn't doing total system architecture a\nfavor.\n\nPersonally, I have little sympathy for helping subpar systems, keeping\nthem on life support while they are in turn trying to squish the better\nsystems.\n\nBut then it is not me doing the actual work, so this is no more than an\nidle reflection.\n\n-- \nDavid Kastrup\n"},{"id":"236416","messageId":"alpine.DEB.2.02.1403100923110.16215@nftneq.ynat.uz","threadId":"36109","inReplyTo":"87k3c2820l.fsf@fencepost.gnu.org","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2014-03-10T16:28:10Z","receivedAt":"2014-03-10T16:28:10Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 10 Mar 2014, David Kastrup wrote:\n\n> Jeff King <peff@peff.net> writes:\n>\n>> On Mon, Mar 10, 2014 at 07:30:45AM -0700, Shawn Pearce wrote:\n>>\n>>>> * Store references in a SQLite database, to get correct transaction\n>>>>   handling.\n>>>\n>>> No to SQLLite in git-core. Using it from JGit requires building\n>>> SQLLite and a JNI wrapper, which makes JGit significantly less\n>>> portable. I know SQLLite is pretty amazing, but implementing\n>>> compatibility with it from JGit will be a big nightmare for us.\n>>\n>> That seems like a poor reason not to implement a pluggable feature for\n>> git-core. If we implement it, then a site using only git-core can take\n>> advantage of it. Sites with JGit cannot, and would use a different\n>> pluggable storage mechanism that's supported by both. But if we don't\n>> implement, it hurts people using only git-core, and it does not help\n>> sites using JGit at all.\n>\n> Of course, the basic premise for this feature is \"let's assume that our\n> file and/or operating system suck at providing file system functionality\n> at file name granularity\".  There have been two historically approaches\n> to that problem that are not independent: a) use Linux b) kick Linus.\n\nAs a note, if this is done properly, it could allow for plugins that connect to \nthe underlying storage system (similar to the Facebook Mecurial change)\n\nEven for those who don't have the $$$$$ storage arrays, there may be other \nstorage specific hacks that can be done to detect that files haven't changed.\n\nFor example, with btrfs and you compile into a different directory thatn your \nsource, you may be able to detect that things didn't change by the fact that the \nfilesystem didn't have to do a rewrite of the parent node.\n\nDavid Lang\n"},{"id":"236424","messageId":"xmqqwqg2q752.fsf@gitster.dls.corp.google.com","threadId":"36109","inReplyTo":"20140310155230.GA29801@sigill.intra.peff.net","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-10T17:46:01Z","receivedAt":"2014-03-10T17:46:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Mon, Mar 10, 2014 at 07:30:45AM -0700, Shawn Pearce wrote:\n>\n>> > * Store references in a SQLite database, to get correct transaction\n>> >   handling.\n>> \n>> No to SQLLite in git-core. Using it from JGit requires building\n>> SQLLite and a JNI wrapper, which makes JGit significantly less\n>> portable. I know SQLLite is pretty amazing, but implementing\n>> compatibility with it from JGit will be a big nightmare for us.\n>\n> That seems like a poor reason not to implement a pluggable feature for\n> git-core. If we implement it, then a site using only git-core can take\n> advantage of it. Sites with JGit cannot, and would use a different\n> pluggable storage mechanism that's supported by both. But if we don't\n> implement, it hurts people using only git-core, and it does not help\n> sites using JGit at all.\n\nWe would need to eventually have at least one backend that we know\nwill play well with different Git implementations that matter\n(namely, git-core, Jgit and libgit2) before the feature can be\nwidely adopted.\n\nThe first backend that is used while the plugging-interface is in\ndevelopment can be anything and does not have to be one that\neventual ubiquitous one, however; as long as it is something that we\ndo not mind carrying it forever, along with that final reference\nbackend.  I take the objection from Shawn only as against making the\nsqlite that final one.\n"},{"id":"236425","messageId":"20140310175658.GA23255@sigill.intra.peff.net","threadId":"36109","inReplyTo":"xmqqwqg2q752.fsf@gitster.dls.corp.google.com","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2014-03-10T17:56:58Z","receivedAt":"2014-03-10T17:56:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 10, 2014 at 10:46:01AM -0700, Junio C Hamano wrote:\n\n> >> No to SQLLite in git-core. Using it from JGit requires building\n> >> SQLLite and a JNI wrapper, which makes JGit significantly less\n> >> portable. I know SQLLite is pretty amazing, but implementing\n> >> compatibility with it from JGit will be a big nightmare for us.\n> >\n> > That seems like a poor reason not to implement a pluggable feature for\n> > git-core. If we implement it, then a site using only git-core can take\n> > advantage of it. Sites with JGit cannot, and would use a different\n> > pluggable storage mechanism that's supported by both. But if we don't\n> > implement, it hurts people using only git-core, and it does not help\n> > sites using JGit at all.\n> \n> We would need to eventually have at least one backend that we know\n> will play well with different Git implementations that matter\n> (namely, git-core, Jgit and libgit2) before the feature can be\n> widely adopted.\n\nI assumed that the current refs/ and logs/ code, massaged into pluggable\nbackend form, would be the first such. And I wouldn't be surprised to\nsee some iteration on that once it is easier to move from scheme to\nscheme (e.g., to use some encoding of the names on the filesystem to\navoid D/F conflicts, and thus allow reflogs for deleted refs).\n\n> The first backend that is used while the plugging-interface is in\n> development can be anything and does not have to be one that\n> eventual ubiquitous one, however; as long as it is something that we\n> do not mind carrying it forever, along with that final reference\n> backend.  I take the objection from Shawn only as against making the\n> sqlite that final one.\n\nSure, I'd agree with that. I'd think something like an sqlite interface\nwould be mainly of interest to people running busy servers. I don't know\nthat it would make a good default.\n\n-Peff\n"},{"id":"236443","messageId":"20140310194247.GA24568@sigill.intra.peff.net","threadId":"36109","inReplyTo":"87k3c2820l.fsf@fencepost.gnu.org","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2014-03-10T19:42:47Z","receivedAt":"2014-03-10T19:42:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 10, 2014 at 05:14:02PM +0100, David Kastrup wrote:\n\n> [storing refs in sqlite]\n>\n> Of course, the basic premise for this feature is \"let's assume that our\n> file and/or operating system suck at providing file system functionality\n> at file name granularity\".  There have been two historically approaches\n> to that problem that are not independent: a) use Linux b) kick Linus.\n\nYou didn't define \"suck\" here, but there are a number of issues with the\ncurrent ref storage system. Here is a sampling:\n\n  1. The filesystem does not present an atomic view of the data (e.g.,\n     you read \"a\", then while you are reading \"b\", somebody else updates\n     \"a\"; your view is one that never existed at any point in time).\n\n  2. Using the filesystem creates D/F conflicts between branches \"foo\"\n     and \"foo/bar\". Because this name is a primary key even for the\n     reflogs, we cannot easily persist reflogs after the ref is removed.\n\n  3. We use packed-refs in conjunction with loose ones to achieve\n     reasonable performance when there are a large number of refs. The\n     scheme for determining the current value of a ref is complicated\n     and error-prone (we had several race conditions that caused real\n     data loss).\n\nThose things can be solved through better support from the filesystem.\nBut they were also solved decades ago by relational databases.\n\nI generally avoid databases where possible. They lock your data up in a\nbinary format that you can't easily touch with standard unix tools. And\nthey introduce complexity and opportunity for bugs.\n\nBut they are also a proven technology for solving exactly the sorts of\nproblems that some people are having with git. I do not see a reason not\nto consider them as an option for a pluggable refs system. But I also do\nnot see a reason to inflict their costs on people who do not have those\nproblems. And that is why Michael's email is about _pluggable_ ref\nbackends, and not \"let's convert git to sqlite\".\n\nI do not even know if sqlite is going to end up as an interesting\noption. But it will be nice to be able to experiment with it easily due\nto git's ref code becoming more modular.\n\n-Peff\n"},{"id":"236447","messageId":"87bnxd96ar.fsf@fencepost.gnu.org","threadId":"36109","inReplyTo":"20140310194247.GA24568@sigill.intra.peff.net","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2014-03-10T19:56:12Z","receivedAt":"2014-03-10T19:56:12Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Mon, Mar 10, 2014 at 05:14:02PM +0100, David Kastrup wrote:\n>\n>> [storing refs in sqlite]\n>>\n>> Of course, the basic premise for this feature is \"let's assume that our\n>> file and/or operating system suck at providing file system functionality\n>> at file name granularity\".  There have been two historically approaches\n>> to that problem that are not independent: a) use Linux b) kick Linus.\n>\n> You didn't define \"suck\" here, but there are a number of issues with the\n> current ref storage system. Here is a sampling:\n>\n>   1. The filesystem does not present an atomic view of the data (e.g.,\n>      you read \"a\", then while you are reading \"b\", somebody else updates\n>      \"a\"; your view is one that never existed at any point in time).\n\nIf there are no system calls suitable for addressing this problem that\nfundamentally concerns the use of the file system as a file-name\naddressed data store, I don't see why \"kick Linus\" would not apply here.\n\n>   2. Using the filesystem creates D/F conflicts between branches \"foo\"\n>      and \"foo/bar\". Because this name is a primary key even for the\n>      reflogs, we cannot easily persist reflogs after the ref is\n>      removed.\n\nThat actually sounds more like \"kick Junio\" territory (the wonderful\ntimes when \"kick Linus\" could achieve almost anything are over).  To\nwit: this sounds like a design shortcoming in Git's use of filesystems,\nnot something that is actually inherent in the use of files.\n\n>   3. We use packed-refs in conjunction with loose ones to achieve\n>      reasonable performance when there are a large number of refs. The\n>      scheme for determining the current value of a ref is complicated\n>      and error-prone (we had several race conditions that caused real\n>      data loss).\n\nAgain, that sounds like we are talking about a scenario that is not a\nproblem of files inherently but rather of Git's ways of managing them.\n\n> Those things can be solved through better support from the filesystem.\n> But they were also solved decades ago by relational databases.\n\nRelational databases that are not implemented on raw storage managed by\ndatabase servers will still map their operations to file operations.\n\n> But they are also a proven technology for solving exactly the sorts of\n> problems that some people are having with git. I do not see a reason\n> not to consider them as an option for a pluggable refs system.\n\nBut I think it would be wrong to try solving \"2.\" above at the database\nlevel when its actual problem lies with the reference->filename mapping\nscheme.\n\n-- \nDavid Kastrup\n"},{"id":"236460","messageId":"531E2986.8050604@alum.mit.edu","threadId":"36109","inReplyTo":"20140310155230.GA29801@sigill.intra.peff.net","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2014-03-10T21:07:18Z","receivedAt":"2014-03-10T21:07:18Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 03/10/2014 04:52 PM, Jeff King wrote:\n> On Mon, Mar 10, 2014 at 07:30:45AM -0700, Shawn Pearce wrote:\n> \n>>> * Store references in a SQLite database, to get correct transaction\n>>>   handling.\n>>\n>> No to SQLLite in git-core. Using it from JGit requires building\n>> SQLLite and a JNI wrapper, which makes JGit significantly less\n>> portable. I know SQLLite is pretty amazing, but implementing\n>> compatibility with it from JGit will be a big nightmare for us.\n> \n> That seems like a poor reason not to implement a pluggable feature for\n> git-core. If we implement it, then a site using only git-core can take\n> advantage of it. Sites with JGit cannot, and would use a different\n> pluggable storage mechanism that's supported by both. But if we don't\n> implement, it hurts people using only git-core, and it does not help\n> sites using JGit at all.\n\nI think it's important to distinguish between two types of backend:\n\n* Exotic backends, optimized for servers, or embedded systems, or other\ncontrolled environments where the person deploying Git can decide about\nthe whole technology stack.  Here I say let a thousand flowers bloom.\nIf user A wants to try an Oracle backend and only uses JGit, there's no\nneed for him to implement the equivalent backend for git-core or libgit2.\n\n* Mainstream backends, intended for use by end-users on their\nworkstations and notebooks.  Such backends will be pretty worthless if\nthey are not supported more or less universally, because one user will\nwant to use the command line and Eclipse, another Visual Studio and\nTortoiseGit, a third will use GitHub for Mac plus a bunch of shell\nscripts written by his IT department.  A backend that is not supported\nby the big three Git implementations (git-core, libgit2, and JGit) will\nprobably be rejected by users.  Realistically there will be at most a\ncouple of mainstream backends--in fact probably usually a single\nestablished one and occasionally a single next-generation one waiting\nfor people to migrate slowly to it.  For mainstream backends I think it\nis important for the implementations to plan and coordinate ahead of\ntime to make sure everybody's concerns are addressed.\n\nIt sounds to me like Shawn is saying \"please don't make a SQLite-based\nbackend the new default git-core backend\" and Peff is saying \"there is\nno reason that a Git hosting service shouldn't experiment with a\nSQLite-based backend\".  I see no contradiction there [1].\n\nAlso, please remember that I'm not advocating a SQLite backend or any\nother at this time.  I'm only refactoring code to open the way for\n*future* flamefests :-)\n\nMichael\n\n[1] There might of course be a technical argument about whether a\nSQLite-based backend would be SO AWESOME for end-users that switching to\nit would be worth the extra inconvenience for the JGit folks.\nPersonally I'm skeptical.\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"236479","messageId":"CAJo=hJt6zoJ=53JNUT6fLXM+5_4Af8enE67z3Ozv4DOz1jU1Eg@mail.gmail.com","threadId":"36109","inReplyTo":"531E2986.8050604@alum.mit.edu","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2014-03-11T02:39:00Z","receivedAt":"2014-03-11T02:39:00Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Mon, Mar 10, 2014 at 2:07 PM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> On 03/10/2014 04:52 PM, Jeff King wrote:\n>> On Mon, Mar 10, 2014 at 07:30:45AM -0700, Shawn Pearce wrote:\n>>\n>>>> * Store references in a SQLite database, to get correct transaction\n>>>>   handling.\n>>>\n>>> No to SQLLite in git-core. Using it from JGit requires building\n>>> SQLLite and a JNI wrapper, which makes JGit significantly less\n>>> portable. I know SQLLite is pretty amazing, but implementing\n>>> compatibility with it from JGit will be a big nightmare for us.\n>>\n>> That seems like a poor reason not to implement a pluggable feature for\n>> git-core. If we implement it, then a site using only git-core can take\n>> advantage of it. Sites with JGit cannot, and would use a different\n>> pluggable storage mechanism that's supported by both. But if we don't\n>> implement, it hurts people using only git-core, and it does not help\n>> sites using JGit at all.\n>\n> I think it's important to distinguish between two types of backend:\n>\n> * Exotic backends, optimized for servers, or embedded systems, or other\n> controlled environments where the person deploying Git can decide about\n> the whole technology stack.  Here I say let a thousand flowers bloom.\n> If user A wants to try an Oracle backend and only uses JGit, there's no\n> need for him to implement the equivalent backend for git-core or libgit2.\n\nFWIW I have been running JGit derived servers using Google Bigtable\nfor reference storage for years. So yes in this sort of environment\nlet people do what they think is best for them.\n\n> * Mainstream backends, intended for use by end-users on their\n> workstations and notebooks.  Such backends will be pretty worthless if\n> they are not supported more or less universally, because one user will\n> want to use the command line and Eclipse, another Visual Studio and\n> TortoiseGit, a third will use GitHub for Mac plus a bunch of shell\n> scripts written by his IT department.  A backend that is not supported\n> by the big three Git implementations (git-core, libgit2, and JGit) will\n> probably be rejected by users.  Realistically there will be at most a\n> couple of mainstream backends--in fact probably usually a single\n> established one and occasionally a single next-generation one waiting\n> for people to migrate slowly to it.  For mainstream backends I think it\n> is important for the implementations to plan and coordinate ahead of\n> time to make sure everybody's concerns are addressed.\n\nYes, this was my real concern. Eclipse users using EGit expect EGit to\nbe compatible with git-core at the filesystem level so they can do\nsomething in EGit then switch to a shell and bang out a command, or\nrun a script provided by their project or co-worker. Build systems\noften integrate with Git to e.g. embed `git describe` output into the\nbinary. In mainstream use cross compatibility of the tools within a\nsingle working directory is something that I think users have come to\nexpect.\n\n> It sounds to me like Shawn is saying \"please don't make a SQLite-based\n> backend the new default git-core backend\" and Peff is saying \"there is\n> no reason that a Git hosting service shouldn't experiment with a\n> SQLite-based backend\".  I see no contradiction there [1].\n\nYes. :-)\n\n> Also, please remember that I'm not advocating a SQLite backend or any\n> other at this time.  I'm only refactoring code to open the way for\n> *future* flamefests :-)\n>\n> Michael\n>\n> [1] There might of course be a technical argument about whether a\n> SQLite-based backend would be SO AWESOME for end-users that switching to\n> it would be worth the extra inconvenience for the JGit folks.\n> Personally I'm skeptical.\n\nIf it was really that amazing, yes, we would probably support it in\nJGit for those that need that amazing.\n\nBut I tend to think we can (usually) find a simpler format that would\nprovide many of the same benefits with less of the drawbacks of\nlocking the data up into SQLLite's file format. I'm with Peff, I kind\nof like the fact that most of the Git data is easy to inspect by hand,\nor with some simple tools written in Git's source tree. Starting with\n\"go get this other SQLLite tool first then write this code\" is a lot\nless fun.\n"},{"id":"236494","messageId":"531EEBCC.10409@gmail.com","threadId":"36109","inReplyTo":"531D9B50.5030404@alum.mit.edu","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2014-03-11T10:56:12Z","receivedAt":"2014-03-11T10:56:12Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 10.03.2014 12:00, schrieb Michael Haggerty:\n> \n> Reference transactions\n> ----------------------\n> \n\nVery cool ideas indeed.\n\nHowever, I'm concerned a bit that transactions are conceptual overkill. How many concurrent updates do you expect in a repository? Wouldn't a single repo-wide lock suffice (and be _much_ simpler to implement with any backend, esp. file-based)?\n\nThe API you posted in [1] doesn't look very much like a transaction API either (rather like batch-updates). E.g. there's no rollback, the queue* methods cannot report failure, and there's no way to read a ref as part of the transaction. So I'm afraid that backends that support transactions out of the box (e.g. RDBMSs) will be hard to adapt to this.\n\nJust my 2cents,\nKarsten\n\n[1] http://article.gmane.org/gmane.comp.version-control.git/243748\n"},{"id":"236569","messageId":"20140312102601.GA26257@inner.h.apk.li","threadId":"36109","inReplyTo":"CAJo=hJt6zoJ=53JNUT6fLXM+5_4Af8enE67z3Ozv4DOz1jU1Eg@mail.gmail.com","subject":"egit vs. git behaviour (was: [RFC/WIP] Pluggable reference backends)","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2014-03-12T10:26:01Z","receivedAt":"2014-03-12T10:26:01Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"On Mon, 10 Mar 2014 19:39:00 +0000, Shawn Pearce wrote:\n> Yes, this was my real concern. Eclipse users using EGit expect EGit to\n> be compatible with git-core at the filesystem level so they can do\n> something in EGit then switch to a shell and bang out a command, or\n> run a script provided by their project or co-worker.\n\nA question: Where to ask/report problems with that?\n\nWe're currently running into problems that egit doesn't push to where\ngit would when the local and remote branches aren't the same name. It\nseems that egit ignores the branch.*.merge settings. Or push.default?\n\nAndreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"236571","messageId":"53204867.4010809@alum.mit.edu","threadId":"36109","inReplyTo":"531EEBCC.10409@gmail.com","subject":"Re: [RFC/WIP] Pluggable reference backends","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2014-03-12T11:43:35Z","receivedAt":"2014-03-12T11:43:35Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Karsten,\n\nThanks for your feedback!\n\nOn 03/11/2014 11:56 AM, Karsten Blees wrote:\n> Am 10.03.2014 12:00, schrieb Michael Haggerty:\n>> \n>> Reference transactions ----------------------\n> \n> Very cool ideas indeed.\n> \n> However, I'm concerned a bit that transactions are conceptual\n> overkill. How many concurrent updates do you expect in a repository?\n> Wouldn't a single repo-wide lock suffice (and be _much_ simpler to\n> implement with any backend, esp. file-based)?\n\nI am mostly thinking about long-running processes, like \"gc\" and\n\"prune-refs\", which need to be made race-free without blocking other\nprocesses for the whole time they are running (whereas it might be quite\ntolerable to have them fail or only complete part of their work in any\ngiven invocation).  Also, I work at GitHub, where we have quite a few\nrepositories, some of which are quite active :-)\n\nRemember that I'm not yet proposing anything like hard-core ACID\nreference transactions.  I'm just clearing the way for various possible\nchanges in reference handling.  I listed the ideas only to whet people's\nappetites and motivate the refactoring, which will take a while before\nit bears any real fruit.\n\n> The API you posted in [1] doesn't look very much like a transaction\n> API either (rather like batch-updates). E.g. there's no rollback, the\n> queue* methods cannot report failure, and there's no way to read a\n> ref as part of the transaction. So I'm afraid that backends that\n> support transactions out of the box (e.g. RDBMSs) will be hard to\n> adapt to this.\n\nGmane is down at the moment but I assume you are referring to my patch\nseries and the ref_transaction implementation therein.\n\nNo explicit rollback is necessary at this stage, because the \"commit\"\nfunction first locks all of the references that it wants to change\n(first verifying that they have the expected values), and then modifies\nthem all.  By the time the references are locked, the whole transaction\nis guaranteed to succeed [1].  If the locks can't all be acquired, then\nany locks that were obtained are released.\n\nIf a caller wants to rollback a transaction, it only needs to free the\ntransaction instead of committing.  I should probably make that clearer\nby renaming free_ref_transaction() to rollback_ref_transaction().  By\nthe time we start implementing other reference backends, that function\nwill of course have to do more.  For that matter, maybe\ncreate_ref_transaction() should be renamed to begin_ref_transaction().\nNow would be a good time for concrete bikeshedding suggestions about\nfunction names or other details of the API :-)\n\nYes, the queue_*() methods should probably later make a preliminary\ncheck of the reference's old value and return an error if the expected\nvalue is already incorrect.  This would allow callers to fail fast if\nthe transaction is doomed to failure.  But that wasn't needed yet for\nthe one existing caller, which builds up a transaction and commits it\nimmediately, so I didn't implement it yet.  And the early checks would\nadd overhead for this caller, so maybe they should be optional anyway.\nMaybe these functions should already be declared to return an error\nstatus, but there should be an option passed to create_ref_transaction()\nthat selects whether fast checks should be performed or not for that\ntransaction.\n\nReally, all that this first patch series does is put a different API\naround the mechanism that was already there, in update_refs().  There\nwill be a lot more steps before we see anything approaching real\nreference transactions.  But I think your (implied) suggestion, to make\nthe API more reminiscent of something like database transactions, is a\ngood one and I will work on it.\n\nCheers,\nMichael\n\n[1] \"Guaranteed\" here is of course relative.  The commit could still\nfail due to the process being killed, disk errors, etc.  But it can't\nfail due to lock contention with another git process.\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"236583","messageId":"CAJo=hJvab468-PRMQJZ690A-ek8p01cwUOCkuy9KOQTRtt0FWQ@mail.gmail.com","threadId":"36109","inReplyTo":"20140312102601.GA26257@inner.h.apk.li","subject":"Re: egit vs. git behaviour (was: [RFC/WIP] Pluggable reference backends)","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2014-03-12T16:48:08Z","receivedAt":"2014-03-12T16:48:08Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Mar 12, 2014 at 3:26 AM, Andreas Krey <a.krey@gmx.de> wrote:\n> On Mon, 10 Mar 2014 19:39:00 +0000, Shawn Pearce wrote:\n>> Yes, this was my real concern. Eclipse users using EGit expect EGit to\n>> be compatible with git-core at the filesystem level so they can do\n>> something in EGit then switch to a shell and bang out a command, or\n>> run a script provided by their project or co-worker.\n>\n> A question: Where to ask/report problems with that?\n\nEGit developers have a bug tracker, from:\n\n  http://eclipse.org/egit/support/\n\nWe see File a bug with a link to:\n\n  https://bugs.eclipse.org/bugs/enter_bug.cgi?product=EGit&rep_platform=All&op_sys=All\n\n> We're currently running into problems that egit doesn't push to where\n> git would when the local and remote branches aren't the same name. It\n> seems that egit ignores the branch.*.merge settings. Or push.default?\n\nI think this is just missing code in EGit. Its probable they already\nknow about it, or many of them don't use these features in .git/config\nand thus don't realize they are missing.\n"}]}