{"thread":{"id":"57868","subject":"[FR] supporting submodules with alternate version control systems (new contributor)","startedAt":"2022-05-10T16:12:18Z","lastAt":"2022-06-06T14:53:28Z","messageCount":13,"participants":["Addison Klinke","Junio C Hamano","Jason Pyeron","rsbecker@nexbridge.com","Philip Oakley"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"455084","messageId":"CAE9CXuhvqfhARrqz2=oS1=9BF=iNhGbJv7y3HmYs1tddn8ndiQ@mail.gmail.com","threadId":"57868","inReplyTo":null,"subject":"[FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Addison Klinke","fromEmail":"addison@baller.tv","sentAt":"2022-05-10T16:11:48Z","receivedAt":"2022-05-10T16:12:18Z","isPatch":false,"sender":{"key":"addison@baller.tv","avatar":null},"body":"Hello all,\n\nI'm familiar with opensource software development through Github, but\nhave not contributed to git before so apologies if I'm using the wrong\navenues. Please point me in the right direction if that is the case. I\nsaw this mailing list mentioned on the\n[mirror](https://github.com/git/git) repository, so it seemed like the\nright place to start.\n\nI have a feature request I'd like some feedback on. The core idea is\nto support submodules with alternate (i.e. non-git based) version\ncontrol systems.\n\n* **Why:** Git is excellent for versioning code and I don't need\nanother VCS for that purpose. However, in machine learning (ML)\nworkflows it has become more\n[standard](https://opendatascience.com/how-data-versioning-can-be-used-in-machine-learning/)\nto version your datasets, and for this purpose many git-like tools\nhave been developed. See [Dolt](https://www.dolthub.com/),\n[LakeFS](https://lakefs.io/), and [DVC](https://dvc.org/) for a few\nexamples. Currently, ML practitioners have to bifurcate their\ndevelopment process - code is committed/managed with git and datasets\nare committed/managed with a 3rd party VCS (and often cloned in a\ndifferent folder outside the git repository). My proposal is to unify\nthe data versioning tools with git submodules so that they can act as\nany other 3rd party library inside a parent repository\n\n* **How:** Most data versioning tools already define a git-like CLI.\nFor instance, you have \"dolt commit\", \"dvc push\", \"lakectl diff\", etc.\nThe set of commands and options is usually a subset of the full list\navailable in git, but the important ones are there. My approach would\nrequire a few steps\n\n1. Git defines an API for configuring 3rd party VCS tools. It's\nessentially a mapping from git command to the equivalent in the 3rd\nparty library. This should also account for which options/flags are\nsupported\n2. Developers from the 3rd party library integrate with this git API\nby maintaining a config file for the mapping that gets installed\nalongside their binaries\n3. The .gitmodules syntax is extended to include a \"type\" field which\ndefaults to git but can be set to other supported values\n4. Then end-users can add submodules with an alternate VCS. Once\nadded, the CLI interaction would appear like normal git but under the\nhood it would be using a different engine (and remote storage)\n\nIs something along these lines feasible? If so, could someone who is\nmore familiar with the code base give me a rough idea how one might go\nabout this? I would like to author the PR to implement this - just\nlooking for some help getting started.\n\nThank you for the help,\n\nAddison\n"},{"id":"455087","messageId":"xmqq4k1x8gqj.fsf@gitster.g","threadId":"57868","inReplyTo":"CAE9CXuhvqfhARrqz2=oS1=9BF=iNhGbJv7y3HmYs1tddn8ndiQ@mail.gmail.com","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-05-10T17:00:52Z","receivedAt":"2022-05-10T17:01:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Addison Klinke <addison@baller.tv> writes:\n\n> Is something along these lines feasible?\n\nOffhand, I only think of one thing that could make it fundamentally\ninfeasible.\n\nWhen you bind an external repository (be it stored in Git or\nsomebody else's system) as a submodule, each commit in the\nsuperproject records which exact commit in the submodule is used\nwith the rest of the superproject tree.  And that is done by\nrecording the object name of the commit in the submodule.\n\nWhat it means for the foreign system that wants to \"plug into\" a\nsuperproject in Git as a submodule?  It is required to do two\nthings:\n\n * At the time \"git commit\" is run at the superproject level, the\n   foreign system has to be able to say \"the version I have to be\n   used in the context of this superproject commit is X\", with X\n   that somehow can be stored in the superproject's tree object\n   (which is sized 20-byte for SHA-1 repositories; in SHA-256\n   repositories, it is a bit wider).\n\n * At the time \"git chekcout\" is run at the superproject level, the\n   superproject will learn the above X (i.e. the version of the\n   submodule that goes with the version of the superproject being\n   checked out).  The foreign system has to be able to perform a\n   \"checkout\" given that X.\n\nIf a foreign system cannot do the above two, then it fundamentally\nwould be incapable of participating in such a \"superproject and\nsubmodule\" relationship.\n\nEverything else I think is feasible in the sense that \"it is just a\nmatter of programming\".\n\nIt is a different story how it is implemented, how much it would\ncost to do so, and if it is worth maintaining it as part of Git, so\nI'd stop at \"is it feasible?\" here, not judging \"if it is realistic\"\nat this point ;-).\n\n"},{"id":"455091","messageId":"01e601d86492$43bb70b0$cb325210$@pdinc.us","threadId":"57868","inReplyTo":"xmqq4k1x8gqj.fsf@gitster.g","subject":"RE: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Jason Pyeron","fromEmail":"jpyeron@pdinc.us","sentAt":"2022-05-10T17:20:36Z","receivedAt":"2022-05-10T17:21:03Z","isPatch":false,"sender":{"key":"jpyeron@pdinc.us","avatar":"https://gravatar.com/avatar/c2e53452caa53d940768a1ffc9cf76196d851b9b534b7a39cd39852a70a0508f?d=mp&s=160"},"body":"> -----Original Message-----\n> From: Junio C Hamano\n> Sent: Tuesday, May 10, 2022 1:01 PM\n> To: Addison Klinke <addison@baller.tv>\n> \n> Addison Klinke <addison@baller.tv> writes:\n> \n> > Is something along these lines feasible?\n> \n> Offhand, I only think of one thing that could make it fundamentally\n> infeasible.\n> \n> When you bind an external repository (be it stored in Git or\n> somebody else's system) as a submodule, each commit in the\n> superproject records which exact commit in the submodule is used\n> with the rest of the superproject tree.  And that is done by\n> recording the object name of the commit in the submodule.\n> \n> What it means for the foreign system that wants to \"plug into\" a\n> superproject in Git as a submodule?  It is required to do two\n> things:\n> \n>  * At the time \"git commit\" is run at the superproject level, the\n>    foreign system has to be able to say \"the version I have to be\n>    used in the context of this superproject commit is X\", with X\n>    that somehow can be stored in the superproject's tree object\n>    (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>    repositories, it is a bit wider).\n> \n>  * At the time \"git chekcout\" is run at the superproject level, the\n>    superproject will learn the above X (i.e. the version of the\n>    submodule that goes with the version of the superproject being\n>    checked out).  The foreign system has to be able to perform a\n>    \"checkout\" given that X.\n> \n> If a foreign system cannot do the above two, then it fundamentally\n> would be incapable of participating in such a \"superproject and\n> submodule\" relationship.\n\nThe submodule \"type\" could create an object (hashed and stored) that contains the needed \"translation\" details. The object would be hashed using SHA1 or SHA256 depending on the git config. The format of the object's contents would be defined by the submodule's \"code\".\n\n\n--\nJason Pyeron  | Architect\nPD Inc        | Certified SBA 8(a)\n10 w 24th St  | Certified SBA HUBZone\nBaltimore, MD | CAGE Code: 1WVR6\n \n.mil: jason.j.pyeron.ctr@mail.mil\n.com: jpyeron@pdinc.us\ntel : 202-741-9397\n\n\n\n"},{"id":"455092","messageId":"CAE9CXujPzu3_95pBDVRXKFU_z40j9Y7v5_1y3c+WnFpz1_oY4w@mail.gmail.com","threadId":"57868","inReplyTo":"01e601d86492$43bb70b0$cb325210$@pdinc.us","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Addison Klinke","fromEmail":"addison@baller.tv","sentAt":"2022-05-10T17:26:56Z","receivedAt":"2022-05-10T17:28:05Z","isPatch":false,"sender":{"key":"addison@baller.tv","avatar":null},"body":"Thanks for the quick replies\n\n> Junio Hamano: When you bind an external repository (be it stored in Git or\nsomebody else's system) as a submodule, each commit in the\nsuperproject records which exact commit in the submodule is used\nwith the rest of the superproject tree.\n\nThis should be fine then - at least the data versioning tools I'm\nfamiliar with can all specify their current commit and checkout by\ncommit hash. Does it matter how the hashes are structured/stored\ninternally? For example, I believe Dolt keeps them in a MySQL table\nthat connects to Noms under the hood.\n\n > Junio Hamano: not judging \"if it is realistic\" at this point\n\nWhat would be the best approach for answering this portion?\n\n> Jason Pyeron: The submodule \"type\" could create an object (hashed and stored) that contains the needed \"translation\" details\n\nThat sounds like an interesting idea. Since I'd like to offload the\nburden of maintaining these translation files to the 3rd party\ndevelopers, it would be nice if they got copied to a standard location\n(i.e. ~/.gitmodules/translations/tool_x) during the 3rd party install.\nThen when a submodule is added with \"type = tool_x\", git checks that\nthe appropriate translation file is available, and if so, copies it\ninto the parent repository.\n\nOn Tue, May 10, 2022 at 11:20 AM Jason Pyeron <jpyeron@pdinc.us> wrote:\n>\n> > -----Original Message-----\n> > From: Junio C Hamano\n> > Sent: Tuesday, May 10, 2022 1:01 PM\n> > To: Addison Klinke <addison@baller.tv>\n> >\n> > Addison Klinke <addison@baller.tv> writes:\n> >\n> > > Is something along these lines feasible?\n> >\n> > Offhand, I only think of one thing that could make it fundamentally\n> > infeasible.\n> >\n> > When you bind an external repository (be it stored in Git or\n> > somebody else's system) as a submodule, each commit in the\n> > superproject records which exact commit in the submodule is used\n> > with the rest of the superproject tree.  And that is done by\n> > recording the object name of the commit in the submodule.\n> >\n> > What it means for the foreign system that wants to \"plug into\" a\n> > superproject in Git as a submodule?  It is required to do two\n> > things:\n> >\n> >  * At the time \"git commit\" is run at the superproject level, the\n> >    foreign system has to be able to say \"the version I have to be\n> >    used in the context of this superproject commit is X\", with X\n> >    that somehow can be stored in the superproject's tree object\n> >    (which is sized 20-byte for SHA-1 repositories; in SHA-256\n> >    repositories, it is a bit wider).\n> >\n> >  * At the time \"git chekcout\" is run at the superproject level, the\n> >    superproject will learn the above X (i.e. the version of the\n> >    submodule that goes with the version of the superproject being\n> >    checked out).  The foreign system has to be able to perform a\n> >    \"checkout\" given that X.\n> >\n> > If a foreign system cannot do the above two, then it fundamentally\n> > would be incapable of participating in such a \"superproject and\n> > submodule\" relationship.\n>\n> The submodule \"type\" could create an object (hashed and stored) that contains the needed \"translation\" details. The object would be hashed using SHA1 or SHA256 depending on the git config. The format of the object's contents would be defined by the submodule's \"code\".\n>\n>\n> --\n> Jason Pyeron  | Architect\n> PD Inc        | Certified SBA 8(a)\n> 10 w 24th St  | Certified SBA HUBZone\n> Baltimore, MD | CAGE Code: 1WVR6\n>\n> .mil: jason.j.pyeron.ctr@mail.mil\n> .com: jpyeron@pdinc.us\n> tel : 202-741-9397\n>\n>\n>\n"},{"id":"455099","messageId":"03ca01d8649b$7d6a3310$783e9930$@nexbridge.com","threadId":"57868","inReplyTo":"CAE9CXujPzu3_95pBDVRXKFU_z40j9Y7v5_1y3c+WnFpz1_oY4w@mail.gmail.com","subject":"RE: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-05-10T18:26:33Z","receivedAt":"2022-05-10T18:26:47Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 10, 2022 1:27 PM, Addison Klinke wrote:\n>Thanks for the quick replies\n>\n>> Junio Hamano: When you bind an external repository (be it stored in\n>> Git or\n>somebody else's system) as a submodule, each commit in the superproject\n>records which exact commit in the submodule is used with the rest of the\n>superproject tree.\n>\n>This should be fine then - at least the data versioning tools I'm familiar with can all\n>specify their current commit and checkout by commit hash. Does it matter how\n>the hashes are structured/stored internally? For example, I believe Dolt keeps\n>them in a MySQL table that connects to Noms under the hood.\n>\n> > Junio Hamano: not judging \"if it is realistic\" at this point\n>\n>What would be the best approach for answering this portion?\n\nBasically, answer the following: Can you implement a command like the cvs2git that can be re-executed on an idempotent (repeatedly with the same result) basis?\n\nIf yes, then you can build your own automation to move code into a submodule from your own VCS system into a git repository and the work with the submodule without the git code-base knowing about this.\n\nIf you can go the other way, from git to your other VCS system, repeatedly, then you can go back again. This is likely to be much harder as git has a much richer representation model than is typical of VCS systems.\n\nOne way may be sufficient for your purposes. Research how cvs2git works and see whether you are able to emulate its functions.\n\n>> Jason Pyeron: The submodule \"type\" could create an object (hashed and\n>> stored) that contains the needed \"translation\" details\n>\n>That sounds like an interesting idea. Since I'd like to offload the burden of\n>maintaining these translation files to the 3rd party developers, it would be nice if\n>they got copied to a standard location (i.e. ~/.gitmodules/translations/tool_x)\n>during the 3rd party install.\n>Then when a submodule is added with \"type = tool_x\", git checks that the\n>appropriate translation file is available, and if so, copies it into the parent\n>repository.\n>\n>On Tue, May 10, 2022 at 11:20 AM Jason Pyeron <jpyeron@pdinc.us> wrote:\n>>\n>> > -----Original Message-----\n>> > From: Junio C Hamano\n>> > Sent: Tuesday, May 10, 2022 1:01 PM\n>> > To: Addison Klinke <addison@baller.tv>\n>> >\n>> > Addison Klinke <addison@baller.tv> writes:\n>> >\n>> > > Is something along these lines feasible?\n>> >\n>> > Offhand, I only think of one thing that could make it fundamentally\n>> > infeasible.\n>> >\n>> > When you bind an external repository (be it stored in Git or\n>> > somebody else's system) as a submodule, each commit in the\n>> > superproject records which exact commit in the submodule is used\n>> > with the rest of the superproject tree.  And that is done by\n>> > recording the object name of the commit in the submodule.\n>> >\n>> > What it means for the foreign system that wants to \"plug into\" a\n>> > superproject in Git as a submodule?  It is required to do two\n>> > things:\n>> >\n>> >  * At the time \"git commit\" is run at the superproject level, the\n>> >    foreign system has to be able to say \"the version I have to be\n>> >    used in the context of this superproject commit is X\", with X\n>> >    that somehow can be stored in the superproject's tree object\n>> >    (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>> >    repositories, it is a bit wider).\n>> >\n>> >  * At the time \"git chekcout\" is run at the superproject level, the\n>> >    superproject will learn the above X (i.e. the version of the\n>> >    submodule that goes with the version of the superproject being\n>> >    checked out).  The foreign system has to be able to perform a\n>> >    \"checkout\" given that X.\n>> >\n>> > If a foreign system cannot do the above two, then it fundamentally\n>> > would be incapable of participating in such a \"superproject and\n>> > submodule\" relationship.\n>>\n>> The submodule \"type\" could create an object (hashed and stored) that contains\n>the needed \"translation\" details. The object would be hashed using SHA1 or\n>SHA256 depending on the git config. The format of the object's contents would be\n>defined by the submodule's \"code\".\n\nI would not try to do this inside the git infrastructure. What you may be able to do in my suggestion above, is to restrict how your other VCS system is used and restrict how your team uses git to make the mapping repeatable. This is typical of some environments where there is an SVN repo and a git repo that are mirrored. This does simplify matters particularly if you do not have to modify either system but are building a façade or wrapper around both.\n\nKeep this as simple as possible to meet a minimum viable set of requirements.\n--Randal \n\n"},{"id":"455119","messageId":"271b6a9a-a5f4-0336-51b8-860ad07f2609@iee.email","threadId":"57868","inReplyTo":"01e601d86492$43bb70b0$cb325210$@pdinc.us","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-05-10T20:54:10Z","receivedAt":"2022-05-10T20:54:19Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 10/05/2022 18:20, Jason Pyeron wrote:\n>> -----Original Message-----\n>> From: Junio C Hamano\n>> Sent: Tuesday, May 10, 2022 1:01 PM\n>> To: Addison Klinke <addison@baller.tv>\n>>\n>> Addison Klinke <addison@baller.tv> writes:\n>>\n>>> Is something along these lines feasible?\n>> Offhand, I only think of one thing that could make it fundamentally\n>> infeasible.\n>>\n>> When you bind an external repository (be it stored in Git or\n>> somebody else's system) as a submodule, each commit in the\n>> superproject records which exact commit in the submodule is used\n>> with the rest of the superproject tree.  And that is done by\n>> recording the object name of the commit in the submodule.\n>>\n>> What it means for the foreign system that wants to \"plug into\" a\n>> superproject in Git as a submodule?  It is required to do two\n>> things:\n>>\n>>   * At the time \"git commit\" is run at the superproject level, the\n>>     foreign system has to be able to say \"the version I have to be\n>>     used in the context of this superproject commit is X\", with X\n>>     that somehow can be stored in the superproject's tree object\n>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>>     repositories, it is a bit wider).\n>>\n>>   * At the time \"git chekcout\" is run at the superproject level, the\n>>     superproject will learn the above X (i.e. the version of the\n>>     submodule that goes with the version of the superproject being\n>>     checked out).  The foreign system has to be able to perform a\n>>     \"checkout\" given that X.\n>>\n>> If a foreign system cannot do the above two, then it fundamentally\n>> would be incapable of participating in such a \"superproject and\n>> submodule\" relationship.\n\nThe sub-modules already have that problem if the user forgets publish \ntheir sub-module (see notes in the docs ;-).\n> The submodule \"type\" could create an object (hashed and stored) that contains the needed \"translation\" details. The object would be hashed using SHA1 or SHA256 depending on the git config. The format of the object's contents would be defined by the submodule's \"code\".\n>\nAnother way of looking at the issue is via a variant of Git-LFS with a \nsmudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n\nThe LFS already uses the .gitattributes to define a 'type', while the \nsubmodules don't yet have that capability. There is just a single \nspecial type within a tree object of \"sub-module\"  being a mode 16000 \ncommit (see https://longair.net/blog/2010/06/02/git-submodules-explained/).\n\nOne thought is that one uses a proper sub-module that within it then has \nthe single 'large' file git-lfs style that hosts the hash reference for \nthe data VCS \n(https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It would be \nthe regular sub-modules .gitattributes file that handles the data \nconversion.\n\nIt may be converting an X-Y problem into an X-Y-Z solution, or just \nextending the problem.\n\n--\nPhilip\n\n\n"},{"id":"456441","messageId":"CAE9CXuiTDjbncEzWJpHN5N0CukcmXbhxQJtzDDhuy0er4Se2DA@mail.gmail.com","threadId":"57868","inReplyTo":"271b6a9a-a5f4-0336-51b8-860ad07f2609@iee.email","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Addison Klinke","fromEmail":"addison@baller.tv","sentAt":"2022-06-01T12:44:33Z","receivedAt":"2022-06-01T12:44:52Z","isPatch":false,"sender":{"key":"addison@baller.tv","avatar":null},"body":"> rsbecker: move code into a submodule from your own VCS system\ninto a git repository and the work with the submodule without the git\ncode-base knowing about this\n\n> Philip: uses a proper sub-module that within it then has\nthe single 'large' file git-lfs style that hosts the hash reference for\nthe data VCS\n\nThe downside I see with both of these approaches is that translating\nthe native data VCS to git (or LFS) negates all the benefits of having\na VCS purpose-built for data. That's why the majority of data\nversioning tools exist - because git (or LFS) are not ideal for\nhandling machine learning datasets\n\nOn Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email> wrote:\n>\n> On 10/05/2022 18:20, Jason Pyeron wrote:\n> >> -----Original Message-----\n> >> From: Junio C Hamano\n> >> Sent: Tuesday, May 10, 2022 1:01 PM\n> >> To: Addison Klinke <addison@baller.tv>\n> >>\n> >> Addison Klinke <addison@baller.tv> writes:\n> >>\n> >>> Is something along these lines feasible?\n> >> Offhand, I only think of one thing that could make it fundamentally\n> >> infeasible.\n> >>\n> >> When you bind an external repository (be it stored in Git or\n> >> somebody else's system) as a submodule, each commit in the\n> >> superproject records which exact commit in the submodule is used\n> >> with the rest of the superproject tree.  And that is done by\n> >> recording the object name of the commit in the submodule.\n> >>\n> >> What it means for the foreign system that wants to \"plug into\" a\n> >> superproject in Git as a submodule?  It is required to do two\n> >> things:\n> >>\n> >>   * At the time \"git commit\" is run at the superproject level, the\n> >>     foreign system has to be able to say \"the version I have to be\n> >>     used in the context of this superproject commit is X\", with X\n> >>     that somehow can be stored in the superproject's tree object\n> >>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n> >>     repositories, it is a bit wider).\n> >>\n> >>   * At the time \"git chekcout\" is run at the superproject level, the\n> >>     superproject will learn the above X (i.e. the version of the\n> >>     submodule that goes with the version of the superproject being\n> >>     checked out).  The foreign system has to be able to perform a\n> >>     \"checkout\" given that X.\n> >>\n> >> If a foreign system cannot do the above two, then it fundamentally\n> >> would be incapable of participating in such a \"superproject and\n> >> submodule\" relationship.\n>\n> The sub-modules already have that problem if the user forgets publish\n> their sub-module (see notes in the docs ;-).\n> > The submodule \"type\" could create an object (hashed and stored) that contains the needed \"translation\" details. The object would be hashed using SHA1 or SHA256 depending on the git config. The format of the object's contents would be defined by the submodule's \"code\".\n> >\n> Another way of looking at the issue is via a variant of Git-LFS with a\n> smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n>\n> The LFS already uses the .gitattributes to define a 'type', while the\n> submodules don't yet have that capability. There is just a single\n> special type within a tree object of \"sub-module\"  being a mode 16000\n> commit (see https://longair.net/blog/2010/06/02/git-submodules-explained/).\n>\n> One thought is that one uses a proper sub-module that within it then has\n> the single 'large' file git-lfs style that hosts the hash reference for\n> the data VCS\n> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It would be\n> the regular sub-modules .gitattributes file that handles the data\n> conversion.\n>\n> It may be converting an X-Y problem into an X-Y-Z solution, or just\n> extending the problem.\n>\n> --\n> Philip\n>\n>\n"},{"id":"456645","messageId":"547b245d-bdb2-5833-fe4d-15222ae32b57@iee.email","threadId":"57868","inReplyTo":"CAE9CXuiTDjbncEzWJpHN5N0CukcmXbhxQJtzDDhuy0er4Se2DA@mail.gmail.com","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-06-03T23:06:36Z","receivedAt":"2022-06-03T23:06:43Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 01/06/2022 13:44, Addison Klinke wrote:\n>> rsbecker: move code into a submodule from your own VCS system\n> into a git repository and the work with the submodule without the git\n> code-base knowing about this\n>\n>> Philip: uses a proper sub-module that within it then has\n> the single 'large' file git-lfs style that hosts the hash reference for\n> the data VCS\n>\n> The downside I see with both of these approaches is that translating\n> the native data VCS to git (or LFS) negates all the benefits of having\n> a VCS purpose-built for data. That's why the majority of data\n> versioning tools exist - because git (or LFS) are not ideal for\n> handling machine learning datasets\n\nThe key aspect is deciding which of the two storage systems (the Data &\nthe Code) will be the overall lead system that contains the linked\nreference to the other storage system to ensure the needed integrity.\nThat is not really a technical question. Rather its somewhat of a social\ndiscussion (workflows, trust, style of integration, etc).\n\nIt maybe that one of the systems does have less long-term integrity, as\nhas been seen in many versioning systems over the last century (both\nmanual and computer), but the UI is also important.\n\nIIRC Junio did note that having a suitable API to access the other\nstorage system (to know its status, etc.) is likely to be core to the\nability to combine the two. It may  be that a top level 'gui' is used\ncontrol both systems and ensure synchronisation to hide the complexities\nof both systems.\n\nI'm still thinking that the \"git-lfs like\" style could be the one to\nuse, but that is very dependant on the API that is available for\ncapturing the Data state into the git entry that records that state,\nwhether that is a file (git-lfs like) or a 'sub-module' (directory as\nstate ) style.  Either way it still need reifying (i.e. coded to make\nthe abstract concept into a concrete implementation).\n\nWhich ever route is chosen, it still sounds to me like a worthwhile\nenterprise. It's all still very abstract.\n>\n> On Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email> wrote:\n>> On 10/05/2022 18:20, Jason Pyeron wrote:\n>>>> -----Original Message-----\n>>>> From: Junio C Hamano\n>>>> Sent: Tuesday, May 10, 2022 1:01 PM\n>>>> To: Addison Klinke <addison@baller.tv>\n>>>>\n>>>> Addison Klinke <addison@baller.tv> writes:\n>>>>\n>>>>> Is something along these lines feasible?\n>>>> Offhand, I only think of one thing that could make it fundamentally\n>>>> infeasible.\n>>>>\n>>>> When you bind an external repository (be it stored in Git or\n>>>> somebody else's system) as a submodule, each commit in the\n>>>> superproject records which exact commit in the submodule is used\n>>>> with the rest of the superproject tree.  And that is done by\n>>>> recording the object name of the commit in the submodule.\n>>>>\n>>>> What it means for the foreign system that wants to \"plug into\" a\n>>>> superproject in Git as a submodule?  It is required to do two\n>>>> things:\n>>>>\n>>>>   * At the time \"git commit\" is run at the superproject level, the\n>>>>     foreign system has to be able to say \"the version I have to be\n>>>>     used in the context of this superproject commit is X\", with X\n>>>>     that somehow can be stored in the superproject's tree object\n>>>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>>>>     repositories, it is a bit wider).\n>>>>\n>>>>   * At the time \"git chekcout\" is run at the superproject level, the\n>>>>     superproject will learn the above X (i.e. the version of the\n>>>>     submodule that goes with the version of the superproject being\n>>>>     checked out).  The foreign system has to be able to perform a\n>>>>     \"checkout\" given that X.\n>>>>\n>>>> If a foreign system cannot do the above two, then it fundamentally\n>>>> would be incapable of participating in such a \"superproject and\n>>>> submodule\" relationship.\n>> The sub-modules already have that problem if the user forgets publish\n>> their sub-module (see notes in the docs ;-).\n>>> The submodule \"type\" could create an object (hashed and stored) that contains the needed \"translation\" details. The object would be hashed using SHA1 or SHA256 depending on the git config. The format of the object's contents would be defined by the submodule's \"code\".\n>>>\n>> Another way of looking at the issue is via a variant of Git-LFS with a\n>> smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n>>\n>> The LFS already uses the .gitattributes to define a 'type', while the\n>> submodules don't yet have that capability. There is just a single\n>> special type within a tree object of \"sub-module\"  being a mode 16000\n>> commit (see https://longair.net/blog/2010/06/02/git-submodules-explained/).\n>>\n>> One thought is that one uses a proper sub-module that within it then has\n>> the single 'large' file git-lfs style that hosts the hash reference for\n>> the data VCS\n>> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It would be\n>> the regular sub-modules .gitattributes file that handles the data\n>> conversion.\n>>\n>> It may be converting an X-Y problem into an X-Y-Z solution, or just\n>> extending the problem.\n>>\n>> --\n>> Philip\n>>\n>>\n\n"},{"id":"456652","messageId":"000301d877b7$0fb1ca20$2f155e60$@nexbridge.com","threadId":"57868","inReplyTo":"547b245d-bdb2-5833-fe4d-15222ae32b57@iee.email","subject":"RE: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-06-04T02:01:47Z","receivedAt":"2022-06-04T02:02:09Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On June 3, 2022 7:07 PM, Philip Oakley wrote:\n>On 01/06/2022 13:44, Addison Klinke wrote:\n>>> rsbecker: move code into a submodule from your own VCS system\n>> into a git repository and the work with the submodule without the git\n>> code-base knowing about this\n>>\n>>> Philip: uses a proper sub-module that within it then has\n>> the single 'large' file git-lfs style that hosts the hash reference\n>> for the data VCS\n>>\n>> The downside I see with both of these approaches is that translating\n>> the native data VCS to git (or LFS) negates all the benefits of having\n>> a VCS purpose-built for data. That's why the majority of data\n>> versioning tools exist - because git (or LFS) are not ideal for\n>> handling machine learning datasets\n>\n>The key aspect is deciding which of the two storage systems (the Data & the Code)\n>will be the overall lead system that contains the linked reference to the other\n>storage system to ensure the needed integrity.\n>That is not really a technical question. Rather its somewhat of a social discussion\n>(workflows, trust, style of integration, etc).\n>\n>It maybe that one of the systems does have less long-term integrity, as has been\n>seen in many versioning systems over the last century (both manual and\n>computer), but the UI is also important.\n>\n>IIRC Junio did note that having a suitable API to access the other storage system\n>(to know its status, etc.) is likely to be core to the ability to combine the two. It\n>may  be that a top level 'gui' is used control both systems and ensure\n>synchronisation to hide the complexities of both systems.\n>\n>I'm still thinking that the \"git-lfs like\" style could be the one to use, but that is very\n>dependant on the API that is available for capturing the Data state into the git\n>entry that records that state, whether that is a file (git-lfs like) or a 'sub-module'\n>(directory as state ) style.  Either way it still need reifying (i.e. coded to make the\n>abstract concept into a concrete implementation).\n>\n>Which ever route is chosen, it still sounds to me like a worthwhile enterprise. It's\n>all still very abstract.\n>>\n>> On Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email> wrote:\n>>> On 10/05/2022 18:20, Jason Pyeron wrote:\n>>>>> -----Original Message-----\n>>>>> From: Junio C Hamano\n>>>>> Sent: Tuesday, May 10, 2022 1:01 PM\n>>>>> To: Addison Klinke <addison@baller.tv>\n>>>>>\n>>>>> Addison Klinke <addison@baller.tv> writes:\n>>>>>\n>>>>>> Is something along these lines feasible?\n>>>>> Offhand, I only think of one thing that could make it fundamentally\n>>>>> infeasible.\n>>>>>\n>>>>> When you bind an external repository (be it stored in Git or\n>>>>> somebody else's system) as a submodule, each commit in the\n>>>>> superproject records which exact commit in the submodule is used\n>>>>> with the rest of the superproject tree.  And that is done by\n>>>>> recording the object name of the commit in the submodule.\n>>>>>\n>>>>> What it means for the foreign system that wants to \"plug into\" a\n>>>>> superproject in Git as a submodule?  It is required to do two\n>>>>> things:\n>>>>>\n>>>>>   * At the time \"git commit\" is run at the superproject level, the\n>>>>>     foreign system has to be able to say \"the version I have to be\n>>>>>     used in the context of this superproject commit is X\", with X\n>>>>>     that somehow can be stored in the superproject's tree object\n>>>>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>>>>>     repositories, it is a bit wider).\n>>>>>\n>>>>>   * At the time \"git chekcout\" is run at the superproject level, the\n>>>>>     superproject will learn the above X (i.e. the version of the\n>>>>>     submodule that goes with the version of the superproject being\n>>>>>     checked out).  The foreign system has to be able to perform a\n>>>>>     \"checkout\" given that X.\n>>>>>\n>>>>> If a foreign system cannot do the above two, then it fundamentally\n>>>>> would be incapable of participating in such a \"superproject and\n>>>>> submodule\" relationship.\n>>> The sub-modules already have that problem if the user forgets publish\n>>> their sub-module (see notes in the docs ;-).\n>>>> The submodule \"type\" could create an object (hashed and stored) that\n>contains the needed \"translation\" details. The object would be hashed using SHA1\n>or SHA256 depending on the git config. The format of the object's contents would\n>be defined by the submodule's \"code\".\n>>>>\n>>> Another way of looking at the issue is via a variant of Git-LFS with\n>>> a smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n>>>\n>>> The LFS already uses the .gitattributes to define a 'type', while the\n>>> submodules don't yet have that capability. There is just a single\n>>> special type within a tree object of \"sub-module\"  being a mode 16000\n>>> commit (see https://longair.net/blog/2010/06/02/git-submodules-explained/).\n>>>\n>>> One thought is that one uses a proper sub-module that within it then\n>>> has the single 'large' file git-lfs style that hosts the hash\n>>> reference for the data VCS\n>>> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It would\n>>> be the regular sub-modules .gitattributes file that handles the data\n>>> conversion.\n>>>\n>>> It may be converting an X-Y problem into an X-Y-Z solution, or just\n>>> extending the problem.\n\nThe most salient issue I have with this is that signatures cannot be validated across VCS systems. Within git, a submodule commit can be signed. This ensures that the contents of the commit in the super-project can also be signed. If someone hacks an underlying VCS that is not git, either:\n\na) git can never sign a commit from an underlying VCS, or\n\nb) git can never trust a commit from an underlying VCS.\n\nThis pollutes a fundamental capability of git, being multiple signers the contents of a commit, and invalidates the integrity of the Merkel tree that underlies git contents.\n\nI do not see that this concept contributes positively to the ecosystem. I do feel strongly about this and hope my points are understood.\n\nSincerely,\nRandall\n\n"},{"id":"456680","messageId":"b767ee9f-5c93-0c82-f551-7c1673adcc62@iee.email","threadId":"57868","inReplyTo":"000301d877b7$0fb1ca20$2f155e60$@nexbridge.com","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-06-04T13:27:54Z","receivedAt":"2022-06-04T13:28:06Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Randall,\n\nOn 04/06/2022 03:01, rsbecker@nexbridge.com wrote:\n> On June 3, 2022 7:07 PM, Philip Oakley wrote:\n>> On 01/06/2022 13:44, Addison Klinke wrote:\n>>>> rsbecker: move code into a submodule from your own VCS system\n>>> into a git repository and the work with the submodule without the git\n>>> code-base knowing about this\n>>>\n>>>> Philip: uses a proper sub-module that within it then has\n>>> the single 'large' file git-lfs style that hosts the hash reference\n>>> for the data VCS\n>>>\n>>> The downside I see with both of these approaches is that translating\n>>> the native data VCS to git (or LFS) negates all the benefits of having\n>>> a VCS purpose-built for data. That's why the majority of data\n>>> versioning tools exist - because git (or LFS) are not ideal for\n>>> handling machine learning datasets\n>> The key aspect is deciding which of the two storage systems (the Data & the Code)\n>> will be the overall lead system that contains the linked reference to the other\n>> storage system to ensure the needed integrity.\n>> That is not really a technical question. Rather its somewhat of a social discussion\n>> (workflows, trust, style of integration, etc).\n>>\n>> It maybe that one of the systems does have less long-term integrity, as has been\n>> seen in many versioning systems over the last century (both manual and\n>> computer), but the UI is also important.\n>>\n>> IIRC Junio did note that having a suitable API to access the other storage system\n>> (to know its status, etc.) is likely to be core to the ability to combine the two. It\n>> may  be that a top level 'gui' is used control both systems and ensure\n>> synchronisation to hide the complexities of both systems.\n>>\n>> I'm still thinking that the \"git-lfs like\" style could be the one to use, but that is very\n>> dependant on the API that is available for capturing the Data state into the git\n>> entry that records that state, whether that is a file (git-lfs like) or a 'sub-module'\n>> (directory as state ) style.  Either way it still need reifying (i.e. coded to make the\n>> abstract concept into a concrete implementation).\n>>\n>> Which ever route is chosen, it still sounds to me like a worthwhile enterprise. It's\n>> all still very abstract.\n>>> On Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email> wrote:\n>>>> On 10/05/2022 18:20, Jason Pyeron wrote:\n>>>>>> -----Original Message-----\n>>>>>> From: Junio C Hamano\n>>>>>> Sent: Tuesday, May 10, 2022 1:01 PM\n>>>>>> To: Addison Klinke <addison@baller.tv>\n>>>>>>\n>>>>>> Addison Klinke <addison@baller.tv> writes:\n>>>>>>\n>>>>>>> Is something along these lines feasible?\n>>>>>> Offhand, I only think of one thing that could make it fundamentally\n>>>>>> infeasible.\n>>>>>>\n>>>>>> When you bind an external repository (be it stored in Git or\n>>>>>> somebody else's system) as a submodule, each commit in the\n>>>>>> superproject records which exact commit in the submodule is used\n>>>>>> with the rest of the superproject tree.  And that is done by\n>>>>>> recording the object name of the commit in the submodule.\n>>>>>>\n>>>>>> What it means for the foreign system that wants to \"plug into\" a\n>>>>>> superproject in Git as a submodule?  It is required to do two\n>>>>>> things:\n>>>>>>\n>>>>>>   * At the time \"git commit\" is run at the superproject level, the\n>>>>>>     foreign system has to be able to say \"the version I have to be\n>>>>>>     used in the context of this superproject commit is X\", with X\n>>>>>>     that somehow can be stored in the superproject's tree object\n>>>>>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>>>>>>     repositories, it is a bit wider).\n>>>>>>\n>>>>>>   * At the time \"git chekcout\" is run at the superproject level, the\n>>>>>>     superproject will learn the above X (i.e. the version of the\n>>>>>>     submodule that goes with the version of the superproject being\n>>>>>>     checked out).  The foreign system has to be able to perform a\n>>>>>>     \"checkout\" given that X.\n>>>>>>\n>>>>>> If a foreign system cannot do the above two, then it fundamentally\n>>>>>> would be incapable of participating in such a \"superproject and\n>>>>>> submodule\" relationship.\n>>>> The sub-modules already have that problem if the user forgets publish\n>>>> their sub-module (see notes in the docs ;-).\n>>>>> The submodule \"type\" could create an object (hashed and stored) that\n>> contains the needed \"translation\" details. The object would be hashed using SHA1\n>> or SHA256 depending on the git config. The format of the object's contents would\n>> be defined by the submodule's \"code\".\n>>>> Another way of looking at the issue is via a variant of Git-LFS with\n>>>> a smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n>>>>\n>>>> The LFS already uses the .gitattributes to define a 'type', while the\n>>>> submodules don't yet have that capability. There is just a single\n>>>> special type within a tree object of \"sub-module\"  being a mode 16000\n>>>> commit (see https://longair.net/blog/2010/06/02/git-submodules-explained/).\n>>>>\n>>>> One thought is that one uses a proper sub-module that within it then\n>>>> has the single 'large' file git-lfs style that hosts the hash\n>>>> reference for the data VCS\n>>>> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It would\n>>>> be the regular sub-modules .gitattributes file that handles the data\n>>>> conversion.\n>>>>\n>>>> It may be converting an X-Y problem into an X-Y-Z solution, or just\n>>>> extending the problem.\n> The most salient issue I have with this is that signatures cannot be validated across VCS systems. \n\nI think I disagree, but let's be sure we are talking about the same\n'signature' aspect, I think there are (at least) three different\nsignatures we could be talking about\n\n1. The hash verification 'signature' that can cascade down the trees. We\nverify against a given hash.\n2. The 'Signed-off-by:' legal/copyright signature - important, but I\ndon't think that's the one being discussed.\n3. The (e.g.) PGP signature of a tag or commit. This provides a (web of)\ntrust mechanism for the _given_ hash in 1. Important in 'open systems',\nless so in more closed systems where trust, and the _given_, is via side\nchannels.\n\nNote the shift from using a hash to using the PGP for the 'signature'.\n\n\n> Within git, a submodule commit can be signed. This ensures that the contents of the commit in the super-project can also be signed. If someone hacks an underlying VCS that is not git, either:\nSubmodules are a remote VCS, it just happens to have the same hash\nvalidation software as the super-project, which is nice.\n>\n> a) git can never sign a commit from an underlying VCS, or\nGit-LFS is a similar hand off, though many accept it's capability.\n>\n> b) git can never trust a commit from an underlying VCS.\n>\n> This pollutes a fundamental capability of git, being multiple signers the contents of a commit, and invalidates the integrity of the Merkel tree that underlies git contents.\n\nThe main issue is how to confirm the integrity the other VCS. Many of\nthe Data VCS systems are based on Git and it's hash integrity approach,\nso as long as the DATA VCS has similar integrity guarantees, we maintain\nthe level of trust in the security of the whole system.\n\n>\n> I do not see that this concept contributes positively to the ecosystem. I do feel strongly about this and hope my points are understood.\n\nI'd agree that there is a need to work out how to integrate the code VCS\nand data VCS in a consistent way. Ignoring the Data VCS problem doesn't\nmake it go away.\n\nMaybe if Addison was able to identify one or two lead contenders as the\nData VCS and how it/they offer their levels of security and integrity,\nthen it would be easier to see where in the Git model that may fit. Or\nwhether Git is the underling VCS (because it has programmable API), and\nthe Data VCS (esp because of scale and non-distributed nature) becomes\nthe \"authority\", even if that has less capability!\n>\n> Sincerely,\n> Randall\n>\nPhilip\n"},{"id":"456683","messageId":"003001d8782b$d207c100$76174300$@nexbridge.com","threadId":"57868","inReplyTo":"b767ee9f-5c93-0c82-f551-7c1673adcc62@iee.email","subject":"RE: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-06-04T15:57:35Z","receivedAt":"2022-06-04T15:57:54Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On June 4, 2022 9:28 AM, Philip Oakley wrote:\n>On 04/06/2022 03:01, rsbecker@nexbridge.com wrote:\n>> On June 3, 2022 7:07 PM, Philip Oakley wrote:\n>>> On 01/06/2022 13:44, Addison Klinke wrote:\n>>>>> rsbecker: move code into a submodule from your own VCS system\n>>>> into a git repository and the work with the submodule without the\n>>>> git code-base knowing about this\n>>>>\n>>>>> Philip: uses a proper sub-module that within it then has\n>>>> the single 'large' file git-lfs style that hosts the hash reference\n>>>> for the data VCS\n>>>>\n>>>> The downside I see with both of these approaches is that translating\n>>>> the native data VCS to git (or LFS) negates all the benefits of\n>>>> having a VCS purpose-built for data. That's why the majority of data\n>>>> versioning tools exist - because git (or LFS) are not ideal for\n>>>> handling machine learning datasets\n>>> The key aspect is deciding which of the two storage systems (the Data\n>>> & the Code) will be the overall lead system that contains the linked\n>>> reference to the other storage system to ensure the needed integrity.\n>>> That is not really a technical question. Rather its somewhat of a\n>>> social discussion (workflows, trust, style of integration, etc).\n>>>\n>>> It maybe that one of the systems does have less long-term integrity,\n>>> as has been seen in many versioning systems over the last century\n>>> (both manual and computer), but the UI is also important.\n>>>\n>>> IIRC Junio did note that having a suitable API to access the other\n>>> storage system (to know its status, etc.) is likely to be core to the\n>>> ability to combine the two. It may  be that a top level 'gui' is used\n>>> control both systems and ensure synchronisation to hide the complexities of\n>both systems.\n>>>\n>>> I'm still thinking that the \"git-lfs like\" style could be the one to\n>>> use, but that is very dependant on the API that is available for\n>>> capturing the Data state into the git entry that records that state, whether that\n>is a file (git-lfs like) or a 'sub-module'\n>>> (directory as state ) style.  Either way it still need reifying (i.e.\n>>> coded to make the abstract concept into a concrete implementation).\n>>>\n>>> Which ever route is chosen, it still sounds to me like a worthwhile\n>>> enterprise. It's all still very abstract.\n>>>> On Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email>\n>wrote:\n>>>>> On 10/05/2022 18:20, Jason Pyeron wrote:\n>>>>>>> -----Original Message-----\n>>>>>>> From: Junio C Hamano\n>>>>>>> Sent: Tuesday, May 10, 2022 1:01 PM\n>>>>>>> To: Addison Klinke <addison@baller.tv>\n>>>>>>>\n>>>>>>> Addison Klinke <addison@baller.tv> writes:\n>>>>>>>\n>>>>>>>> Is something along these lines feasible?\n>>>>>>> Offhand, I only think of one thing that could make it\n>>>>>>> fundamentally infeasible.\n>>>>>>>\n>>>>>>> When you bind an external repository (be it stored in Git or\n>>>>>>> somebody else's system) as a submodule, each commit in the\n>>>>>>> superproject records which exact commit in the submodule is used\n>>>>>>> with the rest of the superproject tree.  And that is done by\n>>>>>>> recording the object name of the commit in the submodule.\n>>>>>>>\n>>>>>>> What it means for the foreign system that wants to \"plug into\" a\n>>>>>>> superproject in Git as a submodule?  It is required to do two\n>>>>>>> things:\n>>>>>>>\n>>>>>>>   * At the time \"git commit\" is run at the superproject level, the\n>>>>>>>     foreign system has to be able to say \"the version I have to be\n>>>>>>>     used in the context of this superproject commit is X\", with X\n>>>>>>>     that somehow can be stored in the superproject's tree object\n>>>>>>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>>>>>>>     repositories, it is a bit wider).\n>>>>>>>\n>>>>>>>   * At the time \"git chekcout\" is run at the superproject level, the\n>>>>>>>     superproject will learn the above X (i.e. the version of the\n>>>>>>>     submodule that goes with the version of the superproject being\n>>>>>>>     checked out).  The foreign system has to be able to perform a\n>>>>>>>     \"checkout\" given that X.\n>>>>>>>\n>>>>>>> If a foreign system cannot do the above two, then it\n>>>>>>> fundamentally would be incapable of participating in such a\n>>>>>>> \"superproject and submodule\" relationship.\n>>>>> The sub-modules already have that problem if the user forgets\n>>>>> publish their sub-module (see notes in the docs ;-).\n>>>>>> The submodule \"type\" could create an object (hashed and stored)\n>>>>>> that\n>>> contains the needed \"translation\" details. The object would be hashed\n>>> using SHA1 or SHA256 depending on the git config. The format of the\n>>> object's contents would be defined by the submodule's \"code\".\n>>>>> Another way of looking at the issue is via a variant of Git-LFS\n>>>>> with a smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n>>>>>\n>>>>> The LFS already uses the .gitattributes to define a 'type', while\n>>>>> the submodules don't yet have that capability. There is just a\n>>>>> single special type within a tree object of \"sub-module\"  being a\n>>>>> mode 16000 commit (see https://longair.net/blog/2010/06/02/git-\n>submodules-explained/).\n>>>>>\n>>>>> One thought is that one uses a proper sub-module that within it\n>>>>> then has the single 'large' file git-lfs style that hosts the hash\n>>>>> reference for the data VCS\n>>>>> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It\n>>>>> would be the regular sub-modules .gitattributes file that handles\n>>>>> the data conversion.\n>>>>>\n>>>>> It may be converting an X-Y problem into an X-Y-Z solution, or just\n>>>>> extending the problem.\n>> The most salient issue I have with this is that signatures cannot be validated\n>across VCS systems.\n>\n>I think I disagree, but let's be sure we are talking about the same 'signature'\n>aspect, I think there are (at least) three different signatures we could be talking\n>about\n>\n>1. The hash verification 'signature' that can cascade down the trees. We verify\n>against a given hash.\n>2. The 'Signed-off-by:' legal/copyright signature - important, but I don't think that's\n>the one being discussed.\n>3. The (e.g.) PGP signature of a tag or commit. This provides a (web of) trust\n>mechanism for the _given_ hash in 1. Important in 'open systems', less so in more\n>closed systems where trust, and the _given_, is via side channels.\n\nThe third is more my concern. I do not know of other (D)VCS systems that have the same level of trust allowed in git - simultaneously PGP/SSH signing commits and potentially multiple tags.\n\n>Note the shift from using a hash to using the PGP for the 'signature'.\n>\n>\n>> Within git, a submodule commit can be signed. This ensures that the contents of\n>the commit in the super-project can also be signed. If someone hacks an\n>underlying VCS that is not git, either:\n>Submodules are a remote VCS, it just happens to have the same hash validation\n>software as the super-project, which is nice.\n>>\n>> a) git can never sign a commit from an underlying VCS, or\n>Git-LFS is a similar hand off, though many accept it's capability.\n>>\n>> b) git can never trust a commit from an underlying VCS.\n>>\n>> This pollutes a fundamental capability of git, being multiple signers the contents\n>of a commit, and invalidates the integrity of the Merkel tree that underlies git\n>contents.\n>\n>The main issue is how to confirm the integrity the other VCS. Many of the Data\n>VCS systems are based on Git and it's hash integrity approach, so as long as the\n>DATA VCS has similar integrity guarantees, we maintain the level of trust in the\n>security of the whole system.\n\nThis is exactly my concern and what I was trying to point out - although more briefly. I do not think (an|there are) underlying VCS can provide similar guarantees. It is all too easy to hack most VCS systems if you have an appropriate user id especially most non-distributed ones. We originally moved to git because we had hacks on two different VCS systems underlying files.\n\n>> I do not see that this concept contributes positively to the ecosystem. I do feel\n>strongly about this and hope my points are understood.\n>\n>I'd agree that there is a need to work out how to integrate the code VCS and data\n>VCS in a consistent way. Ignoring the Data VCS problem doesn't make it go away.\n>\n>Maybe if Addison was able to identify one or two lead contenders as the Data VCS\n>and how it/they offer their levels of security and integrity, then it would be easier\n>to see where in the Git model that may fit. Or whether Git is the underling VCS\n>(because it has programmable API), and the Data VCS (esp because of scale and\n>non-distributed nature) becomes the \"authority\", even if that has less capability!\n\nI agree as well. I want to see assurances that this level of integrity can be maintained - or that the user will have to accept the risks that git signatures are no longer usable. It might be appropriate to disable commit.gpgsign if the underlying VCS cannot be an authority.\n\n--Randall\n\n"},{"id":"456701","messageId":"cf143d14-2265-be7d-d7a9-a4b11ff0f6af@iee.email","threadId":"57868","inReplyTo":"003001d8782b$d207c100$76174300$@nexbridge.com","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-06-05T21:52:21Z","receivedAt":"2022-06-05T21:52:32Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 04/06/2022 16:57, rsbecker@nexbridge.com wrote:\n> On June 4, 2022 9:28 AM, Philip Oakley wrote:\n>> On 04/06/2022 03:01, rsbecker@nexbridge.com wrote:\n>>> On June 3, 2022 7:07 PM, Philip Oakley wrote:\n>>>> On 01/06/2022 13:44, Addison Klinke wrote:\n>>>>>> rsbecker: move code into a submodule from your own VCS system\n>>>>> into a git repository and the work with the submodule without the\n>>>>> git code-base knowing about this\n>>>>>\n>>>>>> Philip: uses a proper sub-module that within it then has\n>>>>> the single 'large' file git-lfs style that hosts the hash reference\n>>>>> for the data VCS\n>>>>>\n>>>>> The downside I see with both of these approaches is that translating\n>>>>> the native data VCS to git (or LFS) negates all the benefits of\n>>>>> having a VCS purpose-built for data. That's why the majority of data\n>>>>> versioning tools exist - because git (or LFS) are not ideal for\n>>>>> handling machine learning datasets\n>>>> The key aspect is deciding which of the two storage systems (the Data\n>>>> & the Code) will be the overall lead system that contains the linked\n>>>> reference to the other storage system to ensure the needed integrity.\n>>>> That is not really a technical question. Rather its somewhat of a\n>>>> social discussion (workflows, trust, style of integration, etc).\n>>>>\n>>>> It maybe that one of the systems does have less long-term integrity,\n>>>> as has been seen in many versioning systems over the last century\n>>>> (both manual and computer), but the UI is also important.\n>>>>\n>>>> IIRC Junio did note that having a suitable API to access the other\n>>>> storage system (to know its status, etc.) is likely to be core to the\n>>>> ability to combine the two. It may  be that a top level 'gui' is used\n>>>> control both systems and ensure synchronisation to hide the complexities of\n>> both systems.\n>>>> I'm still thinking that the \"git-lfs like\" style could be the one to\n>>>> use, but that is very dependant on the API that is available for\n>>>> capturing the Data state into the git entry that records that state, whether that\n>> is a file (git-lfs like) or a 'sub-module'\n>>>> (directory as state ) style.  Either way it still need reifying (i.e.\n>>>> coded to make the abstract concept into a concrete implementation).\n>>>>\n>>>> Which ever route is chosen, it still sounds to me like a worthwhile\n>>>> enterprise. It's all still very abstract.\n>>>>> On Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email>\n>> wrote:\n>>>>>> On 10/05/2022 18:20, Jason Pyeron wrote:\n>>>>>>>> -----Original Message-----\n>>>>>>>> From: Junio C Hamano\n>>>>>>>> Sent: Tuesday, May 10, 2022 1:01 PM\n>>>>>>>> To: Addison Klinke <addison@baller.tv>\n>>>>>>>>\n>>>>>>>> Addison Klinke <addison@baller.tv> writes:\n>>>>>>>>\n>>>>>>>>> Is something along these lines feasible?\n>>>>>>>> Offhand, I only think of one thing that could make it\n>>>>>>>> fundamentally infeasible.\n>>>>>>>>\n>>>>>>>> When you bind an external repository (be it stored in Git or\n>>>>>>>> somebody else's system) as a submodule, each commit in the\n>>>>>>>> superproject records which exact commit in the submodule is used\n>>>>>>>> with the rest of the superproject tree.  And that is done by\n>>>>>>>> recording the object name of the commit in the submodule.\n>>>>>>>>\n>>>>>>>> What it means for the foreign system that wants to \"plug into\" a\n>>>>>>>> superproject in Git as a submodule?  It is required to do two\n>>>>>>>> things:\n>>>>>>>>\n>>>>>>>>   * At the time \"git commit\" is run at the superproject level, the\n>>>>>>>>     foreign system has to be able to say \"the version I have to be\n>>>>>>>>     used in the context of this superproject commit is X\", with X\n>>>>>>>>     that somehow can be stored in the superproject's tree object\n>>>>>>>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n>>>>>>>>     repositories, it is a bit wider).\n>>>>>>>>\n>>>>>>>>   * At the time \"git chekcout\" is run at the superproject level, the\n>>>>>>>>     superproject will learn the above X (i.e. the version of the\n>>>>>>>>     submodule that goes with the version of the superproject being\n>>>>>>>>     checked out).  The foreign system has to be able to perform a\n>>>>>>>>     \"checkout\" given that X.\n>>>>>>>>\n>>>>>>>> If a foreign system cannot do the above two, then it\n>>>>>>>> fundamentally would be incapable of participating in such a\n>>>>>>>> \"superproject and submodule\" relationship.\n>>>>>> The sub-modules already have that problem if the user forgets\n>>>>>> publish their sub-module (see notes in the docs ;-).\n>>>>>>> The submodule \"type\" could create an object (hashed and stored)\n>>>>>>> that\n>>>> contains the needed \"translation\" details. The object would be hashed\n>>>> using SHA1 or SHA256 depending on the git config. The format of the\n>>>> object's contents would be defined by the submodule's \"code\".\n>>>>>> Another way of looking at the issue is via a variant of Git-LFS\n>>>>>> with a smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n>>>>>>\n>>>>>> The LFS already uses the .gitattributes to define a 'type', while\n>>>>>> the submodules don't yet have that capability. There is just a\n>>>>>> single special type within a tree object of \"sub-module\"  being a\n>>>>>> mode 16000 commit (see https://longair.net/blog/2010/06/02/git-\n>> submodules-explained/).\n>>>>>> One thought is that one uses a proper sub-module that within it\n>>>>>> then has the single 'large' file git-lfs style that hosts the hash\n>>>>>> reference for the data VCS\n>>>>>> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It\n>>>>>> would be the regular sub-modules .gitattributes file that handles\n>>>>>> the data conversion.\n>>>>>>\n>>>>>> It may be converting an X-Y problem into an X-Y-Z solution, or just\n>>>>>> extending the problem.\n>>> The most salient issue I have with this is that signatures cannot be validated\n>> across VCS systems.\n>>\n>> I think I disagree, but let's be sure we are talking about the same 'signature'\n>> aspect, I think there are (at least) three different signatures we could be talking\n>> about\n>>\n>> 1. The hash verification 'signature' that can cascade down the trees. We verify\n>> against a given hash.\n>> 2. The 'Signed-off-by:' legal/copyright signature - important, but I don't think that's\n>> the one being discussed.\n>> 3. The (e.g.) PGP signature of a tag or commit. This provides a (web of) trust\n>> mechanism for the _given_ hash in 1. Important in 'open systems', less so in more\n>> closed systems where trust, and the _given_, is via side channels.\n> The third is more my concern. I do not know of other (D)VCS systems that have the same level of trust allowed in git - simultaneously PGP/SSH signing commits and potentially multiple tags.\n\nfor reference of other readers, that's as discussed in\nhttps://git-scm.com/book/en/v2/Git-Tools-Signing-Your-Work esp. the\n'Signing Commits' and 'Everyone Must Sign' sections at the end of the Ch 7.4\n>> Note the shift from using a hash to using the PGP for the 'signature'.\n>>\n>>\n>>> Within git, a submodule commit can be signed. This ensures that the contents of\n>> the commit in the super-project can also be signed. If someone hacks an\n>> underlying VCS that is not git, either:\n>> Submodules are a remote VCS, it just happens to have the same hash validation\n>> software as the super-project, which is nice.\n>>> a) git can never sign a commit from an underlying VCS, or\n>> Git-LFS is a similar hand off, though many accept it's capability.\n>>> b) git can never trust a commit from an underlying VCS.\n>>>\n>>> This pollutes a fundamental capability of git, being multiple signers the contents\n>> of a commit, and invalidates the integrity of the Merkel tree that underlies git\n>> contents.\n>>\n>> The main issue is how to confirm the integrity the other VCS. Many of the Data\n>> VCS systems are based on Git and it's hash integrity approach, so as long as the\n>> DATA VCS has similar integrity guarantees, we maintain the level of trust in the\n>> security of the whole system.\n> This is exactly my concern and what I was trying to point out - although more briefly. I do not think (an|there are) underlying VCS can provide similar guarantees. It is all too easy to hack most VCS systems if you have an appropriate user id especially most non-distributed ones. We originally moved to git because we had hacks on two different VCS systems underlying files.\n>\n>>> I do not see that this concept contributes positively to the ecosystem. I do feel\n>> strongly about this and hope my points are understood.\n>>\n>> I'd agree that there is a need to work out how to integrate the code VCS and data\n>> VCS in a consistent way. Ignoring the Data VCS problem doesn't make it go away.\n>>\n>> Maybe if Addison was able to identify one or two lead contenders as the Data VCS\n>> and how it/they offer their levels of security and integrity,\n\nLooking back at Addison's original email, he did suggest:\n\n- [Dolt](https://www.dolthub.com/),\n- [LakeFS](https://lakefs.io/), and\n- [DVC](https://dvc.org/)\n\nas examples. They all imply git hash style validation of the individual\ndata commits, by not mention of [PGP] signing, though it may available\nfor some.\n\nI did see the Dolt issue [ Cryptographic signing of a changeset? #628\n](https://github.com/dolthub/dolt/issues/628), so it looks like it's on\ntheir radar, though it's likely they'll need similar discussions about\nhow to cross integrate with Git..\n\nHowever, we also need to note the shift to the cloud for these very\nlarge immobile data sets, where there maybe concerns as to the security\nand trustworthiness of the compute and storage platforms (cosmic rays,\nrandom glitches, hacks, etc).\n\nWe are no longer importing code to our local machine that we need to be\nsigned, rather we are exporting our code to their compute\ninfrastructure, so the verification has to happen 'over-there'. So the\nintegrity question is still very pertinent.\n\n>>  then it would be easier\n>> to see where in the Git model that may fit. Or whether Git is the underling VCS\n>> (because it has programmable API), and the Data VCS (esp because of scale and\n>> non-distributed nature) becomes the \"authority\", even if that has less capability!\n> I agree as well. I want to see assurances that this level of integrity can be maintained - or that the user will have to accept the risks that git signatures are no longer usable. It might be appropriate to disable commit.gpgsign if the underlying VCS cannot be an authority.\n>\n>\nI'd also worry, like yourself, about the cloud data sets, and how the\ndata selection subsets are captured (e.g. if multiple individuals have\nused their right to be forgotten to make the old selection no longer\naccessible, then how to validate?). Interesting times.\n--\nPhilip\n"},{"id":"456714","messageId":"CAE9CXugMtYmq4XW+NZeMVYM2-7is7EECBQADSzGUwzAwgzQvUA@mail.gmail.com","threadId":"57868","inReplyTo":"cf143d14-2265-be7d-d7a9-a4b11ff0f6af@iee.email","subject":"Re: [FR] supporting submodules with alternate version control systems (new contributor)","fromName":"Addison Klinke","fromEmail":"addison@baller.tv","sentAt":"2022-06-06T14:53:11Z","receivedAt":"2022-06-06T14:53:28Z","isPatch":false,"sender":{"key":"addison@baller.tv","avatar":null},"body":"> The key aspect is deciding which of the two storage systems (the Data &\nthe Code) will be the overall lead system that contains the linked\nreference to the other storage system\n\nI'd prefer git as the lead system since it's a standard everyone is\nalready used to. With so many variations on data VCS out there, I\nthink it would be difficult to find consensus\n\n> I do not know of other (D)VCS systems that have the same level of trust allowed in git - simultaneously PGP/SSH signing commits and potentially multiple tags\n\nI have not used signed commits/tags with git before since the majority\nof machine learning work in industry is on private repositories with\ninternal teams. The Dolt issue thread that Philip referenced seems\nquite interesting in this regard.\n\n> It might be appropriate to disable commit.gpgsign if the underlying VCS cannot be an authority\n\nWould it be reasonable to start working on submodule integrations and\ndesign a way for signing to be added later on as it (hopefully)\nbecomes supported by each data VCS?\n\nOn Sun, Jun 5, 2022 at 3:52 PM Philip Oakley <philipoakley@iee.email> wrote:\n>\n> On 04/06/2022 16:57, rsbecker@nexbridge.com wrote:\n> > On June 4, 2022 9:28 AM, Philip Oakley wrote:\n> >> On 04/06/2022 03:01, rsbecker@nexbridge.com wrote:\n> >>> On June 3, 2022 7:07 PM, Philip Oakley wrote:\n> >>>> On 01/06/2022 13:44, Addison Klinke wrote:\n> >>>>>> rsbecker: move code into a submodule from your own VCS system\n> >>>>> into a git repository and the work with the submodule without the\n> >>>>> git code-base knowing about this\n> >>>>>\n> >>>>>> Philip: uses a proper sub-module that within it then has\n> >>>>> the single 'large' file git-lfs style that hosts the hash reference\n> >>>>> for the data VCS\n> >>>>>\n> >>>>> The downside I see with both of these approaches is that translating\n> >>>>> the native data VCS to git (or LFS) negates all the benefits of\n> >>>>> having a VCS purpose-built for data. That's why the majority of data\n> >>>>> versioning tools exist - because git (or LFS) are not ideal for\n> >>>>> handling machine learning datasets\n> >>>> The key aspect is deciding which of the two storage systems (the Data\n> >>>> & the Code) will be the overall lead system that contains the linked\n> >>>> reference to the other storage system to ensure the needed integrity.\n> >>>> That is not really a technical question. Rather its somewhat of a\n> >>>> social discussion (workflows, trust, style of integration, etc).\n> >>>>\n> >>>> It maybe that one of the systems does have less long-term integrity,\n> >>>> as has been seen in many versioning systems over the last century\n> >>>> (both manual and computer), but the UI is also important.\n> >>>>\n> >>>> IIRC Junio did note that having a suitable API to access the other\n> >>>> storage system (to know its status, etc.) is likely to be core to the\n> >>>> ability to combine the two. It may  be that a top level 'gui' is used\n> >>>> control both systems and ensure synchronisation to hide the complexities of\n> >> both systems.\n> >>>> I'm still thinking that the \"git-lfs like\" style could be the one to\n> >>>> use, but that is very dependant on the API that is available for\n> >>>> capturing the Data state into the git entry that records that state, whether that\n> >> is a file (git-lfs like) or a 'sub-module'\n> >>>> (directory as state ) style.  Either way it still need reifying (i.e.\n> >>>> coded to make the abstract concept into a concrete implementation).\n> >>>>\n> >>>> Which ever route is chosen, it still sounds to me like a worthwhile\n> >>>> enterprise. It's all still very abstract.\n> >>>>> On Tue, May 10, 2022 at 2:54 PM Philip Oakley <philipoakley@iee.email>\n> >> wrote:\n> >>>>>> On 10/05/2022 18:20, Jason Pyeron wrote:\n> >>>>>>>> -----Original Message-----\n> >>>>>>>> From: Junio C Hamano\n> >>>>>>>> Sent: Tuesday, May 10, 2022 1:01 PM\n> >>>>>>>> To: Addison Klinke <addison@baller.tv>\n> >>>>>>>>\n> >>>>>>>> Addison Klinke <addison@baller.tv> writes:\n> >>>>>>>>\n> >>>>>>>>> Is something along these lines feasible?\n> >>>>>>>> Offhand, I only think of one thing that could make it\n> >>>>>>>> fundamentally infeasible.\n> >>>>>>>>\n> >>>>>>>> When you bind an external repository (be it stored in Git or\n> >>>>>>>> somebody else's system) as a submodule, each commit in the\n> >>>>>>>> superproject records which exact commit in the submodule is used\n> >>>>>>>> with the rest of the superproject tree.  And that is done by\n> >>>>>>>> recording the object name of the commit in the submodule.\n> >>>>>>>>\n> >>>>>>>> What it means for the foreign system that wants to \"plug into\" a\n> >>>>>>>> superproject in Git as a submodule?  It is required to do two\n> >>>>>>>> things:\n> >>>>>>>>\n> >>>>>>>>   * At the time \"git commit\" is run at the superproject level, the\n> >>>>>>>>     foreign system has to be able to say \"the version I have to be\n> >>>>>>>>     used in the context of this superproject commit is X\", with X\n> >>>>>>>>     that somehow can be stored in the superproject's tree object\n> >>>>>>>>     (which is sized 20-byte for SHA-1 repositories; in SHA-256\n> >>>>>>>>     repositories, it is a bit wider).\n> >>>>>>>>\n> >>>>>>>>   * At the time \"git chekcout\" is run at the superproject level, the\n> >>>>>>>>     superproject will learn the above X (i.e. the version of the\n> >>>>>>>>     submodule that goes with the version of the superproject being\n> >>>>>>>>     checked out).  The foreign system has to be able to perform a\n> >>>>>>>>     \"checkout\" given that X.\n> >>>>>>>>\n> >>>>>>>> If a foreign system cannot do the above two, then it\n> >>>>>>>> fundamentally would be incapable of participating in such a\n> >>>>>>>> \"superproject and submodule\" relationship.\n> >>>>>> The sub-modules already have that problem if the user forgets\n> >>>>>> publish their sub-module (see notes in the docs ;-).\n> >>>>>>> The submodule \"type\" could create an object (hashed and stored)\n> >>>>>>> that\n> >>>> contains the needed \"translation\" details. The object would be hashed\n> >>>> using SHA1 or SHA256 depending on the git config. The format of the\n> >>>> object's contents would be defined by the submodule's \"code\".\n> >>>>>> Another way of looking at the issue is via a variant of Git-LFS\n> >>>>>> with a smudge/clean style filter. I.e. the DataVCS would be treated as a 'file'.\n> >>>>>>\n> >>>>>> The LFS already uses the .gitattributes to define a 'type', while\n> >>>>>> the submodules don't yet have that capability. There is just a\n> >>>>>> single special type within a tree object of \"sub-module\"  being a\n> >>>>>> mode 16000 commit (see https://longair.net/blog/2010/06/02/git-\n> >> submodules-explained/).\n> >>>>>> One thought is that one uses a proper sub-module that within it\n> >>>>>> then has the single 'large' file git-lfs style that hosts the hash\n> >>>>>> reference for the data VCS\n> >>>>>> (https://github.com/git-lfs/git-lfs/blob/main/docs/spec.md). It\n> >>>>>> would be the regular sub-modules .gitattributes file that handles\n> >>>>>> the data conversion.\n> >>>>>>\n> >>>>>> It may be converting an X-Y problem into an X-Y-Z solution, or just\n> >>>>>> extending the problem.\n> >>> The most salient issue I have with this is that signatures cannot be validated\n> >> across VCS systems.\n> >>\n> >> I think I disagree, but let's be sure we are talking about the same 'signature'\n> >> aspect, I think there are (at least) three different signatures we could be talking\n> >> about\n> >>\n> >> 1. The hash verification 'signature' that can cascade down the trees. We verify\n> >> against a given hash.\n> >> 2. The 'Signed-off-by:' legal/copyright signature - important, but I don't think that's\n> >> the one being discussed.\n> >> 3. The (e.g.) PGP signature of a tag or commit. This provides a (web of) trust\n> >> mechanism for the _given_ hash in 1. Important in 'open systems', less so in more\n> >> closed systems where trust, and the _given_, is via side channels.\n> > The third is more my concern. I do not know of other (D)VCS systems that have the same level of trust allowed in git - simultaneously PGP/SSH signing commits and potentially multiple tags.\n>\n> for reference of other readers, that's as discussed in\n> https://git-scm.com/book/en/v2/Git-Tools-Signing-Your-Work esp. the\n> 'Signing Commits' and 'Everyone Must Sign' sections at the end of the Ch 7.4\n> >> Note the shift from using a hash to using the PGP for the 'signature'.\n> >>\n> >>\n> >>> Within git, a submodule commit can be signed. This ensures that the contents of\n> >> the commit in the super-project can also be signed. If someone hacks an\n> >> underlying VCS that is not git, either:\n> >> Submodules are a remote VCS, it just happens to have the same hash validation\n> >> software as the super-project, which is nice.\n> >>> a) git can never sign a commit from an underlying VCS, or\n> >> Git-LFS is a similar hand off, though many accept it's capability.\n> >>> b) git can never trust a commit from an underlying VCS.\n> >>>\n> >>> This pollutes a fundamental capability of git, being multiple signers the contents\n> >> of a commit, and invalidates the integrity of the Merkel tree that underlies git\n> >> contents.\n> >>\n> >> The main issue is how to confirm the integrity the other VCS. Many of the Data\n> >> VCS systems are based on Git and it's hash integrity approach, so as long as the\n> >> DATA VCS has similar integrity guarantees, we maintain the level of trust in the\n> >> security of the whole system.\n> > This is exactly my concern and what I was trying to point out - although more briefly. I do not think (an|there are) underlying VCS can provide similar guarantees. It is all too easy to hack most VCS systems if you have an appropriate user id especially most non-distributed ones. We originally moved to git because we had hacks on two different VCS systems underlying files.\n> >\n> >>> I do not see that this concept contributes positively to the ecosystem. I do feel\n> >> strongly about this and hope my points are understood.\n> >>\n> >> I'd agree that there is a need to work out how to integrate the code VCS and data\n> >> VCS in a consistent way. Ignoring the Data VCS problem doesn't make it go away.\n> >>\n> >> Maybe if Addison was able to identify one or two lead contenders as the Data VCS\n> >> and how it/they offer their levels of security and integrity,\n>\n> Looking back at Addison's original email, he did suggest:\n>\n> - [Dolt](https://www.dolthub.com/),\n> - [LakeFS](https://lakefs.io/), and\n> - [DVC](https://dvc.org/)\n>\n> as examples. They all imply git hash style validation of the individual\n> data commits, by not mention of [PGP] signing, though it may available\n> for some.\n>\n> I did see the Dolt issue [ Cryptographic signing of a changeset? #628\n> ](https://github.com/dolthub/dolt/issues/628), so it looks like it's on\n> their radar, though it's likely they'll need similar discussions about\n> how to cross integrate with Git..\n>\n> However, we also need to note the shift to the cloud for these very\n> large immobile data sets, where there maybe concerns as to the security\n> and trustworthiness of the compute and storage platforms (cosmic rays,\n> random glitches, hacks, etc).\n>\n> We are no longer importing code to our local machine that we need to be\n> signed, rather we are exporting our code to their compute\n> infrastructure, so the verification has to happen 'over-there'. So the\n> integrity question is still very pertinent.\n>\n> >>  then it would be easier\n> >> to see where in the Git model that may fit. Or whether Git is the underling VCS\n> >> (because it has programmable API), and the Data VCS (esp because of scale and\n> >> non-distributed nature) becomes the \"authority\", even if that has less capability!\n> > I agree as well. I want to see assurances that this level of integrity can be maintained - or that the user will have to accept the risks that git signatures are no longer usable. It might be appropriate to disable commit.gpgsign if the underlying VCS cannot be an authority.\n> >\n> >\n> I'd also worry, like yourself, about the cloud data sets, and how the\n> data selection subsets are captured (e.g. if multiple individuals have\n> used their right to be forgotten to make the old selection no longer\n> accessible, then how to validate?). Interesting times.\n> --\n> Philip\n"}]}