{"thread":{"id":"58229","subject":"Question: What's the best way to implement directory permission control in git?","startedAt":"2022-07-27T08:56:37Z","lastAt":"2022-08-01T10:15:01Z","messageCount":13,"participants":["ZheNing Hu","Ævar Arnfjörð Bjarmason","Thomas Guyot","Elijah Newren","rsbecker@nexbridge.com","Emily Shaffer","Han-Wen Nienhuys"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"460010","messageId":"CAOLTT8QusNzdO1mHqQFPz84pznYSpFWJunroRGXQ7qk6sJjeYg@mail.gmail.com","threadId":"58229","inReplyTo":null,"subject":"Question: What's the best way to implement directory permission control in git?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2022-07-27T08:56:12Z","receivedAt":"2022-07-27T08:56:37Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"if there is a monorepo such as\ngit@github.com:derrickstolee/sparse-checkout-example.git\n\nThere are many files and directories:\n\nclient/\n    android/\n    electron/\n    iOS/\nservice/\n    common/\n    identity/\n    list/\n    photos/\nweb/\n    browser/\n    editor/\n    friends/\nboostrap.sh\nLICENSE.md\nREADME.md\n\nNow we can use partial-clone + sparse-checkout to reduce\nthe network overhead, and reduce disk storage space size, that's good.\n\nBut I also need a ACL to control what directory or file people can fetch/push.\ne.g. I don't want a client fetch the code in \"service\" or \"web\".\n\nNow if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\nor other git command, git which will  download them by\n\"git fetch --filter=blob:none --stdin <oid>\" automatically.\n\nThis means that the git client and server interact with git objects\n(and don't care about path) we cannot simply ban someone download\na \"path\" on the server side.\n\nWhat should I do? You may recommend me to use submodule,\nbut due to its complexity, I don't really want to use it :-(\n\nZheNing Hu\n"},{"id":"460015","messageId":"220727.86mtculxnz.gmgdl@evledraar.gmail.com","threadId":"58229","inReplyTo":"CAOLTT8QusNzdO1mHqQFPz84pznYSpFWJunroRGXQ7qk6sJjeYg@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-07-27T09:17:44Z","receivedAt":"2022-07-27T09:20:30Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Jul 27 2022, ZheNing Hu wrote:\n\n> if there is a monorepo such as\n> git@github.com:derrickstolee/sparse-checkout-example.git\n>\n> There are many files and directories:\n>\n> client/\n>     android/\n>     electron/\n>     iOS/\n> service/\n>     common/\n>     identity/\n>     list/\n>     photos/\n> web/\n>     browser/\n>     editor/\n>     friends/\n> boostrap.sh\n> LICENSE.md\n> README.md\n>\n> Now we can use partial-clone + sparse-checkout to reduce\n> the network overhead, and reduce disk storage space size, that's good.\n>\n> But I also need a ACL to control what directory or file people can fetch/push.\n> e.g. I don't want a client fetch the code in \"service\" or \"web\".\n>\n> Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> or other git command, git which will  download them by\n> \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n>\n> This means that the git client and server interact with git objects\n> (and don't care about path) we cannot simply ban someone download\n> a \"path\" on the server side.\n>\n> What should I do? You may recommend me to use submodule,\n> but due to its complexity, I don't really want to use it :-(\n\nThere isn't a way to do this in git.\n\nIt's theoretically possible, i.e. a client could be told that the SHA-1\nof a directory is XYZ, and construct a commit object with a reference to\nit.\n\nBut currently a *lot* of things in the client code assume that these\nthings will be available in one way or another.\n\nThe state-of-the-art in the \"sparse\" code may differ from the above, I\ndon't know.\n\nAlso note that there's a well-known edge case in the git protocol where\nit's really incompatible with the notion of \"secret\" data, i.e. even if\nyou hide a ref you'll be able to \"guess\" it by seeing what delta(s) the\nserver will produce or accept etc.\n"},{"id":"460017","messageId":"80dd46c5-f9ff-d2b3-2d7f-4b80e00494b8@gmail.com","threadId":"58229","inReplyTo":"CAOLTT8QusNzdO1mHqQFPz84pznYSpFWJunroRGXQ7qk6sJjeYg@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"Thomas Guyot","fromEmail":"tguyot@gmail.com","sentAt":"2022-07-27T09:24:55Z","receivedAt":"2022-07-27T09:27:23Z","isPatch":false,"sender":{"key":"tguyot@gmail.com","avatar":"https://avatars.githubusercontent.com/u/403890?v=4"},"body":"On 2022-07-27 04:56, ZheNing Hu wrote:\n> if there is a monorepo such as\n> git@github.com:derrickstolee/sparse-checkout-example.git\n>\n> There are many files and directories:\n>\n> client/\n>      android/\n>      electron/\n>      iOS/\n> service/\n>      common/\n>      identity/\n>      list/\n>      photos/\n> web/\n>      browser/\n>      editor/\n>      friends/\n> boostrap.sh\n> LICENSE.md\n> README.md\n>\n> Now we can use partial-clone + sparse-checkout to reduce\n> the network overhead, and reduce disk storage space size, that's good.\n>\n> But I also need a ACL to control what directory or file people can fetch/push.\n> e.g. I don't want a client fetch the code in \"service\" or \"web\".\n\nPushes can easily be blocked with a pre-receive or update hook on the \nserver side. That covers the case where you want to prevenr users to \nupdate certain paths in the repo.\n> Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> or other git command, git which will  download them by\n> \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n>\n> This means that the git client and server interact with git objects\n> (and don't care about path) we cannot simply ban someone download\n> a \"path\" on the server side.\n\nIndeed - core devs can correct me if I'm wrong but afaik even in the \ncase of sparse checkouts and partial clones the packs may include other \nobjects. I have no ideas how git selects objects and packs on sent and \nwhen it decides to repack objects... What I know is it can pack entire \nrepos in just a few files using delta compression and it would probably \nmake sense to sent these pack if there is no real benefit in repacking \njust the requested objects.\n> What should I do? You may recommend me to use submodule,\n> but due to its complexity, I don't really want to use it :-(\n\nSubmodules is definitively an option for read ACLs, and considering git \nwas not originally designed to hide information from a single store it's \nprobably your only option. Moreover, if the git client is able to fetch \ndirectly blobs and trees (the later includes partial trees as a tree \nobject is a single \"directory\" that can contain other blobs and trees), \nthen even the server has no knowledge of where a tree hook into, or even \nhow it's named. All that information would have to be mapped elsewhere.\n\nTo take your example above, the \"common\" subtree of \"service/\" could be \nin multiple top level directories (i,e, the same tree with same \ncontents), and each top level dirs could have a different \"common\" \nsubtree. So git would have to find where each tree object (one per \ndirectory) is accessible from for *each revision* before deciding if a \nclient should be authorized to fetch an object, and the same would be \nrequired for blobs (and tree objects don't even know their own name, \nthat comes from the reference in the parent tree or commit object for \nthe top-level tree).\n\nSo even before solving the client/server protocol issue you mentioned, \nyou can't just hide part of a repo in git right now and changing that is \ndefinitively not trivial.\n\n--\nThomas\n"},{"id":"460073","messageId":"CAOLTT8QpYzoKDq6Pf8+YegCWngogy=3hUf-SyV180kntgucMpQ@mail.gmail.com","threadId":"58229","inReplyTo":"220727.86mtculxnz.gmgdl@evledraar.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2022-07-28T14:54:04Z","receivedAt":"2022-07-28T14:55:59Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2022年7月27日周三 17:20写道：\n>\n>\n> On Wed, Jul 27 2022, ZheNing Hu wrote:\n>\n> > if there is a monorepo such as\n> > git@github.com:derrickstolee/sparse-checkout-example.git\n> >\n> > There are many files and directories:\n> >\n> > client/\n> >     android/\n> >     electron/\n> >     iOS/\n> > service/\n> >     common/\n> >     identity/\n> >     list/\n> >     photos/\n> > web/\n> >     browser/\n> >     editor/\n> >     friends/\n> > boostrap.sh\n> > LICENSE.md\n> > README.md\n> >\n> > Now we can use partial-clone + sparse-checkout to reduce\n> > the network overhead, and reduce disk storage space size, that's good.\n> >\n> > But I also need a ACL to control what directory or file people can fetch/push.\n> > e.g. I don't want a client fetch the code in \"service\" or \"web\".\n> >\n> > Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> > or other git command, git which will  download them by\n> > \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n> >\n> > This means that the git client and server interact with git objects\n> > (and don't care about path) we cannot simply ban someone download\n> > a \"path\" on the server side.\n> >\n> > What should I do? You may recommend me to use submodule,\n> > but due to its complexity, I don't really want to use it :-(\n>\n> There isn't a way to do this in git.\n>\n> It's theoretically possible, i.e. a client could be told that the SHA-1\n> of a directory is XYZ, and construct a commit object with a reference to\n> it.\n>\n\nI guess you mean use a special reference to hold the restricted path which\nthe client can access, and pre-receive-hook can ban the client from downloading\nother references. But this method is a little weird... How can this reference\nsync with main branches? If we have changed client permission to access\nserver directory, how to get the \"history\" of the server directory?\n\nI believe this approach is not very appropriate and is not maintainable.\n\n> But currently a *lot* of things in the client code assume that these\n> things will be available in one way or another.\n>\n> The state-of-the-art in the \"sparse\" code may differ from the above, I\n> don't know.\n>\n> Also note that there's a well-known edge case in the git protocol where\n> it's really incompatible with the notion of \"secret\" data, i.e. even if\n> you hide a ref you'll be able to \"guess\" it by seeing what delta(s) the\n> server will produce or accept etc.\n\nYeah, there are data security issues... Unless we need to isolate objects\nbetween directories. Or in this case we disable the delta object.....\nOkay, this seems a little strange.\n\nAnyway, thanks for the answer!\n\nZheNing Hu\n"},{"id":"460075","messageId":"220728.861qu5kz2c.gmgdl@evledraar.gmail.com","threadId":"58229","inReplyTo":"CAOLTT8QpYzoKDq6Pf8+YegCWngogy=3hUf-SyV180kntgucMpQ@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-07-28T15:50:14Z","receivedAt":"2022-07-28T16:00:03Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Jul 28 2022, ZheNing Hu wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2022年7月27日周三 17:20写道：\n>>\n>>\n>> On Wed, Jul 27 2022, ZheNing Hu wrote:\n>>\n>> > if there is a monorepo such as\n>> > git@github.com:derrickstolee/sparse-checkout-example.git\n>> >\n>> > There are many files and directories:\n>> >\n>> > client/\n>> >     android/\n>> >     electron/\n>> >     iOS/\n>> > service/\n>> >     common/\n>> >     identity/\n>> >     list/\n>> >     photos/\n>> > web/\n>> >     browser/\n>> >     editor/\n>> >     friends/\n>> > boostrap.sh\n>> > LICENSE.md\n>> > README.md\n>> >\n>> > Now we can use partial-clone + sparse-checkout to reduce\n>> > the network overhead, and reduce disk storage space size, that's good.\n>> >\n>> > But I also need a ACL to control what directory or file people can fetch/push.\n>> > e.g. I don't want a client fetch the code in \"service\" or \"web\".\n>> >\n>> > Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n>> > or other git command, git which will  download them by\n>> > \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n>> >\n>> > This means that the git client and server interact with git objects\n>> > (and don't care about path) we cannot simply ban someone download\n>> > a \"path\" on the server side.\n>> >\n>> > What should I do? You may recommend me to use submodule,\n>> > but due to its complexity, I don't really want to use it :-(\n>>\n>> There isn't a way to do this in git.\n>>\n>> It's theoretically possible, i.e. a client could be told that the SHA-1\n>> of a directory is XYZ, and construct a commit object with a reference to\n>> it.\n>>\n>\n> I guess you mean use a special reference to hold the restricted path which\n> the client can access, and pre-receive-hook can ban the client from downloading\n> other references. But this method is a little weird... How can this reference\n> sync with main branches? If we have changed client permission to access\n> server directory, how to get the \"history\" of the server directory?\n>\n> I believe this approach is not very appropriate and is not maintainable.\n\nIt's not maintainable at all, and I don't believe any current git client\nsupports this.\n\nBut due to git's commits referring to a Merkle tree I can tell you that\na subdirectory \"secret\" has a current tree SHA-1 of XYZ, without giving\nyou any of that content.\n\nYou *could* then manually construct a commit like:\n\n\ttree <NEW_TREE>\n\t...\n\nWhere the \"<NEW_TREE>\" would be a tree like:\n\n\t100644 blob <NEW-BLOB-SHA1>\tUPDATED.md\n\t040000 tree <XYZ>\tsecret-stuff\n\nAnd send you a PACK with my new two three new objects (commit, blob &\nnew top-level NEW_TREE). To the remote end & protocol it wouldn't be\ndistinguishable from a \"normal\" push.\n\nBut nothing supports this already, as a practical matter most of git\neither hard dies if content is missing, or has other odd edge-case\nsemantics (and I'm not up-to-date on the state of the art).\n\nAnyway, just saying that for the longer term I'm not aware of an\n*intrinsic* reason for why we couldn't support this sort of thing, in\ncase anyone's interested in putting in a *lot* of leg work to make it\nhappen.\n\n>> But currently a *lot* of things in the client code assume that these\n>> things will be available in one way or another.\n>>\n>> The state-of-the-art in the \"sparse\" code may differ from the above, I\n>> don't know.\n>>\n>> Also note that there's a well-known edge case in the git protocol where\n>> it's really incompatible with the notion of \"secret\" data, i.e. even if\n>> you hide a ref you'll be able to \"guess\" it by seeing what delta(s) the\n>> server will produce or accept etc.\n>\n> Yeah, there are data security issues... Unless we need to isolate objects\n> between directories. Or in this case we disable the delta object.....\n> Okay, this seems a little strange.\n\nYou can't really just \"disable the delta(s)\". Well, you can in\nprinciple, but like what I outlined above it's one of those things\nthat's a far way off, and it's one thing to e.g. have a client that's\nable to craft a commit referring to data it doesn't have.\n\nIt's quite another to secure a server in such a way that it can serve up\nsecret data from the repo to some clients, but not to others.\n\nI can imagine some hacks to make that happen, but I won't go into that\nhere...\n"},{"id":"460169","messageId":"CABPp-BH8BYMaG=VK_OpfX3QKBLAOiLH9sTDdTWq5=4C6-663HA@mail.gmail.com","threadId":"58229","inReplyTo":"220728.861qu5kz2c.gmgdl@evledraar.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2022-07-29T01:48:21Z","receivedAt":"2022-07-29T01:48:42Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Thu, Jul 28, 2022 at 9:28 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n> On Thu, Jul 28 2022, ZheNing Hu wrote:\n>\n> > Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2022年7月27日周三 17:20写道：\n> >>\n> >>\n> >> On Wed, Jul 27 2022, ZheNing Hu wrote:\n> >>\n> >> > if there is a monorepo such as\n> >> > git@github.com:derrickstolee/sparse-checkout-example.git\n> >> >\n> >> > There are many files and directories:\n> >> >\n> >> > client/\n> >> >     android/\n> >> >     electron/\n> >> >     iOS/\n> >> > service/\n> >> >     common/\n> >> >     identity/\n> >> >     list/\n> >> >     photos/\n> >> > web/\n> >> >     browser/\n> >> >     editor/\n> >> >     friends/\n> >> > boostrap.sh\n> >> > LICENSE.md\n> >> > README.md\n> >> >\n> >> > Now we can use partial-clone + sparse-checkout to reduce\n> >> > the network overhead, and reduce disk storage space size, that's good.\n> >> >\n> >> > But I also need a ACL to control what directory or file people can fetch/push.\n> >> > e.g. I don't want a client fetch the code in \"service\" or \"web\".\n> >> >\n> >> > Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> >> > or other git command, git which will  download them by\n> >> > \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n> >> >\n> >> > This means that the git client and server interact with git objects\n> >> > (and don't care about path) we cannot simply ban someone download\n> >> > a \"path\" on the server side.\n> >> >\n> >> > What should I do? You may recommend me to use submodule,\n> >> > but due to its complexity, I don't really want to use it :-(\n> >>\n> >> There isn't a way to do this in git.\n> >>\n> >> It's theoretically possible, i.e. a client could be told that the SHA-1\n> >> of a directory is XYZ, and construct a commit object with a reference to\n> >> it.\n> >>\n> >\n> > I guess you mean use a special reference to hold the restricted path which\n> > the client can access, and pre-receive-hook can ban the client from downloading\n> > other references. But this method is a little weird... How can this reference\n> > sync with main branches? If we have changed client permission to access\n> > server directory, how to get the \"history\" of the server directory?\n> >\n> > I believe this approach is not very appropriate and is not maintainable.\n>\n> It's not maintainable at all, and I don't believe any current git client\n> supports this.\n\nI agree it's not maintainable and a bad idea.  But I did want to\ncorrect one small thing, and I do have an alternative suggestion at\nthe end...\n\n> But due to git's commits referring to a Merkle tree I can tell you that\n> a subdirectory \"secret\" has a current tree SHA-1 of XYZ, without giving\n> you any of that content.\n>\n> You *could* then manually construct a commit like:\n>\n>         tree <NEW_TREE>\n>         ...\n>\n> Where the \"<NEW_TREE>\" would be a tree like:\n>\n>         100644 blob <NEW-BLOB-SHA1>     UPDATED.md\n>         040000 tree <XYZ>       secret-stuff\n>\n> And send you a PACK with my new two three new objects (commit, blob &\n> new top-level NEW_TREE). To the remote end & protocol it wouldn't be\n> distinguishable from a \"normal\" push.\n>\n> But nothing supports this already, as a practical matter most of git\n> either hard dies if content is missing, or has other odd edge-case\n> semantics (and I'm not up-to-date on the state of the art).\n\nActually, this is what sparse-index (as a sub-option in\nsparse-checkout) already basically does.  See\nDocumentation/technical/sparse-index.txt for details, and note that\nwe're basically in Phase IV of that document.  In short, the\nsparse-index makes it so that common operations based on the index do\nnot need and do not use information about some subtrees, so if someone\nhas a partial clone starting with no blobs, they will only have to\ndownload a small subset of the repository blobs in order to handle\nmost Git operations, and many operations become much faster since the\nindex is so much smaller.\n\nHowever:\n\n* Users can run `git sparse-checkout reapply --no-sparse-index` at any\ntime to force the index to be full again.  This is documented, and\neven suggested that users remember in case they attempt to use\nexternal tools (jgit? libgit2? others?) that don't understand sparse\ndirectory entries.  So, removing this ability would be problematic.\n\n* It makes no guarantee whatsoever that the sparse directory entries\nare not expanded by less frequently used Git commands.  Notice the\n\"ensure_full_index()\" calls sprinkled throughout the code.  Some have\nbeen removed, one by one, as commands have been modified to better\noperate with a sparse index.  The odds they'll all be removed in the\nfuture may well be close to 0%.\n\n* The `ort` merge strategy ignores the index altogether during\noperation.  If it needs to walk into a tree to complete a\nmerge/rebase/revert/cherry-pick/etc., it will.  Further, it doesn't\njust look into those paths, it intentionally de-sparsifies paths\ninvolved in conflicts, so it can display it to the user.\n\n* Just because the index is sparse does not mean other commands can't\nwalk into those directories.  So `git grep` (when given a revision),\n`git diff`, `git log`, etc. will look in (old versions of) those\npaths.\n\n> Anyway, just saying that for the longer term I'm not aware of an\n> *intrinsic* reason for why we couldn't support this sort of thing, in\n> case anyone's interested in putting in a *lot* of leg work to make it\n> happen.\n\nAnd on top of the technical leg work required, they would also need to\nsomehow convince everyone else that it's worth accepting the increased\nmaintenance effort.  Right now, even if someone had already done the\nwork to implement it, I'd say it's not worth the maintenance costs.\n\nHowever, there are two alternative choices I can think of here: You\ncan use submodules if you want a fixed part of the repository to only\nbe available to a subset of folks, or use josh\n(https://github.com/josh-project/josh) if you need it to be more\ndynamic.\n"},{"id":"460208","messageId":"CAOLTT8R0C_j1o6bXLtve0kCV2eUUgwLUk8+fUzwEYKgzPLPj3Q@mail.gmail.com","threadId":"58229","inReplyTo":"80dd46c5-f9ff-d2b3-2d7f-4b80e00494b8@gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2022-07-29T12:49:29Z","receivedAt":"2022-07-29T12:49:45Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Thomas Guyot <tguyot@gmail.com> 于2022年7月27日周三 17:27写道：\n>\n> On 2022-07-27 04:56, ZheNing Hu wrote:\n> > if there is a monorepo such as\n> > git@github.com:derrickstolee/sparse-checkout-example.git\n> >\n> > There are many files and directories:\n> >\n> > client/\n> >      android/\n> >      electron/\n> >      iOS/\n> > service/\n> >      common/\n> >      identity/\n> >      list/\n> >      photos/\n> > web/\n> >      browser/\n> >      editor/\n> >      friends/\n> > boostrap.sh\n> > LICENSE.md\n> > README.md\n> >\n> > Now we can use partial-clone + sparse-checkout to reduce\n> > the network overhead, and reduce disk storage space size, that's good.\n> >\n> > But I also need a ACL to control what directory or file people can fetch/push.\n> > e.g. I don't want a client fetch the code in \"service\" or \"web\".\n>\n> Pushes can easily be blocked with a pre-receive or update hook on the\n> server side. That covers the case where you want to prevenr users to\n> update certain paths in the repo.\n\nAgree. pre-receive-hook/update-hook may be the only way git does ACL\nwhen client push.\n\n> > Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> > or other git command, git which will  download them by\n> > \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n> >\n> > This means that the git client and server interact with git objects\n> > (and don't care about path) we cannot simply ban someone download\n> > a \"path\" on the server side.\n>\n> Indeed - core devs can correct me if I'm wrong but afaik even in the\n> case of sparse checkouts and partial clones the packs may include other\n> objects. I have no ideas how git selects objects and packs on sent and\n> when it decides to repack objects... What I know is it can pack entire\n> repos in just a few files using delta compression and it would probably\n> make sense to sent these pack if there is no real benefit in repacking\n> just the requested objects.\n\nYeah, here we need to consider the tradeoff between repacking all\nobjects in path\nand sending objects one by one.\n\n> > What should I do? You may recommend me to use submodule,\n> > but due to its complexity, I don't really want to use it :-(\n>\n> Submodules is definitively an option for read ACLs, and considering git\n> was not originally designed to hide information from a single store it's\n> probably your only option. Moreover, if the git client is able to fetch\n> directly blobs and trees (the later includes partial trees as a tree\n> object is a single \"directory\" that can contain other blobs and trees),\n> then even the server has no knowledge of where a tree hook into, or even\n> how it's named. All that information would have to be mapped elsewhere.\n>\n\nAn association: how does the linux file system do ACL?\n\n1. The uid, gid of the file or sub directory record in directory entry.\n2. When a user wants to access the file or sub directory, filesystem\ncheck if the user has the same permission of uid/gid.\n\nThis is just a casual thought:\n\nLet git imitate the linux file system, we may need to record user\nsignatures in entry of tree objects.\n\nThen we let the server just download the objects which match\nthe user signature.\n\n> To take your example above, the \"common\" subtree of \"service/\" could be\n> in multiple top level directories (i,e, the same tree with same\n> contents), and each top level dirs could have a different \"common\"\n> subtree. So git would have to find where each tree object (one per\n> directory) is accessible from for *each revision* before deciding if a\n> client should be authorized to fetch an object, and the same would be\n> required for blobs (and tree objects don't even know their own name,\n> that comes from the reference in the parent tree or commit object for\n> the top-level tree).\n>\n\nIf the client tells the server what the objects it wants , yes, it\nneeds to check\nall the paths which may \"hold\" these objects... But if user tell\nserver what's the\npath it wants,  that will be easier for the server to check...\n\n> So even before solving the client/server protocol issue you mentioned,\n> you can't just hide part of a repo in git right now and changing that is\n> definitively not trivial.\n>\n\nAgree. It's hard to change...\n\n> --\n> Thomas\n\nZheNing Hu\n"},{"id":"460211","messageId":"CAOLTT8SQsxW4WqwVcE951sW7vqP+YUPauLpMhz8jpRYsmv0bzA@mail.gmail.com","threadId":"58229","inReplyTo":"220728.861qu5kz2c.gmgdl@evledraar.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2022-07-29T13:15:05Z","receivedAt":"2022-07-29T13:15:23Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2022年7月28日周四 23:59写道：\n>\n>\n> On Thu, Jul 28 2022, ZheNing Hu wrote:\n>\n> > Ævar Arnfjörð Bjarmason <avarab@gmail.com> 于2022年7月27日周三 17:20写道：\n> >>\n> >>\n> >> On Wed, Jul 27 2022, ZheNing Hu wrote:\n> >>\n> >> > if there is a monorepo such as\n> >> > git@github.com:derrickstolee/sparse-checkout-example.git\n> >> >\n> >> > There are many files and directories:\n> >> >\n> >> > client/\n> >> >     android/\n> >> >     electron/\n> >> >     iOS/\n> >> > service/\n> >> >     common/\n> >> >     identity/\n> >> >     list/\n> >> >     photos/\n> >> > web/\n> >> >     browser/\n> >> >     editor/\n> >> >     friends/\n> >> > boostrap.sh\n> >> > LICENSE.md\n> >> > README.md\n> >> >\n> >> > Now we can use partial-clone + sparse-checkout to reduce\n> >> > the network overhead, and reduce disk storage space size, that's good.\n> >> >\n> >> > But I also need a ACL to control what directory or file people can fetch/push.\n> >> > e.g. I don't want a client fetch the code in \"service\" or \"web\".\n> >> >\n> >> > Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> >> > or other git command, git which will  download them by\n> >> > \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n> >> >\n> >> > This means that the git client and server interact with git objects\n> >> > (and don't care about path) we cannot simply ban someone download\n> >> > a \"path\" on the server side.\n> >> >\n> >> > What should I do? You may recommend me to use submodule,\n> >> > but due to its complexity, I don't really want to use it :-(\n> >>\n> >> There isn't a way to do this in git.\n> >>\n> >> It's theoretically possible, i.e. a client could be told that the SHA-1\n> >> of a directory is XYZ, and construct a commit object with a reference to\n> >> it.\n> >>\n> >\n> > I guess you mean use a special reference to hold the restricted path which\n> > the client can access, and pre-receive-hook can ban the client from downloading\n> > other references. But this method is a little weird... How can this reference\n> > sync with main branches? If we have changed client permission to access\n> > server directory, how to get the \"history\" of the server directory?\n> >\n> > I believe this approach is not very appropriate and is not maintainable.\n>\n> It's not maintainable at all, and I don't believe any current git client\n> supports this.\n>\n> But due to git's commits referring to a Merkle tree I can tell you that\n> a subdirectory \"secret\" has a current tree SHA-1 of XYZ, without giving\n> you any of that content.\n>\n> You *could* then manually construct a commit like:\n>\n>         tree <NEW_TREE>\n>         ...\n>\n> Where the \"<NEW_TREE>\" would be a tree like:\n>\n>         100644 blob <NEW-BLOB-SHA1>     UPDATED.md\n>         040000 tree <XYZ>       secret-stuff\n>\n> And send you a PACK with my new two three new objects (commit, blob &\n> new top-level NEW_TREE). To the remote end & protocol it wouldn't be\n> distinguishable from a \"normal\" push.\n>\n> But nothing supports this already, as a practical matter most of git\n> either hard dies if content is missing, or has other odd edge-case\n> semantics (and I'm not up-to-date on the state of the art).\n>\n> Anyway, just saying that for the longer term I'm not aware of an\n> *intrinsic* reason for why we couldn't support this sort of thing, in\n> case anyone's interested in putting in a *lot* of leg work to make it\n> happen.\n>\n\nAs Newren said, this is just like what sparse-index does. I use\npartial clone + sparse-checkout + sparse-index to do git add/git commit,\ngit can add and commit correctly without fetching any excess objects.\nBut we can't prevent users from downloading other directories or files.\n\n> >> But currently a *lot* of things in the client code assume that these\n> >> things will be available in one way or another.\n> >>\n> >> The state-of-the-art in the \"sparse\" code may differ from the above, I\n> >> don't know.\n> >>\n> >> Also note that there's a well-known edge case in the git protocol where\n> >> it's really incompatible with the notion of \"secret\" data, i.e. even if\n> >> you hide a ref you'll be able to \"guess\" it by seeing what delta(s) the\n> >> server will produce or accept etc.\n> >\n> > Yeah, there are data security issues... Unless we need to isolate objects\n> > between directories. Or in this case we disable the delta object.....\n> > Okay, this seems a little strange.\n>\n> You can't really just \"disable the delta(s)\". Well, you can in\n> principle, but like what I outlined above it's one of those things\n> that's a far way off, and it's one thing to e.g. have a client that's\n> able to craft a commit referring to data it doesn't have.\n>\n> It's quite another to secure a server in such a way that it can serve up\n> secret data from the repo to some clients, but not to others.\n>\n\nAll right... I might have to think of something else.\n\n> I can imagine some hacks to make that happen, but I won't go into that\n> here...\n\nZheNing Hu\n"},{"id":"460216","messageId":"CAOLTT8R1WQyqLNfymJgxTgCuoOKEe0Vu+7k9D+85u-x53FYJiQ@mail.gmail.com","threadId":"58229","inReplyTo":"CABPp-BH8BYMaG=VK_OpfX3QKBLAOiLH9sTDdTWq5=4C6-663HA@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2022-07-29T14:22:00Z","receivedAt":"2022-07-29T14:22:15Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Elijah Newren <newren@gmail.com> 于2022年7月29日周五 09:48写道：\n\n> > But due to git's commits referring to a Merkle tree I can tell you that\n> > a subdirectory \"secret\" has a current tree SHA-1 of XYZ, without giving\n> > you any of that content.\n> >\n> > You *could* then manually construct a commit like:\n> >\n> >         tree <NEW_TREE>\n> >         ...\n> >\n> > Where the \"<NEW_TREE>\" would be a tree like:\n> >\n> >         100644 blob <NEW-BLOB-SHA1>     UPDATED.md\n> >         040000 tree <XYZ>       secret-stuff\n> >\n> > And send you a PACK with my new two three new objects (commit, blob &\n> > new top-level NEW_TREE). To the remote end & protocol it wouldn't be\n> > distinguishable from a \"normal\" push.\n> >\n> > But nothing supports this already, as a practical matter most of git\n> > either hard dies if content is missing, or has other odd edge-case\n> > semantics (and I'm not up-to-date on the state of the art).\n>\n> Actually, this is what sparse-index (as a sub-option in\n> sparse-checkout) already basically does.  See\n> Documentation/technical/sparse-index.txt for details, and note that\n> we're basically in Phase IV of that document.  In short, the\n> sparse-index makes it so that common operations based on the index do\n> not need and do not use information about some subtrees, so if someone\n> has a partial clone starting with no blobs, they will only have to\n> download a small subset of the repository blobs in order to handle\n> most Git operations, and many operations become much faster since the\n> index is so much smaller.\n>\n\nI think this is mainly due to sparse-checkout instead of sparse-index.\nWithout the sparse-index, we also can do git add, git commit without fetching\nother blob objects.\n\nBut sparse-index can help reduce the size of indexes.\n\n> However:\n>\n> * Users can run `git sparse-checkout reapply --no-sparse-index` at any\n> time to force the index to be full again.  This is documented, and\n> even suggested that users remember in case they attempt to use\n> external tools (jgit? libgit2? others?) that don't understand sparse\n> directory entries.  So, removing this ability would be problematic.\n>\n\nOr `git sparse-checkout disable`? Whatever, when git finds other objects\nmissing, it will fetch the objects from remote, and we may do ACL check here.\nJust let jgit/libgit2/others fail to fetch objects (in this special case?)\n\n> * It makes no guarantee whatsoever that the sparse directory entries\n> are not expanded by less frequently used Git commands.  Notice the\n> \"ensure_full_index()\" calls sprinkled throughout the code.  Some have\n> been removed, one by one, as commands have been modified to better\n> operate with a sparse index.  The odds they'll all be removed in the\n> future may well be close to 0%.\n>\n\nThat's good...\n\n> * The `ort` merge strategy ignores the index altogether during\n> operation.  If it needs to walk into a tree to complete a\n> merge/rebase/revert/cherry-pick/etc., it will.  Further, it doesn't\n> just look into those paths, it intentionally de-sparsifies paths\n> involved in conflicts, so it can display it to the user.\n>\n\nSo the user has to care and deal with a merge conflict in a directory\nthat he \"doesn't have access to\"...\n\nIt would be nice to have the user only care about conflicts in directories/files\nto which he has permissions. I don't know if it would be very\ndifficult to design.\n\n> * Just because the index is sparse does not mean other commands can't\n> walk into those directories.  So `git grep` (when given a revision),\n> `git diff`, `git log`, etc. will look in (old versions of) those\n> paths.\n>\n\nAgree.\n\n> > Anyway, just saying that for the longer term I'm not aware of an\n> > *intrinsic* reason for why we couldn't support this sort of thing, in\n> > case anyone's interested in putting in a *lot* of leg work to make it\n> > happen.\n>\n> And on top of the technical leg work required, they would also need to\n> somehow convince everyone else that it's worth accepting the increased\n> maintenance effort.  Right now, even if someone had already done the\n> work to implement it, I'd say it's not worth the maintenance costs.\n>\n> However, there are two alternative choices I can think of here: You\n> can use submodules if you want a fixed part of the repository to only\n> be available to a subset of folks, or use josh\n> (https://github.com/josh-project/josh) if you need it to be more\n> dynamic.\n\nThanks, I will take a look.\n\nZheNing Hu\n"},{"id":"460221","messageId":"00b901d8a35b$8ebbe6a0$ac33b3e0$@nexbridge.com","threadId":"58229","inReplyTo":"CAOLTT8R1WQyqLNfymJgxTgCuoOKEe0Vu+7k9D+85u-x53FYJiQ@mail.gmail.com","subject":"RE: Question: What's the best way to implement directory permission control in git?","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-07-29T14:57:38Z","receivedAt":"2022-07-29T14:57:47Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On July 29, 2022 10:22 AM, ZheNing Hu wrote:\n>Elijah Newren <newren@gmail.com> 于2022年7月29日周五 09:48写道：\n>\n>> > But due to git's commits referring to a Merkle tree I can tell you\n>> > that a subdirectory \"secret\" has a current tree SHA-1 of XYZ,\n>> > without giving you any of that content.\n>> >\n>> > You *could* then manually construct a commit like:\n>> >\n>> >         tree <NEW_TREE>\n>> >         ...\n>> >\n>> > Where the \"<NEW_TREE>\" would be a tree like:\n>> >\n>> >         100644 blob <NEW-BLOB-SHA1>     UPDATED.md\n>> >         040000 tree <XYZ>       secret-stuff\n>> >\n>> > And send you a PACK with my new two three new objects (commit, blob\n>> > & new top-level NEW_TREE). To the remote end & protocol it wouldn't\n>> > be distinguishable from a \"normal\" push.\n>> >\n>> > But nothing supports this already, as a practical matter most of git\n>> > either hard dies if content is missing, or has other odd edge-case\n>> > semantics (and I'm not up-to-date on the state of the art).\n>>\n>> Actually, this is what sparse-index (as a sub-option in\n>> sparse-checkout) already basically does.  See\n>> Documentation/technical/sparse-index.txt for details, and note that\n>> we're basically in Phase IV of that document.  In short, the\n>> sparse-index makes it so that common operations based on the index do\n>> not need and do not use information about some subtrees, so if someone\n>> has a partial clone starting with no blobs, they will only have to\n>> download a small subset of the repository blobs in order to handle\n>> most Git operations, and many operations become much faster since the\n>> index is so much smaller.\n>>\n>\n>I think this is mainly due to sparse-checkout instead of sparse-index.\n>Without the sparse-index, we also can do git add, git commit without fetching\n>other blob objects.\n>\n>But sparse-index can help reduce the size of indexes.\n>\n>> However:\n>>\n>> * Users can run `git sparse-checkout reapply --no-sparse-index` at any\n>> time to force the index to be full again.  This is documented, and\n>> even suggested that users remember in case they attempt to use\n>> external tools (jgit? libgit2? others?) that don't understand sparse\n>> directory entries.  So, removing this ability would be problematic.\n>>\n>\n>Or `git sparse-checkout disable`? Whatever, when git finds other objects missing,\n>it will fetch the objects from remote, and we may do ACL check here.\n>Just let jgit/libgit2/others fail to fetch objects (in this special case?)\n>\n>> * It makes no guarantee whatsoever that the sparse directory entries\n>> are not expanded by less frequently used Git commands.  Notice the\n>> \"ensure_full_index()\" calls sprinkled throughout the code.  Some have\n>> been removed, one by one, as commands have been modified to better\n>> operate with a sparse index.  The odds they'll all be removed in the\n>> future may well be close to 0%.\n>>\n>\n>That's good...\n>\n>> * The `ort` merge strategy ignores the index altogether during\n>> operation.  If it needs to walk into a tree to complete a\n>> merge/rebase/revert/cherry-pick/etc., it will.  Further, it doesn't\n>> just look into those paths, it intentionally de-sparsifies paths\n>> involved in conflicts, so it can display it to the user.\n>>\n>\n>So the user has to care and deal with a merge conflict in a directory that he\n>\"doesn't have access to\"...\n>\n>It would be nice to have the user only care about conflicts in directories/files to\n>which he has permissions. I don't know if it would be very difficult to design.\n>\n>> * Just because the index is sparse does not mean other commands can't\n>> walk into those directories.  So `git grep` (when given a revision),\n>> `git diff`, `git log`, etc. will look in (old versions of) those\n>> paths.\n>>\n>\n>Agree.\n>\n>> > Anyway, just saying that for the longer term I'm not aware of an\n>> > *intrinsic* reason for why we couldn't support this sort of thing,\n>> > in case anyone's interested in putting in a *lot* of leg work to\n>> > make it happen.\n>>\n>> And on top of the technical leg work required, they would also need to\n>> somehow convince everyone else that it's worth accepting the increased\n>> maintenance effort.  Right now, even if someone had already done the\n>> work to implement it, I'd say it's not worth the maintenance costs.\n>>\n>> However, there are two alternative choices I can think of here: You\n>> can use submodules if you want a fixed part of the repository to only\n>> be available to a subset of folks, or use josh\n>> (https://github.com/josh-project/josh) if you need it to be more\n>> dynamic.\n>\n>Thanks, I will take a look.\n\nAs a completely side perspective on this, I had to integrate security management with five separate security subsystems/mechanisms (not joking) on the NonStop platform that included Unix-style Access Control Lists (ACLs), non-inode ACLs on the NonStop side of the platform, and some recent new thing called XOS - I don't know it yet but provisioned for it. The solution I ended up with was writing a full Workflow wrapper around git that does things similar to GitHub Actions, so after an operation like checkout/switch, merge, pull, etc., specific rules specified in YAML in the repo (if enabled by the user) are run that apply the ACLs. It is a very heavy-weight solution to the problem but works pretty well on this \"exotic\" platform - Workflows were needed for other reasons as well, so I just piggybacked the security handling into my Workflow structure. Again, not built into git but wrapped around it. I could have used hooks for some of it but needed support for more operations than hooks had.\n--Randall\n\n"},{"id":"460291","messageId":"CAJoAoZmsuwYCA8XGziEA-qwghg9h22Af98JQE1AuHHBRfQgrDA@mail.gmail.com","threadId":"58229","inReplyTo":"CAOLTT8QusNzdO1mHqQFPz84pznYSpFWJunroRGXQ7qk6sJjeYg@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2022-07-29T23:50:38Z","receivedAt":"2022-07-29T23:50:53Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, Jul 27, 2022 at 1:56 AM ZheNing Hu <adlternative@gmail.com> wrote:\n>\n> if there is a monorepo such as\n> git@github.com:derrickstolee/sparse-checkout-example.git\n>\n> There are many files and directories:\n>\n> client/\n>     android/\n>     electron/\n>     iOS/\n> service/\n>     common/\n>     identity/\n>     list/\n>     photos/\n> web/\n>     browser/\n>     editor/\n>     friends/\n> boostrap.sh\n> LICENSE.md\n> README.md\n>\n> Now we can use partial-clone + sparse-checkout to reduce\n> the network overhead, and reduce disk storage space size, that's good.\n>\n> But I also need a ACL to control what directory or file people can fetch/push.\n> e.g. I don't want a client fetch the code in \"service\" or \"web\".\n>\n> Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> or other git command, git which will  download them by\n> \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n>\n> This means that the git client and server interact with git objects\n> (and don't care about path) we cannot simply ban someone download\n> a \"path\" on the server side.\n>\n> What should I do? You may recommend me to use submodule,\n> but due to its complexity, I don't really want to use it :-(\n\nAs a quick note, there is some effort on making submodules less\ncomplex, at least from the user perspective. My team and I have been\nactively working on improvements in that area for the past year or so.\nPlease feel free to read and examine the design doc[1] to see if the\nfuture looks brighter in that direction than you thought - or, even\nbetter, if there's something missing from that design that would be\ncompelling in allowing you to use submodules to solve your use case.\n\nAs for differing ACLs within a single repository... Google has had\nsome attempts at it and has only found pain, at least where Git is\ninvolved. As others have mentioned elsewhere downthread, it doesn't\nreally match Git's data model.\n\nGerrit has tried to support something sort of similar to this -\nper-branch read permissions. They were really painful! So much so that\nour Gerrit team is actively discouraging their use, and in the process\nof deprecating them. It turns out that on the server side, calculating\npermissions for which commit should be visible is very expensive,\nbecause you are not just saying \"is commit abcdef on\nforbidden-branch?\" but rather are saying \"is commit abcdef on\nforbidden-branch *and not on any branches $user is allowed to see*?\"\nThe same calculation woes would be true of per-object or per-tree\npermissions, because Git will treat 'everyone/can/see/.linter.config'\nand 'very/secret/dir/.linter.config' as a single object with a single\nID if the contents of each '.linter.config' are identical. It is still\nvery expensive for the server to decide whether or not it's okay to\nsend a certain object. Part of the reason the branch ACL calculation\nis so painful is that we have some repositories with many many\nbranches (100,000+); if you're using a very large monorepo you will\nprobably find similarly expensive and complex calculations even in a\nsingle repository.\n\nGenerally, this isn't something I'd like to see Git support - I think\nit would by necessity be kludgey and has some very pointy edge cases\nfor the user (what if I'm trying to merge from another branch and\nthere is a conflict in very/secret/dir/, but I'm not allowed to see\nit?). But of course Git is open source, and my opinion is only one of\nmany; I just wanted to share some past pain that we've had in this\narea.\n\n - Emily\n\n1: https://lore.kernel.org/git/YHofmWcIAidkvJiD@google.com/\n"},{"id":"460329","messageId":"CAOLTT8RNnbmnckidVtCbfuSymjvPeMnU_uj7bqGj-XUuL+W_mg@mail.gmail.com","threadId":"58229","inReplyTo":"CAJoAoZmsuwYCA8XGziEA-qwghg9h22Af98JQE1AuHHBRfQgrDA@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2022-07-31T16:15:31Z","receivedAt":"2022-07-31T16:16:01Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> 于2022年7月30日周六 07:50写道：\n>\n> On Wed, Jul 27, 2022 at 1:56 AM ZheNing Hu <adlternative@gmail.com> wrote:\n> >\n> > if there is a monorepo such as\n> > git@github.com:derrickstolee/sparse-checkout-example.git\n> >\n> > There are many files and directories:\n> >\n> > client/\n> >     android/\n> >     electron/\n> >     iOS/\n> > service/\n> >     common/\n> >     identity/\n> >     list/\n> >     photos/\n> > web/\n> >     browser/\n> >     editor/\n> >     friends/\n> > boostrap.sh\n> > LICENSE.md\n> > README.md\n> >\n> > Now we can use partial-clone + sparse-checkout to reduce\n> > the network overhead, and reduce disk storage space size, that's good.\n> >\n> > But I also need a ACL to control what directory or file people can fetch/push.\n> > e.g. I don't want a client fetch the code in \"service\" or \"web\".\n> >\n> > Now if the user client use \"git log -p\" or \"git sparse-checkout add service\"...\n> > or other git command, git which will  download them by\n> > \"git fetch --filter=blob:none --stdin <oid>\" automatically.\n> >\n> > This means that the git client and server interact with git objects\n> > (and don't care about path) we cannot simply ban someone download\n> > a \"path\" on the server side.\n> >\n> > What should I do? You may recommend me to use submodule,\n> > but due to its complexity, I don't really want to use it :-(\n>\n> As a quick note, there is some effort on making submodules less\n> complex, at least from the user perspective. My team and I have been\n> actively working on improvements in that area for the past year or so.\n> Please feel free to read and examine the design doc[1] to see if the\n> future looks brighter in that direction than you thought - or, even\n> better, if there's something missing from that design that would be\n> compelling in allowing you to use submodules to solve your use case.\n>\n\nThanks, I think submodules’ improvement may shift my perception.\nBut the problem I'm having is whether I should give permission control\nto all \"subdirectories\" (if and when I find out that this is not necessary,\nthen submodules might be an option)\n\n> As for differing ACLs within a single repository... Google has had\n> some attempts at it and has only found pain, at least where Git is\n> involved. As others have mentioned elsewhere downthread, it doesn't\n> really match Git's data model.\n>\n\nThat's so sad :(\n\n> Gerrit has tried to support something sort of similar to this -\n> per-branch read permissions. They were really painful! So much so that\n> our Gerrit team is actively discouraging their use, and in the process\n> of deprecating them. It turns out that on the server side, calculating\n> permissions for which commit should be visible is very expensive,\n> because you are not just saying \"is commit abcdef on\n> forbidden-branch?\" but rather are saying \"is commit abcdef on\n> forbidden-branch *and not on any branches $user is allowed to see*?\"\n> The same calculation woes would be true of per-object or per-tree\n> permissions, because Git will treat 'everyone/can/see/.linter.config'\n> and 'very/secret/dir/.linter.config' as a single object with a single\n> ID if the contents of each '.linter.config' are identical. It is still\n> very expensive for the server to decide whether or not it's okay to\n> send a certain object. Part of the reason the branch ACL calculation\n> is so painful is that we have some repositories with many many\n> branches (100,000+); if you're using a very large monorepo you will\n> probably find similarly expensive and complex calculations even in a\n> single repository.\n>\n\nAgree, as Avar said that there are delta data too (so data cannot easily\nhidden)\n\n> Generally, this isn't something I'd like to see Git support - I think\n> it would by necessity be kludgey and has some very pointy edge cases\n> for the user (what if I'm trying to merge from another branch and\n> there is a conflict in very/secret/dir/, but I'm not allowed to see\n> it?). But of course Git is open source, and my opinion is only one of\n> many; I just wanted to share some past pain that we've had in this\n> area.\n>\n\nTo summarize (your and other answers' ideas), I have reasons to believe\nthat git itself cannot easily solve this directory permissions problem:\n1. Files with the same object id can be in different directories\n(data cannot be isolated).\n2. DELTA data can share data between multiple objects\n(data cannot be isolated).\n3. Permission management is very cumbersome and time consuming,\nespecially on large repositories.\n4. The directories that are not accessible should be or not see merge\nconflict is a big problem.\n\n>  - Emily\n>\n> 1: https://lore.kernel.org/git/YHofmWcIAidkvJiD@google.com/\n\nThanks.\n\nZheNing Hu\n"},{"id":"460339","messageId":"CAFQ2z_PMZJ0CeEsruhQ_dAna1yTc+z1+p0SaeGg6+XsiKnZ=xQ@mail.gmail.com","threadId":"58229","inReplyTo":"CAJoAoZmsuwYCA8XGziEA-qwghg9h22Af98JQE1AuHHBRfQgrDA@mail.gmail.com","subject":"Re: Question: What's the best way to implement directory permission control in git?","fromName":"Han-Wen Nienhuys","fromEmail":"hanwen@google.com","sentAt":"2022-08-01T10:14:44Z","receivedAt":"2022-08-01T10:15:01Z","isPatch":false,"sender":{"key":"hanwen@google.com","avatar":"https://avatars.githubusercontent.com/u/31547?v=4"},"body":"On Sat, Jul 30, 2022 at 1:50 AM Emily Shaffer <emilyshaffer@google.com> wrote:\n> Gerrit has tried to support something sort of similar to this -\n> per-branch read permissions. They were really painful! So much so that\n> our Gerrit team is actively discouraging their use, and in the process\n> of deprecating them. It turns out that on the server side, calculating\n> permissions for which commit should be visible is very expensive,\n> because you are not just saying \"is commit abcdef on\n> forbidden-branch?\" but rather are saying \"is commit abcdef on\n> forbidden-branch *and not on any branches $user is allowed to see*?\"\n> The same calculation woes would be true of per-object or per-tree\n> permissions, because Git will treat 'everyone/can/see/.linter.config'\n> and 'very/secret/dir/.linter.config' as a single object with a single\n> ID if the contents of each '.linter.config' are identical. It is still\n> very expensive for the server to decide whether or not it's okay to\n> send a certain object. Part of the reason the branch ACL calculation\n> is so painful is that we have some repositories with many many\n> branches (100,000+); if you're using a very large monorepo you will\n> probably find similarly expensive and complex calculations even in a\n> single repository.\n\n\nThanks Emily,\n\nI agree with your points, but as the manager of Google's Gerrit team,\nI just wanted to add a few clarifications:\n\n* The max number of branches we have on repositories is O(1000s). IIRC\nour Android repositories are the worst offenders, because there is a\ncombinatorial explosion of {major release, minor release, target\ndevice}. Pending reviews number in the millions, but we usually don't\nhave to evaluate ACLs fully, as the review refs aren't downloaded\ncommonly.\n\n* The read ACLs are assigned to {branch-regexp, group} tuples. This\nmeans that you can't precompute visibility either, because each\nindividual user may be in a different set of groups.\n\n* Even disconsidering that, you can still do optimizations if updates\nare FF (because each update only increases the visibility of each\ncommit). However, non-FF branch updates preclude such precomputations.\n(Gerrit has non-FF updates in a number of places).\n\n* The Gerrit team isn't actively deprecating read ACLs: the problem is\nhard, because removing read ACLs on branches means that the read ACLs\nmove to repository level, which implies setting up complex ACL\nconfiguration and replication infrastructure for repositories to\naddress existing use cases. It's currently just one of these features\nthat we wish hadn't been added, but now that it's there, we suffer\nthrough it.\n\nMore generally, read permissions are hard to get right in a monorepo:\neven if you stop developers from accessing the code through Git fetch,\nthe permissions must also be enforced throughout the entire dev stack,\nincluding code browsing, code search, viewing CI artifacts etc.\n\n-- \nHan-Wen Nienhuys - Google Munich\nI work 80%. Don't expect answers from me on Fridays.\n--\n\nGoogle Germany GmbH, Erika-Mann-Strasse 33, 80636 Munich\n\nRegistergericht und -nummer: Hamburg, HRB 86891\n\nSitz der Gesellschaft: Hamburg\n\nGeschäftsführer: Paul Manicle, Liana Sebastian\n"}]}