{"thread":{"id":"59561","subject":"[Question] Can git cat-file have a type filtering option?","startedAt":"2023-04-07T14:24:26Z","lastAt":"2023-04-16T12:43:32Z","messageCount":23,"participants":["ZheNing Hu","Junio C Hamano","Taylor Blau","Jeff King","Linus Torvalds","Felipe Contreras"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"475000","messageId":"CAOLTT8RTB7kpabN=Rv1nHvKTaYh6pLR6moOJhfC2wdtUG_xahQ@mail.gmail.com","threadId":"59561","inReplyTo":null,"subject":"[Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-07T14:24:22Z","receivedAt":"2023-04-07T14:24:26Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Sometimes when we use `git cat-file --batch-all-objects`, we only want\ndata of type \"blob\". In order to filter them out, we may need to use\nsome additional processes (such as `git rev-list --objects\n--filter=blob:none --filter-provided-objects`) to obtain the SHA of\nall blobs, and then use `git cat-file --batch` to retrieve them. This\nis not very elegant, or in other words, it might be better to have an\ninternal implementation of filtering within `git cat-file\n--batch-all-objects`.\n\nHowever, `git cat-file` already has a `--filters` option, which is\nused to \"show content as transformed by filters\". I'm not sure if\nthere is a better word to implement the functionality of filtering by\ntype? For example, `--type-filter`?\n\nThanks,\n--\nZheNing Hu\n"},{"id":"475008","messageId":"xmqqy1n3k63p.fsf@gitster.g","threadId":"59561","inReplyTo":"CAOLTT8RTB7kpabN=Rv1nHvKTaYh6pLR6moOJhfC2wdtUG_xahQ@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-04-07T16:30:18Z","receivedAt":"2023-04-07T16:30:23Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"ZheNing Hu <adlternative@gmail.com> writes:\n\n> all blobs, and then use `git cat-file --batch` to retrieve them. This\n> is not very elegant, or in other words, it might be better to have an\n> internal implementation of filtering within `git cat-file\n> --batch-all-objects`.\n\nIt does sound prominently elegant to have each tool does one task\nand does it well, and being able to flexibly combine them to achieve\na larger task.\n\nOnce that approach is working well, it may still make sense to give\na special case codepath that bundles a specific combination of these\nprimitive features, if use cases for the specific combination appear\noften.  But I do not know if the particular one, \"we do not want to\nfeed specific list of objects to check to 'cat-file --batch'\",\nqualifies as one.\n\n> For example, `--type-filter`?\n\nIs the object type the only thing that people often would want to\nbase their filtering decision on?  Will we then see somebody else\nrequest a \"--size-filter\", and then somebody else realizes that the\nfiltering criteria based on size need to be different between blobs\n(most likely counted in bytes) and trees (it may be more convenient\nto count the tree entries, not byes)?  It sounds rather messy and\nwe may be better off having such an extensible logic in one place.\n\nLike rev-list's object list filtering, that is.\n\nIs the logic that implements rev-list's object list filtering\nsomething that is easily called from the side, as if it were a\nlibrary routine?  Refactoring that and teaching cat-file an option\nto activate that logic might be more palatable.\n\nThanks.\n"},{"id":"475032","messageId":"CAOLTT8SXXKG3uEd8Q=uh3zx7XeUDUWezGgNUSCd1Fpq-Kyy-2A@mail.gmail.com","threadId":"59561","inReplyTo":"xmqqy1n3k63p.fsf@gitster.g","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-08T06:27:53Z","receivedAt":"2023-04-08T06:27:53Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Junio C Hamano <gitster@pobox.com> 于2023年4月8日周六 00:30写道：\n>\n> ZheNing Hu <adlternative@gmail.com> writes:\n>\n> > all blobs, and then use `git cat-file --batch` to retrieve them. This\n> > is not very elegant, or in other words, it might be better to have an\n> > internal implementation of filtering within `git cat-file\n> > --batch-all-objects`.\n>\n> It does sound prominently elegant to have each tool does one task\n> and does it well, and being able to flexibly combine them to achieve\n> a larger task.\n>\n\nOkay, you're right. It's not \"ungraceful\" to have each task do its own thing.\nI should clarify that for a command like `git cat-file --batch-all-objects`,\nwhich traverses all objects, it would be better to have a filter. It might be\nmore performant than using `git rev-list --filter | git cat-file --batch`?\n\n\n> Once that approach is working well, it may still make sense to give\n> a special case codepath that bundles a specific combination of these\n> primitive features, if use cases for the specific combination appear\n> often.  But I do not know if the particular one, \"we do not want to\n> feed specific list of objects to check to 'cat-file --batch'\",\n> qualifies as one.\n>\n> > For example, `--type-filter`?\n>\n> Is the object type the only thing that people often would want to\n> base their filtering decision on?  Will we then see somebody else\n> request a \"--size-filter\", and then somebody else realizes that the\n> filtering criteria based on size need to be different between blobs\n> (most likely counted in bytes) and trees (it may be more convenient\n> to count the tree entries, not byes)?  It sounds rather messy and\n> we may be better off having such an extensible logic in one place.\n>\n\nYes, having a generic filter for `git cat-file` would be better.\n\n> Like rev-list's object list filtering, that is.\n>\n> Is the logic that implements rev-list's object list filtering\n> something that is easily called from the side, as if it were a\n> library routine?  Refactoring that and teaching cat-file an option\n> to activate that logic might be more palatable.\n>\n\nI don't think so. While `git rev-list` traverses objects and performs\nfiltering within a revision, `git cat-file --batch-all-objects` traverses\nall loose and packed objects. It might be difficult to perfectly\nextract the filtering from `git rev-list` and apply it to `git cat-file`.\n\n> Thanks.\n"},{"id":"475043","messageId":"ZDITjrSU7jADX/Jw@nand.local","threadId":"59561","inReplyTo":"CAOLTT8RTB7kpabN=Rv1nHvKTaYh6pLR6moOJhfC2wdtUG_xahQ@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-04-09T01:23:26Z","receivedAt":"2023-04-09T01:23:34Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Apr 07, 2023 at 10:24:22PM +0800, ZheNing Hu wrote:\n> However, `git cat-file` already has a `--filters` option, which is\n> used to \"show content as transformed by filters\". I'm not sure if\n> there is a better word to implement the functionality of filtering by\n> type? For example, `--type-filter`?\n\nThere is the `--filter='object:type=blob'` that should do what you're\nlooking for.\n\nIn other words, if you wanted to dump the contents of all blobs in your\nrepository, this should do the trick:\n\n  $ git rev-list --all --objects --filter='object:type=blob' |\n    git cat-file --batch[=<format>]\n\nThanks,\nTaylor\n"},{"id":"475044","messageId":"ZDIUKqPatP+FX8dM@nand.local","threadId":"59561","inReplyTo":"xmqqy1n3k63p.fsf@gitster.g","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-04-09T01:26:02Z","receivedAt":"2023-04-09T01:26:07Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Fri, Apr 07, 2023 at 09:30:18AM -0700, Junio C Hamano wrote:\n> ZheNing Hu <adlternative@gmail.com> writes:\n>\n> > all blobs, and then use `git cat-file --batch` to retrieve them. This\n> > is not very elegant, or in other words, it might be better to have an\n> > internal implementation of filtering within `git cat-file\n> > --batch-all-objects`.\n>\n> It does sound prominently elegant to have each tool does one task\n> and does it well, and being able to flexibly combine them to achieve\n> a larger task.\n\nYeah, agreed. It may be *convenient* to have an easy-to-reach option in\ncat-file like '--exclude-type=tree,commit,tag' or something. But the\nargument falls on a pretty slippery slope, as I think you note below.\n\n> Is the object type the only thing that people often would want to\n> base their filtering decision on?  Will we then see somebody else\n> request a \"--size-filter\", and then somebody else realizes that the\n> filtering criteria based on size need to be different between blobs\n> (most likely counted in bytes) and trees (it may be more convenient\n> to count the tree entries, not byes)?  It sounds rather messy and\n> we may be better off having such an extensible logic in one place.\n>\n> Like rev-list's object list filtering, that is.\n\nYes, exactly. This definitely feels like a \"do one thing and do it\nwell\". `rev-list` is the tool we have for listing revisions and objects,\nand it can produce output that is compatible with the kind of input that\nother tools (like `cat-file`) can interpret.\n\nThanks,\nTaylor\n"},{"id":"475045","messageId":"ZDIUvK/bF7BFqX5q@nand.local","threadId":"59561","inReplyTo":"CAOLTT8SXXKG3uEd8Q=uh3zx7XeUDUWezGgNUSCd1Fpq-Kyy-2A@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-04-09T01:28:28Z","receivedAt":"2023-04-09T01:28:33Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sat, Apr 08, 2023 at 02:27:53PM +0800, ZheNing Hu wrote:\n> Okay, you're right. It's not \"ungraceful\" to have each task do its own thing.\n> I should clarify that for a command like `git cat-file --batch-all-objects`,\n> which traverses all objects, it would be better to have a filter. It might be\n> more performant than using `git rev-list --filter | git cat-file --batch`?\n\nPerhaps slightly so, since there is naturally going to be some\nduplicated effort spawning processes, loading any shared libraries,\ninitializing the repository and reading its configuration, etc.\n\nBut I'd wager that these are all a negligible cost when compared to the\ntime we'll have to spend reading, inflating, and printing out all of the\nobjects in your repository.\n\nHopefully any task(s) where that cost *wouldn't* be negligible relative\nto the rest of the job would be small enough that they could fit into a\nsingle process.\n\n> I don't think so. While `git rev-list` traverses objects and performs\n> filtering within a revision, `git cat-file --batch-all-objects` traverses\n> all loose and packed objects. It might be difficult to perfectly\n> extract the filtering from `git rev-list` and apply it to `git cat-file`.\n\n`rev-list`'s `--all` option does exactly the former: it looks at all\nloose and packed objects instead of doing a traditional object walk.\n\nThanks,\nTaylor\n"},{"id":"475046","messageId":"ZDIgyKDQ2rJT2YEI@nand.local","threadId":"59561","inReplyTo":"ZDIUvK/bF7BFqX5q@nand.local","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-04-09T02:19:52Z","receivedAt":"2023-04-09T02:19:58Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sat, Apr 08, 2023 at 09:28:28PM -0400, Taylor Blau wrote:\n> > I don't think so. While `git rev-list` traverses objects and performs\n> > filtering within a revision, `git cat-file --batch-all-objects` traverses\n> > all loose and packed objects. It might be difficult to perfectly\n> > extract the filtering from `git rev-list` and apply it to `git cat-file`.\n>\n> `rev-list`'s `--all` option does exactly the former: it looks at all\n> loose and packed objects instead of doing a traditional object walk.\n\nSorry, this isn't right: --all pretends as if you passed all references\nto it over argv, not to just look at the individual loose and packed\nobjects.\n\nThanks,\nTaylor\n"},{"id":"475047","messageId":"ZDIiO1HMjej+rnMk@nand.local","threadId":"59561","inReplyTo":"ZDIgyKDQ2rJT2YEI@nand.local","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-04-09T02:26:03Z","receivedAt":"2023-04-09T02:26:08Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sat, Apr 08, 2023 at 10:19:52PM -0400, Taylor Blau wrote:\n> On Sat, Apr 08, 2023 at 09:28:28PM -0400, Taylor Blau wrote:\n> > > I don't think so. While `git rev-list` traverses objects and performs\n> > > filtering within a revision, `git cat-file --batch-all-objects` traverses\n> > > all loose and packed objects. It might be difficult to perfectly\n> > > extract the filtering from `git rev-list` and apply it to `git cat-file`.\n> >\n> > `rev-list`'s `--all` option does exactly the former: it looks at all\n> > loose and packed objects instead of doing a traditional object walk.\n>\n> Sorry, this isn't right: --all pretends as if you passed all references\n> to it over argv, not to just look at the individual loose and packed\n> objects.\n\nThe right thing to do here if you wanted to get a listing of all blobs\nin your repository regardless of their reachability or whether they are\nloose or packed is:\n\n    git cat-file --batch-check='%(objectname)' --batch-all-objects |\n    git rev-list --objects --stdin --no-walk --filter='object:type=blob'\n\nOr, if your filter is as straightforward as \"is this object a blob or\nnot\", you could write something like:\n\n    git cat-file --batch-check --batch-all-objects | awk '\n      if ($2 == \"blob\") { print $0 }'\n\nOr you could tighten up the AWK expression by doing something like:\n\n    git cat-file --batch-check='%(objecttype) %(objectname)' \\\n      --batch-all-objects | awk '/^blob / { print $2 }'\n\nSorry for the brain fart.\n\nThanks,\nTaylor\n"},{"id":"475049","messageId":"CAOLTT8RbU6G67BtE9fSv4gEn10dtR7cT-jf+dcEfhvNhvcwETQ@mail.gmail.com","threadId":"59561","inReplyTo":"ZDIUvK/bF7BFqX5q@nand.local","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-09T06:47:30Z","receivedAt":"2023-04-09T06:47:28Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Taylor Blau <me@ttaylorr.com> 于2023年4月9日周日 09:28写道：\n>\n> On Sat, Apr 08, 2023 at 02:27:53PM +0800, ZheNing Hu wrote:\n> > Okay, you're right. It's not \"ungraceful\" to have each task do its own thing.\n> > I should clarify that for a command like `git cat-file --batch-all-objects`,\n> > which traverses all objects, it would be better to have a filter. It might be\n> > more performant than using `git rev-list --filter | git cat-file --batch`?\n>\n> Perhaps slightly so, since there is naturally going to be some\n> duplicated effort spawning processes, loading any shared libraries,\n> initializing the repository and reading its configuration, etc.\n>\n> But I'd wager that these are all a negligible cost when compared to the\n> time we'll have to spend reading, inflating, and printing out all of the\n> objects in your repository.\n>\n\n\"What you said makes sense. I implemented the --type-filter option for\ngit cat-file and compared the performance of outputting all blobs in the\ngit repository with and without using the type-filter. I found that the\ndifference was not significant.\n\ntime git  cat-file --batch-all-objects --batch-check=\"%(objectname)\n%(objecttype)\" |\nawk '{ if ($2 == \"blob\") print $1 }' | git cat-file --batch > /dev/null\n17.10s user 0.27s system 102% cpu 16.987 total\n\ntime git cat-file --batch-all-objects --batch --type-filter=blob >/dev/null\n16.74s user 0.19s system 95% cpu 17.655 total\n\nAt first, I thought the processes that provide all blob oids by using\ngit rev-list or git cat-file --batch-all-objects --batch-check might waste\ncpu, io, memory resources because they need to read a large number\nof objects, and then they are read again by git cat-file --batch.\nHowever, it seems that this is not actually the bottleneck in performance.\n\n> Hopefully any task(s) where that cost *wouldn't* be negligible relative\n> to the rest of the job would be small enough that they could fit into a\n> single process.\n>\n> > I don't think so. While `git rev-list` traverses objects and performs\n> > filtering within a revision, `git cat-file --batch-all-objects` traverses\n> > all loose and packed objects. It might be difficult to perfectly\n> > extract the filtering from `git rev-list` and apply it to `git cat-file`.\n>\n> `rev-list`'s `--all` option does exactly the former: it looks at all\n> loose and packed objects instead of doing a traditional object walk.\n>\n> Thanks,\n> Taylor\n"},{"id":"475050","messageId":"CAOLTT8TFiXG1hABFVLp_TOEZ4__s2k4+nvcG3Ax867=LJxOi_g@mail.gmail.com","threadId":"59561","inReplyTo":"ZDIiO1HMjej+rnMk@nand.local","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-09T06:51:34Z","receivedAt":"2023-04-09T06:51:34Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Taylor Blau <me@ttaylorr.com> 于2023年4月9日周日 10:26写道：\n>\n> On Sat, Apr 08, 2023 at 10:19:52PM -0400, Taylor Blau wrote:\n> > On Sat, Apr 08, 2023 at 09:28:28PM -0400, Taylor Blau wrote:\n> > > > I don't think so. While `git rev-list` traverses objects and performs\n> > > > filtering within a revision, `git cat-file --batch-all-objects` traverses\n> > > > all loose and packed objects. It might be difficult to perfectly\n> > > > extract the filtering from `git rev-list` and apply it to `git cat-file`.\n> > >\n> > > `rev-list`'s `--all` option does exactly the former: it looks at all\n> > > loose and packed objects instead of doing a traditional object walk.\n> >\n> > Sorry, this isn't right: --all pretends as if you passed all references\n> > to it over argv, not to just look at the individual loose and packed\n> > objects.\n>\n> The right thing to do here if you wanted to get a listing of all blobs\n> in your repository regardless of their reachability or whether they are\n> loose or packed is:\n>\n>     git cat-file --batch-check='%(objectname)' --batch-all-objects |\n>     git rev-list --objects --stdin --no-walk --filter='object:type=blob'\n>\n\nThis looks like a mistake. Try passing a tree oid to git rev-list:\n\ngit rev-list --objects --stdin --no-walk --filter='object:type=blob'\n<<< HEAD^{tree}\n27f9fa75c6d8cdae7834f38006b631522c6a5ac3\n4860bebd32f8d3f34c2382f097ac50c0b972d3a0 .cirrus.yml\nc592dda681fecfaa6bf64fb3f539eafaf4123ed8 .clang-format\nf9d819623d832113014dd5d5366e8ee44ac9666a .editorconfig\nb0044cf272fec9b987e99c600d6a95bc357261c3 .gitattributes\n...\n\n> Or, if your filter is as straightforward as \"is this object a blob or\n> not\", you could write something like:\n>\n>     git cat-file --batch-check --batch-all-objects | awk '\n>       if ($2 == \"blob\") { print $0 }'\n>\n> Or you could tighten up the AWK expression by doing something like:\n>\n>     git cat-file --batch-check='%(objecttype) %(objectname)' \\\n>       --batch-all-objects | awk '/^blob / { print $2 }'\n>\n> Sorry for the brain fart.\n>\n> Thanks,\n> Taylor\n"},{"id":"475079","messageId":"20230410200141.GB104097@coredump.intra.peff.net","threadId":"59561","inReplyTo":"CAOLTT8TFiXG1hABFVLp_TOEZ4__s2k4+nvcG3Ax867=LJxOi_g@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-04-10T20:01:41Z","receivedAt":"2023-04-10T20:01:45Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Apr 09, 2023 at 02:51:34PM +0800, ZheNing Hu wrote:\n\n> > The right thing to do here if you wanted to get a listing of all blobs\n> > in your repository regardless of their reachability or whether they are\n> > loose or packed is:\n> >\n> >     git cat-file --batch-check='%(objectname)' --batch-all-objects |\n> >     git rev-list --objects --stdin --no-walk --filter='object:type=blob'\n> >\n> \n> This looks like a mistake. Try passing a tree oid to git rev-list:\n> \n> git rev-list --objects --stdin --no-walk --filter='object:type=blob'\n> <<< HEAD^{tree}\n> 27f9fa75c6d8cdae7834f38006b631522c6a5ac3\n> 4860bebd32f8d3f34c2382f097ac50c0b972d3a0 .cirrus.yml\n> c592dda681fecfaa6bf64fb3f539eafaf4123ed8 .clang-format\n> f9d819623d832113014dd5d5366e8ee44ac9666a .editorconfig\n> b0044cf272fec9b987e99c600d6a95bc357261c3 .gitattributes\n> ...\n\nThis is the expected behavior. The filter options are meant to support\npartial clones, and the behavior is really \"filter things we'd traverse\nto\". It is intentional that objects the caller directly asks for will\nalways be included in the output.\n\nI certainly found that convention confusing, but I imagine it solves\nsome problems with the lazy-fetch requests themselves. Regardless,\nthat's how it works and it's not going to change anytime soon. :)\n\nFor that reason, and just for general flexibility, I think you are\nmostly better off piping cat-file through an external filter program\n(and then back to cat-file to get more data on each object).\n\n-Peff\n"},{"id":"475080","messageId":"20230410201414.GC104097@coredump.intra.peff.net","threadId":"59561","inReplyTo":"CAOLTT8RbU6G67BtE9fSv4gEn10dtR7cT-jf+dcEfhvNhvcwETQ@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-04-10T20:14:14Z","receivedAt":"2023-04-10T20:14:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Apr 09, 2023 at 02:47:30PM +0800, ZheNing Hu wrote:\n\n> > Perhaps slightly so, since there is naturally going to be some\n> > duplicated effort spawning processes, loading any shared libraries,\n> > initializing the repository and reading its configuration, etc.\n> >\n> > But I'd wager that these are all a negligible cost when compared to the\n> > time we'll have to spend reading, inflating, and printing out all of the\n> > objects in your repository.\n> \n> \"What you said makes sense. I implemented the --type-filter option for\n> git cat-file and compared the performance of outputting all blobs in the\n> git repository with and without using the type-filter. I found that the\n> difference was not significant.\n> \n> time git  cat-file --batch-all-objects --batch-check=\"%(objectname)\n> %(objecttype)\" |\n> awk '{ if ($2 == \"blob\") print $1 }' | git cat-file --batch > /dev/null\n> 17.10s user 0.27s system 102% cpu 16.987 total\n> \n> time git cat-file --batch-all-objects --batch --type-filter=blob >/dev/null\n> 16.74s user 0.19s system 95% cpu 17.655 total\n> \n> At first, I thought the processes that provide all blob oids by using\n> git rev-list or git cat-file --batch-all-objects --batch-check might waste\n> cpu, io, memory resources because they need to read a large number\n> of objects, and then they are read again by git cat-file --batch.\n> However, it seems that this is not actually the bottleneck in performance.\n\nYeah, I think most of your time there is spent on the --batch command\nitself, which is just putting through a lot of bytes. You might also try\nwith \"--unordered\". The default ordering for --batch-all-objects is in\nsha1 order, which has pretty bad locality characteristics for delta\ncaching. Using --unordered goes in pack-order, which should be optimal.\n\nE.g., in git.git, running:\n\n  time \\\n    git cat-file --batch-all-objects --batch-check='%(objecttype) %(objectname)' |\n    perl -lne 'print $1 if /^blob (.*)/' |\n    git cat-file --batch >/dev/null\n\ntakes:\n\n  real\t0m29.961s\n  user\t0m29.128s\n  sys\t0m1.461s\n\nAdding \"--unordered\" to the initial cat-file gives:\n\n  real\t0m1.970s\n  user\t0m2.170s\n  sys\t0m0.126s\n\nSo reducing the size of the actual --batch printing may make the\nrelative cost of using multiple processes much higher (I didn't apply\nyour --type-filter patches to test myself).\n\nIn general, I do think having a processing pipeline like this is OK, as\nit's pretty flexible. But especially for smaller queries (even ones that\ndon't ask for the whole object contents), the per-object lookup costs\ncan start to dominate (especially in a repository that hasn't been\nrecently packed). Right now, even your \"--batch --type-filter\" example\nis probably making at least two lookups per object, because we don't\nhave a way to open a \"handle\" to an object to check its type, and then\nextract the contents conditionally. And of course with multiple\nprocesses, we're naturally doing a separate lookup in each one.\n\nSo a nice thing about being able to do the filtering in one process is\nthat we could _eventually_ do it all with one object lookup. But I'd\nprobably wait on adding something like --type-filter until we have an\ninternal single-lookup API, and then we could time it to see how much\nspeedup we can get.\n\n-Peff\n"},{"id":"475099","messageId":"ZDSZs8LgZzDLf5m1@nand.local","threadId":"59561","inReplyTo":"20230410200141.GB104097@coredump.intra.peff.net","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-04-10T23:20:19Z","receivedAt":"2023-04-10T23:20:37Z","isPatch":false,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Mon, Apr 10, 2023 at 04:01:41PM -0400, Jeff King wrote:\n> For that reason, and just for general flexibility, I think you are\n> mostly better off piping cat-file through an external filter program\n> (and then back to cat-file to get more data on each object).\n\nYeah, agreed. The convention of printing objects listed on the\ncommand-line regardless of whether they would pass through the object\nfilter is confusing to me, too.\n\nBut using `rev-list --no-walk` to accomplish the same job for a filter\nas trivial as the type-level one feels overkill anyway, so I agree that\njust relying on `cat-file` to produce the list of objects, filtering it\nyourself, and then handing it back to `cat-file` is the easiest thing to\ndo.\n\nThanks,\nTaylor\n"},{"id":"475147","messageId":"CAOLTT8T9pJFr94acvUo-8EYriST1gOAkXaDZBxHk54o=Zm5=Sg@mail.gmail.com","threadId":"59561","inReplyTo":"20230410201414.GC104097@coredump.intra.peff.net","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-11T14:09:33Z","receivedAt":"2023-04-11T14:10:05Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Jeff King <peff@peff.net> 于2023年4月11日周二 04:14写道：\n>\n> On Sun, Apr 09, 2023 at 02:47:30PM +0800, ZheNing Hu wrote:\n>\n> > > Perhaps slightly so, since there is naturally going to be some\n> > > duplicated effort spawning processes, loading any shared libraries,\n> > > initializing the repository and reading its configuration, etc.\n> > >\n> > > But I'd wager that these are all a negligible cost when compared to the\n> > > time we'll have to spend reading, inflating, and printing out all of the\n> > > objects in your repository.\n> >\n> > \"What you said makes sense. I implemented the --type-filter option for\n> > git cat-file and compared the performance of outputting all blobs in the\n> > git repository with and without using the type-filter. I found that the\n> > difference was not significant.\n> >\n> > time git  cat-file --batch-all-objects --batch-check=\"%(objectname)\n> > %(objecttype)\" |\n> > awk '{ if ($2 == \"blob\") print $1 }' | git cat-file --batch > /dev/null\n> > 17.10s user 0.27s system 102% cpu 16.987 total\n> >\n> > time git cat-file --batch-all-objects --batch --type-filter=blob >/dev/null\n> > 16.74s user 0.19s system 95% cpu 17.655 total\n> >\n> > At first, I thought the processes that provide all blob oids by using\n> > git rev-list or git cat-file --batch-all-objects --batch-check might waste\n> > cpu, io, memory resources because they need to read a large number\n> > of objects, and then they are read again by git cat-file --batch.\n> > However, it seems that this is not actually the bottleneck in performance.\n>\n> Yeah, I think most of your time there is spent on the --batch command\n> itself, which is just putting through a lot of bytes. You might also try\n> with \"--unordered\". The default ordering for --batch-all-objects is in\n> sha1 order, which has pretty bad locality characteristics for delta\n> caching. Using --unordered goes in pack-order, which should be optimal.\n>\n> E.g., in git.git, running:\n>\n>   time \\\n>     git cat-file --batch-all-objects --batch-check='%(objecttype) %(objectname)' |\n>     perl -lne 'print $1 if /^blob (.*)/' |\n>     git cat-file --batch >/dev/null\n>\n> takes:\n>\n>   real  0m29.961s\n>   user  0m29.128s\n>   sys   0m1.461s\n>\n> Adding \"--unordered\" to the initial cat-file gives:\n>\n>   real  0m1.970s\n>   user  0m2.170s\n>   sys   0m0.126s\n>\n> So reducing the size of the actual --batch printing may make the\n> relative cost of using multiple processes much higher (I didn't apply\n> your --type-filter patches to test myself).\n>\n\nYou are right. Adding the --unordered option can avoid the\ntime-consuming sorting process from affecting the test results.\n\ntime git cat-file --unordered --batch-all-objects \\\n--batch-check=\"%(objectname) %(objecttype)\" | \\\nawk '{ if ($2 == \"blob\") print $1 }' | git cat-file --batch > /dev/null\n\n4.17s user 0.23s system 109% cpu 4.025 total\n\ntime git cat-file --unordered --batch-all-objects --batch\n--type-filter=blob >/dev/null\n\n3.84s user 0.17s system 97% cpu 4.099 total\n\nIt looks like the difference is not significant either.\n\nAfter all, the truly time-consuming process is reading\nthe entire data of the blob, whereas git cat-file --batch-check\nonly reads the first few bytes of the object in comparison.\n\n> In general, I do think having a processing pipeline like this is OK, as\n> it's pretty flexible. But especially for smaller queries (even ones that\n> don't ask for the whole object contents), the per-object lookup costs\n> can start to dominate (especially in a repository that hasn't been\n> recently packed). Right now, even your \"--batch --type-filter\" example\n> is probably making at least two lookups per object, because we don't\n> have a way to open a \"handle\" to an object to check its type, and then\n> extract the contents conditionally. And of course with multiple\n> processes, we're naturally doing a separate lookup in each one.\n>\n\nYes, the type of the object is encapsulated in the header of the loose\nobject file or the object entry header of the pack file. We have to read\nit to get the object type. This may be a lingering question I have had:\nwhy does git put the type/size in the file data instead of storing it as some\nkind of metadata elsewhere?\n\n> So a nice thing about being able to do the filtering in one process is\n> that we could _eventually_ do it all with one object lookup. But I'd\n> probably wait on adding something like --type-filter until we have an\n> internal single-lookup API, and then we could time it to see how much\n> speedup we can get.\n>\n\nI am highly skeptical of this \"internal single-lookup API\". Do we really\nneed an extra metadata table to record all objects?\nSomething like: metadata: {oid: type, size}?\n\n> -Peff\n\nZheNing Hu\n"},{"id":"475200","messageId":"20230412074309.GB1695531@coredump.intra.peff.net","threadId":"59561","inReplyTo":"CAOLTT8T9pJFr94acvUo-8EYriST1gOAkXaDZBxHk54o=Zm5=Sg@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-04-12T07:43:09Z","receivedAt":"2023-04-12T07:43:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 11, 2023 at 10:09:33PM +0800, ZheNing Hu wrote:\n\n> > So reducing the size of the actual --batch printing may make the\n> > relative cost of using multiple processes much higher (I didn't apply\n> > your --type-filter patches to test myself).\n> >\n> \n> You are right. Adding the --unordered option can avoid the\n> time-consuming sorting process from affecting the test results.\n\nJust to be clear: it's not the cost of sorting, but rather that\naccessing the object contents in a sub-optimal order is much worse (and\nthat sub-optimal order happens to be \"sorted by sha1\", since that is\neffectively random with respect to the contents).\n\n> time git cat-file --unordered --batch-all-objects \\\n> --batch-check=\"%(objectname) %(objecttype)\" | \\\n> awk '{ if ($2 == \"blob\") print $1 }' | git cat-file --batch > /dev/null\n> \n> 4.17s user 0.23s system 109% cpu 4.025 total\n> \n> time git cat-file --unordered --batch-all-objects --batch\n> --type-filter=blob >/dev/null\n> \n> 3.84s user 0.17s system 97% cpu 4.099 total\n> \n> It looks like the difference is not significant either.\n\nOK, good, that means we can probably not worry about it. :)\n\n> > In general, I do think having a processing pipeline like this is OK, as\n> > it's pretty flexible. But especially for smaller queries (even ones that\n> > don't ask for the whole object contents), the per-object lookup costs\n> > can start to dominate (especially in a repository that hasn't been\n> > recently packed). Right now, even your \"--batch --type-filter\" example\n> > is probably making at least two lookups per object, because we don't\n> > have a way to open a \"handle\" to an object to check its type, and then\n> > extract the contents conditionally. And of course with multiple\n> > processes, we're naturally doing a separate lookup in each one.\n> >\n> \n> Yes, the type of the object is encapsulated in the header of the loose\n> object file or the object entry header of the pack file. We have to read\n> it to get the object type. This may be a lingering question I have had:\n> why does git put the type/size in the file data instead of storing it as some\n> kind of metadata elsewhere?\n\nIt's not just metadata; it's actually part of what we hash to get the\nobject id (though of course it doesn't _have_ to be stored in a linear\nbuffer, as the pack storage shows). But for loose objects, where would\nsuch metadata be? And accessing it isn't too expensive; we only zlib\ninflate the first few bytes (the main cost is in the syscalls to find\nand open the file).\n\nFor packed object, it effectively is metadata, just stuck at the front\nof the object contents, rather than in a separate table. That lets us\nuse the same .idx file for finding that metadata as we do for the\ncontents themselves (at the slight cost that if you're _just_ accessing\nmetadata, the results are sparser within the file, which has worse\nbehavior for cold-cache disks).\n\nBut when I say that lookup costs dominate, what I mean is that we'd\nspend a lot of our time binary searching within the pack .idx file, or\nfalling back to syscalls to look for loose objects.\n\n> > So a nice thing about being able to do the filtering in one process is\n> > that we could _eventually_ do it all with one object lookup. But I'd\n> > probably wait on adding something like --type-filter until we have an\n> > internal single-lookup API, and then we could time it to see how much\n> > speedup we can get.\n> \n> I am highly skeptical of this \"internal single-lookup API\". Do we really\n> need an extra metadata table to record all objects?\n> Something like: metadata: {oid: type, size}?\n\nNo, I don't mean changing the storage at all. I mean that rather than\ndoing this:\n\n  /* get type, size, etc, for --batch format */\n  type = oid_object_info(&oid, &size);\n\n  /* now get the contents for --batch to write them itself; but note\n   * that this searches for the entry again within all packs, etc */\n  contents = read_object_file(oid, &type, &size);\n\nas the cat-file code now does (because the first call is in\nbatch_object_write(), and the latter in print_object_or_die()), they\ncould be a single call that does the lookup once.\n\nWe could actually do that today, since the object contents are\neventually fed from oid_object_info_extended(), and we know ahead of\ntime that we want both the metadata and the contents. But that wouldn't\nwork if we filtered by type, etc.\n\nI'm not sure how much of a speedup it would yield in practice, though.\nIf you're printing the object contents, then the extra lookup is\nprobably not that expensive by comparison.\n\n-Peff\n"},{"id":"475206","messageId":"CAOLTT8Rw796zxMYxg5+nx8+YoQVnfy=nPXH8Aq0j0Cw+GLT1rA@mail.gmail.com","threadId":"59561","inReplyTo":"20230412074309.GB1695531@coredump.intra.peff.net","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-12T09:57:02Z","receivedAt":"2023-04-12T09:57:00Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Jeff King <peff@peff.net> 于2023年4月12日周三 15:43写道：\n>\n> On Tue, Apr 11, 2023 at 10:09:33PM +0800, ZheNing Hu wrote:\n>\n> > > So reducing the size of the actual --batch printing may make the\n> > > relative cost of using multiple processes much higher (I didn't apply\n> > > your --type-filter patches to test myself).\n> > >\n> >\n> > You are right. Adding the --unordered option can avoid the\n> > time-consuming sorting process from affecting the test results.\n>\n> Just to be clear: it's not the cost of sorting, but rather that\n> accessing the object contents in a sub-optimal order is much worse (and\n> that sub-optimal order happens to be \"sorted by sha1\", since that is\n> effectively random with respect to the contents).\n>\n\nOkay, thanks for correcting me. Reading the packfile in SHA1 order is\nactually a type of random read, and it should cause additional overhead.\n\n> > time git cat-file --unordered --batch-all-objects \\\n> > --batch-check=\"%(objectname) %(objecttype)\" | \\\n> > awk '{ if ($2 == \"blob\") print $1 }' | git cat-file --batch > /dev/null\n> >\n> > 4.17s user 0.23s system 109% cpu 4.025 total\n> >\n> > time git cat-file --unordered --batch-all-objects --batch\n> > --type-filter=blob >/dev/null\n> >\n> > 3.84s user 0.17s system 97% cpu 4.099 total\n> >\n> > It looks like the difference is not significant either.\n>\n> OK, good, that means we can probably not worry about it. :)\n>\n> > > In general, I do think having a processing pipeline like this is OK, as\n> > > it's pretty flexible. But especially for smaller queries (even ones that\n> > > don't ask for the whole object contents), the per-object lookup costs\n> > > can start to dominate (especially in a repository that hasn't been\n> > > recently packed). Right now, even your \"--batch --type-filter\" example\n> > > is probably making at least two lookups per object, because we don't\n> > > have a way to open a \"handle\" to an object to check its type, and then\n> > > extract the contents conditionally. And of course with multiple\n> > > processes, we're naturally doing a separate lookup in each one.\n> > >\n> >\n> > Yes, the type of the object is encapsulated in the header of the loose\n> > object file or the object entry header of the pack file. We have to read\n> > it to get the object type. This may be a lingering question I have had:\n> > why does git put the type/size in the file data instead of storing it as some\n> > kind of metadata elsewhere?\n>\n> It's not just metadata; it's actually part of what we hash to get the\n> object id (though of course it doesn't _have_ to be stored in a linear\n> buffer, as the pack storage shows).\n\nI'm still puzzled why git calculated the object id based on {type, size, data}\n together instead of just {data}?\n\n> But for loose objects, where would\n> such metadata be? And accessing it isn't too expensive; we only zlib\n> inflate the first few bytes (the main cost is in the syscalls to find\n> and open the file).\n>\n\nI may not have a lot of experience with this here. It looks like I should\ngo ahead and do some performance testing to compare the cost of searching\nand opening loose objects v.s reading and inflating loose objects.\n\n> For packed object, it effectively is metadata, just stuck at the front\n> of the object contents, rather than in a separate table. That lets us\n> use the same .idx file for finding that metadata as we do for the\n> contents themselves (at the slight cost that if you're _just_ accessing\n> metadata, the results are sparser within the file, which has worse\n> behavior for cold-cache disks).\n>\n\nAgree. But what if there is a metadata table in the .idx file?\nWe can even know the type and size of the object without accessing\nthe packfile.\n\n> But when I say that lookup costs dominate, what I mean is that we'd\n> spend a lot of our time binary searching within the pack .idx file, or\n> falling back to syscalls to look for loose objects.\n>\n\nAlright, binary search in .idx may indeed be more time-consuming than\nreading type and size from the packfile.\n\n> > > So a nice thing about being able to do the filtering in one process is\n> > > that we could _eventually_ do it all with one object lookup. But I'd\n> > > probably wait on adding something like --type-filter until we have an\n> > > internal single-lookup API, and then we could time it to see how much\n> > > speedup we can get.\n> >\n> > I am highly skeptical of this \"internal single-lookup API\". Do we really\n> > need an extra metadata table to record all objects?\n> > Something like: metadata: {oid: type, size}?\n>\n> No, I don't mean changing the storage at all. I mean that rather than\n> doing this:\n>\n>   /* get type, size, etc, for --batch format */\n>   type = oid_object_info(&oid, &size);\n>\n>   /* now get the contents for --batch to write them itself; but note\n>    * that this searches for the entry again within all packs, etc */\n>   contents = read_object_file(oid, &type, &size);\n>\n> as the cat-file code now does (because the first call is in\n> batch_object_write(), and the latter in print_object_or_die()), they\n> could be a single call that does the lookup once.\n>\n> We could actually do that today, since the object contents are\n> eventually fed from oid_object_info_extended(), and we know ahead of\n> time that we want both the metadata and the contents. But that wouldn't\n> work if we filtered by type, etc.\n>\n\nSo what you mentioned earlier about single read refers to combining\nthe two read operations of getting type size and getting content into one,\nwhen we know exactly that we need to retrieve the content.(in order to reduce\nthe overhead of the binary search once).\n\n> I'm not sure how much of a speedup it would yield in practice, though.\n> If you're printing the object contents, then the extra lookup is\n> probably not that expensive by comparison.\n>\n\nI feel like this solution may not be feasible. After we get the type and size\nfor the first time, we go through different output processes for different types\nof objects: use `stream_blob()` for blobs, and `read_object_file()` with\n`batch_write()` for other objects. If we obtain the content of a blob in one\nsingle read operation, then the performance optimization provided by\n`stream_blob()` would be invalidated.\n\n> -Peff\n"},{"id":"475345","messageId":"20230414073035.GB540206@coredump.intra.peff.net","threadId":"59561","inReplyTo":"CAOLTT8Rw796zxMYxg5+nx8+YoQVnfy=nPXH8Aq0j0Cw+GLT1rA@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2023-04-14T07:30:35Z","receivedAt":"2023-04-14T07:30:40Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 12, 2023 at 05:57:02PM +0800, ZheNing Hu wrote:\n\n> > It's not just metadata; it's actually part of what we hash to get the\n> > object id (though of course it doesn't _have_ to be stored in a linear\n> > buffer, as the pack storage shows).\n> \n> I'm still puzzled why git calculated the object id based on {type, size, data}\n>  together instead of just {data}?\n\nYou'd have to ask Linus for the original reasoning. ;)\n\nBut one nice thing about including these, especially the type, in the\nhash, is that the object id gives the complete context for an object.\nSo if another object claims to point to a tree, say, and points to a blob\ninstead, we can detect that problem immediately.\n\nOr worse, think about something like \"git show 1234abcd\". If the\nmetadata was not part of the object, then how would we know if you\nwanted to show a commit, or a blob (that happens to look like a commit),\netc? That metadata could be carried outside the hash, but then it has to\nbe stored somewhere, and is subject to ending up mismatched to the\ncontents. Hashing all of it (including the size) makes consistency\nchecking much easier.\n\n> > For packed object, it effectively is metadata, just stuck at the front\n> > of the object contents, rather than in a separate table. That lets us\n> > use the same .idx file for finding that metadata as we do for the\n> > contents themselves (at the slight cost that if you're _just_ accessing\n> > metadata, the results are sparser within the file, which has worse\n> > behavior for cold-cache disks).\n> \n> Agree. But what if there is a metadata table in the .idx file?\n> We can even know the type and size of the object without accessing\n> the packfile.\n\nI'm not sure it would be any faster than accessing the packfile. If you\nstick the metadata in the .idx file's oid lookup table, then lookups\nperform a bit worse because you're wasting memory cache. If you make a\nseparate table in the .idx file that's OK, but I'm not sure it's\nconsistently better than finding the data in the packfile.\n\nThe oid lookup table gives you a way to index the table in\nconstant-time (if you store the table as fixed-size entries in sha1\norder), but we can also access the packfile in constant-time (the idx\ntable gives us offsets). The idx metadata table would have better cache\nbehavior if you're only looking at metadata, and not contents. But\notherwise it's worse (since you have to hit the packfile, too). And I\ncheated a bit to say \"fixed-size\" above; the packfile metadata is in a\nvariable-length encoding, so in some ways it's more efficient.\n\nSo I doubt it would make any operations appreciably faster, and even if\nit did, you'd possibly be trading off versus other operations. I think\nthe more interesting metadata is not type/size, but properties such as\nthose stored by the commit graph. And there we do have separate tables\nfor fast access (and it's a _lot_ faster, because it's helping us avoid\ninflating the object contents).\n\n> > I'm not sure how much of a speedup it would yield in practice, though.\n> > If you're printing the object contents, then the extra lookup is\n> > probably not that expensive by comparison.\n> >\n> \n> I feel like this solution may not be feasible. After we get the type and size\n> for the first time, we go through different output processes for different types\n> of objects: use `stream_blob()` for blobs, and `read_object_file()` with\n> `batch_write()` for other objects. If we obtain the content of a blob in one\n> single read operation, then the performance optimization provided by\n> `stream_blob()` would be invalidated.\n\nGood point. So yeah, even to use it in today's code you'd need something\nconditional. A few years ago I played with an option for object_info\nthat would let the caller say \"please give me the object contents if\nthey are smaller than N bytes, otherwise don't\".\n\nAnd that would let many call-sites get type, size, and content together\nmost of the time (for small objects), and then stream only when\nnecessary. I still have the patches, and running them now it looks like\nthere's about a 10% speedup running:\n\n  git cat-file --unordered --batch-all-objects --batch >/dev/null\n\nOther code paths dealing with blobs would likewise get a small speedup,\nI'd think. I don't remember why I didn't send it. I think there was some\nugly refactoring that I needed to double-check, and my attention just\ngot pulled elsewhere. The messy patches are at:\n\n  https://github.com/peff/git jk/object-info-round-trip\n\nif you're interested.\n\n-Peff\n"},{"id":"475355","messageId":"CAOLTT8SEeY1tfU39xHPJ21F7o3dmgEFwNCny=Z2F4Y2HFR3DzA@mail.gmail.com","threadId":"59561","inReplyTo":"20230414073035.GB540206@coredump.intra.peff.net","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-14T12:17:34Z","receivedAt":"2023-04-14T12:17:25Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Jeff King <peff@peff.net> 于2023年4月14日周五 15:30写道：\n>\n> On Wed, Apr 12, 2023 at 05:57:02PM +0800, ZheNing Hu wrote:\n>\n> > > It's not just metadata; it's actually part of what we hash to get the\n> > > object id (though of course it doesn't _have_ to be stored in a linear\n> > > buffer, as the pack storage shows).\n> >\n> > I'm still puzzled why git calculated the object id based on {type, size, data}\n> >  together instead of just {data}?\n>\n> You'd have to ask Linus for the original reasoning. ;)\n>\n> But one nice thing about including these, especially the type, in the\n> hash, is that the object id gives the complete context for an object.\n> So if another object claims to point to a tree, say, and points to a blob\n> instead, we can detect that problem immediately.\n>\n> Or worse, think about something like \"git show 1234abcd\". If the\n> metadata was not part of the object, then how would we know if you\n> wanted to show a commit, or a blob (that happens to look like a commit),\n> etc? That metadata could be carried outside the hash, but then it has to\n> be stored somewhere, and is subject to ending up mismatched to the\n> contents. Hashing all of it (including the size) makes consistency\n> checking much easier.\n>\n\nOh, you are right, this could be to prevent conflicts between Git objects\nwith identical content but different types. However, I always associate\nGit with the file system, where metadata such as file type and size is\nstored in the inode, while the file data is stored in separate chunks.\n\n> > > For packed object, it effectively is metadata, just stuck at the front\n> > > of the object contents, rather than in a separate table. That lets us\n> > > use the same .idx file for finding that metadata as we do for the\n> > > contents themselves (at the slight cost that if you're _just_ accessing\n> > > metadata, the results are sparser within the file, which has worse\n> > > behavior for cold-cache disks).\n> >\n> > Agree. But what if there is a metadata table in the .idx file?\n> > We can even know the type and size of the object without accessing\n> > the packfile.\n>\n> I'm not sure it would be any faster than accessing the packfile. If you\n> stick the metadata in the .idx file's oid lookup table, then lookups\n> perform a bit worse because you're wasting memory cache. If you make a\n> separate table in the .idx file that's OK, but I'm not sure it's\n> consistently better than finding the data in the packfile.\n>\n\nYes, but it maybe be very convenient if we need to filter by object\ntype or size.\n\n> The oid lookup table gives you a way to index the table in\n> constant-time (if you store the table as fixed-size entries in sha1\n> order), but we can also access the packfile in constant-time (the idx\n> table gives us offsets). The idx metadata table would have better cache\n> behavior if you're only looking at metadata, and not contents. But\n> otherwise it's worse (since you have to hit the packfile, too). And I\n> cheated a bit to say \"fixed-size\" above; the packfile metadata is in a\n> variable-length encoding, so in some ways it's more efficient.\n>\n\nYes, if we only use git cat-file --batch-check, we may be able to improve\nperformance by avoiding access to the pack file. Additionally, I think\nthis metadata table is very suitable for filtering and aggregating operations.\n\n> So I doubt it would make any operations appreciably faster, and even if\n> it did, you'd possibly be trading off versus other operations. I think\n> the more interesting metadata is not type/size, but properties such as\n> those stored by the commit graph. And there we do have separate tables\n> for fast access (and it's a _lot_ faster, because it's helping us avoid\n> inflating the object contents).\n>\n\nYeah, optimizing the retrieval of metadata such as type/size may not provide\nas much benefit as recording the commit properties in the metadata table,\nlike the commit graph optimization does.\n\n> > > I'm not sure how much of a speedup it would yield in practice, though.\n> > > If you're printing the object contents, then the extra lookup is\n> > > probably not that expensive by comparison.\n> > >\n> >\n> > I feel like this solution may not be feasible. After we get the type and size\n> > for the first time, we go through different output processes for different types\n> > of objects: use `stream_blob()` for blobs, and `read_object_file()` with\n> > `batch_write()` for other objects. If we obtain the content of a blob in one\n> > single read operation, then the performance optimization provided by\n> > `stream_blob()` would be invalidated.\n>\n> Good point. So yeah, even to use it in today's code you'd need something\n> conditional. A few years ago I played with an option for object_info\n> that would let the caller say \"please give me the object contents if\n> they are smaller than N bytes, otherwise don't\".\n>\n> And that would let many call-sites get type, size, and content together\n> most of the time (for small objects), and then stream only when\n> necessary. I still have the patches, and running them now it looks like\n> there's about a 10% speedup running:\n>\n>   git cat-file --unordered --batch-all-objects --batch >/dev/null\n>\n> Other code paths dealing with blobs would likewise get a small speedup,\n> I'd think. I don't remember why I didn't send it. I think there was some\n> ugly refactoring that I needed to double-check, and my attention just\n> got pulled elsewhere. The messy patches are at:\n>\n>   https://github.com/peff/git jk/object-info-round-trip\n>\n> if you're interested.\n>\n\nAlright, this does feel a bit hackish, allowing most objects to fetch the\ncontent when first read and allowing blobs larger than N to be\nstreamed via stream_blob().\n\nI feel like this optimization for single-reads is a bit off-topic, I quote\nyour previous sentence:\n\n> So a nice thing about being able to do the filtering in one process is\n> that we could _eventually_ do it all with one object lookup. But I'd\n> probably wait on adding something like --type-filter until we have an\n> internal single-lookup API, and then we could time it to see how much\n> speedup we can get.\n\nThis optimization for single-reads doesn't seem to provide much benefit\nfor implementing object filters, because we have already read the content\nof the object in advance?\n\nZheNing Hu\n"},{"id":"475387","messageId":"xmqqh6titpzk.fsf@gitster.g","threadId":"59561","inReplyTo":"CAOLTT8SEeY1tfU39xHPJ21F7o3dmgEFwNCny=Z2F4Y2HFR3DzA@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-04-14T15:58:39Z","receivedAt":"2023-04-14T15:58:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"ZheNing Hu <adlternative@gmail.com> writes:\n\n> Oh, you are right, this could be to prevent conflicts between Git objects\n> with identical content but different types. However, I always associate\n> Git with the file system, where metadata such as file type and size is\n> stored in the inode, while the file data is stored in separate chunks.\n\nI am afraid the presentation order Peff used caused a bit of\nconfusion.  The true reason is what Peff brought up as \"Or worse\".\nWe need to be able to tell, given only the name of an object,\neverything that we need to know about the object, and for that, we\nneed the type information when we ask for an object by its name.\nHaving size embedded in the data that comes back to us when we\nconsult object database with an object name helps the implementation\nto pre-allocate a buffer and then inflate into it--there is no\nfundamental reason why it should be there.\n\nIt is a secondary problem created by the design choice that we store\ntype together with contents, that the object type recorded in a tree\nentry may contradict the actual type of the object recorded in the\ntree entry.  We could have declared that the object type found in a\ntree entry is to be trusted, if we didn't record the type in the\nobject database together with the object contents.\n\nI think your original question was not \"why do we store type and\nsize together with the contents?\", but was \"why do we include in the\nhash computation?\", and all of the above discuss related tangent\nwithout touching the original question.\n\nThe need to have type or size available when we ask the object\ndatabase for data associated with the object does not necessarily\nmean they must be hashed together with the contents.  It was done\nmerely because \"why not? that way, we do not have to worry about\ncatching corrupt values for type and size information we want to\nstore together with the contents\".  IOW, we could have checksummed\nthese two pieces of information separately, but why bother?\n"},{"id":"475391","messageId":"CAHk-=wjr-CMLX2Jo2++rwcv0VNr+HmZqXEVXNsJGiPRUwNxzBQ@mail.gmail.com","threadId":"59561","inReplyTo":"CAOLTT8SEeY1tfU39xHPJ21F7o3dmgEFwNCny=Z2F4Y2HFR3DzA@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2023-04-14T17:04:59Z","receivedAt":"2023-04-14T17:05:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Apr 14, 2023 at 5:17 AM ZheNing Hu <adlternative@gmail.com> wrote:\n>\n> Jeff King <peff@peff.net> 于2023年4月14日周五 15:30写道：\n> >\n> > On Wed, Apr 12, 2023 at 05:57:02PM +0800, ZheNing Hu wrote:\n> > >\n> > > I'm still puzzled why git calculated the object id based on {type, size, data}\n> > >  together instead of just {data}?\n> >\n> > You'd have to ask Linus for the original reasoning. ;)\n\nI originally thought of the git object store as \"tagged pointers\".\n\nThat actually caused confusion initially when I tried to explain this\nto SCM people, because \"tag\" means something very different in an SCM\nenvironment than it means in computer architecture.\n\nAnd the implication of a tagged pointer is that you have two parts of\nit - the \"tag\" and the \"address\". Both are relevant at all points.\n\nThis isn't quite as obvious in everyday moden git usage, because a lot\nof uses end up _only_ using the \"address\" (aka SHA1), but it's very\nmuch part of the object store design. Internally, the object layout\nnever uses just the SHA1, it's all \"type:SHA1\", even if sometimes the\ntypes are implied (ie the tree object doesn't spell out \"blob\", but\nit's still explicit in the mode bits).\n\nThis is very very obvious in \"git cat-file\", which was one of the\noriginal scripts in the first commit (but even there the tag/type has\nchanged meaning over time: the very first version didn't use it as\ninput at all, then it started verifying it, and then later it got the\nmore subtle context of \"peel the tags until you find this type\").\n\nYou can also see this in the original README (again, go look at that\nfirst git commit): the README talks about the \"tag of their type\".\n\nOf course, in practice git then walked away from having to specify the\ntype all the time. It started even in that original release, in that\nthe HEAD file never contained the type - because it was implicit (a\nHEAD is always a commit).\n\nSo we ended up having a lot of situations like that where the actual\ntag part was implicit from context, and these days people basically\nnever refer to the \"full\" object name with tag, but only the SHA1\naddress.\n\nSo now we have situations where the type really has to be looked up\ndynamically, because it's not explicitly encoded anywhere. While HEAD\nis supposed to always be a commit, other refs can be pretty much\nanything, and can point to a tag object, a commit, a tree or a blob.\nSo then you actually have to look up the type based on the address.\n\nEnd result: these days people don't even think of git objects as\n\"tagged pointers\".  Even internally in git, lots of code just passes\nthe \"object name\" along without any tag/type, just the raw SHA1 / OID.\n\nSo that originally \"everything is a tagged pointer\" is much less true\nthan it used to be, and now, instead of having tagged pointers, you\nmostly end up with just \"bare pointers\" and look up the type\ndynamically from there.\n\nAnd that \"look up the type in the object\" is possible because even\noriginally, I did *not* want any kind of \"object type aliasing\".\n\nSo even when looking up the object with the full \"tag:pointer\", the\nencoding of the object itself then also contains that object type, so\nthat you can cross-check that you used the right tag.\n\nThat said, you *can* see some of the effects of this \"tagged pointers\"\nin how the internals do things like\n\n    struct commit *commit = lookup_commit(repo, &oid);\n\nwhich conceptually very much is about tagged pointers. And the fact\nthat two objects cannot alias is actually somewhat encoded in that: a\n\"struct commit\" contains a \"struct object\" as a member. But so does\n\"struct blob\" - and the two \"struct object\" cases are never the same\n\"object\".\n\nSo there's never any worry about \"could blob.object be the same object\nas commit.object\"?\n\nThat is actually inherent in the code, in how \"lookup_commit()\"\nactually does lookup_object() and then does object_as_type(OBJ_COMMIT)\non the result.\n\n> Oh, you are right, this could be to prevent conflicts between Git objects\n> with identical content but different types. However, I always associate\n> Git with the file system, where metadata such as file type and size is\n> stored in the inode, while the file data is stored in separate chunks.\n\nSee above: yes, git design was *also* influenced heavily by\nfilesystems, but that was mostly in the sense of \"this is how to\nencode these things without undue pain\".\n\nThe object database being immutable was partly a security and safety\nmeasure, but it was also very much partly a \"rewriting files is going\nto be a major pain from a filesystem consistency standpoint - don't do\nit\".\n\nBut even more than a filesystem design, it's an \"computer\narchitecture\" design. Think of the git object store as a very abstract\ncomputer architecture that has tagged pointers, stable storage, and no\naliasing - and where the tag is actually verified at each lookup.\n\nThe \"no aliasing\" means that no two distinct pointers can point to the\nsame data. So a tagged pointer of type \"commit\" can not point to the\nsame object as a tagged pointer of type \"blob\". They are distinct\npointers, even if (maybe) the commit object encoding ends up then\nbeing identical to a blob object.\n\nAnd as mentioned, that \"verified at each lookup\" has mostly gone away,\nand \"each lookup\" has become more of a \"can be verified by fsck\", but\nit's probably still a good thing to think that way.\n\nYou still have \"lookup_object_by_type()\" internally in git that takes\nthe full tagged pointer, but almost nobody uses it any more. The\nclosest you get is those \"lookup_commit()\" things (which are fairly\ncommon, still).\n\n              Linus\n"},{"id":"475484","messageId":"CAOLTT8QS8VzepLid7V4FXMfGJpiQL6P5_Bd+2=YygfNoZrPU7w@mail.gmail.com","threadId":"59561","inReplyTo":"xmqqh6titpzk.fsf@gitster.g","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-16T11:15:47Z","receivedAt":"2023-04-16T11:15:36Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Junio C Hamano <gitster@pobox.com> 于2023年4月14日周五 23:58写道：\n>\n> ZheNing Hu <adlternative@gmail.com> writes:\n>\n> > Oh, you are right, this could be to prevent conflicts between Git objects\n> > with identical content but different types. However, I always associate\n> > Git with the file system, where metadata such as file type and size is\n> > stored in the inode, while the file data is stored in separate chunks.\n>\n> I am afraid the presentation order Peff used caused a bit of\n> confusion.  The true reason is what Peff brought up as \"Or worse\".\n> We need to be able to tell, given only the name of an object,\n> everything that we need to know about the object, and for that, we\n> need the type information when we ask for an object by its name.\n> Having size embedded in the data that comes back to us when we\n> consult object database with an object name helps the implementation\n> to pre-allocate a buffer and then inflate into it--there is no\n> fundamental reason why it should be there.\n>\n\nYes, I think I understand the point now. Since Git addresses objects\nbased on their content, if type information is not included in the object,\nwe cannot easily understand what type of Git object corresponds to\na given object ID. Moreover, if we don't include type and size information\nin Git objects, We would need to maintain a large number of external tables\nto record this information, in order to inflate and identify the type.\n\n> It is a secondary problem created by the design choice that we store\n> type together with contents, that the object type recorded in a tree\n> entry may contradict the actual type of the object recorded in the\n> tree entry.  We could have declared that the object type found in a\n> tree entry is to be trusted, if we didn't record the type in the\n> object database together with the object contents.\n>\n\nYes, that may not be crucial, but including type information\nin Git objects can help validate the correctness of tree entries better.\n\n> I think your original question was not \"why do we store type and\n> size together with the contents?\", but was \"why do we include in the\n> hash computation?\", and all of the above discuss related tangent\n> without touching the original question.\n>\n\nYes, but I think these two problems should be similar.\n\n> The need to have type or size available when we ask the object\n> database for data associated with the object does not necessarily\n> mean they must be hashed together with the contents.  It was done\n> merely because \"why not? that way, we do not have to worry about\n> catching corrupt values for type and size information we want to\n> store together with the contents\".  IOW, we could have checksummed\n> these two pieces of information separately, but why bother?\n\n Thank you. I think I roughly understand.\n"},{"id":"475486","messageId":"643be4aa2bdf9_f43d294d9@chronos.notmuch","threadId":"59561","inReplyTo":"CAHk-=wjr-CMLX2Jo2++rwcv0VNr+HmZqXEVXNsJGiPRUwNxzBQ@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2023-04-16T12:06:02Z","receivedAt":"2023-04-16T12:06:07Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Linus Torvalds wrote:\n> On Fri, Apr 14, 2023 at 5:17 AM ZheNing Hu <adlternative@gmail.com> wrote:\n> >\n> > Jeff King <peff@peff.net> 于2023年4月14日周五 15:30写道：\n> > >\n> > > On Wed, Apr 12, 2023 at 05:57:02PM +0800, ZheNing Hu wrote:\n> > > >\n> > > > I'm still puzzled why git calculated the object id based on {type, size, data}\n> > > >  together instead of just {data}?\n> > >\n> > > You'd have to ask Linus for the original reasoning. ;)\n> \n> I originally thought of the git object store as \"tagged pointers\".\n> \n> That actually caused confusion initially when I tried to explain this\n> to SCM people, because \"tag\" means something very different in an SCM\n> environment than it means in computer architecture.\n> \n> And the implication of a tagged pointer is that you have two parts of\n> it - the \"tag\" and the \"address\". Both are relevant at all points.\n> \n> This isn't quite as obvious in everyday moden git usage, because a lot\n> of uses end up _only_ using the \"address\" (aka SHA1), but it's very\n> much part of the object store design. Internally, the object layout\n> never uses just the SHA1, it's all \"type:SHA1\", even if sometimes the\n> types are implied (ie the tree object doesn't spell out \"blob\", but\n> it's still explicit in the mode bits).\n> \n> This is very very obvious in \"git cat-file\", which was one of the\n> original scripts in the first commit (but even there the tag/type has\n> changed meaning over time: the very first version didn't use it as\n> input at all, then it started verifying it, and then later it got the\n> more subtle context of \"peel the tags until you find this type\").\n> \n> You can also see this in the original README (again, go look at that\n> first git commit): the README talks about the \"tag of their type\".\n> \n> Of course, in practice git then walked away from having to specify the\n> type all the time. It started even in that original release, in that\n> the HEAD file never contained the type - because it was implicit (a\n> HEAD is always a commit).\n> \n> So we ended up having a lot of situations like that where the actual\n> tag part was implicit from context, and these days people basically\n> never refer to the \"full\" object name with tag, but only the SHA1\n> address.\n> \n> So now we have situations where the type really has to be looked up\n> dynamically, because it's not explicitly encoded anywhere. While HEAD\n> is supposed to always be a commit, other refs can be pretty much\n> anything, and can point to a tag object, a commit, a tree or a blob.\n> So then you actually have to look up the type based on the address.\n> \n> End result: these days people don't even think of git objects as\n> \"tagged pointers\".  Even internally in git, lots of code just passes\n> the \"object name\" along without any tag/type, just the raw SHA1 / OID.\n> \n> So that originally \"everything is a tagged pointer\" is much less true\n> than it used to be, and now, instead of having tagged pointers, you\n> mostly end up with just \"bare pointers\" and look up the type\n> dynamically from there.\n> \n> And that \"look up the type in the object\" is possible because even\n> originally, I did *not* want any kind of \"object type aliasing\".\n> \n> So even when looking up the object with the full \"tag:pointer\", the\n> encoding of the object itself then also contains that object type, so\n> that you can cross-check that you used the right tag.\n> \n> That said, you *can* see some of the effects of this \"tagged pointers\"\n> in how the internals do things like\n> \n>     struct commit *commit = lookup_commit(repo, &oid);\n> \n> which conceptually very much is about tagged pointers. And the fact\n> that two objects cannot alias is actually somewhat encoded in that: a\n> \"struct commit\" contains a \"struct object\" as a member. But so does\n> \"struct blob\" - and the two \"struct object\" cases are never the same\n> \"object\".\n> \n> So there's never any worry about \"could blob.object be the same object\n> as commit.object\"?\n> \n> That is actually inherent in the code, in how \"lookup_commit()\"\n> actually does lookup_object() and then does object_as_type(OBJ_COMMIT)\n> on the result.\n\nThis explains rather well why the object type is used in the calculation, and\nit makes sense.\n\nBut I don't see anything about the object size. Isn't that unnecessary?\n\n-- \nFelipe Contreras"},{"id":"475487","messageId":"CAOLTT8S7i2CRz0Oo6u6qMD2BD9=3UaJ0C0iNeY1B+hF6H5S94A@mail.gmail.com","threadId":"59561","inReplyTo":"CAHk-=wjr-CMLX2Jo2++rwcv0VNr+HmZqXEVXNsJGiPRUwNxzBQ@mail.gmail.com","subject":"Re: [Question] Can git cat-file have a type filtering option?","fromName":"ZheNing Hu","fromEmail":"adlternative@gmail.com","sentAt":"2023-04-16T12:43:22Z","receivedAt":"2023-04-16T12:43:32Z","isPatch":false,"sender":{"key":"adlternative@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58138461?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> 于2023年4月15日周六 01:05写道：\n>\n> On Fri, Apr 14, 2023 at 5:17 AM ZheNing Hu <adlternative@gmail.com> wrote:\n> >\n> > Jeff King <peff@peff.net> 于2023年4月14日周五 15:30写道：\n> > >\n> > > On Wed, Apr 12, 2023 at 05:57:02PM +0800, ZheNing Hu wrote:\n> > > >\n> > > > I'm still puzzled why git calculated the object id based on {type, size, data}\n> > > >  together instead of just {data}?\n> > >\n> > > You'd have to ask Linus for the original reasoning. ;)\n>\n> I originally thought of the git object store as \"tagged pointers\".\n>\n> That actually caused confusion initially when I tried to explain this\n> to SCM people, because \"tag\" means something very different in an SCM\n> environment than it means in computer architecture.\n>\n> And the implication of a tagged pointer is that you have two parts of\n> it - the \"tag\" and the \"address\". Both are relevant at all points.\n>\n> This isn't quite as obvious in everyday moden git usage, because a lot\n> of uses end up _only_ using the \"address\" (aka SHA1), but it's very\n> much part of the object store design. Internally, the object layout\n> never uses just the SHA1, it's all \"type:SHA1\", even if sometimes the\n> types are implied (ie the tree object doesn't spell out \"blob\", but\n> it's still explicit in the mode bits).\n>\n> This is very very obvious in \"git cat-file\", which was one of the\n> original scripts in the first commit (but even there the tag/type has\n> changed meaning over time: the very first version didn't use it as\n> input at all, then it started verifying it, and then later it got the\n> more subtle context of \"peel the tags until you find this type\").\n>\n\nYes, in the initial commit of Git, \"git cat-file\" only needs to pass the\nobject ID to obtain both the content and type of the object. However,\nmodern \"git cat-file\" requires specifying both the expected object type\nand  object ID by default. e.g. \"git cat-file commit v2.9.1\". This should\nbe where you mentioned the simultaneous appearance of \"tag\" and\n\"address\". This design model may not be very user-friendly for users\nto use, so nowadays, people prefer to use \"git cat-file -p\", this may be\nvery similar to the initial version of git cat-file.\n\n> You can also see this in the original README (again, go look at that\n> first git commit): the README talks about the \"tag of their type\".\n>\n> Of course, in practice git then walked away from having to specify the\n> type all the time. It started even in that original release, in that\n> the HEAD file never contained the type - because it was implicit (a\n> HEAD is always a commit).\n>\n> So we ended up having a lot of situations like that where the actual\n> tag part was implicit from context, and these days people basically\n> never refer to the \"full\" object name with tag, but only the SHA1\n> address.\n>\n> So now we have situations where the type really has to be looked up\n> dynamically, because it's not explicitly encoded anywhere. While HEAD\n> is supposed to always be a commit, other refs can be pretty much\n> anything, and can point to a tag object, a commit, a tree or a blob.\n> So then you actually have to look up the type based on the address.\n>\n> End result: these days people don't even think of git objects as\n> \"tagged pointers\".  Even internally in git, lots of code just passes\n> the \"object name\" along without any tag/type, just the raw SHA1 / OID.\n>\n> So that originally \"everything is a tagged pointer\" is much less true\n> than it used to be, and now, instead of having tagged pointers, you\n> mostly end up with just \"bare pointers\" and look up the type\n> dynamically from there.\n>\n\nI feel that if type was not included in the objects initially, people would\nneed to specify both \"tag\" and \"address\" at the same time to explain\nthe objects. Otherwise, this \"tag\" can only be used for checking the type,\nand this is not necessary in most cases.\n\n> And that \"look up the type in the object\" is possible because even\n> originally, I did *not* want any kind of \"object type aliasing\".\n>\n> So even when looking up the object with the full \"tag:pointer\", the\n> encoding of the object itself then also contains that object type, so\n> that you can cross-check that you used the right tag.\n>\n> That said, you *can* see some of the effects of this \"tagged pointers\"\n> in how the internals do things like\n>\n>     struct commit *commit = lookup_commit(repo, &oid);\n>\n> which conceptually very much is about tagged pointers. And the fact\n> that two objects cannot alias is actually somewhat encoded in that: a\n> \"struct commit\" contains a \"struct object\" as a member. But so does\n> \"struct blob\" - and the two \"struct object\" cases are never the same\n> \"object\".\n>\n> So there's never any worry about \"could blob.object be the same object\n> as commit.object\"?\n>\n\nYes, if an object can be interpreted as multiple types, it will certainly\nmake it very difficult for git higher-level logic to handle it.\n\n> That is actually inherent in the code, in how \"lookup_commit()\"\n> actually does lookup_object() and then does object_as_type(OBJ_COMMIT)\n> on the result.\n>\n> > Oh, you are right, this could be to prevent conflicts between Git objects\n> > with identical content but different types. However, I always associate\n> > Git with the file system, where metadata such as file type and size is\n> > stored in the inode, while the file data is stored in separate chunks.\n>\n> See above: yes, git design was *also* influenced heavily by\n> filesystems, but that was mostly in the sense of \"this is how to\n> encode these things without undue pain\".\n>\n> The object database being immutable was partly a security and safety\n> measure, but it was also very much partly a \"rewriting files is going\n> to be a major pain from a filesystem consistency standpoint - don't do\n> it\".\n>\n\nYou're right. Git objects are immutable, while data in a file system\nis mutable, so Git's design doesn't need to follow the file system completely.\n\n> But even more than a filesystem design, it's an \"computer\n> architecture\" design. Think of the git object store as a very abstract\n> computer architecture that has tagged pointers, stable storage, and no\n> aliasing - and where the tag is actually verified at each lookup.\n>\n> The \"no aliasing\" means that no two distinct pointers can point to the\n> same data. So a tagged pointer of type \"commit\" can not point to the\n> same object as a tagged pointer of type \"blob\". They are distinct\n> pointers, even if (maybe) the commit object encoding ends up then\n> being identical to a blob object.\n>\n> And as mentioned, that \"verified at each lookup\" has mostly gone away,\n> and \"each lookup\" has become more of a \"can be verified by fsck\", but\n> it's probably still a good thing to think that way.\n>\n> You still have \"lookup_object_by_type()\" internally in git that takes\n> the full tagged pointer, but almost nobody uses it any more. The\n> closest you get is those \"lookup_commit()\" things (which are fairly\n> common, still).\n>\n\nWell, now I understand: everything in Git's architecture is \"tag:pointer\".\nTags are used to verify object types (although it's not necessary now),\nand pointers are used for addressing. This is also one of the reasons\nwhy Git initially included the type in its objects.\n\n>               Linus\n\nThank you for your wonderful answer regarding the design concept\nof \"tagged pointers\"! This deepens my understanding of Git's design. :-)\n\nZheNing Hu\n"}]}