git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [Question] Can git cat-file have a type filtering option?

From
ZheNing Hu <adlternative@gmail.com>
Date
Apr 16, 2023, 11:15 UTC
Message-ID
<CAOLTT8QS8VzepLid7V4FXMfGJpiQL6P5_Bd+2=YygfNoZrPU7w@mail.gmail.com>
In-Reply-To
<xmqqh6titpzk.fsf@gitster.g>
Junio C Hamano <gitster@pobox.com> 于2023年4月14日周五 23:58写道:
Show 18 quoted lines
>
> ZheNing Hu <adlternative@gmail.com> writes:
>
> > Oh, you are right, this could be to prevent conflicts between Git objects
> > with identical content but different types. However, I always associate
> > Git with the file system, where metadata such as file type and size is
> > stored in the inode, while the file data is stored in separate chunks.
>
> I am afraid the presentation order Peff used caused a bit of
> confusion.  The true reason is what Peff brought up as "Or worse".
> We need to be able to tell, given only the name of an object,
> everything that we need to know about the object, and for that, we
> need the type information when we ask for an object by its name.
> Having size embedded in the data that comes back to us when we
> consult object database with an object name helps the implementation
> to pre-allocate a buffer and then inflate into it--there is no
> fundamental reason why it should be there.
>

Yes, I think I understand the point now. Since Git addresses objects based on their content, if type information is not included in the object, we cannot easily understand what type of Git object corresponds to a given object ID. Moreover, if we don't include type and size information in Git objects, We would need to maintain a large number of external tables to record this information, in order to inflate and identify the type.

Show 7 quoted lines
> It is a secondary problem created by the design choice that we store
> type together with contents, that the object type recorded in a tree
> entry may contradict the actual type of the object recorded in the
> tree entry.  We could have declared that the object type found in a
> tree entry is to be trusted, if we didn't record the type in the
> object database together with the object contents.
>

Yes, that may not be crucial, but including type information in Git objects can help validate the correctness of tree entries better.

Show 5 quoted lines
> I think your original question was not "why do we store type and
> size together with the contents?", but was "why do we include in the
> hash computation?", and all of the above discuss related tangent
> without touching the original question.
>
Yes, but I think these two problems should be similar.
Show 7 quoted lines
> The need to have type or size available when we ask the object
> database for data associated with the object does not necessarily
> mean they must be hashed together with the contents.  It was done
> merely because "why not? that way, we do not have to worry about
> catching corrupt values for type and size information we want to
> store together with the contents".  IOW, we could have checksummed
> these two pieces of information separately, but why bother?
 Thank you. I think I roughly understand.
Previous: Junio C HamanoNext: Linus Torvalds
Message 18 of 23 in “[Question] Can git cat-file have a type filtering option?”
  1. ZheNing HuApr 7, 2023
  2. Junio C HamanoApr 7, 2023
  3. ZheNing HuApr 8, 2023
  4. Taylor BlauApr 9, 2023
  5. Taylor BlauApr 9, 2023
  6. Taylor BlauApr 9, 2023
  7. ZheNing HuApr 9, 2023
  8. Jeff KingApr 10, 2023
  9. Taylor BlauApr 10, 2023
  10. ZheNing HuApr 9, 2023
  11. Jeff KingApr 10, 2023
  12. ZheNing HuApr 11, 2023
  13. Jeff KingApr 12, 2023
  14. ZheNing HuApr 12, 2023
  15. Jeff KingApr 14, 2023
  16. ZheNing HuApr 14, 2023
  17. Junio C HamanoApr 14, 2023
  18. ZheNing HuApr 16, 2023
  19. Linus TorvaldsApr 14, 2023
  20. Felipe ContrerasApr 16, 2023
  21. ZheNing HuApr 16, 2023
  22. Taylor BlauApr 9, 2023
  23. Taylor BlauApr 9, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.