git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Is GIT_DEFAULT_HASH flawed?

From
Felipe Contreras <felipe.contreras@gmail.com>
Date
May 3, 2023, 15:44 UTC
Message-ID
<64528158cdd1e_68229498@chronos.notmuch>
In-Reply-To
<70103746-6980-baed-13d9-afeae6cee464@zombino.com>
Adam Majer wrote:
Show 6 quoted lines
> On 5/3/23 01:46, Felipe Contreras wrote:
> > To be honest this whole approach seems to be completely flawed to me and
> > against the whole design of git in the first place.
> 
> The discussion above is mostly moot now since this has been fixed in 
> later patches in this thread, AFAIK.

That particular isssue might be fixed, but that issue should never have happened in the first place if the design was correct.

A bad design makes certain errors prone to happen, a good design makes the same errors happen rarely, a great design makes those errors impossible.

Git was designed to make it *impossible* to confuse two commits with similar data.

The symptom might have been fixed, that doesn't mean there's no underlying problem.

> It's also moot for other reasons, like the hash function transition plan is
> not really implemented, yet.
The implemention of the plan isn't the problem, it's the plan itself.
> Also, this was about corner-case, like it often is.
A corner-case that should be impossible.
Show 7 quoted lines
> > In a recent email Linus Torvalds explained why object ids were
> > calculated based {type, size, data} [1], and he explained very clearly
> > that two objects with exactly the same data are not supposed to have the
> > same id if the type is different.
> 
> This is different. But aside, type + size + data are not really much 
> different from just having data in a hash function.
It's completely different.
Show 6 quoted lines
> There are plenty of hash collisions where
> 
>      HASH(type + size + data) == HASH(type + size + data')
> 
> by definition of how these functions work. The problem is always in 
> finding these collisions. But anyway...
I don't think you understand why Linus Torvalds chose to hash objects.
> > In my view one repository should be able to have part SHA-1 history,
> > part SHA3-256 history, and part BLAKE2b history.
> 
> Yes, that would be great. Please provide patch series for this :-)

I have hundreds of patches being ignored, why would I write yet another patch series that will be ignored?

Show 7 quoted lines
> > I have not been following the SHA-1 -> OID discussions, but I
> > distinctively recall Linus Torvalds mentioning that the choice of using
> > SHA-1 wasn't even for security purposes, it was to ensure integrity.
> 
> These are different sides of the same coin. Hashes are used to provide 
> integrity. Hashes like MD4, MD5, SHA1, SHA256 are there for integrity. 
> Some of these are no longer recommended and some are completely broken.

There are different philosophical views of what "security" means, and it seems pretty clear to me that your view does not align with the view of Linus Torvalds.

> > Better the SHA-1 you know, than the SHA-256 you don't.
> 
> Wrong conclusion ;) Also, we know SHA-256
You don't understand what is being said.
Which hash is more trustworthy?
 a. 69c786637d7a7fe3b2b8f7d989af095f5f49c3a8
 b. d891b12414e1d9331f8cbb15acfe690671974f27ba76e2b423294cfb7a055f2f
If you answer b just beacuse it's SHA-256 you don't understand security.
b is a random commit I generated, a is the current git.git master.

A SHA-1 hash from a source you trust is inifinitely more trustworthy than a random SHA-256 hash. Even a known MD5 hash is better in this respect.

> Keep in mind -- hashes are there for object reference.
No. I don't think you understand why Linus Torvalds used hashes.
> If you have SHA1 repo, you can calculate a SHA256 or whatever hash for any
> type object.

I know it *can* be done, I understand how hash algorithms work, but just because something *can* be done doesn't mean it *should*.

You *can* generate a SHA-1 of a blob's data, instead of a SHA-1 of a blob's `type + size + data`, does that mean we should? No.

> Finally, let not have a "bike shed" discussion about this.

Discussing the original design of git's object storage which has withstood the test of time for 18 years is not "bike sheding".

I don't even think you understand what I'm trying to say.
---
Why do you think these commands generate different hashes?
  git hash-object -t blob /dev/null
  git hash-object -t tree /dev/null
-- 
Felipe Contreras
Previous: Adam MajerNext: Adam Majer
Message 35 of 58 in “git clone of empty repositories doesn't preserve hash”
  1. Adam MajerApr 5, 2023
  2. Junio C HamanoApr 5, 2023
  3. Adam MajerApr 5, 2023
  4. Jeff KingApr 5, 2023
  5. Junio C HamanoApr 5, 2023
  6. Junio C HamanoApr 5, 2023
  7. Jeff KingApr 5, 2023
  8. brian m. carlsonApr 5, 2023
  9. Adam MajerApr 6, 2023
  10. brian m. carlsonApr 25, 2023
  11. Junio C HamanoApr 25, 2023
  12. Junio C HamanoApr 25, 2023
  13. brian m. carlsonApr 26, 2023
  14. Jeff KingApr 26, 2023
  15. Junio C HamanoApr 26, 2023
  16. doc: GIT_DEFAULT_HASH is and will be ignored during "clone"Junio C Hamano, Apr 26, 2023
  17. brian m. carlsonApr 26, 2023
  18. Jeff KingApr 27, 2023
  19. Jeff KingApr 26, 2023
  20. Junio C HamanoApr 26, 2023
  21. brian m. carlsonApr 26, 2023
  22. 0/2 Fix empty SHA-256 clones with v0 and v1brian m. carlson, Apr 26, 2023
  23. 1/2 http: advertise capabilities when cloning empty reposbrian m. carlson, Apr 26, 2023
  24. Junio C HamanoApr 26, 2023
  25. brian m. carlsonApr 26, 2023
  26. Jeff KingApr 27, 2023
  27. Jeff KingApr 27, 2023
  28. Junio C HamanoApr 27, 2023
  29. 2/2 Honor GIT_DEFAULT_HASH for empty clones without remote algobrian m. carlson, Apr 26, 2023
  30. Junio C HamanoApr 26, 2023
  31. Junio C HamanoApr 26, 2023
  32. Jeff KingApr 27, 2023
  33. Is GIT_DEFAULT_HASH flawed?Felipe Contreras, May 2, 2023
  34. Adam MajerMay 3, 2023
  35. Felipe ContrerasMay 3, 2023
  36. Adam MajerMay 3, 2023
  37. Felipe ContrerasMay 8, 2023
  38. demerphqMay 3, 2023
  39. Felipe ContrerasMay 3, 2023
  40. brian m. carlsonMay 3, 2023
  41. Felipe ContrerasMay 8, 2023
  42. brian m. carlsonMay 8, 2023
  43. Oswald BuddenhagenMay 9, 2023
  44. Junio C HamanoMay 9, 2023
  45. Junio C HamanoApr 26, 2023
  46. Jeff KingApr 27, 2023
  47. 0/1 Fix empty SHA-256 clones with v0 and v1brian m. carlson, May 1, 2023
  48. 1/1 upload-pack: advertise capabilities when cloning empty reposbrian m. carlson, May 1, 2023
  49. Jeff KingMay 1, 2023
  50. Junio C HamanoMay 1, 2023
  51. Junio C HamanoMay 1, 2023
  52. 0/1 Fix empty SHA-256 clones with v0 and v1brian m. carlson, May 17, 2023
  53. 1/1 upload-pack: advertise capabilities when cloning empty reposbrian m. carlson, May 17, 2023
  54. Junio C HamanoMay 17, 2023
  55. brian m. carlsonMay 17, 2023
  56. Jeff KingMay 18, 2023
  57. brian m. carlsonMay 19, 2023
  58. Jeff KingApr 5, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.