git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [doc] User Manual Suggestion

From
Björn Steinbrink <b.steinbrink@gmx.de>
Date
May 2, 2009, 21:11 UTC
Message-ID
<20090502211110.GC6135@atjola.homenet>
In-Reply-To
<b4087cc50905021136l5209777bs2209bab385deeef6@mail.gmail.com>
On 2009.05.02 13:36:35 -0500, Michael Witten wrote:
Show 9 quoted lines
> 2009/5/2 Björn Steinbrink <B.Steinbrink@gmx.de>:
> >> As I've stated: "address", "pointer", and "handle" are an analogy to
> >> terminology that has been around for ages. In fact, another name for
> >> "pointer" is "reference".
> >
> > AFAIK a pointer is just one kind of reference. C++ references are
> > another kind...
> 
> Actually, a C++ reference is a pointer with restrictions (AFAIK).

I'm not really aware of what the C++ standard says about it, but from a usage point of view, they're IMHO different enough to consider them as truly different types of references.

Show 5 quoted lines
> > And there are probably plenty of examples where you could apply that
> > analogy, yet nobody (I know) does. Arrays, database tables, ...
> 
> Well, this terminology is certainly used with arrays in C, because
> array elements can be accessed with pointers.

But when you apply the analogy, then the array is the memory, and an integer is an address and an index variable is a pointer.

> Also, databases use a much different scheme for addressing information
> than does memory.

I don't see any inherent problem in saying that the primary key determines the address of a row. (It just gets funny when you have a table schema without a primary key *g*)

Show 7 quoted lines
> > And "memory" usually means "RAM" to me, not "WORM"-memory (well,
> > actually, you can also delete and then rewrite, but not modify).
> 
> Well, I don't see how Random Access Memory really conflicts. One
> certainly can access objects in the object memory/store randomly. The
> main difference is that the computer store is addressed by location,
> wheras the git store is addressed by content.

When I have a (non const) pointer in C I can write to the memory location it references. With git, I can't do that. ("RWRAM" would have been more correct, I'm damaged by the common usage of RAM as meaning RWRAM).

> Also, I would say that conceptually deletion is an implementation
> detail.

Yeah, thus I put it in parentheses, just to show that, in practise, we don't even have WORM-memory (but still taking the hash collision problem into account, so we need to write once).

Show 8 quoted lines
> Because git's object store is content addressable, one could
> think of it as already containing all possible objects (of course, I'm
> assuming that the 160-bit hash is also an implementation detail; an
> infinite number of objects implies infinitely large addresses, though
> the nonsignificant zeros could be disregarded as with real numbers or
> something. I don't know, I'm making this up as I go :-D). That the git
> tools ever complain no such object exists is an implementation detail
> resulting from our finite storage in reality.

I prefer to take the hash collision into account when looking at things like that, but yeah, one could look at it like that.

Show 9 quoted lines
> > So the analogy would even hurt my mental model (just like the
> > "commit --amend" command might be consider harmful, because it
> > actually creates a new commit, but some users actually think the
> > original commit is modified).
> 
> Actually, this is why it's so important to have the underlying
> concepts at hand. Understanding that objects are simply addressed by
> content (that is, objects are immutable) completely extirpates this
> kind of confusion.

I never disagreed with that, though I put more emphasis on the plain object relationships and their immutability than on the fact that hashes are used. Having that part right (how objects work together to form history) is a large part of what you need to understand all the rest.

Show 18 quoted lines
> >> >> So, a pointer variable's value is an object address that is the
> >> >> location of an object in git 'memory'. I think using this approach
> >> >> would make things significantly more transparent.
> >> >
> >> > But then HEAD would be a pointer pointer variable (symbolic ref), unless
> >> > you have a detached HEAD.
> >>
> >> We call those handles.
> >
> > Isn't a handle basically an opaque/abstract reference, at least in
> > "modern" usage? Symvolic references aren't. The user is free to create
> > and manipulate them, and gets full access to the things referenced by
> > them. And saying that HEAD is a reference, that might be symbolic is
> > IMHO by far easier to understand than saying that HEAD might be a
> > pointer or a handle.
> 
> Fair enough. Call them symbolic pointers; however, I don't really see
> the problem with pointer pointers.

You called them handles anyway ;-) But seriously, it's that "pointer" triggers C for me. And having an entity that can switch between being a pointer and a pointer pointer needs casting or a union (or a struct if you want to), not something I'd like to have to think about in my mental model of git.

Show 8 quoted lines
> In any case, I *think* my point is that it's important to understand
> that git uses content addressing; at first I was emphatic about the
> idea of 'addressing', so I went with pointer terminology (which works
> quite well, in my opinion). However, I think the 'content' part is
> more important, which is why 'object hash' is loads better than
> 'object name' or 'object id'. Also, at least the documentation could
> say that 'objects are addressed by their hashes', which says a whole
> lot in one quick sentence about how git works.
Hm, like chapter 7 "Git concepts"?
>>>>>>
The Object Database

We already saw in the section called “Understanding History: Commits” that all commits are stored under a 40-digit "object name". In fact, all the information needed to represent the history of a project is stored in objects with such names. In each case the name is calculated by taking the SHA-1 hash of the contents of the object. The SHA-1 hash is a cryptographic hash function. What that means to us is that it is impossible to find two different objects with the same name. This has a number of advantages; among others:

    * Git can quickly determine whether two objects are identical or
      not, just by comparing names.
    * Since object names are computed the same way in every repository,
      the same content stored in two repositories will always be stored
      under the same name.
    * Git can detect errors when it reads an object, by checking that
      the object's name is still the SHA-1 hash of its contents.
<<<<<<
Björn
Previous: Michael WittenNext: Michael Witten
Message 66 of 90 in “[doc] User Manual Suggestion”
  1. David AbrahamsApr 22, 2009
  2. J. Bruce FieldsApr 23, 2009
  3. Michael WittenApr 23, 2009
  4. Jeff KingApr 23, 2009
  5. Michael WittenApr 23, 2009
  6. David AbrahamsApr 23, 2009
  7. Michael WittenApr 24, 2009
  8. Jeff KingApr 24, 2009
  9. J. Bruce FieldsApr 24, 2009
  10. David AbrahamsApr 24, 2009
  11. Jeff KingApr 24, 2009
  12. David AbrahamsApr 24, 2009
  13. Jeff KingApr 24, 2009
  14. David AbrahamsApr 24, 2009
  15. Björn SteinbrinkApr 24, 2009
  16. David AbrahamsApr 25, 2009
  17. Björn SteinbrinkApr 26, 2009
  18. Jeff KingApr 24, 2009
  19. Michael WittenApr 24, 2009
  20. Michael WittenApr 24, 2009
  21. Jeff KingApr 24, 2009
  22. Michael WittenApr 24, 2009
  23. J. Bruce FieldsApr 24, 2009
  24. Jeff KingApr 24, 2009
  25. J. Bruce FieldsApr 24, 2009
  26. Michael WittenApr 24, 2009
  27. David AbrahamsApr 23, 2009
  28. Johan HerlandApr 23, 2009
  29. Michael WittenApr 24, 2009
  30. Johan HerlandApr 24, 2009
  31. Daniel BarkalowApr 24, 2009
  32. Jeff KingApr 24, 2009
  33. Michael WittenApr 24, 2009
  34. Michael WittenApr 24, 2009
  35. Daniel BarkalowApr 24, 2009
  36. Jeff KingApr 24, 2009
  37. Michael WittenApr 24, 2009
  38. Michael WittenApr 24, 2009
  39. Jeff KingApr 24, 2009
  40. Michael WittenApr 25, 2009
  41. Felipe ContrerasApr 25, 2009
  42. Michael WittenApr 24, 2009
  43. Daniel BarkalowApr 25, 2009
  44. Michael WittenApr 25, 2009
  45. Felipe ContrerasApr 25, 2009
  46. David AbrahamsApr 25, 2009
  47. Felipe ContrerasApr 25, 2009
  48. Björn SteinbrinkApr 26, 2009
  49. David AbrahamsApr 26, 2009
  50. Björn SteinbrinkApr 26, 2009
  51. David AbrahamsApr 26, 2009
  52. Björn SteinbrinkApr 26, 2009
  53. David AbrahamsApr 27, 2009
  54. David AbrahamsApr 27, 2009
  55. Michael WittenApr 27, 2009
  56. Michael WittenApr 26, 2009
  57. Björn SteinbrinkApr 26, 2009
  58. David AbrahamsApr 26, 2009
  59. David AbrahamsApr 25, 2009
  60. Björn SteinbrinkApr 24, 2009
  61. Michael WittenApr 25, 2009
  62. David AbrahamsApr 25, 2009
  63. Björn SteinbrinkApr 26, 2009
  64. Björn SteinbrinkMay 2, 2009
  65. Michael WittenMay 2, 2009
  66. Björn SteinbrinkMay 2, 2009
  67. Michael WittenMay 2, 2009
  68. Björn SteinbrinkMay 2, 2009
  69. Michael WittenMay 3, 2009
  70. Björn SteinbrinkMay 3, 2009
  71. Mark LodatoMay 3, 2009
  72. Michael WittenMay 3, 2009
  73. Daniel BarkalowApr 24, 2009
  74. Jeff KingApr 24, 2009
  75. Björn SteinbrinkApr 26, 2009
  76. Michael WittenApr 24, 2009
  77. Björn SteinbrinkApr 27, 2009
  78. David AbrahamsApr 25, 2009
  79. Michael WittenApr 25, 2009
  80. Jeff KingApr 25, 2009
  81. David AbrahamsApr 25, 2009
  82. Jeff KingApr 29, 2009
  83. David AbrahamsApr 29, 2009
  84. Jeff KingApr 29, 2009
  85. J. Bruce FieldsApr 24, 2009
  86. Michael WittenApr 24, 2009
  87. David AbrahamsApr 24, 2009
  88. J. Bruce FieldsApr 24, 2009
  89. J. Bruce FieldsApr 24, 2009
  90. Felipe ContrerasApr 25, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.