threads / discuss / 1556

First stab at glossary

Subject: First stab at glossary

## tl;dr

17 messages between Aug 17, 2005 and Sep 4, 2005.

replies: 16people: 4as markdown or json

Johannes Schindelin· Aug 17, 2005, 14:56 UTC · lore
Hi,

long, long time. Here´s my first stab at the glossary, attached the alphabetically sorted, asciidoc marked up txt file (Comments? Suggestions? Pizzas?):

object::
	The unit of storage in GIT. It is uniquely identified by
	the SHA1 of its contents. Consequently, an object can not
	be changed.
SHA1::
	A 20-byte sequence (or 41-byte file containing the hex
	representation and a newline). It is calculated from the
	contents of an object by the Secure Hash Algorithm 1.
object database::
	Stores a set of "objects", and an individial object is identified
	by its SHA1 (its ref). The objects are either stored as single
	files, or live inside of packs.
object name::
	Synonym for SHA1.
blob object::
	Untyped object, i.e. the contents of a file.
tree object::
	An object containing a list of blob and/or tree objects.
	(A tree usually corresponds to a directory without
	subdirectories).
tree::
	Either a working tree, or a tree object together with the
	dependent blob and tree objects (i.e. a stored representation
	of a working tree).
cache::
	A collection of files whose contents are stored as objects.
	The cache is a stored version of your working tree. Well, can
	also contain a second, and even a third version of a working
	tree, which are used when merging.
cache entry::
	The information regarding a particular file, stored in the index.
	A cache entry can be unmerged, if a merge was started, but not
	yet finished (i.e. if the cache contains multiple versions of
	that file).
index::
	Contains information about the cache contents, in particular
	timestamps and mode flags ("stat information") for the files
	stored in the cache. An unmerged index is an index which contains
	unmerged cache entries.
working tree::
	The set of files and directories currently being worked on.
	Think "ls -laR"
directory::
	The list you get with "ls" :-)
checkout::
	The action of updating the working tree to a revision which was
	stored in the object database.
revision::
	A particular state of files and directories which was stored in
	the object database. It is referenced by a commit object.
commit::
	The action of storing the current state of the cache in the
	object database. The result is a revision.
commit object::
	An object which contains the information about a particular
	revision, such as parents, committer, author, date and the
	tree object which corresponds to the top directory of the
	stored revision.
changeset::
	BitKeeper/cvsps speak for "commit". Since git does not store
	changes, but states, it really does not make sense to use
	the term "changesets" with git.
ent::
	Favorite synonym to "tree-ish" by some total geeks.
clean::
	A working tree is clean, if it corresponds to the revision
	referenced by the current head.
dirty::
	A working tree is said to be dirty if it contains modifications
	which have not been committed to the current branch.
head::
	The top of a branch. It contains a ref to the corresponding
	commit object.
branch::
	A non-cyclical graph of revisions, i.e. the complete history of
	a particular revision, which does not (yet) have children, which
	is called the branch head. The branch heads are stored in
	$GIT_DIR/refs/heads/.
ref::
	A 40-byte hex representation of a SHA1 pointing to a particular
	object. These are stored in $GIT_DIR/refs/.
head ref::
	A ref pointing to a head. Often, this is abbreviated to "head".
	Head refs are stored in $GIT_DIR/refs/heads/.
tree-ish::
	A ref pointing to either a commit object, a tree object, or a
	tag object pointing to a commit or tree object.
tag object::
	An object containing a ref pointing to another object. It can
	contain a (PGP) signature, in which case it is called "signed
	tag object".
tag::
	A ref pointing to a tag or commit object. In contrast to a head,
	a tag is not changed by a commit. Tags (not tag objects) are
	stored in $GIT_DIR/refs/tags/. A git tag has nothing to do with
	a Lisp tag (which is called object type in git's context).
merge::
	To merge branches means to try to accumulate the changes since a
	common ancestor and apply them to the first branch. An automatic
	merge uses heuristics to accomplish that. Evidently, an automatic
	merge can fail.
resolve::
	The action of fixing up manually what a failed automatic merge
	left behind.
repository::
	A collection of refs together with an object database containing
	all objects, which are reachable from the refs. A repository can
	share an object database with other repositories.
alternate object database::
	Via the alternates mechanism, a repository can inherit part of its
	object database from another object database, which is called
	"alternate".
reachable::
	An object is reachable from a ref/commit/tree/tag, if there is a
	chain leading from the latter to the former.
chain::
	A list of objects, where each object in the list contains a
	reference to its successor (for example, the successor of a commit
	could be one of its parents).
parent::
	A commit object contains a (possibly empty) list of the logical
	predecessor(s) in the line of development, i.e. its parents.
fetch::
	Fetching a branch means to get the branch's head ref from a
	remote repository, to find out which objects are missing from
	the local object database, and to get them, too.
pull::
	Pulling a branch means to fetch it and merge it.
push::
	Pushing a branch means to get the branch's head ref from a remote
	repository, find out if it is an ancestor to the branch's local
	head ref is a direct, and in that case, putting all objects, which
	are reachable from the local head ref, and which are missing from
	the remote repository, into the remote object database, and updating
	the remote head ref. If the remote head is not an ancestor to the
	local head, the push fails.
pack::
	A set of objects which have been compressed into one file (to save
	space or to transmit them efficiently).
pack index::
	Contains offsets into a pack, so the pack can be used instead of
	the unpacked objects.
plumbing::
	Cute name for core git.
porcelain::
	Cute name for programs and program suites depending on core git,
	presenting a high level access to core git. Porcelains expose
	more of a SCM interface than the plumbing.
object type:
	One of the identifiers "commit","tree","tag" and "blob" describing
	the type of an object.
SCM::
	Source code management (tool).
dircache::
	You are *waaaaay* behind.

GIT Glossary ============ Aug 2005

[[ref_SCM]]SCM::
	Source code management (tool). 
[[ref_SHA1]]SHA1::
	A 20-byte sequence (or 41-byte file containing the hex representation
	and a newline). It is calculated from the contents of an
	<<ref_object,object>> by the Secure Hash Algorithm 1. 
[[ref_alternate_object_database]]alternate object database::
	Via the alternates mechanism, a <<ref_repository,repository>> can
	inherit part of its <<ref_object_database,object database>> from another
	<<ref_object_database,object database>>, which is called "alternate". 
[[ref_blob_object]]blob object::
	Untyped <<ref_object,object>>, i.e. the contents of a file. 
[[ref_branch]]branch::
	A non-cyclical graph of revisions, i.e. the complete history of a
	particular <<ref_revision,revision>>, which does not (yet) have
	children, which is called the <<ref_branch,branch>> <<ref_head,head>>.
	The <<ref_branch,branch>> heads are stored in $GIT_DIR/refs/heads/. 
[[ref_cache]]cache::
	A collection of files whose contents are stored as objects. The
	<<ref_cache,cache>> is a stored version of your
	<<ref_working_tree,working tree>>. Well, can also contain a second, and
	even a third version of a working <<ref_tree,tree>>, which are used when
	merging. 
[[ref_cache_entry]]cache entry::
	The information regarding a particular file, stored in the
	<<ref_index,index>>. A <<ref_cache_entry,cache entry>> can be unmerged,
	if a <<ref_merge,merge>> was started, but not yet finished (i.e. if the
	<<ref_cache,cache>> contains multiple versions of that file). 
[[ref_chain]]chain::
	A list of objects, where each <<ref_object,object>> in the list contains
	a reference to its successor (for example, the successor of a
	<<ref_commit,commit>> could be one of its parents). 
[[ref_changeset]]changeset::
	BitKeeper/cvsps speak for "<<ref_commit,commit>>". Since git does not
	store changes, but states, it really does not make sense to use the term
	"changesets" with git. 
[[ref_checkout]]checkout::
	The action of updating the <<ref_working_tree,working tree>> to a
	<<ref_revision,revision>> which was stored in the
	<<ref_object_database,object database>>. 
[[ref_clean]]clean::
	A <<ref_working_tree,working tree>> is <<ref_clean,clean>>, if it
	corresponds to the <<ref_revision,revision>> referenced by the current
	<<ref_head,head>>. 
[[ref_commit]]commit::
	The action of storing the current state of the <<ref_cache,cache>> in
	the <<ref_object_database,object database>>. The result is a
	<<ref_revision,revision>>. 
[[ref_commit_object]]commit object::
	An <<ref_object,object>> which contains the information about a
	particular <<ref_revision,revision>>, such as parents, committer,
	author, date and the <<ref_tree_object,tree object>> which corresponds
	to the top <<ref_directory,directory>> of the stored
	<<ref_revision,revision>>. 
[[ref_dircache]]dircache::
	You are *waaaaay* behind. 
[[ref_directory]]directory::
	The list you get with "ls" :-) 
[[ref_dirty]]dirty::
	A <<ref_working_tree,working tree>> is said to be <<ref_dirty,dirty>> if
	it contains modifications which have not been committed to the current
	<<ref_branch,branch>>. 
[[ref_ent]]ent::
	Favorite synonym to "<<ref_tree-ish,tree-ish>>" by some total geeks. 
[[ref_fetch]]fetch::
	Fetching a <<ref_branch,branch>> means to get the
	<<ref_branch,branch>>'s <<ref_head_ref,head ref>> from a remote
	<<ref_repository,repository>>, to find out which objects are missing
	from the local <<ref_object_database,object database>>, and to get them,
	too. 
[[ref_head]]head::
	The top of a <<ref_branch,branch>>. It contains a <<ref_ref,ref>> to the
	corresponding <<ref_commit_object,commit object>>. 
[[ref_head_ref]]head ref::
	A <<ref_ref,ref>> pointing to a <<ref_head,head>>. Often, this is
	abbreviated to "<<ref_head,head>>". Head refs are stored in
	$GIT_DIR/refs/heads/. 
[[ref_index]]index::
	Contains information about the <<ref_cache,cache>> contents, in
	particular timestamps and mode flags ("stat information") for the files
	stored in the <<ref_cache,cache>>. An unmerged <<ref_index,index>> is an
	<<ref_index,index>> which contains unmerged <<ref_cache,cache>> entries.
[[ref_merge]]merge::
	To <<ref_merge,merge>> branches means to try to accumulate the changes
	since a common ancestor and apply them to the first
	<<ref_branch,branch>>. An automatic <<ref_merge,merge>> uses heuristics
	to accomplish that. Evidently, an automatic <<ref_merge,merge>> can
	fail. 
[[ref_object]]object::
	The unit of storage in GIT. It is uniquely identified by the
	<<ref_SHA1,SHA1>> of its contents. Consequently, an
	<<ref_object,object>> can not be changed. 
[[ref_object_database]]object database::
	Stores a set of "objects", and an individial <<ref_object,object>> is
	identified by its <<ref_SHA1,SHA1>> (its <<ref_ref,ref>>). The objects
	are either stored as single files, or live inside of packs. 
[[ref_object_name]]object name::
	Synonym for <<ref_SHA1,SHA1>>. 
[[ref_pack]]pack::
	A set of objects which have been compressed into one file (to save space
	or to transmit them efficiently). 
[[ref_pack_index]]pack index::
	Contains offsets into a <<ref_pack,pack>>, so the <<ref_pack,pack>> can
	be used instead of the unpacked objects. 
[[ref_parent]]parent::
	A <<ref_commit_object,commit object>> contains a (possibly empty) list
	of the logical predecessor(s) in the line of development, i.e. its
	parents. 
[[ref_plumbing]]plumbing::
	Cute name for core git. 
[[ref_porcelain]]porcelain::
	Cute name for programs and program suites depending on core git,
	presenting a high level access to core git. Porcelains expose more of a
	<<ref_SCM,SCM>> interface than the <<ref_plumbing,plumbing>>. 
[[ref_pull]]pull::
	Pulling a <<ref_branch,branch>> means to <<ref_fetch,fetch>> it and
	<<ref_merge,merge>> it. 
[[ref_push]]push::
	Pushing a <<ref_branch,branch>> means to get the <<ref_branch,branch>>'s
	<<ref_head_ref,head ref>> from a remote <<ref_repository,repository>>,
	find out if it is an ancestor to the <<ref_branch,branch>>'s local
	<<ref_head_ref,head ref>> is a direct, and in that case, putting all
	objects, which are <<ref_reachable,reachable>> from the local
	<<ref_head_ref,head ref>>, and which are missing from the remote
	<<ref_repository,repository>>, into the remote
	<<ref_object_database,object database>>, and updating the remote
	<<ref_head_ref,head ref>>. If the remote <<ref_head,head>> is not an
	ancestor to the local <<ref_head,head>>, the <<ref_push,push>> fails. 
[[ref_reachable]]reachable::
	An <<ref_object,object>> is <<ref_reachable,reachable>> from a
	<<ref_ref,ref>>/<<ref_commit,commit>>/<<ref_tree,tree>>/<<ref_tag,tag>>,
	if there is a <<ref_chain,chain>> leading from the latter to the former.
[[ref_ref]]ref::
	A 40-byte hex representation of a <<ref_SHA1,SHA1>> pointing to a
	particular <<ref_object,object>>. These are stored in $GIT_DIR/refs/. 
[[ref_repository]]repository::
	A collection of refs together with an <<ref_object_database,object
	database>> containing all objects, which are <<ref_reachable,reachable>>
	from the refs. A <<ref_repository,repository>> can share an
	<<ref_object_database,object database>> with other repositories. 
[[ref_resolve]]resolve::
	The action of fixing up manually what a failed automatic
	<<ref_merge,merge>> left behind. 
[[ref_revision]]revision::
	A particular state of files and directories which was stored in the
	<<ref_object_database,object database>>. It is referenced by a
	<<ref_commit_object,commit object>>. 
[[ref_tag]]tag::
	A <<ref_ref,ref>> pointing to a <<ref_tag,tag>> or
	<<ref_commit_object,commit object>>. In contrast to a <<ref_head,head>>,
	a <<ref_tag,tag>> is not changed by a <<ref_commit,commit>>. Tags (not
	<<ref_tag,tag>> objects) are stored in $GIT_DIR/refs/tags/. A git
	<<ref_tag,tag>> has nothing to do with a Lisp <<ref_tag,tag>> (which is
	called <<ref_object,object>> type in git's context). 
[[ref_tag_object]]tag object::
	An <<ref_object,object>> containing a <<ref_ref,ref>> pointing to
	another <<ref_object,object>>. It can contain a (PGP) signature, in
	which case it is called "signed <<ref_tag_object,tag object>>". 
[[ref_tree]]tree::
	Either a <<ref_working_tree,working tree>>, or a <<ref_tree_object,tree
	object>> together with the dependent blob and <<ref_tree,tree>> objects
	(i.e. a stored representation of a <<ref_working_tree,working tree>>). 
[[ref_tree_object]]tree object::
	An <<ref_object,object>> containing a list of blob and/or
	<<ref_tree,tree>> objects. (A <<ref_tree,tree>> usually corresponds to a
	<<ref_directory,directory>> without subdirectories). 
[[ref_tree-ish]]tree-ish::
	A <<ref_ref,ref>> pointing to either a <<ref_commit_object,commit
	object>>, a <<ref_tree_object,tree object>>, or a <<ref_tag_object,tag
	object>> pointing to a <<ref_commit,commit>> or <<ref_tree_object,tree
	object>>. 
[[ref_working_tree]]working tree::
	The set of files and directories currently being worked on. Think "ls
	-laR" 
Daniel Barkalow· Aug 17, 2005, 19:13 UTC · re: Johannes Schindelin · lore

Re: First stab at glossary

On Wed, 17 Aug 2005, Johannes Schindelin wrote:
Show 15 quoted lines
> Hi,
>
> long, long time. Here�s my first stab at the glossary, attached the
> alphabetically sorted, asciidoc marked up txt file (Comments?
> Suggestions? Pizzas?):
>
> object::
> 	The unit of storage in GIT. It is uniquely identified by
> 	the SHA1 of its contents. Consequently, an object can not
> 	be changed.
>
> SHA1::
> 	A 20-byte sequence (or 41-byte file containing the hex
> 	representation and a newline). It is calculated from the
> 	contents of an object by the Secure Hash Algorithm 1.

It's also often 40-character string (with whatever termination) in places like commit objects, tag objects, command-line arguments, listings, and so forth.

Show 7 quoted lines
> object database::
> 	Stores a set of "objects", and an individial object is identified
> 	by its SHA1 (its ref). The objects are either stored as single
> 	files, or live inside of packs.
>
> object name::
> 	Synonym for SHA1.

Have we killed the use of the third term "hash" for this? I'd say that "object name" is the standard term, and "SHA1" is a nickname, if only because "object name" is more descriptive of the particular use of the term.

> blob object::
> 	Untyped object, i.e. the contents of a file.

This "i.e." should be "e.g.", since symlink targets are also stored as blobs, and any other bulk data stored by itself would be. (IIRC, Junio has a tagged blob to hold his public key, for example)

Show 27 quoted lines
> tree object::
> 	An object containing a list of blob and/or tree objects.
> 	(A tree usually corresponds to a directory without
> 	subdirectories).
>
> tree::
> 	Either a working tree, or a tree object together with the
> 	dependent blob and tree objects (i.e. a stored representation
> 	of a working tree).
>
> cache::
> 	A collection of files whose contents are stored as objects.
> 	The cache is a stored version of your working tree. Well, can
> 	also contain a second, and even a third version of a working
> 	tree, which are used when merging.
>
> cache entry::
> 	The information regarding a particular file, stored in the index.
> 	A cache entry can be unmerged, if a merge was started, but not
> 	yet finished (i.e. if the cache contains multiple versions of
> 	that file).
>
> index::
> 	Contains information about the cache contents, in particular
> 	timestamps and mode flags ("stat information") for the files
> 	stored in the cache. An unmerged index is an index which contains
> 	unmerged cache entries.

I think we might want to entirely kill the "cache" term, and talk only about the "index" and "index entries". Of course, a bunch of the code will have to be renamed to make this completely successful, but we could change the glossary and documentation, and mention "cache" and "cache entry" as old names for "index" and "index entry" respectively.

> working tree::
> 	The set of files and directories currently being worked on.
> 	Think "ls -laR"

This is where the data is actually in the filesystem, and you can edit and compile it (as opposed to a tree object or the index, which semantically have the same contents, but aren't presented in the filesystem that way).

Show 6 quoted lines
> directory::
> 	The list you get with "ls" :-)
>
> checkout::
> 	The action of updating the working tree to a revision which was
> 	stored in the object database.
Move after "revision"?
Show 13 quoted lines
> revision::
> 	A particular state of files and directories which was stored in
> 	the object database. It is referenced by a commit object.
>
> commit::
> 	The action of storing the current state of the cache in the
> 	object database. The result is a revision.
>
> commit object::
> 	An object which contains the information about a particular
> 	revision, such as parents, committer, author, date and the
> 	tree object which corresponds to the top directory of the
> 	stored revision.
Move "parent" around here.
Show 7 quoted lines
> changeset::
> 	BitKeeper/cvsps speak for "commit". Since git does not store
> 	changes, but states, it really does not make sense to use
> 	the term "changesets" with git.
>
> ent::
> 	Favorite synonym to "tree-ish" by some total geeks.
Move after "tree-ish".
Show 9 quoted lines
> head::
> 	The top of a branch. It contains a ref to the corresponding
> 	commit object.
>
> branch::
> 	A non-cyclical graph of revisions, i.e. the complete history of
> 	a particular revision, which does not (yet) have children, which
> 	is called the branch head. The branch heads are stored in
> 	$GIT_DIR/refs/heads/.

A branch head might have children, if they're in another branch. (E.g., I pull mainline, make a new branch based on it, and commit a change; the head of mainline is still a branch head, even though it's the parent of my new commit, because my new commit isn't in mainline.)

Show 22 quoted lines
> ref::
> 	A 40-byte hex representation of a SHA1 pointing to a particular
> 	object. These are stored in $GIT_DIR/refs/.
>
> head ref::
> 	A ref pointing to a head. Often, this is abbreviated to "head".
> 	Head refs are stored in $GIT_DIR/refs/heads/.
>
> tree-ish::
> 	A ref pointing to either a commit object, a tree object, or a
> 	tag object pointing to a commit or tree object.
>
> tag object::
> 	An object containing a ref pointing to another object. It can
> 	contain a (PGP) signature, in which case it is called "signed
> 	tag object".
>
> tag::
> 	A ref pointing to a tag or commit object. In contrast to a head,
> 	a tag is not changed by a commit. Tags (not tag objects) are
> 	stored in $GIT_DIR/refs/tags/. A git tag has nothing to do with
> 	a Lisp tag (which is called object type in git's context).

As above, only the head for the branch being committed to is changed by a commit. A tag, not being the head of a branch, is therefore never changed by a commit.

Show 9 quoted lines
> merge::
> 	To merge branches means to try to accumulate the changes since a
> 	common ancestor and apply them to the first branch. An automatic
> 	merge uses heuristics to accomplish that. Evidently, an automatic
> 	merge can fail.
>
> resolve::
> 	The action of fixing up manually what a failed automatic merge
> 	left behind.

"Resolve" is also used for the automatic case (e.g., in "git-resolve-script", which goes from having two commits and a message to having a new commit). I'm not sure what the distinction is supposed to be.

	-Daniel
*This .sig left intentionally blank*
Johannes Schindelin· Aug 17, 2005, 20:05 UTC · re: Daniel Barkalow · lore

Re: First stab at glossary

Hi,
On Wed, 17 Aug 2005, Daniel Barkalow wrote:
Show 10 quoted lines
> On Wed, 17 Aug 2005, Johannes Schindelin wrote:
> 
> > SHA1::
> > 	A 20-byte sequence (or 41-byte file containing the hex
> > 	representation and a newline). It is calculated from the
> > 	contents of an object by the Secure Hash Algorithm 1.
> 
> It's also often 40-character string (with whatever termination) in places
> like commit objects, tag objects, command-line arguments, listings, and so
> forth.
Okay.
Show 7 quoted lines
> > object name::
> > 	Synonym for SHA1.
> 
> Have we killed the use of the third term "hash" for this? I'd say that
> "object name" is the standard term, and "SHA1" is a nickname, if only
> because "object name" is more descriptive of the particular use of the
> term.

Okay for "hash". What is the consensus on "object name" being more standard than "SHA1"?

Show 6 quoted lines
> > blob object::
> > 	Untyped object, i.e. the contents of a file.
> 
> This "i.e." should be "e.g.", since symlink targets are also stored as
> blobs, and any other bulk data stored by itself would be. (IIRC, Junio has
> a tagged blob to hold his public key, for example)
Agree.
Show 5 quoted lines
> I think we might want to entirely kill the "cache" term, and talk only
> about the "index" and "index entries". Of course, a bunch of the code will
> have to be renamed to make this completely successful, but we could change
> the glossary and documentation, and mention "cache" and "cache entry" as
> old names for "index" and "index entry" respectively.

For me, "index" is just the file named "index" (holding stat data and a ref for each cache entry). That is why I say an "index" contains "cache entries", not "index entries" (wee, that sounds wrong :-).

Show 7 quoted lines
> > working tree::
> > 	The set of files and directories currently being worked on.
> > 	Think "ls -laR"
> 
> This is where the data is actually in the filesystem, and you can edit and
> compile it (as opposed to a tree object or the index, which semantically
> have the same contents, but aren't presented in the filesystem that way).

Maybe I was too cautious. Linus very new idea was to think of the lowest level of an SCM as a file system. But I did not want to mention that. Thinking of it again, maybe I should.

> > checkout::
>
> Move after "revision"?

Ultimately, the glossary terms will be sorted alphabetically. If you look at the file attached to my original mail, this is already sorted and marked up using asciidoc. However, I wanted you and the list to understand how I grouped terms. The asciidoc'ed file is generated by a perl script.

> Move "parent" around here.
See above.
> Move after "tree-ish".
Ditto.
Show 10 quoted lines
> > branch::
> > 	A non-cyclical graph of revisions, i.e. the complete history of
> > 	a particular revision, which does not (yet) have children, which
> > 	is called the branch head. The branch heads are stored in
> > 	$GIT_DIR/refs/heads/.
> 
> A branch head might have children, if they're in another branch. (E.g., I
> pull mainline, make a new branch based on it, and commit a change; the
> head of mainline is still a branch head, even though it's the parent of my
> new commit, because my new commit isn't in mainline.)
Well noted! I'll just delete that part.
Show 9 quoted lines
> > tag::
> > 	A ref pointing to a tag or commit object. In contrast to a head,
> > 	a tag is not changed by a commit. Tags (not tag objects) are
> > 	stored in $GIT_DIR/refs/tags/. A git tag has nothing to do with
> > 	a Lisp tag (which is called object type in git's context).
> 
> As above, only the head for the branch being committed to is changed by a
> commit. A tag, not being the head of a branch, is therefore never changed
> by a commit.
I tried to say that.
Show 7 quoted lines
> > resolve::
> > 	The action of fixing up manually what a failed automatic merge
> > 	left behind.
> 
> "Resolve" is also used for the automatic case (e.g., in
> "git-resolve-script", which goes from having two commits and a message to
> having a new commit). I'm not sure what the distinction is supposed to be.

I did not like that naming anyway. In reality, git-resolve-script does not resolve anything, but it merges two revisions, possibly leaving something to resolve.

Ciao, Dscho

Junio C Hamano· Aug 17, 2005, 20:57 UTC · re: Johannes Schindelin · lore

Re: First stab at glossary

Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:
> Okay for "hash". What is the consensus on "object name" being more 
> standard than "SHA1"?

The tutorial uses the term "object name", so does README (implicitly, by saying "All objects are named by their content, which is approximated by the SHA1 hash of the object itself"). I think it is pretty safe to assume the list agrees with this term.

> For me, "index" is just the file named "index" (holding stat data and a 
> ref for each cache entry). That is why I say an "index" contains "cache 
> entries", not "index entries" (wee, that sounds wrong :-).

I think Linus already commented on using "index file" and "index entries" as the canonical terms. It would be a good idea to mention "cache" as a historical synonym in the documentation, so that we do not have to rename the symbols in the code.

> Ultimately, the glossary terms will be sorted alphabetically. If you look 
> at the file attached to my original mail, this is already sorted and 
> marked up using asciidoc. However, I wanted you and the list to understand 
> how I grouped terms. The asciidoc'ed file is generated by a perl script.

Then we should put the text version under Documentation, along with that script and a Makefile entry to do asciidoc and another to go to html. No rush for the script and Makefile entries, but it would make things easier to manage if we put the text version in the tree soonish. I've pushed out the one from your original "First stab" message.

Show 5 quoted lines
>> > branch::
>> > 	A non-cyclical graph of revisions, i.e. the complete history of
>> > 	a particular revision, which does not (yet) have children, which
>> > 	is called the branch head. The branch heads are stored in
>> > 	$GIT_DIR/refs/heads/.

I wonder if there is a math term for a non-cyclical graph that has a single "greater than anything else in the graph" node (but not necessarily a single but possibly more "lesser than anything else in the graph" nodes)?

Show 5 quoted lines
>> > tag::
>> > 	A ref pointing to a tag or commit object. In contrast to a head,
>> > 	a tag is not changed by a commit. Tags (not tag objects) are
>> > 	stored in $GIT_DIR/refs/tags/. A git tag has nothing to do with
>> > 	a Lisp tag (which is called object type in git's context).

I think this is good already, but maybe mention why you would use a tag in a sentence? "Most typically used to mark a particular point in the commit ancestry chain," or something.

Show 11 quoted lines
>> > resolve::
>> > 	The action of fixing up manually what a failed automatic merge
>> > 	left behind.
>> 
>> "Resolve" is also used for the automatic case (e.g., in
>> "git-resolve-script", which goes from having two commits and a message to
>> having a new commit). I'm not sure what the distinction is supposed to be.
>
> I did not like that naming anyway. In reality, git-resolve-script does not 
> resolve anything, but it merges two revisions, possibly leaving something 
> to resolve.

I am sure this would break people's script, but I am not against renaming git-resolve-script to say git-merge-script.

Anyway, thanks for doing this less-fun and not-so-glorious job.
Johannes Schindelin· Aug 17, 2005, 21:24 UTC · re: Junio C Hamano · lore

Re: First stab at glossary

Hi,
On Wed, 17 Aug 2005, Junio C Hamano wrote:
Show 10 quoted lines
> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:
> 
> > Okay for "hash". What is the consensus on "object name" being more 
> > standard than "SHA1"?
> 
> The tutorial uses the term "object name", so does README
> (implicitly, by saying "All objects are named by their content,
> which is approximated by the SHA1 hash of the object itself").
> I think it is pretty safe to assume the list agrees with this
> term.
Okay, I'll give in, then.
Show 10 quoted lines
> 
> > For me, "index" is just the file named "index" (holding stat data and a 
> > ref for each cache entry). That is why I say an "index" contains "cache 
> > entries", not "index entries" (wee, that sounds wrong :-).
> 
> I think Linus already commented on using "index file" and "index
> entries" as the canonical terms.  It would be a good idea to
> mention "cache" as a historical synonym in the documentation, so
> that we do not have to rename the symbols in the code.
> 
If the king penguin speaketh, the little blue penguin followeth.
Show 11 quoted lines
> > Ultimately, the glossary terms will be sorted alphabetically. If you look 
> > at the file attached to my original mail, this is already sorted and 
> > marked up using asciidoc. However, I wanted you and the list to understand 
> > how I grouped terms. The asciidoc'ed file is generated by a perl script.
> 
> Then we should put the text version under Documentation, along
> with that script and a Makefile entry to do asciidoc and another
> to go to html.  No rush for the script and Makefile entries, but
> it would make things easier to manage if we put the text version
> in the tree soonish.  I've pushed out the one from your original
> "First stab" message.

Okay. Then I follow the advice of the large and angry Saucer Crunching Monster, and shuffle the entries into a more logical order.

Show 10 quoted lines
> >> > branch::
> >> > 	A non-cyclical graph of revisions, i.e. the complete history of
> >> > 	a particular revision, which does not (yet) have children, which
> >> > 	is called the branch head. The branch heads are stored in
> >> > 	$GIT_DIR/refs/heads/.
> 
> I wonder if there is a math term for a non-cyclical graph that
> has a single "greater than anything else in the graph" node (but
> not necessarily a single but possibly more "lesser than anything
> else in the graph" nodes)?

Yes, there is. Although git itself is an example that there are two "greater than almost anything else in the graph" nodes.

Also, let's not be overzealous with our math degrees, okay? :-)
Show 9 quoted lines
> >> > tag::
> >> > 	A ref pointing to a tag or commit object. In contrast to a head,
> >> > 	a tag is not changed by a commit. Tags (not tag objects) are
> >> > 	stored in $GIT_DIR/refs/tags/. A git tag has nothing to do with
> >> > 	a Lisp tag (which is called object type in git's context).
> 
> I think this is good already, but maybe mention why you would
> use a tag in a sentence?  "Most typically used to mark a
> particular point in the commit ancestry chain," or something.
Done.
Show 14 quoted lines
> >> > resolve::
> >> > 	The action of fixing up manually what a failed automatic merge
> >> > 	left behind.
> >> 
> >> "Resolve" is also used for the automatic case (e.g., in
> >> "git-resolve-script", which goes from having two commits and a message to
> >> having a new commit). I'm not sure what the distinction is supposed to be.
> >
> > I did not like that naming anyway. In reality, git-resolve-script does not 
> > resolve anything, but it merges two revisions, possibly leaving something 
> > to resolve.
> 
> I am sure this would break people's script, but I am not against
> renaming git-resolve-script to say git-merge-script.

I do not mind changing the description if the consensus is to keep git-resolve-script.

> Anyway, thanks for doing this less-fun and not-so-glorious job.
The little blue penguin says: Thanks for all the fish!
Daniel Barkalow· Aug 17, 2005, 22:09 UTC · re: Johannes Schindelin · lore

Re: First stab at glossary

On Wed, 17 Aug 2005, Johannes Schindelin wrote:
Show 15 quoted lines
> Hi,
>
> On Wed, 17 Aug 2005, Daniel Barkalow wrote:
>
> > On Wed, 17 Aug 2005, Johannes Schindelin wrote:
> >
> > > object name::
> > > 	Synonym for SHA1.
> >
> > Have we killed the use of the third term "hash" for this? I'd say that
> > "object name" is the standard term, and "SHA1" is a nickname, if only
> > because "object name" is more descriptive of the particular use of the
> > term.
>
> Okay for "hash".

I think we only need at most two names for this, so this is more a matter of fixing old usage than documenting it.

Show 9 quoted lines
> > I think we might want to entirely kill the "cache" term, and talk only
> > about the "index" and "index entries". Of course, a bunch of the code will
> > have to be renamed to make this completely successful, but we could change
> > the glossary and documentation, and mention "cache" and "cache entry" as
> > old names for "index" and "index entry" respectively.
>
> For me, "index" is just the file named "index" (holding stat data and a
> ref for each cache entry). That is why I say an "index" contains "cache
> entries", not "index entries" (wee, that sounds wrong :-).

Well, it often contains information not present anywhere else (the status of a merge; the set of files being committed, added, or removed), so it isn't really a cache at all.

Show 11 quoted lines
> > > working tree::
> > > 	The set of files and directories currently being worked on.
> > > 	Think "ls -laR"
> >
> > This is where the data is actually in the filesystem, and you can edit and
> > compile it (as opposed to a tree object or the index, which semantically
> > have the same contents, but aren't presented in the filesystem that way).
>
> Maybe I was too cautious. Linus very new idea was to think of the lowest
> level of an SCM as a file system. But I did not want to mention that.
> Thinking of it again, maybe I should.

You probably don't need to mention that tree objects and index files can be thought of as filesystems, but you should specify that the working tree really is in the Unix filesystem, in case people have heard of the idea.

It should be clear to say 'You can "cd" there and "ls" to list your files.', rather than 'Think "ls -laR"' which makes my think of the output, which is like the output from git-ls-files.

Show 8 quoted lines
> > > checkout::
> >
> > Move after "revision"?
>
> Ultimately, the glossary terms will be sorted alphabetically. If you look
> at the file attached to my original mail, this is already sorted and
> marked up using asciidoc. However, I wanted you and the list to understand
> how I grouped terms. The asciidoc'ed file is generated by a perl script.
Ah, okay.
Show 11 quoted lines
> > > resolve::
> > > 	The action of fixing up manually what a failed automatic merge
> > > 	left behind.
> >
> > "Resolve" is also used for the automatic case (e.g., in
> > "git-resolve-script", which goes from having two commits and a message to
> > having a new commit). I'm not sure what the distinction is supposed to be.
>
> I did not like that naming anyway. In reality, git-resolve-script does not
> resolve anything, but it merges two revisions, possibly leaving something
> to resolve.
Right; I think we should change the name of the script.
	-Daniel
*This .sig left intentionally blank*
Johannes Schindelin· Aug 17, 2005, 22:19 UTC · re: Daniel Barkalow · lore

Re: First stab at glossary

Hi,
On Wed, 17 Aug 2005, Daniel Barkalow wrote:
Show 8 quoted lines
> On Wed, 17 Aug 2005, Johannes Schindelin wrote:
> 
> > On Wed, 17 Aug 2005, Daniel Barkalow wrote:
> > > [...]
> > Okay for "hash".
> 
> I think we only need at most two names for this, so this is more a matter
> of fixing old usage than documenting it.

It's short enough to keep it in the glossary _and_ fix the old documentation.

Show 5 quoted lines
> > [blabla] index [blable] cache [bliblo]
>
> Well, it often contains information not present anywhere else (the status
> of a merge; the set of files being committed, added, or removed), so it
> isn't really a cache at all.
Okay, okay. I stand corrected.
Show 11 quoted lines
> > Maybe I was too cautious. Linus very new idea was to think of the lowest
> > level of an SCM as a file system. But I did not want to mention that.
> > Thinking of it again, maybe I should.
> 
> You probably don't need to mention that tree objects and index files can
> be thought of as filesystems, but you should specify that the working tree
> really is in the Unix filesystem, in case people have heard of the idea.
> 
> It should be clear to say 'You can "cd" there and "ls" to list your
> files.', rather than 'Think "ls -laR"' which makes my think of the output,
> which is like the output from git-ls-files.
How about this:
working tree::
        The set of files and directories currently being worked on,
        i.e. you can work in your working tree without using git at all.
Show 10 quoted lines
> > > > checkout::
> > >
> > > Move after "revision"?
> >
> > Ultimately, the glossary terms will be sorted alphabetically. If you look
> > at the file attached to my original mail, this is already sorted and
> > marked up using asciidoc. However, I wanted you and the list to understand
> > how I grouped terms. The asciidoc'ed file is generated by a perl script.
> 
> Ah, okay.

Sorry, I attributed these "moving suggestions" to the large and angry SCM, while those were your comments. Since Junio decided to keep the "topic ordered" form in his repository, I moved them around according to your mail.

Show 13 quoted lines
> > > > resolve::
> > > > 	The action of fixing up manually what a failed automatic merge
> > > > 	left behind.
> > >
> > > "Resolve" is also used for the automatic case (e.g., in
> > > "git-resolve-script", which goes from having two commits and a message to
> > > having a new commit). I'm not sure what the distinction is supposed to be.
> >
> > I did not like that naming anyway. In reality, git-resolve-script does not
> > resolve anything, but it merges two revisions, possibly leaving something
> > to resolve.
> 
> Right; I think we should change the name of the script.

How many users are there? Probably many call git-pull-script anyway, right?

Ciao, Dscho

Tim Ottinger· Aug 24, 2005, 15:03 UTC · re: Johannes Schindelin · lore

Tool renames? was Re: First stab at glossary

So when this gets all settled, will we see a lot of tool renaming? 

While it would cause me and my team some personal effort (we have a special-purpose porcelain), it would be welcome to have a lexicon that is sane and consistent, and in tune with all the documentation.

Others may feel differently, I understand.
-- 
                             ><>
... either 'way ahead of the game, or 'way out in left field.
Junio C Hamano· Aug 25, 2005, 01:16 UTC · re: Tim Ottinger · lore

Re: Tool renames? was Re: First stab at glossary

Tim Ottinger <tottinge@progeny.com> writes:
> So when this gets all settled, will we see a lot of tool renaming? 

I personally do not see it coming. Any particular one you have in mind?

Tim Ottinger· Sep 1, 2005, 17:55 UTC · re: Junio C Hamano · lore

Re: Tool renames? was Re: First stab at glossary

Junio C Hamano wrote:
Show 13 quoted lines
>Tim Ottinger <tottinge@progeny.com> writes:
>
>  
>
>>So when this gets all settled, will we see a lot of tool renaming? 
>>    
>>
>
>I personally do not see it coming.  Any particular one you have
>in mind?
>
>  
>

git-update-cache for instance? I am not sure which 'cache' commands need to be 'index' now.

-- 
                             ><>
... either 'way ahead of the game, or 'way out in left field.
Junio C Hamano· Sep 2, 2005, 00:38 UTC · re: Tim Ottinger · lore

Re: Tool renames? was Re: First stab at glossary

Tim Ottinger <tottinge@progeny.com> writes:
> git-update-cache for instance?
> I am not sure which 'cache' commands need to be 'index' now.

Logically you are right, but I suspect that may not fly well in practice. Too many of us have already got our fingers wired to type cache, and the glossary is there to describe both cache and index.

 
Daniel Barkalow· Sep 2, 2005, 18:09 UTC · re: Junio C Hamano · lore

Re: Tool renames? was Re: First stab at glossary

On Thu, 1 Sep 2005, Junio C Hamano wrote:
Show 9 quoted lines
> Tim Ottinger <tottinge@progeny.com> writes:
> 
> > git-update-cache for instance?
> > I am not sure which 'cache' commands need to be 'index' now.
> 
> Logically you are right, but I suspect that may not fly well in
> practice.  Too many of us have already got our fingers wired to
> type cache, and the glossary is there to describe both cache and
> index.

My vote's for changing the official names, but keeping symlinks for the old names. As far as I know, there aren't any actual conflicts, and we might as well have new users pick up the logical names. I particularly think "git merge" would be really good to have.

	-Daniel
*This .sig left intentionally blank*
Junio C Hamano· Sep 2, 2005, 18:33 UTC · re: Daniel Barkalow · lore

Re: Tool renames? was Re: First stab at glossary

Daniel Barkalow <barkalow@iabervon.org> writes:
Show 15 quoted lines
> On Thu, 1 Sep 2005, Junio C Hamano wrote:
>
>> Tim Ottinger <tottinge@progeny.com> writes:
>> 
>> > git-update-cache for instance?
>> 
>> Logically you are right, but I suspect that may not fly well in
>> practice.  Too many of us have already got our fingers wired to
>> type cache, and the glossary is there to describe both cache and
>> index.
>
> My vote's for changing the official names, but keeping symlinks for the 
> old names. As far as I know, there aren't any actual conflicts, and we 
> might as well have new users pick up the logical names. I particularly 
> think "git merge" would be really good to have.
OK.  As Horst also says, we should do this before 1.0.
0.99.6::
	This hopefully will be done on Sep 7th.  Tool renames
	will not happen in this release, but the set of cleaned
	up names will be discussed on the list during this
	timeperiod.  I'll draw up a strawman tonight unless
	somebody else does it first.
0.99.7::
	We install symbolic links for the old names.  For the
	documentation, we do not bother --- just install under
	new names.  Also remove support for ancient environment
	variable names from gitenv().  Aim for Sep 17th.
0.99.8::
	Aim for Oct 1st; we do not install symbolic links
	anymore and supply "clean-old-install" target in the
	Makefile that removes symlinks installed by 0.99.7 from
	DESTDIR.  This target is not run automatically from other
	usual make targets; it is just there for your
	convenience.
Junio C Hamano· Sep 3, 2005, 06:05 UTC · re: Junio C Hamano · lore

Re: Tool renames? was Re: First stab at glossary

I said:
> 	I'll draw up a strawman tonight unless somebody else
> 	does it first.
1. Say 'index' when you are tempted to say 'cache'.
        git-checkout-cache      git-checkout-index
        git-convert-cache       git-convert-index
        git-diff-cache          git-diff-index
        git-fsck-cache          git-fsck-index
        git-merge-cache         git-merge-index
        git-update-cache        git-update-index
2. The act of combining two or more heads is called 'merging';
   fetching immediately followed by merging is called 'pulling'.
        git-resolve-script      git-merge-script
   The commit walkers are called *-pull, but this is probably
   confusing.  They are not pulling.
        git-http-pull           git-http-walk
        git-local-pull          git-local-walk
        git-ssh-pull            git-ssh-walk
3. Non-binaries are called '*-scripts'.
   In earlier discussions some people seem to like the
   distinction between *-script and others; I did not
   particularly like it, but I am throwing this in for
   discussion.
        git-applymbox           git-applymbox-script
        git-applypatch          git-applypatch-script
        git-cherry              git-cherry-script
        git-shortlog            git-shortlog-script
        git-whatchanged         git-whatchanged-script
4. To be removed shortly.
        git-clone-dumb-http     should be folded into git-clone-script
Daniel Barkalow· Sep 3, 2005, 06:54 UTC · re: Junio C Hamano · lore

Re: Tool renames? was Re: First stab at glossary

On Fri, 2 Sep 2005, Junio C Hamano wrote:
Show 13 quoted lines
> I said:
> 
> > 	I'll draw up a strawman tonight unless somebody else
> > 	does it first.
> 
> 1. Say 'index' when you are tempted to say 'cache'.
> 
>         git-checkout-cache      git-checkout-index
>         git-convert-cache       git-convert-index
>         git-diff-cache          git-diff-index
>         git-fsck-cache          git-fsck-index
>         git-merge-cache         git-merge-index
>         git-update-cache        git-update-index

Agreed, except that git-convert-cache and git-fsck-cache actually have nothing to do this the index by any name, and should probably be git-convert-objects and git-fsck-objects.

Show 11 quoted lines
> 2. The act of combining two or more heads is called 'merging';
>    fetching immediately followed by merging is called 'pulling'.
> 
>         git-resolve-script      git-merge-script
> 
>    The commit walkers are called *-pull, but this is probably
>    confusing.  They are not pulling.
> 
>         git-http-pull           git-http-walk
>         git-local-pull          git-local-walk
>         git-ssh-pull            git-ssh-walk
I think "fetch" is more applicable to what they do.
Show 12 quoted lines
> 3. Non-binaries are called '*-scripts'.
> 
>    In earlier discussions some people seem to like the
>    distinction between *-script and others; I did not
>    particularly like it, but I am throwing this in for
>    discussion.
> 
>         git-applymbox           git-applymbox-script
>         git-applypatch          git-applypatch-script
>         git-cherry              git-cherry-script
>         git-shortlog            git-shortlog-script
>         git-whatchanged         git-whatchanged-script

I don't think it matters very much whether something is a script or not; on the other hand, it would be good to have "git" list a reasonable set of commands to use through the interface, which would exclude, for example, git-merge-one-file-script, and include the above commands.

> 4. To be removed shortly.
> 
>         git-clone-dumb-http     should be folded into git-clone-script
Agreed.
	-Daniel
*This .sig left intentionally blank*
Junio C Hamano· Sep 3, 2005, 08:29 UTC · re: Daniel Barkalow · lore

Re: Tool renames? was Re: First stab at glossary

Daniel Barkalow <barkalow@iabervon.org> writes:
> Agreed, except that git-convert-cache and git-fsck-cache actually have 
> nothing to do this the index by any name, and should probably be 
> git-convert-objects and git-fsck-objects.
You are right.
> I think "fetch" is more applicable to what they do.

OK. then they are git-http-fetch and friends. How about git-ssh-push? The counterpart of fetch-pack/clone-pack is called upload-pack, so would git-ssh-upload make things more consistent? I dunno.

> I don't think it matters very much whether something is a script or not; 
> on the other hand, it would be good to have "git" list a reasonable set of 
> commands to use through the interface, which would exclude, for example, 
> git-merge-one-file-script, and include the above commands.

Are you suggesting to drop -script from git-merge-one-file? Then git-cherry should perhaps keep its current name.

Daniel Barkalow· Sep 4, 2005, 17:23 UTC · re: Junio C Hamano · lore

Re: Tool renames? was Re: First stab at glossary

On Sat, 3 Sep 2005, Junio C Hamano wrote:
Show 8 quoted lines
> Daniel Barkalow <barkalow@iabervon.org> writes:
> 
> > I think "fetch" is more applicable to what they do.
> 
> OK.  then they are git-http-fetch and friends.  How about
> git-ssh-push?  The counterpart of fetch-pack/clone-pack is
> called upload-pack, so would git-ssh-upload make things more
> consistent?  I dunno.
I like that idea.
Show 7 quoted lines
> > I don't think it matters very much whether something is a script or not; 
> > on the other hand, it would be good to have "git" list a reasonable set of 
> > commands to use through the interface, which would exclude, for example, 
> > git-merge-one-file-script, and include the above commands.
> 
> Are you suggesting to drop -script from git-merge-one-file?
> Then git-cherry should perhaps keep its current name.

I'd suggest it get a different ending, like .sh or -helper. That way, it's distinct both from binaries and from scripts that people run directly.

	-Daniel
*This .sig left intentionally blank*

← back to recent threads