threads / discuss / 6207

Re: [DRAFT] Branching and merging with git

Subject: Re: [DRAFT] Branching and merging with git

## tl;dr

37 messages between Nov 16, 2006 and Jan 10, 2007.

replies: 36people: 13as markdown or json

linux@horizon.com· Nov 16, 2006, 22:17 UTC · lore

[DRAFT] Branching and merging with git

I know it took me a while to get used to playing with branches, and I still get nervous when doing something creative. So I've been trying to get more comfortable, and wrote the following to document what I've learned.

It's a first draft - I just finished writing it, so there are probably some glaring errors - but I thought it might be of interest anyway.

* Branching and merging in git

In CVS, branches are difficult and awkward to use, and generally considered an advanced technique. Many people use CVS for a long time without departing from the trunk.

Git is very different. Branching and merging are central to effective use of git, and if you aren't comfortable with them, you won't be comfortable with git. In particular, they are required to share work with other people.

The only things that are a bit confusing are some of the names.
In particular, at least when beginning:
- You create new branches with "git checkout -b".
  "git branch" should only be used to list and delete branches.
- You share work with "git fetch" and "git push".  These are opposites.
- You merge with "git pull", not "git merge".  "git pull" can
  also do a "git fetch", but that's optional.  What's not optional
  is the merge.
* A brief digression on command names.

Originally, all git commands were named "git-foo". When there got to be over a hundred, people started complaining about the clutter in /usr/bin. After some discussion, the following solution was reached:

- It's now possible to place all of the git-foo commands into a separate
  directory.  (Despite the complaints, not too many people are doing it
  yet.)
- One option for git users is to add that directory to their $PATH.
- Another is provided by a wrapper called just "git".  It's intended to
  live in a public directory like /usr/bin, and knows the location of
  the separate directory.  When you type "git foo", it finds and executes
  "git-foo".
- Some simple commands are built into the git wrapper.  When you type
  "git add", it just does it internally.  (On the git mailing list,
  you will see patches like "make git diff a builtin"; this is what
  they're talking about.)
- For compatibility, for each builtin, there is a "git-add" file,
  which is just a link to the "git" wrapper.  It looks at the name it
  was invoked as to figure out what it should do.

The one confusing thing is that, although people usually type "git foo" in examples, they're interchangeable in practice. I go back and forth for no good reason. The main caveat is that to get the man page, you still need to type "man git-foo". Fortunately, there are two other ways to get the man page:

	1) "git help foo"
	2) "git foo --help"

Git doesn't have a specialized built-in help system; it just shows you the man pages.

One outstanding problem with git's man pages is that often the most detail is in the command page that was written first, not the user-friendly one that you should use. For example, there are a number of special cases of the "git diff" command that were written first, and the man pages for these commands (git-diff-index, git-diff-files, git-diff-tree, and git-diff-stages) are considerably more informative than the page for plain git-diff, even though that's the command that you should use 99% of the time.

* Git's representation of history

As you recall from Git 101, there are exactly four kinds of objects in Git's object database. All of them have globally unique 40-character hex names made by hashing their type and contents. Blob objects record file contents; they contain bytes. Tree objects record directory contents; they contain file names, permissions, and the associated tree or blob object names. Tag objects are shareable pointers to other objects; they're generally used to store a digital signature.

And then, we come to commit objects. Every commit points to (contains the name of) an associated tree object which records the state of the source code at the time of the commit, and some descriptive data (time, author, committer, commit comment) about the commit.

And most importantly, it contains a list of "parent commits", older commits from which this one is derived. These pointers are what produce the history graph.

Typically only one commit (the initial commit) has zero parents. It's possible to have more than one such commit (if you merge two projects with different history), but that's unusual.

Many commits have exactly one parent. These are made by a normal commit after editing. From a branching and merging point of view, they're not too exciting.

And then there are commits which have multiple parents. Two is most common, but git allows many more. (There's a limit of sixteen in the source code, and the most anyone's ever used in real life is 12, and that was generally regarded as overdoing it. Google on "doedecapus" for discussion of it.)

Finally, there are references, stored in the .git/refs directory. These are the human-readable names associated with commits, and the "root set" from which all other commits should be reachable.

These references are generally divided into two types, although
there is no fundamental difference:
- Tags are references that are intended to be immutable.
  The "v1.2" tag is a historical record.  A tag may point to
  a tag object (which will hold a signature), or just to a commit
  directly.  The latter isn't cryptographically authenticated, but
  works just fine for everyday use.
- Heads are references that are intended to be updated.  "Head"
  is actually synonymous with "branch", although one emphasizes the
  tip more, while the other directs your attention to the entire
  path that got there.
Either way, they're just a 41-byte file that contains a 40-byte hex
object ID, plus a newline.  Tags are stored in .git/refs/tags, and heads
are stored in .git/refs/heads.  Creating a new branch is literally just
picking a file name and writing the ID of an existing commit into it.

The git programs enforce the immutability of tags, but that's a safety feature, not something fundamental. You can rename a tag to the heads directory and go wild.

The only limit on branches is clutter. A number of git commands have ways to operate on "all heads", and if you have too many, it can get annoying. If you're not using a branch, either delete it, or move it somewhere (like the tags directory) where it won't clutter up the list of "currently active heads".

(Note that CVS doesn't have this all-heads default, so people tend to use longer branch names and keep them around after they've been merged into the trunk. Old CVS repositories converted to git generally need an old-branch cleanup.)

Another thing that's worth mentioning is that head and tag names can contain slashes; i.e. you're allowed to make subdirectories in the .git/refs/heads and .git/refs/tags directories. See the name page for "git-check-ref-format" for full details of legal names.

* Naming revisions

CVS encourages you to tag like crazy, because the only other way to find a given revision is by date. Git makes it a lot easier, so most revisions don't need names.

You can find a full description in the git-rev-parse man page, but here's a summary.

First of all, every commit has a globally unique name, its 40-digit hex object ID. It's a bit long and awkward, but always works. This is useful for talking about a specific commit on a mailing list. You can abbreviate it to a unique prefix; most people find about 8 digits sufficient.

(Subversion is easier yet, because it assigns a sequential number to each commit. However, that isn't possible in a distributed system like git.)

Second, you can refer to a head or tag name.  Git looks in the
following places, in order, for a head:
	1) .git
	2) .git/refs
	3) .git/refs/heads
	4) .git/refs/tags

You should avoid having e.g. a head and a tag with the same name, but if you do, you can specify one or the other with heads/foo and tags/foo.

Third, you can specify a commit relative to another. The simplest one is "the parent", specified by appending ^ to a name. E.g. HEAD^ or deadbeef^. If there are multiple parents, then ^ is the same as ^1, and the others are ^2, ^3, etc.

So the last few commits you've made are HEAD, HEAD^, HEAD^^, HEAD^^^, etc. After a while, counting carets becomes annoying, so you can abbreviate ^^^^ as ~4. Note that this only lets you specify the first parent. If you want to follow a side branch, you have to specify something like "master~305^2~22".

* Converting between names

Git has two helpers (programs designed mainly for use in shell scripts) to convert between global object IDs and human-readable names.

The first is git-rev-parse. This is a general git shell script helper, which validates the command line and converts object names to absolute object IDs. Its man page has a detailed description of the object name syntax.

The second is git-name-rev, which converts the other way around. It's particularly useful for seeing which tags a given commit falls between.

* Working with branches, the trivial cases.

By convention, the local "trunk" of git development is called "master". This is just the name of the branch it creates when you start an empty repository. You can delete it if you don't like the name.

If you create your repository by cloning someone else's repository, the remote "master" branch is copied to a local branch named "origin". You get your own "master" branch which is not tied to the remote repository.

There is always a current head, known as HEAD. (This is actually a symbolic link, .git/HEAD, to a file like refs/heads/master.) Git requires that this always point to the refs/heads directory.

	Minor technical details:
	1) HEAD used to be a Unix symlink, and can still be though of that
	   way, but for Microsoft support, this is now what's called a
	   "symbolic reference" or symref, and is a plain file containing
	   "ref: refs/heads/master".  Git treats it just like a symlink.
	   There's a git-update-ref helper which writes these.
	2) While HEAD must point to refs/heads, it's legal for it to
	   point to a file that doesn't exist.	This is what happens
	   before the first commit in a brand new repository.

When you do "git commit", a new commit object is created with the old HEAD as a parent, and the new commit is written to the current head (pointed to by HEAD).

* The three uses of "git checkout"
Git checkout can do three separate things:
1) Change to a new head
	git checkout [-f|-m] <branch>
   This makes <branch> the new HEAD, and copies its state to the index
   and the working directory.
   If a file has unsaved changes in the working directory, this tries
   to preserve them.  This is a simple attempt, and requires that the
   modified files(s) are not altered between the old and new HEADs.
   In that case, the version in the working directory is left untouched.
   A more aggressive option is -m, which will try to do a three-way
   (intra-file) merge.  This can fail, leaving unmerged files in the
   index.
   An alternative is to use -f, which will overwrite any unsaved changes
   in the working directory.  This option can be used with no <branch>
   specified (defaults to HEAD) to undo local edits.
2) Revert changes to a small number of files.
	git checkout [<revision>] [--] <paths>
   will copy the version of the <paths> from the index to the working
   directory.  If a <revision> is given, the index for those paths will
   be updated from the given revision before copying from the index to
   the working tree.
   Unlike the version with no <paths> specified, this does NOT update
   HEAD, even if <paths> is ".".
3) Create a branch.
	git checkout [-f|-m] -b <branch> [revision]
   will create, and switch to, a new branch with the given name.
   This is equivalent to
	git branch <branch> [<revision>]
	git checkout [-f|-m] <branch>
   If <revision> is omitted, it defaults to the current HEAD, in which
   case no working directory files are altered.
   This is the usual way that one checks out a revision that does not
   have an existing head pointing to it.
* Deleting branches

"git branch -d <head>" is safe. It deletes the given <head>, but first it checks that the commit is reachable some other way. That is, you merged the branch in somewhere, or you never did any edits on that branch.

It's a good idea to create a "topic branch" when you're working on anything bigger than a one-liner, but it's also a good idea to delete them when you're done. It's still there in the history.

* Doing rude things to heads: git reset

If you need to overwrite the current HEAD for some reason, the tool to do it with is "git reset". There are three levels of reset:

git reset --soft <head>
	This overwrites the current HEAD with the contents of <head>.
	If you omit <head>, it defaults to HEAD, so this does nothing.
git reset [<head>]
git reset --mixed [<head>]
	These overwrite the current HEAD, and copy it to the index,
	undoing any git-update-index commands you may have executed.
	If you omit <head>, it default to HEAD, so there is no change
	to the current branch, but all index changes are undone.
git reset --hard [<head>]
	This does everything mentioned above, and updates the
	working directory.  This throws away all of your in-progress
	edits and gets you a clean copy.  This is also commonly
	used without an explicit <head>, in which case the current
	HEAD is used.
* Using git-reset to fix mistakes
"Oh, no!  I didn't mean to commit *that*!  How do I undo it?"

If you just want to undo a commit, then you can use "git reset HEAD^" to return the current HEAD to the previous version. If you want to leave the commit in the index (this only applies to you if you are familiar with using the index; see below), then you can use "git reset --soft HEAD^".

And if you want to blow away every record of the changes you made, you can use "git reset --hard HEAD^"

If you just want a stupid trivial mistake and want to replace the most recent commit with a corrected one, "git commit --amend" is your friend. It makes a new commit with HEAD^ rather than HEAD as its ancestor.

* Fixing mistakes without git-reset

git-reset has the problem that it doesn't preserve hacking in progress in the working directory. It can leave the working directory alone (making everything a "hack in progress"), but it can't merge changes like git checkout.

So, suppose you've been trying something that should have been simple, and made three commits before realizing that the problem is harder than you thought and you want your work so far to be on a new branch of its own; committing them on the current HEAD (I'll call it "old") was a mistake.

You don't want to erase anything, just rename it. Make "new" a copy of the current "old" and move old back to HEAD^^^ (three commits ago).

While there are ways to do that using git-reset, but far better is to use "git branch -f":

	git checkout -b new
		Create (and switch to) the "new" branch.
	git branch -f old HEAD^^^
		Forcibly move "old" back three versions.
		(You could also use old~3 or new^^^ or any synonymous name.)

You can use a similar trick to rename a branch. If it's the current HEAD, then:

	git checkout -b newname
	git branch -d oldname
and if it's not, then
	git branch newname oldname
	git branch -d oldname

An alternative in the latter case is to just use mv on the raw .git/refs/heads/oldname file.

* How do I check out an old version?

A very common beginning question is how to check out an old version. Say you need to compile an old release for test purposes. "git checkout v1.2" gives a funny error message. What's going on?

Well, "git checkout" makes the current HEAD point to the head that you specify. And, as previously mentioned, git requires that it point to something in the .git/refs/heads directory. So you can't do that.

If you're busy doing things in your working directory, and don't want to overwrite your work with an old version, then you can get a snapshot with the (old) git-tar-tree or (new) git-archive commands. These produce a tar file (git-archive can also produce a zip file) which is a snapshot of any version you like. You can then unpack this file in a different directory and build it.

However, if you haven't got any edits in progress, and want to check out the old version into your working directory, just create a temp branch!

	git checkout -b temp v1.2

Will do what you want. This will also do what you want if you have a local edit (like the "#define DEBUG 1" mentioned above) that you want to preserve while working on the old version.

You'll see this in use if you ever use the (highly recommended) git-bisect tool. It creates a branch called "bisect" for the duration of the bisect.

(Yes, I have to confess, I sometimes wish that git would enforce the "HEAD must point to .git/refs/heads" rule when committing (checking in) rather than when checking out, but that's the way git has grown up.)

Note that if you want *exactly* an old version, with no local hacks, make sure there are none (with "git status") when doing this. It's more convenient if you do it before the checkout, but you'll get the same answer if you ask afterwards.

Now, what about the complex case: you have local hacks that you want to keep, but not have polluting the old version?

Well, one way of the other, you'll have to commit it. If you don't mind committing your changes to the current branch ("git commit -a"), do that.

If they're not ready to commit, you can commit them anyway, and back them out when you're done:

	git commit -a -m "Temp commit"
	git checkout -b temp v1.2
	make ; make test ; whatever
	git checkout master
	git branch -d temp
	git reset HEAD^

This leaves both the working directory and the master head in the states they were in at the beginning.

If you don't like committing to the master branch, you can make a new one. In this example, it's "work in progress", a.k.a. "wip":

	git checkout -b wip
	git commit -a -m "Temp commit"
	git checkout -b temp v1.2
	make ; make test ; whatever
	git checkout wip
	git branch -d temp
	git reset master
	git checkout master	# Won't change working directory
	git branch -d wip
* Examining history: git-log and git-rev-list

In another example of docs being better on the first command written, the all-purpose utility for examining history is "git log", but all of the examples of clever ways to use it are in the git-rev-list man page. And git-log also has most of git-diff's options.

Other utilities, notably the gitk and qgit GUIs, also use the git-rev-list command-line options, so it's well worth learning them.

git-rev-list gives you a filtered subset of the repository history. There are two basic ways that you can do the filtering:

1) By ancestry.  You specify a set of commits to include all the
   ancestors of, and another set to exclude all the ancestors of.
   (For this purpose, a commit is considered an ancestor of itself.)
   So if you want to see all commits between v1.1 and v1.2, you
   can specify
   	git log ^v1.1 v1.2
   or, with a more convenient syntax
   	git log v1.1..v1.2
  However, there are times when you want to specify something more
  complex.  For example, if a big branch that had been in progress since
  v1.0.7 was merged between v1.1 and v1.2, but you don't want to see it,
  you could specify any of:
   	git log v1.2 ^v1.1 ^bigbranch
   	git log ^bigbranch v1.1..v1.2
	git log ^v1.1 bigbranch..v1.2
  They're all equivalent.  Another special syntax that's sometimes
  handy is
	git log branch1...branch2
   Note the three dots.  This generates the symmetric difference between
   the two; basically it's a diff between the commits that went into
   each of them.
   "git log" by default pipes its output through less(1), and generates
   its output from newest to oldest on the fly, so there's no great
   speed penalty to not specifying a starting place.  It'll generate a
   few screen fulls more than you look at, but not waste any more effort
   than that.
2) By path name.  This is a feature which appears to be unique to git.
   If you give git-rev-list (or git-log, or gitk, or qgit) a list of
   pathname prefixes, it will list only commits which touch those
   paths. So "git log drivers/scsi include/scsi" will list only
   commits which alters a file whose name begins with drivers/scsi
   or include/scsi.
   (If there's any possible ambiguity between a path name and a commit
   name, git-rev-list will refuse to proceed.  You can resolve it by
   including "--" on the command line.  Everything before that is a
   commit name; everything after is a path.)
   This filter is in addition to the ancestry filter.  It's also rather
   clever about omitting unnecessary detail.  In particular, if there's
   a side branch which does touch drivers/scsi, then the entire branch,
   and the merge at the end, will be removed from the log.

You can additionally limit the commits to a certain number, or by date, author, committer, and so on.

By default, "git log" only shows the commit messages, so it's important to write good ones. Other tools compress commit messages down to the first line, so try to make that as informative as possible.

* History diagrams
When talking about various situations involving multiple branches,
people often find it handy to draw pictures.  Gitk draws nice pictures
vertically, but for e-mail, ASCII art drawn horizontally is often easier.
Commits are shown as "o", and the links between them with lines drawn with
- / and \.  Time goes left to right, and heads may be labelled with names.
For example:
         o--o--o <-- Branch A
        /
 o--o--o <-- master
        \
         o--o--o <-- Branch B

If someone needs to talk about a particular commit, the character "o" may be replaced with another letter or number.

* Trivial merges: fast-forward and already up-to-date.

There are two kinds of merge that are particularly simple, and you will encounter them in git a great deal. They are mirror images.

Suppose that you are working on branch A and merge in branch B, but no work has been done to branch B since the last time you merged, or since you spawned branch A from it. That is, the history looks like

 o--o--o--o <-- B
           \
	    o--o--o <-- A
or
 o--o--o--o--o--o <-- B
           \     \
	    o--o--o--o--o <-- A

If you then merge B into A, A is described as "already up to date". It is already a strict superset of B, and the merge does nothing.

In particular, git will not create a dummy commit to record the fact that a merge was done. It turns out that are a number of bad things that would happen if you did this, but for now, I'll just say that git doesn't do it.

Now, the opposite scenario is the "fast-forward" merge. Suppose you merge A into B. Again, A is a strict superset of B.

In this case, git will simply change the head B to point to the same commit as A and say that it did a "fast-forward" merge. Again, no commit object is created to reflect this fact.

The effect is to unclutter the git history. If I create a topic branch to work on a feature, do some hacking, and then merge the result back into the (untouched!) master, the history will look just like I did all the work on the master directly. If I then delete the topic branch (because I'm done using it), the repository state is truly indistinguishable.

While the topic branch existed, you could have done something to the master branch, in which case the final merge would have been non-trivial, but if that didn't happen, git produces a simple, easy-to-follow linear history.

Some people used to heavyweight branches find this confusing; they think a merge is a big deal and it should be memorialized, but there are actually excellent reasons for doing this.

The most important one is that a fit of merging back and forth will eventually end. Suppose that branches A and B are maintained by separate developers who like to track each other's work closely.

If the fast-forward case did create a commit, then merging A into B would produce

 o--o--o--o---------o <-- B
           \       /
	    o--o--o <-- A
then merging B into A would produce:
 o--o--o--o---------o <-- B
           \       / \
	    o--o--o---o <-- A

and further merges would produce more and more dummy commits, all without ever reaching a steady state, and without making it obvious that the two heads are actually identical.

Since history lasts forever, cluttering it up with unimportant stuff is a burden to all future users, and not a good idea. Allowing the merge of a branch to be seamless in the simple case encourages lightweight branches. If you _might_ need a separate branch, create it. If it turned out that you didn't, it won't make a difference.

* Exchanging work with other repositories

The basic tools for exchanging work with other repositories are "git fetch" and "git push". The fact that "git pull" is not the opposite of "git push" is often confusing to beginners (it's a superset of git fetch), but that's the terminology that has grown up.

The unit of sharing in git is the branch. If you've used branches in CVS, you'll be familiar with using "CVS update" to pull changes from your "current branch" in the repository into your working directory.

In Git, you don't pull into the working directory, but rather into a tracking branch. You set up a branch in your repository which will be a copy of the branch in the remote repository. For example, if you use "git clone", then the remote "master" branch is tracked by the local "origin" branch.

Then, when you do a "git fetch", git fetches all of the new commits and sets the origin head to point to the newly fetched head of the remote branch.

By default, git checks that this is a trivial fast-forward merge, that is not throwing away history. If it finds something like:

o--o--o--o--o--o <-- remote master
       \
        o <-- Local origin

It will complain and abort the fetch. This is usually a warning that something has gone wrong - in particular, you forgot that this was supposed to be a tracking branch and committed some work to it - and it aborts before throwing your work away.

However, sometimes the remote git user will have a branch name that they delete and re-create frequently. There are plenty of reasons to do this. The most common is doing a "test merge" between various branches in progress. They're all unfinished, so the developer of branch A doesn't want to merge in all the new bugs in branch B, but a tester might want to create a merged version with both sets of bugs for testing.

The merged version is not intended to be a permanent part of history - it'll get deleted after the test - but it can still be useful to have a draft copy.

In this case, you can mark the source branch with a leading "+", to disable this sanity check. (See the git-fetch man page for details.)

Note that in this case, you should specifically avoid merging from such a branch into any non-test branches of your own. It is, as mentioned, not intended to be a permanent part of history, so don't make it part of your permanent history. (You still might want to test-merge it with your work in progress, of course.)

The fact that you should know to treat such branches specially is why git doesn't try to automatically cope with them.

* Alternate branch naming

The original git scheme mixes tracking branches with all the other heads. This requires that you remember which branches are tracking branches and which aren't. Hopefully, you remember what all your branches are for, but if you track a lot of remote repositories, you might not remember what every remote branch is for and what you called it locally.

* Remotes files

You can specify what to fetch on the git-fetch command line. However, if you intend to monitor another repository on an ongoing basis, it's generally easier to set up a short-cut by placing the options in .git/remotes/<name>.

The syntax is explained in the git-fetch man page. When this is st up, "git fetch <name>" will retrieve all the branches listed in the .git/remotes/<name> file. The ability to fetch multiple branches at once (such as release, beta, and development) is an advantage of using a remotes file.

You can also create the remotes file "origin" (not necessarily any relation to the branch named "origin"), which is the default for git-fetch. If you have a single primary "upstream" repository that you sync to, place it in the origin remotes file, and you can just type "git fetch" to get all the latest changes.

Note that branches to fetch are identified by "Pull: " lines in the remotes file. This is another example of the fetch/pull confusion. git-pull will be explained eventually.

* Remote tags
TODO: Figure out how remote tags work, under what circumstances
they are fetched, and what git does if there are conflicts.
* Exchanging work with other repositories, part II: git-push

It's simpler to set up git sharing on a pull basis. If your source code isn't secret, you can set up a public read-only server very easily (see the git-daemon man page for details), and have other fetch from that.

However, N developers all pulling from each other is an N^2 mess. Some centralization helps.

One way is to have a central coordinator (like Linus) who pulls from all of the developers, and who they in turn pull from.

The other is to have a central repository that people can push to. This generally requires an ssh login on the server. You can use git-shell as the login shell if all you want to allow the account to do is git fetch and push. (You can use the hook scripts to enforce rules about who's allowed to do what to which branch.)

Git-push to the remote machine works exactly like git-fetch from the remote machine. The objects are moved over, and the branches pushed to are fast-forwarded. If fast-forward is impossible, you get an error.

So if you have multiple people committing to a branch on the server, you will not be allowed to push if someone has pushed more to that branch since last time you fetched it.

You have to merge the changes locally, and re-try the push when you've got a new head that includes the most recently pushed work as an ancestor.

This is exactly like "cvs commit" not working if your recent checkout wasn't the (current) tip of the branch, but git can upload more than one commit.

The simplest way to resolve the conflict is to merge the remote head with your local head. This is easiest if you have different local branches for fetching the remote repository and for pushing to it.

That is, you have one head that just tracks the master repository's main branch, and another that you add your work to, and push from. This makes merging simpler when there are conflicts.

Another use for git-push, even for a solo developer, is sharing your work with the world. You can set up a public git server on a high-bandwidth machine (possibly rented from a hosting service) and then push to it to publish something.

* Merging (finally!)

I went through everything else first because the most common merge case is local changes with remote changes. Not that you can't merge two branches of your own, but you don't need to do that nearly as often.

The primitive that does the merging is called (guess what?) git-merge. And you can use that if you want. If you want to create a so-called octopus merge, with more than two parents, you have to.

However, it's usually easier to use the git-pull wrapper. This merges the changes from some other branch into the current HEAD and generates a commit message automatically.

git-merge lets you specify the commit message (rather than generating it automatically) and use a non-HEAD destination branch, but those options are usually more annoying than useful.

The basic git-pull syntax is
	git-pull <repository> <branch>

The repository can be any URL that git supports. Including, particularly, a local file. So to do a simple local merge, you just type

	git-pull . <branch>
So after doing some hacking on branch "foo", you would
	git checkout master
	git pull . foo
and ba-boom, all is done.

Now, you can also specify a remote repository to merge from, using a git://, http:// or git+ssh:// URL. This is what Linus does all day long, and why the git-pull tool is optimized to allow that. It uses git-fetch to fetch the remote branch without assigning it a branch name (it gets the special name FETCH_HEAD temporarily), and them merges it into the current HEAD directly.

There is absolutely nothing wrong with doing that, but beginners often find it confusing to have a single short command do quite so much. And if you are working closely with someone, it's often more convenient and less confusing to keep local tracking branches. Then you can

	git fetch upstream	# Fetches 'origin'
	git pull . origin
It's also possible to give just a single remotes file name to git-pull:
	git pull upstream

That does a git fetch, updating all of the listed branches as usual, then merges the _first_ listed branch into HEAD.

By the way: don't blink, you might miss it! As I mentioned, pulling is a very big part of Linus's daily routine, and he's made sure it's fast. (Actually, it produces a fair bit of output, so you'll see.)

Just to clarify, because people often get confused:

git-pull is a MERGING tool. It always does a merge, as well as an optional fetch. If you just want to LOOK at a remote branch, use git-fetch.

* Undoing a merge

If you discover that a merge was a mistake, it can be undone just like any other commit. The HEAD you merged to is the first parent, so just do

	git reset --hard HEAD^

This is why Linus likes a git-pull command that does so much in one shot - if he doesn't like what he pulls, it's easy to undo.

* How merging operates

Git uses the basic three-way merge. First, it applies it to whole files, and then to lines within files.

To do a three-way merge, you need three versions of a file. The versions A and B you want to merge, and a common ancestor, commonly called O. That is, history proceeds something like:

         o--o--A
        /
 o--o--O
        \
	 o--B

The basic idea is "I want the file O, plus all the changes made from O to A, plus all the changes made from O to B." Since the cases where one of A or B is a direct ancestor of the other have already been disposed of, the three commits must be different.

For each file, there are a few cases that are trivial, and git gets these out of the way immediately:

- If A and B are identical, the merged result is obvious.
- If O and A are the same, then the result should be B.
- If O and B are the same, then the result should be A.

In the completely trivial case when O, A and B are the same, then all three rules apply, they all produce the same obvious result.

The "merge base" version O is generally the most recent common ancestor of A and B. The only problem is, that's not necessarily unique!

The classic confusing case is called a "criss-cross merge", and looks like this:

         o--b-o-o--B
        /    \ /
 o--o--o      X
        \    / \
	 o--a-o-o--A

There are two common ancestors of A and B, marked a and b in the graph above. And they're not the same. You could use either one and get reasonable results, but how to choose?

The details are too advanced for this discussion, but the default "recursive" merge strategy that git uses solves the answer by merging a and b into a temporary commit and using *that* as the merge base.

Of course, a and b could have the same problem, so merging them could require another merge of still-older commits. This is why the algorithm is called "recursive." It's been tested with pathological conditions, but multiply nested criss-cross merges are very rare, so the recursion isn't a performance limit in practice.

If all three of a given file in O, A, B are different, then the three versions are pulled into the index file, called "stage 1", "stage 2", and "stage 3", and a merge strategy driver is called to resolve the mess. Git then uses the classic line-based three-way merge, looking for isolated changes and applying the same rules as for files when two of the source files are the same in some range.

* Alternate merge strategies

In every version control system prior to git, the merging algorithm was buried deep in the bowels of the software, and very difficult to change. One of particularly nice things that git did was allow for easily replaceable "merge strategies". Indeed, you can try multiple merge strategies, and the fallback - print an error message and let the user sort it out - can be thought of as just another merge strategy.

Enabling this is why the index is so important to git. It provides a place to store an unfinished merge, so you can try various strategies (including hand-editing) to finish it.

Generally, git's default merge strategies are just fine. There is, however, one special case that is occasionally useful, specified with the "-s ours" strategy.

That strategy instructs git that the merged result should be the same as the current HEAD. Any other branches are recorded as parents, but their contents are ignored.

What the heck is the use of that? Well, it lets you record the fact that some work has been done in the history, and that it shouldn't be merged again. For example, say you write and share a popular patch set. People are always merging it in to their local source trees. But then you discover a much better way to achieve the goal of that patch set, and you want to publish the fact that the new patch supersedes the old one.

If you developed the new set starting from the old one, that would happen automatically. But another way to achieve the same goal is to merge the old branch it in using the "ours" strategy. Everyone else's git will notice that the patch is already included, and stop trying to merge it in.

* When merging goes wrong

This is the fun part. Git's default recursive-merge strategy is pretty clever, but sometimes changes truly do conflict and need manual fix-up.

When git is unable to complete a merge, it leaves the three different versions in the index and places a file with CVS-style conflict markers in the working directory.

As long as there is a "staged" file in the index, you will not be able to commit. You must resolve the conflict, and update the index with the resolved versions. You can do this one at a time with git-update-index, or at the end by giving the files as arguments to git-commit.

Doing them one at a time is probably safest; checking in a file which still has conflict markers makes a bit of a mess. Note that git will still use the automatically generated commit message when you finally commit. (It's in .git/MERGE_MSG, if you care.)

Note that "git diff" knows how to be useful with a staged file. By default, it displays a multi-way diff. For example, suppose I take a (slightly buggy) hello.c:

--- hello.c --- #include <stdio.h>

int main(void)
{
	printf("Hello, world!");
}
--- end ---

Now, suppose that in branch A, I fix some bugs - add the missing newline and "return 0;". In branch B, I display my angst and change it to "Goodbye, cruel world!". When I try to merge A into B, obviously I'll get a conflict. The resultant file, with conflict markers, looks like:

--- hello.c --- #include <stdio.h>

int
main(void)
{
<<<<<<< HEAD/hello.c
        printf("Goodbye, cruel world!");
=======
	printf("Hello, world!\n");
	return 0;
>>>>>>> edadc53fc7a8aef2a672a4fa9d09aa16f4e14706/hello.c

} --- end ---

and the result of "git diff" is

diff --cc hello.c index 4b7f550,948a5f8..0000000

--- a/hello.c
+++ b/hello.c
@@@ -3,5 -3,6 +3,10 @@@
  int
  main(void)
  {
++<<<<<<< HEAD/hello.c
 +      printf("Goodbye, cruel world!");
++=======
+       printf("Hello, world!\n");
+       return 0;
++>>>>>>> edadc53fc7a8aef2a672a4fa9d09aa16f4e14706/hello.c
  }

Notice how this is not a standard diff!  It has two columns of diff
symbols, and shows the difference from each of the ancestors to the
current hello.c contents.  I can also use "git diff -1" to compare
against the common ancestor, or "-2" or "-3" to compare against each of
the merged copies individually.


* Alternatives to merging

The bigger and more active your source tree, the more important it is to
keep the history reasonably clean.  Just because git can do a merge in
under a second doesn't mean that you should do one daily.  When you look
back at a feature's development history, you'd like to see meaningful
changes recorded and not a lot of meaningless ones.

Now, once you have shared a commit with others, and they have incorporated
it into their development, it becomes impossible to undo.  But git
provides tools that are useful for "rewriting history" before public
release.  These can be used to edit a commit for publication.

* Test merging

One way to keep the history clean is to simply not merge other branches
into your development branch.  If you want to use your new features and
other people's code changes, make a test merge and use that, but don't
make that merge part of your branch.

This is slightly more work (you have to change to a test branch and do
your merging there), but not very much.

Sometimes, when doing this, a conflict appears between your changes and
someone else's development.  If you get tired of fixing the same conflict
every time you do a test merge, have a look at the git-rerere tool.
This remembers resolved conflicts and tries to apply the same resolution
patch the next time.

It's written specifically to help you not do an extra merge unnecessarily.
Although its man page is well worth reading, you never invoke git-rerere
explicitly; it's invoked automatically by the merge and patch tools if
you create a .git/rr-cache directory.

* Cherry picking

If you have a series of patches on a branch, but you want a subset
of them, or in a different order, there's a handy utility called
"git-cherry-pick" which will find the diff and apply it as a patch to
the current HEAD.  It automatically recycles the commit message from
the original commit.

If the patch can't be applied, it leaves the versions in the index and
conflict markers in the working directory just like a failed merge.
And just like a merge, it remembers the commit message and provides it
as a default when I finally commit.

Note that this can only work on a chain of single-parent commits.
If a commit has multiple parents, there's no single patch to apply.


You van get the list of commits on a branch with git-log or git-rev-list,
but for more complex cases, the git-cherry tool is designed to generate
the list of commits to merge.  It has a rather neat approximate-match
function built in which identifies patches that appear to already be
present in the target branch.

* Rebasing

A special case of cherry-picking is if you want to move a whole branch
to a newer "base" commit.  This is done by git-rebase.  You specify
the branch to move (default HEAD) and where to move it to (no default),
and git cherry-picks every patch out of that branch, applies it on top
of the target, and moves the refs/heads/<branch> pointer to the newly
created commits.

By default, "the branch" is every commit back to the last common
ancestor of the branch head and the target, but you can override that
with command-line arguments.

If you want to avoid merge conflicts due to the master code changing out
from under your edits, but not have "cleanup" merges in your history,
git-rebase is the tool to use.

Git-rebase will also use git-rerere if enabled ("mkdir .git/rr-cache").


If rebasing encounters a conflict it can't resolve, it will stop halfway
and ask you to resolve the problem by hand.  However, it still knows it
has a job to finish!  The unapplied patches are remembered until you do
one of

	git-rebase --continue
		This will check in the current index.  You should
		do git-update-index <files> in the conflicts that
		you resolve, but NOT do an actual git-commit.
		git-rebase --continue will do the commit.
	git-rebase --skip
		This will skip the conflicting patch.  You
		don't have to resolve the conflicts; git will
		just back up and try the next patch in the series.
	git-rebase --abort
		This will abandon the whole rebase operation (including
		any half-done work) and return you to where you began.


Git-rebase can also help you divide up work.  Suppose you've mixed up
development of two features in the current HEAD, a branch called "dev".
You want to divide them up into "dev1" and "dev2".  Assuming that HEAD
is a branch off master, then you can either look through

	git log master..HEAD
or just get a raw list of the commits with
	git rev-list master..HEAD

Either way, suppose you figure out a list of commits that you want in
dev1 and create that branch:

	git checkout -b dev1 master
	for i in `cat commit_list`; do
		git-cherry-pick $i
	done

You can use the other half of the list you edited to generate the dev2
branch, but if you're not sure if you forgot something, or just don't
feel like doing that manual work, then you can use git-rebase to do it
for you...

	git checkout -b dev2 dev	# Create dev2 branch
	git-rebase --onto master dev1	# Subreact dev1 and rebase

This will find all patches that are in dev and not in dev1,
apply them on top of master, and call the result dev2.


* Experimenting with merging

To play with non-trivial merging, get an existing git repository of
a non-trivial project (git itself and the Linux kernel are readily
available.  Fire up gitk to look at history, find some interesting-looking
merges, and redo them yourself on a test branch.

As long as you do everything on test branches, you aren't going to screw
anything up.  So play!

You can use gitk to search for "Conflicts:" in the commit comments to
find merges that didn't go smoothly and see what happens.  (Or you can
search in "git log" output.  gitk just draws prettier pictures.)

You can also set up two repositories on the same machine and try pulling
and pushing between them.

To identify arbitrary commits, the 40-byte raw hex ID is probably easiest;
you can cut-and-paste them from the gitk window.

For example, in the git repository,
3f69d405d749742945afd462bff6541604ecd420

looks like an interesting merge.  Its parents are
Parent: 7d55561986ffe94ca7ca22dc0a6846f698893226
Parent: 097dc3d8c32f4b85bf9701d5e1de98999ac25c1c

Let's try doing that manually:

$ git checkout -b test 7d55561986ffe94ca7ca22dc0a6846f698893226
$ git pull . 097dc3d8c32f4b85bf9701d5e1de98999ac25c1c
error: no such remote ref refs/heads/097dc3d8c32f4b85bf9701d5e1de98999ac25c1c
Fetch failure: .

Cool!  I didn't know that wasn't allowed.  (I'll have to ask why it's
not; perhaps it's because it uses the branch name in the automatic
commit message.)  I could do it by hand with git-merge, but I'll just
give it a branch name:

$ git branch test2 097dc3d8c32f4b85bf9701d5e1de98999ac25c1c
$ git pull . test2
Merging HEAD with 097dc3d8c32f4b85bf9701d5e1de98999ac25c1c
Merging:
7d55561986ffe94ca7ca22dc0a6846f698893226 Merge branch 'jc/dirwalk-n-cache-tree' into jc/cache-tree
097dc3d8c32f4b85bf9701d5e1de98999ac25c1c Remove "tree->entries" tree-entry list from tree parser
found 2 common ancestor(s):
d9b814cc97f16daac06566a5340121c446136d22 Add builtin "git rm" command
288c0384505e6c25cc1a162242919a0485d50a74 Merge branch 'js/fetchconfig'
  Merging:
  d9b814cc97f16daac06566a5340121c446136d22 Add builtin "git rm" command
  288c0384505e6c25cc1a162242919a0485d50a74 Merge branch 'js/fetchconfig'
  found 1 common ancestor(s):
  63dffdf03da65ddf1a02c3215ad15ba109189d42 Remove old "git-grep.sh" remnants
  Auto-merging Makefile
merge: warning: conflicts during merge
  CONFLICT (content): Merge conflict in Makefile
  Auto-merging builtin.h
merge: warning: conflicts during merge
  CONFLICT (content): Merge conflict in builtin.h
  Auto-merging cache.h
  Removing check-ref-format.c
  Auto-merging git.c
merge: warning: conflicts during merge
  CONFLICT (content): Merge conflict in git.c
  Auto-merging read-cache.c
  Auto-merging update-index.c
merge: warning: conflicts during merge
  CONFLICT (content): Merge conflict in update-index.c
Renaming apply.c => builtin-apply.c
Auto-merging builtin-apply.c
Renaming read-tree.c => builtin-read-tree.c
Auto-merging builtin-read-tree.c
Auto-merging .gitignore
Auto-merging Makefile
merge: warning: conflicts during merge
CONFLICT (content): Merge conflict in Makefile
Auto-merging builtin.h
merge: warning: conflicts during merge
CONFLICT (content): Merge conflict in builtin.h
Auto-merging cache.h
Auto-merging fsck-objects.c
Removing git-format-patch.sh
Auto-merging git.c
merge: warning: conflicts during merge
CONFLICT (content): Merge conflict in git.c
Auto-merging update-index.c
Automatic merge failed; fix conflicts and then commit the result.

$ git status
Hey, look, lots of interesting stuff.  Particularly, see
# Changed but not updated:
#   (use git-update-index to mark for commit)
#
#       unmerged: Makefile
#       modified: Makefile
#       unmerged: builtin.h
#       modified: builtin.h
#       unmerged: git.c
#       modified: git.c

The "unmerged" (a.k.a. "staged") files are ones that need manual resolution.

(I notice that update-index.c isn't listed, despite being mentioned
as a conflict in the message.  Can someone explain that?)

Fixing those is easy, but as you can see from the original commit comment
and diffs, there were some additional changes that were necessary to
make that compile.

You can test before committing the change, or do it the git way - commit
anyway, then test and "git commit --amend" with the fixes, of any.

Unlike a centralized VCS, committing is not the same as pushing upstream.
You can use test branches in the repository to save as much work as
you like.  While it's still nice to keep the public repository clean,
you don't have to worry about "breaking the tree" every time you commit.
You can do all kinds of stuff in test branches, and clean it up later.

This is why all the git merge tools do the commit without waiting for
you to test it.  The merge is usually okay, and it saves time.  If not,
Jakub Narebski· Nov 17, 2006, 09:37 UTC · re: linux@horizon.com · lore
linux@horizon.com wrote:
> Either way, they're just a 41-byte file that contains a 40-byte hex
> object ID, plus a newline.  Tags are stored in .git/refs/tags, and heads
> are stored in .git/refs/heads.  Creating a new branch is literally just
> picking a file name and writing the ID of an existing commit into it.

This is an implementation detail, and is not true in repository with packed refs. Although usually (by default) only tags are packed.

But it remains true that ref (be it branch or tag) is just name and ID.
> The git programs enforce the immutability of tags, but that's a safety
> feature, not something fundamental.  You can rename a tag to the heads
> directory and go wild.

You can have only refs to commit objects in heads directory (and I hope this is verified by fsck-objects), you can have refs to tag objects (heavyweight tags), to commits (lightweight tags), to blobs (for example public PGP key used for signing tags), to trees (I guess unused).

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Jakub Narebski· Nov 17, 2006, 09:41 UTC · re: linux@horizon.com · lore
linux@horizon.com wrote:
> There is always a current head, known as HEAD.  (This is actually a
> symbolic link, .git/HEAD, to a file like refs/heads/master.)

Usually this is symref, not symlink, i.e. .git/HEAD (or rather $GIT_DIR/HEAD) is a file which contains single line like this:

  ref: refs/heads/master

There is a talk about relaxing HEAD restriction to allow it to contain ref to tag, or bare SHA1 id for "seeking"; you are forbidden to commit to such state.

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Jakub Narebski· Nov 17, 2006, 10:37 UTC · re: linux@horizon.com · lore
linux@horizon.com wrote:
Show 6 quoted lines
> * Remotes files
> 
> You can specify what to fetch on the git-fetch command line.  However,
> if you intend to monitor another repository on an ongoing basis,
> it's generally easier to set up a short-cut by placing the options in
> .git/remotes/<name>.

You can also set up this in config file (remote and branch sections), in modern git.

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Theodore Tso· Nov 17, 2006, 15:32 UTC · re: linux@horizon.com · lore
On Thu, Nov 16, 2006 at 05:17:01PM -0500, linux@horizon.com wrote:
Show 7 quoted lines
> I know it took me a while to get used to playing with branches, and I
> still get nervous when doing something creative.  So I've been trying
> to get more comfortable, and wrote the following to document what I've
> learned.
> 
> It's a first draft - I just finished writing it, so there are probably
> some glaring errors - but I thought it might be of interest anyway.

This is really, really good stuff that you've written! Have you any thoughts or suggestions about where this text should end up? Personally, I think this information is actually more important to an end-user than the current "part two" of the tutorial, which discusses the object database and the index file. Perhaps this should be "part 2", and the object database and index file should become "part 3"?

It might also be a good to consider moving some of the "discussion" portion the top-level git(7) man page into the object database and index file discussion. Right now, the best way to introduce git's concepts (IMHO), is to start with the part 1 of the tutorial, then go into the your draft branch/merging with git, then the current part 2 of the tutorial, and then direct folks to read the "discussion" section of git(7). Only then do they really have enough background understanding of the fundamental concepts of git that they won't get confused when they start talking to other git users, on the git mailing list, for example.

It would be nice if there was an easy way to direct users through the documentation in a way which makes good pedagogical sense. Right now, one of the reasons why life gets hard for new users is that the current tutorials aren't enough for them to really undersatnd what's going on at a conceptual level. And if users start using "everyday git" as a crutch, without the right background concepts, the human brain naturally tries to intuit what's happening in the background, but without reading the background docs, git is different enough that they will probably get it wrong, which means more stuff that they have to unlearn later.

Show 9 quoted lines
> * Git's representation of history
> 
> As you recall from Git 101, there are exactly four kinds of objects in
> Git's object database.  All of them have globally unique 40-character hex
> names made by hashing their type and contents.  Blob objects record file
> contents; they contain bytes.  Tree objects record directory contents;
> they contain file names, permissions, and the associated tree or blob
> object names.  Tag objects are shareable pointers to other objects;
> they're generally used to store a digital signature.

Hmm... this assumes that you've read the Git(7) discussion first. There is enough information here though that maybe you don't need to say "as you recall". It might be enough to give a quick summary of the concepts that are needed to understand the rest of your tutorial, and then point to git(7) Discussion section for people who need to learn more details.

Show 5 quoted lines
> * Remotes files
> 
> Note that branches to fetch are identified by "Pull: " lines in the
> remotes file.  This is another example of the fetch/pull confusion.
> git-pull will be explained eventually.

Maybe we should change git so that a "Fetch: " line in the remotes file works the same way as "Pull: ", and then recommend that people use "Fetch: " in order to reduce confusion, as opposed to simply explaining it away as "yet another example of the histororical fetch/pull confusion"?

Thanks,
Sean· Nov 17, 2006, 15:57 UTC · re: Theodore Tso · lore

On Fri, 17 Nov 2006 10:32:46 -0500 Theodore Tso <tytso@mit.edu> wrote:

Show 10 quoted lines
> It would be nice if there was an easy way to direct users through the
> documentation in a way which makes good pedagogical sense.  Right now,
> one of the reasons why life gets hard for new users is that the
> current tutorials aren't enough for them to really undersatnd what's
> going on at a conceptual level.  And if users start using "everyday
> git" as a crutch, without the right background concepts, the human
> brain naturally tries to intuit what's happening in the background,
> but without reading the background docs, git is different enough that
> they will probably get it wrong, which means more stuff that they have
> to unlearn later.  

It would be nice to post this information on the Git website and not have it overshadowed by Cogito examples with paragraphs explaining how Cogito makes things easier. The current website distracts users away from learning Git or ever reading about this kind of information. Maybe we can pass a hat around for some funds for a separate Cogito website. ;o)

Show 5 quoted lines
> Maybe we should change git so that a "Fetch: " line in the remotes
> file works the same way as "Pull: ", and then recommend that people
> use "Fetch: " in order to reduce confusion, as opposed to simply
> explaining it away as "yet another example of the histororical
> fetch/pull confusion"?
 
That's quite a good idea.  The name was fixed when the option to move
this info into the config file was added (remote.<name>.fetch).  So
another option would be to show new users the config file method and
just damn the remotes file to a historical footnote.
Nguyen Thai Ngoc Duy· Nov 17, 2006, 16:19 UTC · re: Sean · lore
On 11/17/06, Sean <seanlkml@sympatico.ca> wrote:
Show 6 quoted lines
> It would be nice to post this information on the Git website and not
> have it overshadowed by Cogito examples with paragraphs explaining how
> Cogito makes things easier.  The current website distracts users away
> from learning Git or ever reading about this kind of information.
> Maybe we can pass a hat around for some funds for a separate Cogito
> website. ;o)

Or.. find a way to merge cogito back to git :-) /me runs into a nearest bush.

Marko Macek· Nov 17, 2006, 16:25 UTC · re: Nguyen Thai Ngoc Duy · lore
Nguyen Thai Ngoc Duy wrote:
Show 10 quoted lines
> On 11/17/06, Sean <seanlkml@sympatico.ca> wrote:
>> It would be nice to post this information on the Git website and not
>> have it overshadowed by Cogito examples with paragraphs explaining how
>> Cogito makes things easier.  The current website distracts users away
>> from learning Git or ever reading about this kind of information.
>> Maybe we can pass a hat around for some funds for a separate Cogito
>> website. ;o)
> 
> Or.. find a way to merge cogito back to git :-)
> /me runs into a nearest bush.

I agree, this would certainly be the best solution. But it would imply hiding the 'index' by default which would probably an incompatible change.

The alternative would be to explain that git is a low level tool suitable mostly for integrators like Linus (that, and that Cogito and/or StGit should be used by developers/contributors).

Petr Baudis· Nov 17, 2006, 16:33 UTC · re: Marko Macek · lore
On Fri, Nov 17, 2006 at 05:25:25PM CET, Marko Macek wrote:
Show 11 quoted lines
> Nguyen Thai Ngoc Duy wrote:
> >On 11/17/06, Sean <seanlkml@sympatico.ca> wrote:
> >>It would be nice to post this information on the Git website and not
> >>have it overshadowed by Cogito examples with paragraphs explaining how
> >>Cogito makes things easier.  The current website distracts users away
> >>from learning Git or ever reading about this kind of information.
> >>Maybe we can pass a hat around for some funds for a separate Cogito
> >>website. ;o)
> >
> >Or.. find a way to merge cogito back to git :-)
> >/me runs into a nearest bush.

I think we are trying to figure that out in the last few days in those mammoth threads. UI-wise with no big breakthroughs so far I guess, though.

> The alternative would be to explain that git is a low level tool suitable 
> mostly for integrators like Linus (that, and that Cogito and/or StGit 
> should be used by developers/contributors).

This is in essence what many people (including Junio) are saying. I'm not saying it's a totally great situation, hence the previous paragraph.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
#!/bin/perl -sp0777i<X+d*lMLa^*lN%0]dsXx++lMlN/dsM0<j]dsj
$/=unpack('H*',$_);$_=`echo 16dio\U$k"SK$/SM$n\EsN0p[lN*1
Sean· Nov 17, 2006, 16:34 UTC · re: Nguyen Thai Ngoc Duy · lore

On Fri, 17 Nov 2006 23:19:23 +0700 "Nguyen Thai Ngoc Duy" <pclouds@gmail.com> wrote:

> Or.. find a way to merge cogito back to git :-)
> /me runs into a nearest bush.

Pasky has already given a lot to Git, and it would be great to see even more merged back into Git where a consensus can be reached. In fact Pasky has said that his plan is to push a lot more towards Git and make Cogito a thinner UI layer. Either way, there's absolutely nothing wrong with people choosing to use Cogito rather than Git. It's just that the separate Cogito tool shouldn't have a place on the Git website any more prominent than say StGit does.

The Git website should be a place where Git makes the best case it can for _itself_, not for its sister tools. It's a distraction and gets in the way of promoting Git as a stand alone tool. At least one new user has complained that it was confusing.

Personally I have nothing against Cogito, I just think Pasky should separate his role as Git webmaster from his role as Cogito author. If people have good ideas for Git documentation, the website would be a natural place for it, and it shouldn't have to compete with Cogito tutorials etc.

J. Bruce Fields· Nov 17, 2006, 17:44 UTC · re: linux@horizon.com · lore

This has some useful material that fills gaps in the existing documentation. We need to think a little more about the intended audience, and about how to fit it in with existing documentation.

On Thu, Nov 16, 2006 at 05:17:01PM -0500, linux@horizon.com wrote:
Show 33 quoted lines
> * A brief digression on command names.
> 
> Originally, all git commands were named "git-foo".  When there got to
> be over a hundred, people started complaining about the clutter in
> /usr/bin.  After some discussion, the following solution was reached:
> 
> - It's now possible to place all of the git-foo commands into a separate
>   directory.  (Despite the complaints, not too many people are doing it
>   yet.)
> - One option for git users is to add that directory to their $PATH.
> - Another is provided by a wrapper called just "git".  It's intended to
>   live in a public directory like /usr/bin, and knows the location of
>   the separate directory.  When you type "git foo", it finds and executes
>   "git-foo".
> - Some simple commands are built into the git wrapper.  When you type
>   "git add", it just does it internally.  (On the git mailing list,
>   you will see patches like "make git diff a builtin"; this is what
>   they're talking about.)
> - For compatibility, for each builtin, there is a "git-add" file,
>   which is just a link to the "git" wrapper.  It looks at the name it
>   was invoked as to figure out what it should do.
> 
> The one confusing thing is that, although people usually type "git foo"
> in examples, they're interchangeable in practice.  I go back and forth
> for no good reason.  The main caveat is that to get the man page, you
> still need to type "man git-foo".  Fortunately, there are two other ways
> to get the man page:
> 
> 	1) "git help foo"
> 	2) "git foo --help"
> 
> Git doesn't have a specialized built-in help system; it just shows you
> the man pages.

Who's the audience for the above? I can see that it's useful for administrators, who may need help deciding how to install stuff, and for developers, who need to know where the heck the code for "git-add" came from. But the case I'm most interested in is the user whose distribution installs git for them, in which case I think the above could be distilled down to:

	- "git-foo" and "git foo" can be used interchangeably.
	- Documentation for the command foo is available from any of
		- man git-foo
		- git help foo
		- git foo --help

Then the additional details above could be postponed to a later part of the documentation.

Show 8 quoted lines
> One outstanding problem with git's man pages is that often the most detail
> is in the command page that was written first, not the user-friendly
> one that you should use.  For example, there are a number of special
> cases of the "git diff" command that were written first, and the man
> pages for these commands (git-diff-index, git-diff-files, git-diff-tree,
> and git-diff-stages) are considerably more informative than the page for
> plain git-diff, even though that's the command that you should use 99%
> of the time.

I agree that that's helpful. Though we should probably also be working on the man pages to make this organization clearer.

> As you recall from Git 101

Obviously a more specific reference would be more useful here--if there's nothing useful to point to among the existing documentation, we should figure out how to fix that problem.

That might also remove the need for some of the recap that follows.
> there are exactly four kinds of objects in
> Git's object database.  All of them have globally unique 40-character hex
....
> Finally, there are references, stored in the .git/refs directory.
> These are the human-readable names associated with commits, and the
> "root set" from which all other commits should be reachable.

This is good; a comprehensive discussion of references will fill a gap in the current documentation.

....
Show 8 quoted lines
> * Naming revisions
> 
> CVS encourages you to tag like crazy, because the only other way to
> find a given revision is by date.  Git makes it a lot easier, so most
> revisions don't need names.
> 
> You can find a full description in the git-rev-parse man page, but here's
> a summary.

This has a lot more overlap with existing documentation. The extra detail is useful, but we need to decide what our audience and goal is here, to decide exactly what niche we're trying to fill between the brief stuff that's in the tutorial part I and the details in "man git-rev-parse".

Show 12 quoted lines
> * Converting between names
> 
> Git has two helpers (programs designed mainly for use in shell scripts)
> to convert between global object IDs and human-readable names.
> 
> The first is git-rev-parse.  This is a general git shell script helper,
> which validates the command line and converts object names to absolute
> object IDs.  Its man page has a detailed description of the object
> name syntax.
> 
> The second is git-name-rev, which converts the other way around.  It's
> particularly useful for seeing which tags a given commit falls between.
Also discuss git-describe?
> * The three uses of "git checkout"

Obviously there's a lot of overlap here with "man git-checkout". What's the goal here? Maybe this should just be worked in to a revision of that man page?

Show 5 quoted lines
> * Deleting branches
> 
> "git branch -d <head>" is safe.  It deletes the given <head>, but first
> it checks that the commit is reachable some other way.  That is, you
> merged the branch in somewhere, or you never did any edits on that branch.

It only checks whether the head of the branch to delete is reachable from the *current* branch. The man page could be clearer here.

....
> * Examining history: git-log and git-rev-list

Yep, we should definitely have a good long chapter just devoted to history examination. Most of it could be just cool examples, so it would be fun.

Note some of this is done in the last half of cvs-migration.txt; we should mine that section for whatever's useful and then replace by a reference to the new chapter.

> * History diagrams
...
> * Trivial merges: fast-forward and already up-to-date.
These two sections are useful, yep.
> * Exchanging work with other repositories, part II: git-push

There's a lot of overlap here with cvs-migration.txt. Maybe some better organization is needed to make that more prominent.

> The details are too advanced for this discussion, but the default
> "recursive" merge strategy that git uses solves the answer by merging
> a and b into a temporary commit and using *that* as the merge base.

I'm tempted to ignore any description of the merge strategy, or postpone it till later; as a first pass I think it's better just to say "obvious cases will be handled automatically, and you'll be prompted for comments." Only other SCM developers are going to wonder how you handle the corner cases.

> * When merging goes wrong
But yes, I think people could use more help on how to resolve merges.
> * Test merging
...
> * Cherry picking
...
> * Rebasing
Yup, I agree that that's good material to cover together.
Jakub Narebski· Nov 17, 2006, 18:16 UTC · re: J. Bruce Fields · lore
J. Bruce Fields wrote:
Show 6 quoted lines
> This has some useful material that fills gaps in the existing
> documentation.  We need to think a little more about the intended
> audience, and about how to fit it in with existing documentation.
> 
> On Thu, Nov 16, 2006 at 05:17:01PM -0500, linux@horizon.com wrote:
>> * A brief digression on command names.
Show 5 quoted lines
> But the case I'm most interested in is the user whose
> distribution installs git for them, in which case I think the above
> could be distilled down to:
> 
>       - "git-foo" and "git foo" can be used interchangeably.

But it is encouraged (also for example by git-completion.bash) to use "git foo" form in command line (because git commands can be not in the PATH, although usually they are), and "git-foo" form in scripts (if possible).

Show 9 quoted lines
>> The details are too advanced for this discussion, but the default
>> "recursive" merge strategy that git uses solves the answer by merging
>> a and b into a temporary commit and using *that* as the merge base.
> 
> I'm tempted to ignore any description of the merge strategy, or postpone
> it till later; as a first pass I think it's better just to say "obvious
> cases will be handled automatically, and you'll be prompted for
> comments."  Only other SCM developers are going to wonder how you handle
> the corner cases.
See below...
 
>> * When merging goes wrong
> 
> But yes, I think people could use more help on how to resolve merges.

It would be useful to cover all non-reductible cases of recursive merge strategy (the default merge strategy for two-head merges) conflicts: contents (covered), add/add, rename/modify etc.

So some info about recirsive merge strategy would be useful.
-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Theodore Tso· Jan 3, 2007, 17:04 UTC · re: linux@horizon.com · lore
On Thu, Nov 16, 2006 at 05:17:01PM -0500, linux@horizon.com wrote:
> I know it took me a while to get used to playing with branches, and I
> still get nervous when doing something creative.  So I've been trying
> to get more comfortable, and wrote the following to document what I've
> learned.

What ever happened to this document? There was some talk of getting this integrated into the git tree as Docmentation/tutorial-3.txt. IMHO it would be really, really good to do this before 1.5.0, since I think a lot of users would find it really useful. Some of the text may need to be moved to other locations, but it might go faster if we get the base document into the tree first, and then we can submit patches to move text around to integrate it into the other documentation files.

I'm certainly willing to help out submitting patches to improve the documentation, and I think this would be a big step towards helping new users to git become much more quickly proficient.

						- Ted
Junio C Hamano· Jan 3, 2007, 17:08 UTC · re: Theodore Tso · lore
Theodore Tso <tytso@mit.edu> writes:
Show 10 quoted lines
> On Thu, Nov 16, 2006 at 05:17:01PM -0500, linux@horizon.com wrote:
>> I know it took me a while to get used to playing with branches, and I
>> still get nervous when doing something creative.  So I've been trying
>> to get more comfortable, and wrote the following to document what I've
>> learned.
>
> What ever happened to this document?  There was some talk of getting
> this integrated into the git tree as Docmentation/tutorial-3.txt.
> IMHO it would be really, really good to do this before 1.5.0, since I
> think a lot of users would find it really useful.
Seconded.  Can I have the latest round?
J. Bruce Fields· Jan 7, 2007, 23:44 UTC · re: Theodore Tso · lore
On Wed, Jan 03, 2007 at 12:04:11PM -0500, Theodore Tso wrote:
> What ever happened to this document?  There was some talk of getting
> this integrated into the git tree as Docmentation/tutorial-3.txt.
Just to throw more fuel on the fire....
I have a draft attempt at a complete "git user's manual" at
	http://www.fieldses.org/~bfields/
The goals are:
	- Readable from beginning to end in order without having read
	  any other git documentation beforehand.
	- Helpful section names and cross-references, so it's not too
	  hard to skip around some if you need to.
	- Organized to allow it to grow much larger (unlike the
	  tutorials)

It's more liesurely than tutorial.txt, but tries to stay focused on practical how-to stuff. It adds a discussion of how to resolve merge conflicts, and partial instructions on setting up and dealing with a public repository.

I've lifted a little bit from "branching and merging" (e.g., some of the discussion of history diagrams), and could probably steal more if that's OK. (Similarly anyone should of course feel free to reuse bits of this if any parts seem more useful than the whole.)

There's a lot of detail on managing branches and using git-fetch, just because those are essential even to people needing read-only access (e.g., kernel testers). I think those sections will be much shorter once the new "git remote" command and the disconnected checkouts are taken into account.

I do feel bad about adding yet another piece of documentation, but I we need something that goes through all the basics in a logical order, and I wasn't seeing how to grow the tutorials into that.

Opinions?
--b.
Junio C Hamano· Jan 8, 2007, 00:24 UTC · re: J. Bruce Fields · lore
"J. Bruce Fields" <bfields@fieldses.org> writes:
Show 5 quoted lines
> I do feel bad about adding yet another piece of documentation, but I we
> need something that goes through all the basics in a logical order, and
> I wasn't seeing how to grow the tutorials into that.
>
> Opinions?

I was having the feeling that we need to start over the documentation from a clean slate by first coming up with a coherent presentation order and then filling sections in it, instead of tweaking existing documents here and there. The existing documents were written in different development stages of git, and each document tries to be more or less independent from others in the area it wants to talk about, and reading all of them in _any_ order is not the best way to learn git because of duplication. Also I suspect some information in older documents, while being still valid and technically correct, predates invention of a better/simpler alternative.

In other words, I think we have enough information in the tutorial documents, but the problem is not the lack of information -- the problem is the lack of organization.

I think this effort of yours is wonderful because it directly tackles that problem.

J. Bruce Fields· Jan 8, 2007, 02:35 UTC · re: Junio C Hamano · lore
On Sun, Jan 07, 2007 at 04:24:08PM -0800, Junio C Hamano wrote:
Show 7 quoted lines
> "J. Bruce Fields" <bfields@fieldses.org> writes:
> In other words, I think we have enough information in the
> tutorial documents, but the problem is not the lack of
> information -- the problem is the lack of organization.
> 
> I think this effort of yours is wonderful because it directly
> tackles that problem.

OK, thanks for the vote of confidence.... My tentative organization (which I'm totally open to argument about) is:

chapters 1 and 2: "Read-only" operations:
	clone, fetch, the commit DAG, etc.; material that could be
	useful to a linux kernel tester, for example.  This also
	includes lots of stuff about branch manipulation and fetching,
	just because that's necessary to keep a repo up to date and
	check out random commits.  Once we have "git remote" and
	disconnected checkouts most of this could be postponed till
	later.
Chapter 3: "Read-write" operations:
	Read-write stuff: creating commits (basic mention of index),
	handling merges, git-gc, ending with distributed stuff:
	importing and exporting patches, pull and push, etc.
Chapter 4 (unwritten): interactions with other VCS's
	cvs, subversion.  Also some of us use track projects with git
	even when all we've got is a sequence of release tarballs to
	track, and that might be worth documenting.
Chapter 5 (unwritten): rewriting history
	rebasing, cherry-picking, managing patch series, etc.
Chapter 6 (unwritten): git internals
	I intend to just do a wholesale import of either tutorial-2.txt,
	core-tutorial.txt, or the README, or some combination thereof,
	but can't decide which.
--b.
David Kågedal· Jan 8, 2007, 13:04 UTC · re: J. Bruce Fields · lore
"J. Bruce Fields" <bfields@fieldses.org> writes:
> OK, thanks for the vote of confidence....  My tentative organization
> (which I'm totally open to argument about) is:
>
> chapters 1 and 2: "Read-only" operations:
> Chapter 3: "Read-write" operations:
> Chapter 4 (unwritten): interactions with other VCS's

I think this should be considered more peripheral, since it is really an independent piece, and nobody needs to read it to learn how git works. So I would probably move it to the end.

> Chapter 5 (unwritten): rewriting history
> Chapter 6 (unwritten): git internals
-- 
David Kågedal
Theodore Tso· Jan 8, 2007, 14:03 UTC · re: J. Bruce Fields · lore
On Sun, Jan 07, 2007 at 09:35:11PM -0500, J. Bruce Fields wrote:
Show 9 quoted lines
> chapters 1 and 2: "Read-only" operations:
> 
> 	clone, fetch, the commit DAG, etc.; material that could be
> 	useful to a linux kernel tester, for example.  This also
> 	includes lots of stuff about branch manipulation and fetching,
> 	just because that's necessary to keep a repo up to date and
> 	check out random commits.  Once we have "git remote" and
> 	disconnected checkouts most of this could be postponed till
> 	later.

I would add a QuickStart Chapter before you start going into the "read-only" oeperations. It would show how to create a completely empty repository, and add a few commits. It would also demonstrate how to clone an example repository (with a fixed set of contents, stored at git://git.kernel.org/pub/scm/git/example and add a commit using "git commit -a".

The basic idea is to show the user that git really isn't that hard, *before* you start diving into a lot of details. If you don't tell a user how to make a commit until Chapter 3, he/she will assume it's because it's Really Hard, and you may end up losing them before that.

Show 5 quoted lines
> Chapter 3: "Read-write" operations:
> 
> 	Read-write stuff: creating commits (basic mention of index),
> 	handling merges, git-gc, ending with distributed stuff:
> 	importing and exporting patches, pull and push, etc.

At least some discussions of branches needs to happen here; it's really important to talk about different workflows, and how you use branches as part of your read-write operations. Some folks might or might not use topic branches, but the concept of using temporary branches to try things out is critical.

Show 11 quoted lines
> Chapter 4 (unwritten): interactions with other VCS's
> 
> 	cvs, subversion.  Also some of us use track projects with git
> 	even when all we've got is a sequence of release tarballs to
> 	track, and that might be worth documenting.
> 
> Chapter 6 (unwritten): git internals
> 
> 	I intend to just do a wholesale import of either tutorial-2.txt,
> 	core-tutorial.txt, or the README, or some combination thereof,
> 	but can't decide which.
You might want to consider putting these two chapters into appendices.
						- Ted
J. Bruce Fields· Jan 9, 2007, 02:41 UTC · re: Theodore Tso · lore
On Mon, Jan 08, 2007 at 09:03:05AM -0500, Theodore Tso wrote:
Show 11 quoted lines
> I would add a QuickStart Chapter before you start going into the
> "read-only" oeperations.  It would show how to create a completely
> empty repository, and add a few commits.  It would also demonstrate
> how to clone an example repository (with a fixed set of contents,
> stored at git://git.kernel.org/pub/scm/git/example and add a commit
> using "git commit -a".
>
> The basic idea is to show the user that git really isn't that hard,
> *before* you start diving into a lot of details.  If you don't tell a
> user how to make a commit until Chapter 3, he/she will assume it's
> because it's Really Hard, and you may end up losing them before that.

Yeah, I agree. I just haven't been able to decide quite what to choose for that purpose. Some choices:

	- We could just pare down the tutorial a bit and drag it in as
	  chapter one.
	- I tried writing something modeled loosely on the hg quick
	  start.  It's a little out of date now, but that could be
	  fixed:
		http://www.fieldses.org/~bfields/git-quick-start.html
	- Or maybe a revised everyday.txt would do the job?
Any opinions?
> At least some discussions of branches needs to happen here;

The basic nuts-and-bolts (how to create and delete branches, etc.) should all be covered, of course, but....

> it's really important to talk about different workflows, and how you
> use branches as part of your read-write operations.  Some folks might
> or might not use topic branches, but the concept of using temporary
> branches to try things out is critical.

.... Maybe it'd be fun to have a section called just "examples" at the end of each chapter. The sort of thing you're describing could fit in well there. I'd need some help collecting interesting examples.

--b.
Andreas Ericsson· Jan 9, 2007, 08:46 UTC · re: J. Bruce Fields · lore
J. Bruce Fields wrote:
Show 25 quoted lines
> On Mon, Jan 08, 2007 at 09:03:05AM -0500, Theodore Tso wrote:
>> I would add a QuickStart Chapter before you start going into the
>> "read-only" oeperations.  It would show how to create a completely
>> empty repository, and add a few commits.  It would also demonstrate
>> how to clone an example repository (with a fixed set of contents,
>> stored at git://git.kernel.org/pub/scm/git/example and add a commit
>> using "git commit -a".
>>
>> The basic idea is to show the user that git really isn't that hard,
>> *before* you start diving into a lot of details.  If you don't tell a
>> user how to make a commit until Chapter 3, he/she will assume it's
>> because it's Really Hard, and you may end up losing them before that.
> 
> Yeah, I agree.  I just haven't been able to decide quite what to choose
> for that purpose.  Some choices:
> 
> 	- We could just pare down the tutorial a bit and drag it in as
> 	  chapter one.
> 
> 	- I tried writing something modeled loosely on the hg quick
> 	  start.  It's a little out of date now, but that could be
> 	  fixed:
> 
> 		http://www.fieldses.org/~bfields/git-quick-start.html
> 

I like this, although fetch should probably have "--force" instead of the "+branch" notation. --force stands out more and users are familiar with --force possibly destroying things (rm -rf, anyone?).

> 	- Or maybe a revised everyday.txt would do the job?
> 
> Any opinions?
> 

I think the document is fine as it is, but could probably start off with a link to the tutorial, quickstart or a revised version of everyday.txt, stating that "here's something you might want to read if you prefer to experiment. If you think something goes wrong, come back here and find out why".

Show 5 quoted lines
>> At least some discussions of branches needs to happen here;
> 
> The basic nuts-and-bolts (how to create and delete branches, etc.)
> should all be covered, of course, but....
> 

I found it quite sufficient. Perhaps it would be nice to include some more advanced examples, like octopus merges and things like that, although I feel such things could well live in an appendix to keep all the easy operations up front. Most people I know will most likely *never* use octopus merges. 90% of the merges we do here at work result in fast-forwards, so a real merge is already considered a bit odd.

Show 9 quoted lines
>> it's really important to talk about different workflows, and how you
>> use branches as part of your read-write operations.  Some folks might
>> or might not use topic branches, but the concept of using temporary
>> branches to try things out is critical.
> 
> .... Maybe it'd be fun to have a section called just "examples" at the
> end of each chapter.  The sort of thing you're describing could fit in
> well there.  I'd need some help collecting interesting examples.
> 
Indeed. I for one like examples that tell me

# type this # this will happen # you can see what you just did with this, this, and this command # this is because...

Not only is it good for learning the how and the why, but it also trains the fingers right from the start. Hopefully the UI is stabilized enough by now that we can reliably tell users how to accomplish a certain thing. UI changes must almost certainly be listed at whatever official site git has. As Junio has already pointed out, the members of the git mailing list are now in minority among the git users, so some other place has to hold the user-visible changes as well and the location of that site must probably be published along with the tools.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231
J. Bruce Fields· Jan 9, 2007, 15:49 UTC · re: Andreas Ericsson · lore
On Tue, Jan 09, 2007 at 09:46:06AM +0100, Andreas Ericsson wrote:
Show 11 quoted lines
> J. Bruce Fields wrote:
> >	- I tried writing something modeled loosely on the hg quick
> >	  start.  It's a little out of date now, but that could be
> >	  fixed:
> >
> >		http://www.fieldses.org/~bfields/git-quick-start.html
> >
> 
> I like this, although fetch should probably have "--force" instead of 
> the "+branch" notation. --force stands out more and users are familiar 
> with --force possibly destroying things (rm -rf, anyone?).

I started out writing it that way (for the reasons you give), then changed it on the theory starting out with the "+" notation would make it simpler explaining how to do the remote configuration.

Now that there's git-remote, and less need to manipulate the remote configuration by hand, maybe that's less important.

Show 5 quoted lines
> I think the document is fine as it is, but could probably start off with 
> a link to the tutorial, quickstart or a revised version of everyday.txt, 
> stating that "here's something you might want to read if you prefer to 
> experiment. If you think something goes wrong, come back here and find 
> out why".
Sounds sensible.
Show 9 quoted lines
> Indeed. I for one like examples that tell me
> 
> # type this
> # this will happen
> # you can see what you just did with this, this, and this command
> # this is because...
> 
> Not only is it good for learning the how and the why, but it also trains 
> the fingers right from the start.
OK.  This is a place where I'd really appreciate any contributions.
--b.
Theodore Tso· Jan 9, 2007, 16:58 UTC · re: Andreas Ericsson · lore
On Tue, Jan 09, 2007 at 09:46:06AM +0100, Andreas Ericsson wrote:
Show 5 quoted lines
> I think the document is fine as it is, but could probably start off with 
> a link to the tutorial, quickstart or a revised version of everyday.txt, 
> stating that "here's something you might want to read if you prefer to 
> experiment. If you think something goes wrong, come back here and find 
> out why".

If what we're going to do is a "git user's manual", I'd recommend keeping the 2-3 pages in the manual, and do it via a link to some other document. One of the issues with the git documentation is that it's *too* branchy, and some the branches go off to some truly scary low-level implementation detail. If we are going to assume that isn't going to change (and I am glad that the low-level details are documented, and am not advocating that they be deleted), then keeping a user-friendly QuickStart in the main document might not be a bad decision.

						- Ted
J. Bruce Fields· Jan 10, 2007, 04:15 UTC · re: Theodore Tso · lore
On Tue, Jan 09, 2007 at 11:58:28AM -0500, Theodore Tso wrote:
Show 9 quoted lines
> If what we're going to do is a "git user's manual", I'd recommend
> keeping the 2-3 pages in the manual, and do it via a link to some
> other document.  One of the issues with the git documentation is that
> it's *too* branchy, and some the branches go off to some truly scary
> low-level implementation detail.  If we are going to assume that isn't
> going to change (and I am glad that the low-level details are
> documented, and am not advocating that they be deleted), then keeping
> a user-friendly QuickStart in the main document might not be a bad
> decision.
Sounds reasonable.

I'll probably set this aside a few days, then do some more work on it this weekend. (Patches welcomed, though--source is in the master branch of git://linux-nfs.org/~bfields/git.git.)

--b.
Theodore Tso· Jan 8, 2007, 00:40 UTC · re: J. Bruce Fields · lore
On Sun, Jan 07, 2007 at 06:44:11PM -0500, J. Bruce Fields wrote:
Show 9 quoted lines
> On Wed, Jan 03, 2007 at 12:04:11PM -0500, Theodore Tso wrote:
> > What ever happened to this document?  There was some talk of getting
> > this integrated into the git tree as Docmentation/tutorial-3.txt.
> 
> Just to throw more fuel on the fire....
> 
> I have a draft attempt at a complete "git user's manual" at
> 
> 	http://www.fieldses.org/~bfields/

Is that the right URL? That gets me to "Not Bruce's Webpage" and I don't see an obvious link to git documentation...

						- Ted
J. Bruce Fields· Jan 8, 2007, 00:46 UTC · re: Theodore Tso · lore
On Sun, Jan 07, 2007 at 07:40:06PM -0500, Theodore Tso wrote:
Show 13 quoted lines
> On Sun, Jan 07, 2007 at 06:44:11PM -0500, J. Bruce Fields wrote:
> > On Wed, Jan 03, 2007 at 12:04:11PM -0500, Theodore Tso wrote:
> > > What ever happened to this document?  There was some talk of getting
> > > this integrated into the git tree as Docmentation/tutorial-3.txt.
> > 
> > Just to throw more fuel on the fire....
> > 
> > I have a draft attempt at a complete "git user's manual" at
> > 
> > 	http://www.fieldses.org/~bfields/
> 
> Is that the right URL?  That gets me to "Not Bruce's Webpage" and I
> don't see an obvious link to git documentation...
Crap:
	http://www.fieldses.org/~bfields/git-user-manual.html
Sorry about that.--b.
Jakub Narebski· Jan 8, 2007, 01:22 UTC · re: J. Bruce Fields · lore
J. Bruce Fields wrote:
Show 13 quoted lines
> On Sun, Jan 07, 2007 at 07:40:06PM -0500, Theodore Tso wrote:
>> On Sun, Jan 07, 2007 at 06:44:11PM -0500, J. Bruce Fields wrote:
>>> 
>>> I have a draft attempt at a complete "git user's manual" at
>>> 
>>>     http://www.fieldses.org/~bfields/
>> 
>> Is that the right URL?  That gets me to "Not Bruce's Webpage" and I
>> don't see an obvious link to git documentation...
> 
> Crap:
> 
>       http://www.fieldses.org/~bfields/git-user-manual.html
Added to
  http://git.or.cz/gitwiki/GitDocumentation
  http://git.or.cz/gitwiki/GitLinks
-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Guilhem Bonnefille· Jan 8, 2007, 12:38 UTC · re: J. Bruce Fields · lore
On 1/8/07, J. Bruce Fields <bfields@fieldses.org> wrote:
Show 18 quoted lines
> On Sun, Jan 07, 2007 at 07:40:06PM -0500, Theodore Tso wrote:
> > On Sun, Jan 07, 2007 at 06:44:11PM -0500, J. Bruce Fields wrote:
> > > On Wed, Jan 03, 2007 at 12:04:11PM -0500, Theodore Tso wrote:
> > > > What ever happened to this document?  There was some talk of getting
> > > > this integrated into the git tree as Docmentation/tutorial-3.txt.
> > >
> > > Just to throw more fuel on the fire....
> > >
> > > I have a draft attempt at a complete "git user's manual" at
> > >
> > >     http://www.fieldses.org/~bfields/
> >
> > Is that the right URL?  That gets me to "Not Bruce's Webpage" and I
> > don't see an obvious link to git documentation...
>
> Crap:
>
>         http://www.fieldses.org/~bfields/git-user-manual.html
Nice work.

My only 2 cents: the SVN book is really a good book, as it contains both simple user and advanced hacker info. As it is in free licence, perhaps it could be possible to "port" the book to Git. I saw that the SVK book is such a port. But it's a DocBook document. http://svnbook.red-bean.com/

-- 
Guilhem BONNEFILLE
-=- #UIN: 15146515 JID: guyou@im.apinc.org MSN: guilhem_bonnefille@hotmail.com
-=- mailto:guilhem.bonnefille@gmail.com
-=- http://nathguil.free.fr/
J. Bruce Fields· Jan 9, 2007, 04:17 UTC · re: Guilhem Bonnefille · lore
On Mon, Jan 08, 2007 at 01:38:19PM +0100, Guilhem Bonnefille wrote:
> Nice work.
Thanks!
Show 5 quoted lines
> My only 2 cents: the SVN book is really a good book, as it contains
> both simple user and advanced hacker info. As it is in free licence,
> perhaps it could be possible to "port" the book to Git. I saw that the
> SVK book is such a port. But it's a DocBook document.
> http://svnbook.red-bean.com/
Thanks, yes, that does look very polished.

If there's any part you'd be particularly interested in seeing "ported", I'd be happy to help incorporate your work.

--b.
Petr Baudis· Nov 17, 2006, 16:53 UTC · lore
On Fri, Nov 17, 2006 at 05:34:04PM CET, Sean wrote:
> It's just that the separate Cogito tool shouldn't have a place on the
> Git website any more prominent than say StGit does.
It doesn't - look at the "Maintaining external patches" crash course.

Porcelains are integral part of the Git environment. I think several people have already tried to explain it before.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
#!/bin/perl -sp0777i<X+d*lMLa^*lN%0]dsXx++lMlN/dsM0<j]dsj
$/=unpack('H*',$_);$_=`echo 16dio\U$k"SK$/SM$n\EsN0p[lN*1
Sean· Nov 17, 2006, 17:01 UTC · re: Petr Baudis · lore

On Fri, 17 Nov 2006 17:53:33 +0100 Petr Baudis <pasky@suse.cz> wrote:

Show 9 quoted lines
> On Fri, Nov 17, 2006 at 05:34:04PM CET, Sean wrote:
> > It's just that the separate Cogito tool shouldn't have a place on the
> > Git website any more prominent than say StGit does.
> 
> It doesn't - look at the "Maintaining external patches" crash course.
> 
> Porcelains are integral part of the Git environment. I think several
> people have already tried to explain it before.
> 

There is enough native Git documentation and hopefully more coming that third party tools should be pushed behind the scenes a bit. At least on the GIT website.

Of course there is nothing wrong with having information there, but the main thrust should be about Git and how to use it directly without porcelains. Especially in the light that people have recently expressed a desire to advocate and document the use of native Git more strongly.

Having a link to Cogito off the front page of the Git website that says... Cogito makes things "easier", no matter how much you personally believe it, isn't the way everyone feels and is at odds with the native-git message and improvement effort.

Petr Baudis· Nov 17, 2006, 21:31 UTC · lore
On Fri, Nov 17, 2006 at 06:01:54PM CET, Sean wrote:
> There is enough native Git documentation and hopefully more coming
> that third party tools should be pushed behind the scenes a bit.
> At least on the GIT website.

It's not about documentation but ease to use. I agree and sympathise very much with the effort of making core Git more easy to use and obsoleting Cogito, but until it gets there we should have what's nicest to the users.

Show 5 quoted lines
> Of course there is nothing wrong with having information there, but
> the main thrust should be about Git and how to use it directly without
> porcelains.  Especially in the light that people have recently
> expressed a desire to advocate and document the use of native Git
> more strongly.

If someone writes a crash course in pure Git covering the same grounds as the current ones (possibly by just extending/retouching the tutorial) (it does not necessarily need to be a "refugee" crash course, it can build up from scratch), I can add it on the web. If it becomes as easy to use and with as mild learning curve as Cogito, it means Cogito got mostly obsolete and I'll happily remove the Cogito crash courses from the web.

> Having a link to Cogito off the front page of the Git website that
> says... Cogito makes things "easier", no matter how much you
> personally believe it, isn't the way everyone feels and is at
> odds with the native-git message and improvement effort.

If you disagree about that fact, can you provide some specific argumentation?

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
#!/bin/perl -sp0777i<X+d*lMLa^*lN%0]dsXx++lMlN/dsM0<j]dsj
$/=unpack('H*',$_);$_=`echo 16dio\U$k"SK$/SM$n\EsN0p[lN*1
Chris Riddoch· Nov 17, 2006, 22:36 UTC · re: Petr Baudis · lore
On 11/17/06, Petr Baudis <pasky@suse.cz> wrote:
Show 7 quoted lines
> If someone writes a crash course in pure Git covering the same grounds
> as the current ones (possibly by just extending/retouching the tutorial)
> (it does not necessarily need to be a "refugee" crash course, it can
> build up from scratch), I can add it on the web. If it becomes as easy
> to use and with as mild learning curve as Cogito, it means Cogito got
> mostly obsolete and I'll happily remove the Cogito crash courses from
> the web.

As a relatively new user myself, I ran into the same confusion when I came to the website for the first time. One of the most prominent things on the front page is the "Git Crash Courses." Clicking on that gives me the crash courses, all of which are about Cogito, not for Git. So why doesn't the front page say "Cogito Crash Courses" instead?

And I don't think it matters much whether Cogito makes things easier or not -- the Git website really should make Git's documentation more prominent than Cogito's. I'd expect the opposite of Cogito's website.

It *is* unnecessarily confusing.
-- 
epistemological humility
Petr Baudis· Nov 17, 2006, 22:50 UTC · re: Chris Riddoch · lore
On Fri, Nov 17, 2006 at 11:36:25PM CET, Chris Riddoch wrote:
Show 19 quoted lines
> On 11/17/06, Petr Baudis <pasky@suse.cz> wrote:
> >If someone writes a crash course in pure Git covering the same grounds
> >as the current ones (possibly by just extending/retouching the tutorial)
> >(it does not necessarily need to be a "refugee" crash course, it can
> >build up from scratch), I can add it on the web. If it becomes as easy
> >to use and with as mild learning curve as Cogito, it means Cogito got
> >mostly obsolete and I'll happily remove the Cogito crash courses from
> >the web.
> 
> As a relatively new user myself, I ran into the same confusion when I
> came to the website for the first time.  One of the most prominent
> things on the front page is the "Git Crash Courses."  Clicking on that
> gives me the crash courses, all of which are about Cogito, not for
> Git.  So why doesn't the front page say "Cogito Crash Courses"
> instead?
> 
> And I don't think it matters much whether Cogito makes things easier
> or not -- the Git website really should make Git's documentation more
> prominent than Cogito's.  I'd expect the opposite of Cogito's website.

I think the difference here is the Git _tool_ vs. the Git version control system. Cogito is an element of the second: To use Git, you can either use the Git tool or the Cogito tool or the StGIT tool or even just the qgit tool (which also lets you inspect the working copy and commit). I believe the tool best suited for general usage by newbies _at this point_ is Cogito, so that's what I use for introduction to Git. I'm not saying this is ideal situation and I and others are/will be working to fix it.

I'm all for making it more obvious what's going on at the website, I think the current wording is better. Also, if people believe that a crash course for core Git would help things, I'm all for it as well.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
#!/bin/perl -sp0777i<X+d*lMLa^*lN%0]dsXx++lMlN/dsM0<j]dsj
$/=unpack('H*',$_);$_=`echo 16dio\U$k"SK$/SM$n\EsN0p[lN*1
Sean· Nov 17, 2006, 23:30 UTC · re: Petr Baudis · lore

On Fri, 17 Nov 2006 22:31:26 +0100 Petr Baudis <pasky@suse.cz> wrote:

> It's not about documentation but ease to use. I agree and sympathise
> very much with the effort of making core Git more easy to use and
> obsoleting Cogito, but until it gets there we should have what's nicest
> to the users.

As some new users have already tried to tell you, it's confusing for _them_ when they're trying to learn Git to be confronted with Cogito documentation.

The way we're going to get Git to be better is to expose new people to it and respond to their comments, complaints and ideas about how to make it better and easier to understand as they get up to speed. Having Cogito plastered all over the Git website as the _easy_ alternative is counterproductive to that effort. We need fresh blood looking at the Git documentation and trying to learn Git.

By using the GIT webpage to promote Cogito as the "easy" alternative you make it look like the entire GIT community is recommending new users should use Cogito instead. That does not represent the views of the entire GIT community. You should be very careful to represent the entire community in your role as GIT webmaster.

If people go to a Cogito website, _that's_ where they should learn about your opinions about why someone should use Cogito in place of Git. Cogito isn't "nicest" for users who don't need its extra functionality, or for getting new users involved in the improvement effort of native Git.

← back to recent threads