{"thread":{"id":"8149","subject":"[PATCH 01/10] Add a birdview-on-the-source-code section to the user manual","startedAt":"2007-05-14T15:21:20Z","lastAt":"2007-05-14T15:21:20Z","messageCount":1,"participants":["J. Bruce Fields"],"isPatch":true,"patchVersion":1,"patchTotal":10},"messages":[{"id":"42127","messageId":"46836.6894325756$1179156135@news.gmane.org","threadId":"8149","inReplyTo":null,"subject":"[PATCH 01/10] Add a birdview-on-the-source-code section to the user manual","fromName":"J. Bruce Fields","fromEmail":"bfields@citi.umich.edu","sentAt":"2007-05-14T15:21:20Z","receivedAt":"2007-05-14T15:21:20Z","isPatch":true,"sender":{"key":"bfields@citi.umich.edu","avatar":null},"body":"From: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n\nIn http://thread.gmane.org/gmane.comp.version-control.git/42479,\na birdview on the source code was requested.\n\nJ. Bruce Fields suggested that my reply should be included in the\nuser manual, and there was nothing of an outcry, so here it is,\nnot 2 months later.\n\nIt includes modifications as suggested by J. Bruce Fields, Karl\nHasselstrÃ¶m and Daniel Barkalow.\n\nSigned-off-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n---\n Documentation/user-manual.txt |  219 +++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 219 insertions(+), 0 deletions(-)\n\ndiff --git a/Documentation/user-manual.txt b/Documentation/user-manual.txt\nindex 13db969..bac9660 100644\n--- a/Documentation/user-manual.txt\n+++ b/Documentation/user-manual.txt\n@@ -3160,6 +3160,225 @@ confusing and scary messages, but it won't actually do anything bad. In\n contrast, running \"git prune\" while somebody is actively changing the \n repository is a *BAD* idea).\n \n+[[birdview-on-the-source-code]]\n+A birdview on Git's source code\n+-----------------------------\n+\n+While Git's source code is quite elegant, it is not always easy for\n+new  developers to find their way through it.  A good idea is to look\n+at the contents of the initial commit:\n+_e83c5163316f89bfbde7d9ab23ca2e25604af290_ (also known as _v0.99~954_).\n+\n+Tip: you can see what files are in there with\n+\n+----------------------------------------------------\n+$ git show e83c5163316f89bfbde7d9ab23ca2e25604af290:\n+----------------------------------------------------\n+\n+and look at those files with something like\n+\n+-----------------------------------------------------------\n+$ git show e83c5163316f89bfbde7d9ab23ca2e25604af290:cache.h\n+-----------------------------------------------------------\n+\n+Be sure to read the README in that revision _after_ you are familiar with\n+the terminology (<<glossary>>), since the terminology has changed a little\n+since then.  For example, we call the things \"commits\" now, which are\n+described in that README as \"changesets\".\n+\n+Actually a lot of the structure as it is now can be explained by that\n+initial commit.\n+\n+For example, we do not call it \"cache\" any more, but \"index\", however, the\n+file is still called `cache.h`.  Remark: Not much reason to change it now,\n+especially since there is no good single name for it anyway, because it is\n+basically _the_ header file which is included by _all_ of Git's C sources.\n+\n+If you grasp the ideas in that initial commit (it is really small and you\n+can get into it really fast, and it will help you recognize things in the\n+much larger code base we have now), you should go on skimming `cache.h`,\n+`object.h` and `commit.h` in the current version.\n+\n+In the early days, Git (in the tradition of UNIX) was a bunch of programs\n+which were extremely simple, and which you used in scripts, piping the\n+output of one into another. This turned out to be good for initial\n+development, since it was easier to test new things.  However, recently\n+many of these parts have become builtins, and some of the core has been\n+\"libified\", i.e. put into libgit.a for performance, portability reasons,\n+and to avoid code duplication.\n+\n+By now, you know what the index is (and find the corresponding data\n+structures in `cache.h`), and that there are just a couple of object types\n+(blobs, trees, commits and tags) which inherit their common structure from\n+`struct object`, which is their first member (and thus, you can cast e.g.\n+`(struct object *)commit` to achieve the _same_ as `&commit->object`, i.e.\n+get at the object name and flags).\n+\n+Now is a good point to take a break to let this information sink in.\n+\n+Next step: get familiar with the object naming.  Read <<naming-commits>>.\n+There are quite a few ways to name an object (and not only revisions!).\n+All of these are handled in `sha1_name.c`. Just have a quick look at\n+the function `get_sha1()`. A lot of the special handling is done by\n+functions like `get_sha1_basic()` or the likes.\n+\n+This is just to get you into the groove for the most libified part of Git:\n+the revision walker.\n+\n+Basically, the initial version of `git log` was a shell script:\n+\n+----------------------------------------------------------------\n+$ git-rev-list --pretty $(git-rev-parse --default HEAD \"$@\") | \\\n+\tLESS=-S ${PAGER:-less}\n+----------------------------------------------------------------\n+\n+What does this mean?\n+\n+`git-rev-list` is the original version of the revision walker, which\n+_always_ printed a list of revisions to stdout.  It is still functional,\n+and needs to, since most new Git programs start out as scripts using\n+`git-rev-list`.\n+\n+`git-rev-parse` is not as important any more; it was only used to filter out\n+options that were relevant for the different plumbing commands that were\n+called by the script.\n+\n+Most of what `git-rev-list` did is contained in `revision.c` and\n+`revision.h`.  It wraps the options in a struct named `rev_info`, which\n+controls how and what revisions are walked, and more.\n+\n+The original job of `git-rev-parse` is now taken by the function\n+`setup_revisions()`, which parses the revisions and the common command line\n+options for the revision walker. This information is stored in the struct\n+`rev_info` for later consumption. You can do your own command line option\n+parsing after calling `setup_revisions()`. After that, you have to call\n+`prepare_revision_walk()` for initialization, and then you can get the\n+commits one by one with the function `get_revision()`.\n+\n+If you are interested in more details of the revision walking process,\n+just have a look at the first implementation of `cmd_log()`; call\n+`git-show v1.3.0~155^2~4` and scroll down to that function (note that you\n+no longer need to call `setup_pager()` directly).\n+\n+Nowadays, `git log` is a builtin, which means that it is _contained_ in the\n+command `git`.  The source side of a builtin is\n+\n+- a function called `cmd_<bla>`, typically defined in `builtin-<bla>.c`,\n+  and declared in `builtin.h`,\n+\n+- an entry in the `commands[]` array in `git.c`, and\n+\n+- an entry in `BUILTIN_OBJECTS` in the `Makefile`.\n+\n+Sometimes, more than one builtin is contained in one source file.  For\n+example, `cmd_whatchanged()` and `cmd_log()` both reside in `builtin-log.c`,\n+since they share quite a bit of code.  In that case, the commands which are\n+_not_ named like the `.c` file in which they live have to be listed in\n+`BUILT_INS` in the `Makefile`.\n+\n+`git log` looks more complicated in C than it does in the original script,\n+but that allows for a much greater flexibility and performance.\n+\n+Here again it is a good point to take a pause.\n+\n+Lesson three is: study the code.  Really, it is the best way to learn about\n+the organization of Git (after you know the basic concepts).\n+\n+So, think about something which you are interested in, say, \"how can I\n+access a blob just knowing the object name of it?\".  The first step is to\n+find a Git command with which you can do it.  In this example, it is either\n+`git show` or `git cat-file`.\n+\n+For the sake of clarity, let's stay with `git cat-file`, because it\n+\n+- is plumbing, and\n+\n+- was around even in the initial commit (it literally went only through\n+  some 20 revisions as `cat-file.c`, was renamed to `builtin-cat-file.c`\n+  when made a builtin, and then saw less than 10 versions).\n+\n+So, look into `builtin-cat-file.c`, search for `cmd_cat_file()` and look what\n+it does.\n+\n+------------------------------------------------------------------\n+        git_config(git_default_config);\n+        if (argc != 3)\n+                usage(\"git-cat-file [-t|-s|-e|-p|<type>] <sha1>\");\n+        if (get_sha1(argv[2], sha1))\n+                die(\"Not a valid object name %s\", argv[2]);\n+------------------------------------------------------------------\n+\n+Let's skip over the obvious details; the only really interesting part\n+here is the call to `get_sha1()`.  It tries to interpret `argv[2]` as an\n+object name, and if it refers to an object which is present in the current\n+repository, it writes the resulting SHA-1 into the variable `sha1`.\n+\n+Two things are interesting here:\n+\n+- `get_sha1()` returns 0 on _success_.  This might surprise some new\n+  Git hackers, but there is a long tradition in UNIX to return different\n+  negative numbers in case of different errors -- and 0 on success.\n+\n+- the variable `sha1` in the function signature of `get_sha1()` is `unsigned\n+  char *`, but is actually expected to be a pointer to `unsigned\n+  char[20]`.  This variable will contain the 160-bit SHA-1 of the given\n+  commit.  Note that whenever a SHA-1 is passed as \"unsigned char *\", it\n+  is the binary representation, as opposed to the ASCII representation in\n+  hex characters, which is passed as \"char *\".\n+\n+You will see both of these things throughout the code.\n+\n+Now, for the meat:\n+\n+-----------------------------------------------------------------------------\n+        case 0:\n+                buf = read_object_with_reference(sha1, argv[1], &size, NULL);\n+-----------------------------------------------------------------------------\n+\n+This is how you read a blob (actually, not only a blob, but any type of\n+object).  To know how the function `read_object_with_reference()` actually\n+works, find the source code for it (something like `git grep\n+read_object_with | grep \":[a-z]\"` in the git repository), and read\n+the source.\n+\n+To find out how the result can be used, just read on in `cmd_cat_file()`:\n+\n+-----------------------------------\n+        write_or_die(1, buf, size);\n+-----------------------------------\n+\n+Sometimes, you do not know where to look for a feature.  In many such cases,\n+it helps to search through the output of `git log`, and then `git show` the\n+corresponding commit.\n+\n+Example: If you know that there was some test case for `git bundle`, but\n+do not remember where it was (yes, you _could_ `git grep bundle t/`, but that\n+does not illustrate the point!):\n+\n+------------------------\n+$ git log --no-merges t/\n+------------------------\n+\n+In the pager (`less`), just search for \"bundle\", go a few lines back,\n+and see that it is in commit 18449ab0...  Now just copy this object name,\n+and paste it into the command line\n+\n+-------------------\n+$ git show 18449ab0\n+-------------------\n+\n+Voila.\n+\n+Another example: Find out what to do in order to make some script a\n+builtin:\n+\n+-------------------------------------------------\n+$ git log --no-merges --diff-filter=A builtin-*.c\n+-------------------------------------------------\n+\n+You see, Git is actually the best tool to find out about the source of Git\n+itself!\n+\n [[glossary]]\n include::glossary.txt[]\n \n-- \n1.5.1.4.19.g69e2\n"}]}