{"thread":{"id":"15387","subject":"Git Community Book","startedAt":"2008-09-05T19:08:44Z","lastAt":"2008-09-06T18:26:45Z","messageCount":11,"participants":["Scott Chacon","Thomas Adam","Junio C Hamano","Linus Torvalds","Felipe Contreras","Stephan Beyer","Shawn O. Pearce","Christos Τrochalakis"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"89863","messageId":"d411cc4a0809051208k2a15c4a7te09a6979929e52f7@mail.gmail.com","threadId":"15387","inReplyTo":null,"subject":"Git Community Book","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2008-09-05T19:08:44Z","receivedAt":"2008-09-05T19:08:44Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"Hey all,\n\nI just wanted to let those of you who are interested know that I've\nbeen making a lot of progress on the Git Community Book\n(http://book.git-scm.com)  I was wondering if anyone was interested in\nhelping me with a few parts.  For one, there are some sections that I\npersonally have very little experience with, and was looking for some\nnotes/blog posts/personal experiences on, namely Advanced History\nModification (filter-branch, advanced rebasing, etc), Corruption\nRecovery, Branch Tracking, Subversion Integration, Git with\nPerl/Python/PHP, and Using Git with Editors (especially\nNetBeans/Eclipse).\n\nAlso, the last section of the book is on some of the plumbing - mostly\nstuff I've found difficult to pick up with the existing documentation\nwhile re-implementing stuff in Ruby.  I would really appreciate it if\nsomeone could proofread some of these chapters for errors:\n\nhttp://book.git-scm.com/7_the_packfile.html\nhttp://book.git-scm.com/7_raw_git.html\nhttp://book.git-scm.com/7_transfer_protocols.html\n\nSome of the next things I'm interested in producing is a cookbook\nstyle guide and some searching tools for all the online documentation,\njust to keep everyone up to date on where I'm going with the project.\nAlso, there is now a simple PDF downloadable version of the book\navailable and being kept up to date with the html version.\n\nThanks,\nScott\n"},{"id":"89864","messageId":"18071eea0809051215n6a8e1468gfd28876d7d5b0488@mail.gmail.com","threadId":"15387","inReplyTo":"d411cc4a0809051208k2a15c4a7te09a6979929e52f7@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Thomas Adam","fromEmail":"thomas.adam22@gmail.com","sentAt":"2008-09-05T19:15:55Z","receivedAt":"2008-09-05T19:15:55Z","isPatch":false,"sender":{"key":"thomas.adam22@gmail.com","avatar":"https://gravatar.com/avatar/137f9858bc6bfd5b2f743aefd988c81ce0cbd306248889df80e269519cfc8741?d=mp&s=160"},"body":"2008/9/5 Scott Chacon <schacon@gmail.com>:\n> Hey all,\n>\n> I just wanted to let those of you who are interested know that I've\n> been making a lot of progress on the Git Community Book\n> (http://book.git-scm.com)  I was wondering if anyone was interested in\n\nI'm going to bite and ask the obvious questions:\n\n1.  How does what you're producing differ from the current Git Users' Manual?\n2.  Is this project of yours aiming to obsolete the Git Users' Manual\nwith \"official\" sanctioning from people involved with Git?\n3.  Assuming 2 is a \"no\", patches to the Users' Guide would be nice.  :)\n\n-- Thomas Adam\n"},{"id":"89868","messageId":"7vmyimv0qr.fsf@gitster.siamese.dyndns.org","threadId":"15387","inReplyTo":"d411cc4a0809051208k2a15c4a7te09a6979929e52f7@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-09-05T19:41:48Z","receivedAt":"2008-09-05T19:41:48Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Scott Chacon\" <schacon@gmail.com> writes:\n\n> Also, the last section of the book is on some of the plumbing - mostly\n> stuff I've found difficult to pick up with the existing documentation\n> while re-implementing stuff in Ruby.  I would really appreciate it if\n> someone could proofread some of these chapters for errors:\n>\n> http://book.git-scm.com/7_the_packfile.html\n\nNice pictures.  You might also want to know that code for reading pack idx\nversion 2 was backported to 1.4.4.5 for people who are stuck on 1.4.4\nseries for whatever reason.\n\nWhat is the target audience of this section?  If it is written for a mere\ncurious type, or if it is written to give \"here is the general idea, for\nmore details read the source\", the level of detail here would be Ok.\n\nIf you are writing for people who want to (re)implement something that\nproduces these files, you might want to at least say that offset/sha1[]\ntable is sorted by sha1[] values (this is to allow binary search of this\ntable), and fanout[] table points at the offset/sha1[] table in a specific\nway (so that part of the latter table that covers all hashes that start\nwith a given byte can be found to avoid 8 iterations of the binary\nsearch).\n\n<data> part is just zlib stream for non-delta object types; for the two\ndelta object representations, the <data> portion contains something that\nidentifies which base object this delta representation depends on, and the\ndelta to apply on the base object to resurrect this object.  ref-delta\nuses 20-byte hash of the base object at the beginning of <data>, while\nofs-delta stores an offset within the same packfile to identify the base\nobject.  In either case, two important constraints a reimplementor must\nadhere to are:\n\n * delta representation must be based on some other object within the same\n   packfile;\n\n * the base object must be of the same underlying type (blob, tree, commit\n   or tag);\n\n> http://book.git-scm.com/7_raw_git.html\n\nI am guessing this is for Porcelain writers who use plumbing.  Please\ndon't teach echoing into .git/refs/...  but DO teach using update-ref with\nthe -m option.  We do not want people's random Porcelains flipping the tip\nof branches without leaving trail in reflog for users to use to recover\nfrom mistakes.\n"},{"id":"89870","messageId":"alpine.LFD.1.10.0809051317510.3117@nehalem.linux-foundation.org","threadId":"15387","inReplyTo":"d411cc4a0809051208k2a15c4a7te09a6979929e52f7@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-09-05T20:27:43Z","receivedAt":"2008-09-05T20:27:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Sep 2008, Scott Chacon wrote:\n> \n> http://book.git-scm.com/7_the_packfile.html \n\nThe checksums in the index file \"trailers\" are all claiming to be 4 bytes, \nand that's wrong - they're full SHA1 sums at 20 bytes each.\n\nThe v2 pack-file _also_ has per-object CRC's, and those are indeed just 4 \nbytes each, and are correctly listed as such.\n\nThe pack-file itself also has a few more things there, it's not just the \n\"PACK\" string and then the objects. It has two more 32-bit words: a pack \nfile version number and the number of entries in the pack-file (all \nnetwork byte order). It also has its own checksum at the end (20-byte SHA1 \nagain).\n\nBut looks good otherwise from a quick look.\n\n\t\tLinus\n"},{"id":"89871","messageId":"d411cc4a0809051345p1c1b127epa7a186f49414a1fc@mail.gmail.com","threadId":"15387","inReplyTo":"18071eea0809051215n6a8e1468gfd28876d7d5b0488@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2008-09-05T20:45:51Z","receivedAt":"2008-09-05T20:45:51Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"On Fri, Sep 5, 2008 at 12:15 PM, Thomas Adam <thomas.adam22@gmail.com> wrote:\n> 2008/9/5 Scott Chacon <schacon@gmail.com>:\n>> Hey all,\n>>\n>> I just wanted to let those of you who are interested know that I've\n>> been making a lot of progress on the Git Community Book\n>> (http://book.git-scm.com)  I was wondering if anyone was interested in\n>\n> I'm going to bite and ask the obvious questions:\n\nJust for reference, a lot of this was discussed here a while back:\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/90653\n\nhowever, I would be happy to answer these for you.\n\n>\n> 1.  How does what you're producing differ from the current Git Users' Manual?\n\nI'm going for a different audience with this project.  I'd like for it\nto be a lot more user-friendly, easily digestible, and to include\nimages, diagrams and screencasts.\n\n> 2.  Is this project of yours aiming to obsolete the Git Users' Manual\n> with \"official\" sanctioning from people involved with Git?\n\nI think there will be people who prefer the Users Manual format, who\nthink screencasts are wussy :)\nAlso, I'm not sure an \"official\" sanctioning would do much of anything\n- because of the images and screencasts, this will never be included\nin the git source like the UM is, but it's also open source so if\npeople want to take content from it to improve the UM, that's cool.\n\n> 3.  Assuming 2 is a \"no\", patches to the Users' Guide would be nice.  :)\n\nI would love to do this, but I don't know what exactly the community\nthinks is missing/lacking.  My ideas about what is helpful is rarely\nthe same as the git lists :)  However, if someone pointed to one of\nthe chapters I wrote and said \"that would be great in the UM\", I would\nhappily convert it.\n\nScott\n"},{"id":"89873","messageId":"d411cc4a0809051434g4e92790fsa38d12487630aa9f@mail.gmail.com","threadId":"15387","inReplyTo":"7vmyimv0qr.fsf@gitster.siamese.dyndns.org","subject":"Re: Git Community Book","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2008-09-05T21:34:34Z","receivedAt":"2008-09-05T21:34:34Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"On Fri, Sep 5, 2008 at 12:41 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> \"Scott Chacon\" <schacon@gmail.com> writes:\n>\n>> Also, the last section of the book is on some of the plumbing - mostly\n>> stuff I've found difficult to pick up with the existing documentation\n>> while re-implementing stuff in Ruby.  I would really appreciate it if\n>> someone could proofread some of these chapters for errors:\n>>\n>> http://book.git-scm.com/7_the_packfile.html\n>\n> Nice pictures.  You might also want to know that code for reading pack idx\n> version 2 was backported to 1.4.4.5 for people who are stuck on 1.4.4\n> series for whatever reason.\n>\n> What is the target audience of this section?  If it is written for a mere\n> curious type, or if it is written to give \"here is the general idea, for\n> more details read the source\", the level of detail here would be Ok.\n>\n> If you are writing for people who want to (re)implement something that\n> produces these files, you might want to at least say that offset/sha1[]\n> table is sorted by sha1[] values (this is to allow binary search of this\n> table), and fanout[] table points at the offset/sha1[] table in a specific\n> way (so that part of the latter table that covers all hashes that start\n> with a given byte can be found to avoid 8 iterations of the binary\n> search).\n>\n> <data> part is just zlib stream for non-delta object types; for the two\n> delta object representations, the <data> portion contains something that\n> identifies which base object this delta representation depends on, and the\n> delta to apply on the base object to resurrect this object.  ref-delta\n> uses 20-byte hash of the base object at the beginning of <data>, while\n> ofs-delta stores an offset within the same packfile to identify the base\n> object.  In either case, two important constraints a reimplementor must\n> adhere to are:\n>\n>  * delta representation must be based on some other object within the same\n>   packfile;\n>\n>  * the base object must be of the same underlying type (blob, tree, commit\n>   or tag);\n>\n>> http://book.git-scm.com/7_raw_git.html\n>\n> I am guessing this is for Porcelain writers who use plumbing.  Please\n> don't teach echoing into .git/refs/...  but DO teach using update-ref with\n> the -m option.  We do not want people's random Porcelains flipping the tip\n> of branches without leaving trail in reflog for users to use to recover\n> from mistakes.\n>\n\nI've implemented all of these and Linus's fixes and suggestions.\nThanks for the feedback.\n\nTo answer your earlier question, these docs are basically for people\nworking on bindings/re-implementations in other languages, since there\nis no real linked library available yet, as a primer before they dig\ninto the source, or possibly so they don't have to.\n\nI'm not fantastic at C, so it took me a while in some cases - figuring\nout that the size listed in the object header was not the actual size\nof the data, but the size of it when expanded, for example, was not\nvery easy to do.  I've been doing a lot of work on re-implementations\nin Ruby and ObjC because I can't easily make real bindings, so I\nthought I would add things that I could not easily find in the docs\nfor others that are trying in other languages.\n\nIf you want, I could create a patch for any of this stuff to\nDocumentation/ (that goes for the whole book), but someone will have\nto tell me which parts might be useful to add.\n\nThanks again for taking the time!\nScott\n"},{"id":"89875","messageId":"94a0d4530809051509y5f3cf2f6g209047debc584ed9@mail.gmail.com","threadId":"15387","inReplyTo":"d411cc4a0809051434g4e92790fsa38d12487630aa9f@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2008-09-05T22:09:41Z","receivedAt":"2008-09-05T22:09:41Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Sep 6, 2008 at 12:34 AM, Scott Chacon <schacon@gmail.com> wrote:\n> On Fri, Sep 5, 2008 at 12:41 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>> \"Scott Chacon\" <schacon@gmail.com> writes:\n>>\n>>> Also, the last section of the book is on some of the plumbing - mostly\n>>> stuff I've found difficult to pick up with the existing documentation\n>>> while re-implementing stuff in Ruby.  I would really appreciate it if\n>>> someone could proofread some of these chapters for errors:\n>>>\n>>> http://book.git-scm.com/7_the_packfile.html\n>>\n>> Nice pictures.  You might also want to know that code for reading pack idx\n>> version 2 was backported to 1.4.4.5 for people who are stuck on 1.4.4\n>> series for whatever reason.\n>>\n>> What is the target audience of this section?  If it is written for a mere\n>> curious type, or if it is written to give \"here is the general idea, for\n>> more details read the source\", the level of detail here would be Ok.\n>>\n>> If you are writing for people who want to (re)implement something that\n>> produces these files, you might want to at least say that offset/sha1[]\n>> table is sorted by sha1[] values (this is to allow binary search of this\n>> table), and fanout[] table points at the offset/sha1[] table in a specific\n>> way (so that part of the latter table that covers all hashes that start\n>> with a given byte can be found to avoid 8 iterations of the binary\n>> search).\n>>\n>> <data> part is just zlib stream for non-delta object types; for the two\n>> delta object representations, the <data> portion contains something that\n>> identifies which base object this delta representation depends on, and the\n>> delta to apply on the base object to resurrect this object.  ref-delta\n>> uses 20-byte hash of the base object at the beginning of <data>, while\n>> ofs-delta stores an offset within the same packfile to identify the base\n>> object.  In either case, two important constraints a reimplementor must\n>> adhere to are:\n>>\n>>  * delta representation must be based on some other object within the same\n>>   packfile;\n>>\n>>  * the base object must be of the same underlying type (blob, tree, commit\n>>   or tag);\n>>\n>>> http://book.git-scm.com/7_raw_git.html\n>>\n>> I am guessing this is for Porcelain writers who use plumbing.  Please\n>> don't teach echoing into .git/refs/...  but DO teach using update-ref with\n>> the -m option.  We do not want people's random Porcelains flipping the tip\n>> of branches without leaving trail in reflog for users to use to recover\n>> from mistakes.\n>>\n>\n> I've implemented all of these and Linus's fixes and suggestions.\n> Thanks for the feedback.\n>\n> To answer your earlier question, these docs are basically for people\n> working on bindings/re-implementations in other languages, since there\n> is no real linked library available yet, as a primer before they dig\n> into the source, or possibly so they don't have to.\n>\n> I'm not fantastic at C, so it took me a while in some cases - figuring\n> out that the size listed in the object header was not the actual size\n> of the data, but the size of it when expanded, for example, was not\n> very easy to do.  I've been doing a lot of work on re-implementations\n> in Ruby and ObjC because I can't easily make real bindings, so I\n> thought I would add things that I could not easily find in the docs\n> for others that are trying in other languages.\n\nI have experience mixing C and Ruby code if you are interested, it's\nactually quite easy.\n\nI also think a shared library would make sense.\n\nKeep up the good work ;)\n\n-- \nFelipe Contreras\n"},{"id":"89884","messageId":"20080906004831.GA8984@leksak.fem-net","threadId":"15387","inReplyTo":"d411cc4a0809051208k2a15c4a7te09a6979929e52f7@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Stephan Beyer","fromEmail":"s-beyer@gmx.net","sentAt":"2008-09-06T00:48:31Z","receivedAt":"2008-09-06T00:48:31Z","isPatch":false,"sender":{"key":"s-beyer@gmx.net","avatar":"https://avatars.githubusercontent.com/u/143889?v=4"},"body":"Hi,\n\nScott Chacon wrote:\n> I just wanted to let those of you who are interested know that I've\n> been making a lot of progress on the Git Community Book\n> (http://book.git-scm.com)  I was wondering if anyone was interested in\n> helping me with a few parts.\n\nI just had a very quick look over the PDF, meaning only looking at\npictures and headlines.\n\nJust nitpicking about one thing:\nI was wondering if \"Stash Queue\" is the right headline, because I\nusually use\n\n\tgit stash save\t# oh, an interrupt, have to do something else now\n\nand after this is done:\n\n\tgit stash pop\t# back to the real work\n\nAnd if you are interrupted in an interrupt, you want the last stash\nbeing the first one to pop, which is a stack-like (last in, first out)\nbehavior.\n\nOf course, there may be cases where you want the queuing behavior that\nyou advertise in the book.\nI use it rather seldomly. But perhaps it is just me :-)\n\nRegards,\n  Stephan\n\n-- \nStephan Beyer <s-beyer@gmx.net>, PGP 0x6EDDD207FCC5040F\n"},{"id":"89895","messageId":"20080906063325.GD28035@spearce.org","threadId":"15387","inReplyTo":"d411cc4a0809051434g4e92790fsa38d12487630aa9f@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-09-06T06:33:25Z","receivedAt":"2008-09-06T06:33:25Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Scott Chacon <schacon@gmail.com> wrote:\n> On Fri, Sep 5, 2008 at 12:41 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> > \"Scott Chacon\" <schacon@gmail.com> writes:\n> >\n> >> Also, the last section of the book is on some of the plumbing - mostly\n> >> stuff I've found difficult to pick up with the existing documentation\n> >> while re-implementing stuff in Ruby.  I would really appreciate it if\n> >> someone could proofread some of these chapters for errors:\n> >>\n> >> http://book.git-scm.com/7_the_packfile.html\n\nOK, time for me to throw in comments.  ;-)\n\nI do like this book, its organized and concise.  Thanks for doing it.\n\n\nhttp://book.git-scm.com/7_how_git_stores_objects.html:\n\nThe loose object formatting of\n\n header = \"#{type} #{size}#body\"\n store = header + content\n\nI can't read Ruby so I'm not sure what the header value computes\nout to here.  #body should be a \\0.  I'm also not sure that the\nprior line setting size = content.length.to_s is very clear for\nthe non-Ruby people to understand how a size is formatted.\n\nIf the code shown here is the Ruby implementation I'm a little\nconcerned about it writing directly into the loose object.  If the\nwrite is partial then you have a partial object which is at the\nright name, but is unusable.  That can give you corruption that\nis difficult to track down and fix.  C Git and JGit both write\nto temporary files then atomically move the temporary file into\nposition under its proper name only after it has been fully written.\n\nIf an implementor is implementing they should be offered this advice,\nand probably do so right here in this section of the book.\n\n\"When objects are written to disk, it is often in the loose format,\nsince that format is less expensive to access.\"\n\nI'm not sure that statement is true.  Access from packs tends\nto scream compared to access from loose objects.  The overheads\nof opening and closing the file descriptors, even on Linux, is\nwhat kills performance for data access.  However Git writes to\nloose objects first and packs later for _safety_ not efficiency.\nAlthough it is a lot more efficient to write a 2 KB loose object\nand avoid rewriting a 50 MB pack, but its also less likely to fail\nand make you lose your work.\n\n\nhttp://book.git-scm.com/7_the_git_index.html:\n\nI wouldn't say that the index stores permissions.  More like it\nstores the \"class\" or \"type\" of the thing located at that path.\nThere are 4 major classes:\n\n\t- regular file\n\t- executable file\n\t- symbolic link\n\t- git submodule\n\nThe 5th class is the subtree, but only appears in trees and not\nin the index since the index file is actually flat.\n\n\nhttp://book.git-scm.com/7_the_packfile.html:\n\nYou should probably point out that the .idx file uses network byte\norder for the numeric fields like the version number and the file\noffsets.\n\nI'd also point out that the offsets in index v1 are unsigned and\nfrom the start of the pack file.  The offsets in index v2 are\nalso unsigned, but the 1<<31 is tested in the 32 bit offset to\nsee if a 64 bit offset is used.  The algorithm there is:\n\n\tif offset32 & 1<<31:\n\t\toffset = ofs64_table[offset32 & ~(1<<31)]\n\telse\n\t\toffset = offset32\n\nIts also rather unclear how the fan out table can be used to limit\nthe binary search.  What you are missing is describing that fanout[X]\nholds the number of objects whose first byte of their SHA-1 is <= X.\nHence fanout[0] has the number of objects whose SHA-1 starts with\n\"00\" and fanout[0x15] has the number of objects whose SHA-1 starts\nwith \"15\", \"14\", \"13\", ..., \"00\".  Thus fanout[0xff] has the total\nnumber of objects in the pack.\n\nIn the pack file section I'd also point out the version and entry\ncount are unsigned network byte order.  This is not clear from the\nRuby code, although one can guess at it if one knows the git.git\ncode very very well (like I do).\n\n\"After that, you get a series of packed objects, in order of thier SHAs\"\n\nAside from s/thier/their/ this is not a correct statement _AT ALL_.\n\nThe ordering of objects in the packfile is very carefully planned\nby the packer to maximize data locality from most recent -> least\nrecent information, making the most recent revisions of a project\nthe fastest to access.  This has _NOTHING_ to do with their SHA-1\nnames.\n\nTechnically a pack may store objects in any random order.  Heck,\nyou can wire up an RNG to the packer to always produce a different\nordering each time you pack.  Practically an implementation shouldn't\nbe that stupid and should instead try to order objects by recency,\nlike git.git and JGit both do.\n\n\"At the end of the packfile is a 20-byte SHA1 sum of all the shas\n(in sorted order) in that packfile.\"\n\nAlso incorrect.  The 20-byte checksum at the end of the pack file\nis a checksum of all bytes preceeding the checksum itself.  We use\nit as an end-to-end data integrity check, especially on the network\ntransport to verify that every bit sent by the one side is received\ncorrectly on the other side.\n\nBTW, can I just say, I love the graphics in this book.  They are\nquite well done.  Very worthwhile.\n\n\nhttp://book.git-scm.com/7_transfer_protocols.html:\n\nYou might as well explain that the stream returned by upload-pack\nuses the same 4 byte line length framing to form \"packets\", with\nthe 5th byte (really first byte of the payload) indicating the\n\"stream\":\n\n\t- stream 1 ('\\001') is the PACK data\n\t- stream 2 ('\\002') is progress data/information\n\t- stream 3 ('\\003') is the OH S**T we are aborting, died, dead\n\nYou may also want to explain that the way you know the end of the\npack is to read the header, get the entry count, and then read that\nmany objects from the stream, and then verify the pack checksum.\n\n-- \nShawn.\n"},{"id":"89917","messageId":"d411cc4a0809061114k6b9f01b9sdf479360d4cb4c41@mail.gmail.com","threadId":"15387","inReplyTo":"20080906063325.GD28035@spearce.org","subject":"Re: Git Community Book","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2008-09-06T18:14:22Z","receivedAt":"2008-09-06T18:14:22Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"Thanks a ton for this, I'll incorporate all of this.\n\nOn Fri, Sep 5, 2008 at 11:33 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Scott Chacon <schacon@gmail.com> wrote:\n>> On Fri, Sep 5, 2008 at 12:41 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>> > \"Scott Chacon\" <schacon@gmail.com> writes:\n>> >\n>> >> Also, the last section of the book is on some of the plumbing - mostly\n>> >> stuff I've found difficult to pick up with the existing documentation\n>> >> while re-implementing stuff in Ruby.  I would really appreciate it if\n>> >> someone could proofread some of these chapters for errors:\n>> >>\n>> >> http://book.git-scm.com/7_the_packfile.html\n>\n> OK, time for me to throw in comments.  ;-)\n>\n> I do like this book, its organized and concise.  Thanks for doing it.\n>\n>\n> http://book.git-scm.com/7_how_git_stores_objects.html:\n>\n> The loose object formatting of\n>\n>  header = \"#{type} #{size}#body\"\n>  store = header + content\n>\n> I can't read Ruby so I'm not sure what the header value computes\n> out to here.  #body should be a \\0.  I'm also not sure that the\n> prior line setting size = content.length.to_s is very clear for\n> the non-Ruby people to understand how a size is formatted.\n>\n\nSorry, the markdown thingy is translating all the '\\0's to '#body' for\nsome freaking reason unless I write it as '\\\\0'.  I'll fix this - it's\ndifficult for me to find these sometimes.  As for the rest of the ruby\nstuff, I think I'll add some comments.\n\n> If the code shown here is the Ruby implementation I'm a little\n> concerned about it writing directly into the loose object.  If the\n> write is partial then you have a partial object which is at the\n> right name, but is unusable.  That can give you corruption that\n> is difficult to track down and fix.  C Git and JGit both write\n> to temporary files then atomically move the temporary file into\n> position under its proper name only after it has been fully written.\n\nThat is a good idea - I don't do it that way and I certainly will\nchange the implementation to do so and modify these docs to reflect\nthat advice.\n\n> \"When objects are written to disk, it is often in the loose format,\n> since that format is less expensive to access.\"\n>\n> I'm not sure that statement is true.  Access from packs tends\n> to scream compared to access from loose objects.  The overheads\n> of opening and closing the file descriptors, even on Linux, is\n> what kills performance for data access.  However Git writes to\n> loose objects first and packs later for _safety_ not efficiency.\n> Although it is a lot more efficient to write a 2 KB loose object\n> and avoid rewriting a 50 MB pack, but its also less likely to fail\n> and make you lose your work.\n\nThanks for the clarification.  I write to loose objects first largely\nbecause it's so much easier to do.  But also because I don't mmap\nobjects, so packfile access is not faster for implementations that\ncan't do that very well.  Also, I had originally meant \"less expensive\nto write\", but I can see that is not clear.\n\n\n> http://book.git-scm.com/7_the_git_index.html:\n>\n> I wouldn't say that the index stores permissions.  More like it\n> stores the \"class\" or \"type\" of the thing located at that path.\n> There are 4 major classes:\n>\n>        - regular file\n>        - executable file\n>        - symbolic link\n>        - git submodule\n>\n> The 5th class is the subtree, but only appears in trees and not\n> in the index since the index file is actually flat.\n\nInteresting.  This documentation is actually from the User Manual -\nI'll update this chapter first and if it looks better, I'll submit a\npatch to the UM, too.\n\n> http://book.git-scm.com/7_the_packfile.html:\n>\n> You should probably point out that the .idx file uses network byte\n> order for the numeric fields like the version number and the file\n> offsets.\n\nWill do.\n\n>\n> I'd also point out that the offsets in index v1 are unsigned and\n> from the start of the pack file.  The offsets in index v2 are\n> also unsigned, but the 1<<31 is tested in the 32 bit offset to\n> see if a 64 bit offset is used.  The algorithm there is:\n>\n>        if offset32 & 1<<31:\n>                offset = ofs64_table[offset32 & ~(1<<31)]\n>        else\n>                offset = offset32\n>\n> Its also rather unclear how the fan out table can be used to limit\n> the binary search.  What you are missing is describing that fanout[X]\n> holds the number of objects whose first byte of their SHA-1 is <= X.\n> Hence fanout[0] has the number of objects whose SHA-1 starts with\n> \"00\" and fanout[0x15] has the number of objects whose SHA-1 starts\n> with \"15\", \"14\", \"13\", ..., \"00\".  Thus fanout[0xff] has the total\n> number of objects in the pack.\n>\n> In the pack file section I'd also point out the version and entry\n> count are unsigned network byte order.  This is not clear from the\n> Ruby code, although one can guess at it if one knows the git.git\n> code very very well (like I do).\n>\n> \"After that, you get a series of packed objects, in order of thier SHAs\"\n>\n> Aside from s/thier/their/ this is not a correct statement _AT ALL_.\n>\n> The ordering of objects in the packfile is very carefully planned\n> by the packer to maximize data locality from most recent -> least\n> recent information, making the most recent revisions of a project\n> the fastest to access.  This has _NOTHING_ to do with their SHA-1\n> names.\n>\n> Technically a pack may store objects in any random order.  Heck,\n> you can wire up an RNG to the packer to always produce a different\n> ordering each time you pack.  Practically an implementation shouldn't\n> be that stupid and should instead try to order objects by recency,\n> like git.git and JGit both do.\n>\n> \"At the end of the packfile is a 20-byte SHA1 sum of all the shas\n> (in sorted order) in that packfile.\"\n>\n> Also incorrect.  The 20-byte checksum at the end of the pack file\n> is a checksum of all bytes preceeding the checksum itself.  We use\n> it as an end-to-end data integrity check, especially on the network\n> transport to verify that every bit sent by the one side is received\n> correctly on the other side.\n>\n\nI'm an idiot.  I say this because I actually implemented a bunch of\nthis stuff (in Ruby) and ran into most of these issues when trying to\nimplement it.  So I knew these things not 3 weeks ago, but I still\nwrote it this way.  Dur.  Thanks for the corrections, I'll update\neverything accordingly.\n\n> BTW, can I just say, I love the graphics in this book.  They are\n> quite well done.  Very worthwhile.\n\nThanks.\n\n>\n>\n> http://book.git-scm.com/7_transfer_protocols.html:\n>\n> You might as well explain that the stream returned by upload-pack\n> uses the same 4 byte line length framing to form \"packets\", with\n> the 5th byte (really first byte of the payload) indicating the\n> \"stream\":\n>\n>        - stream 1 ('\\001') is the PACK data\n>        - stream 2 ('\\002') is progress data/information\n>        - stream 3 ('\\003') is the OH S**T we are aborting, died, dead\n>\n> You may also want to explain that the way you know the end of the\n> pack is to read the header, get the entry count, and then read that\n> many objects from the stream, and then verify the pack checksum.\n>\n> --\n> Shawn.\n>\n\nThanks again for all the time it must have taken to review all of this\n- I'll make sure it gets into the book, and where appropriate, back\ninto the UM or other internal git docs.\n\nScott\n"},{"id":"89920","messageId":"f7b87f7c0809061126y17f6b601iaeb9a3661c914391@mail.gmail.com","threadId":"15387","inReplyTo":"d411cc4a0809051208k2a15c4a7te09a6979929e52f7@mail.gmail.com","subject":"Re: Git Community Book","fromName":"Christos Τrochalakis","fromEmail":"yatiohi@ideopolis.gr","sentAt":"2008-09-06T18:26:45Z","receivedAt":"2008-09-06T18:26:45Z","isPatch":false,"sender":{"key":"yatiohi@ideopolis.gr","avatar":"https://gravatar.com/avatar/0a2086bcb916d6c7d3aab6edf1b697bc54d8f376d79d73b6deb96db048f7c64a?d=mp&s=160"},"body":"On Fri, Sep 5, 2008 at 10:08 PM, Scott Chacon <schacon@gmail.com> wrote:\n> Hey all,\n>\n> I just wanted to let those of you who are interested know that I've\n> been making a lot of progress on the Git Community Book\n> (http://book.git-scm.com)\n> ...\n\nHello Scott!\n\nNice book, I just started reading it and I have a recommendation to\nmake, at \"Chapter 4: Git Treeishes\" you write\n\n---------\nhttp://book.git-scm.com/4_git_treeishes.html\nRange\n\nFinally, you can specify a range of commits with the range spec. This\nwill give you all the commits between 7b593b5 and 51bea1 (where 51bea1\nis most recent), excluding 7b593b5 but including 51bea1:\n\n7b593b5..51bea1\n\nThis will include every commit since 7b593b:\n\n7b593b..\n---------\n\nThis in not quite correct. \"commits between A and B\" cannot really\napply here. I believe that \"commits reachable from B and not from A\"\nis more precise. Actually you are already using the \"reachability\"\nexplanation at the start of \"Chapter 3: Basic usage\".\n\nThis issue is also described at the rev-parse man page.\n\nApart from that, you could also include \"a...b\" syntax for completeness.\n\n-christos\n"}]}