{"thread":{"id":"7265","subject":"Libification project (SoC)","startedAt":"2007-03-16T04:24:06Z","lastAt":"2007-03-22T09:51:23Z","messageCount":62,"participants":["Luiz Fernando N. Capitulino","Shawn O. Pearce","Junio C Hamano","Johannes Sixt","Matthieu Moy","Johannes Schindelin","Petr Baudis","Rocco Rutte","Nicolas Pitre","Marco Costalba","Steve Frécinaux","Andy Parkins","Jakub Narebski","Theodore Tso","Linus Torvalds","Andreas Ericsson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"37200","messageId":"20070316042406.7e750ed0@home.brethil","threadId":"7265","inReplyTo":null,"subject":"Libification project (SoC)","fromName":"Luiz Fernando N. Capitulino","fromEmail":"lcapitulino@mandriva.com.br","sentAt":"2007-03-16T04:24:06Z","receivedAt":"2007-03-16T04:24:06Z","isPatch":false,"sender":{"key":"lcapitulino@mandriva.com.br","avatar":null},"body":"\n Hi Shawn,\n\n I'm going to apply for the libification project and, in order to help\nme to get started, would be good to get some feedback regarding the\nproject's goal and your expectations.\n\n I'll just dump some thoughts/question I had, so that we can\nstart some discussion.\n\n 1. This' a more complete todo list, based on the wiki and a\nquick look at the code.\n\n    o Remove static variables\n    o Avoid dying when a function call fails (eg, malloc())\n    o Input parameter checking (plus errno setting)\n    o Documentation (eg, doxygen)\n    o Unit-tests\n    o Add prefix (eg, git_*) to public API functions\n\n Do we agree here? Is there more suggestions?\n\n 2. What's the minimum amount of work that need to be done for\nthe SoC project to be considered successful?\n\n 3. I don't code in Perl, is it a problem? I mean, the project's\ngoal is to have a Perl binding but I think it goes far from\nthat: we could have a python module, a C program, or anything\nthat shows the libgit is useful.\n\n Thanks,\n\n-- \nLuiz Fernando N. Capitulino\n"},{"id":"37203","messageId":"20070316045928.GB31606@spearce.org","threadId":"7265","inReplyTo":"20070316042406.7e750ed0@home.brethil","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-16T04:59:28Z","receivedAt":"2007-03-16T04:59:28Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br> wrote:\n>  I'm going to apply for the libification project and, in order to help\n> me to get started, would be good to get some feedback regarding the\n> project's goal and your expectations.\n\nExcellent!\n \n>  1. This' a more complete todo list, based on the wiki and a\n> quick look at the code.\n> \n>     o Remove static variables\n\nYes.  Removing all of these is not completely necessary in the\nfirst version; in fact I would recommened against it.\n\nFor example the active_cache variable and its related friends\nis referenced a lot. lt contains the index in memory.  I think\nits perfectly OK to say that in the first iteration of a public\nlibgit.a that the process may only use one index at a time, if it\ncan even use the index at all (see below).  But if you eventually\ngot around to even helping the index parts of \"the Git library\",\nthat would certainly be appreciated!\n\nOn the other hand, many of the variables declared in environment.c\nare repository specific configuration variables.  These probably\nshould be abstracted into some sort of wrapper, so that multiple\nrepositories can be accessed from within the same process.  Why?\na future mod_perl running gitweb.cgi accessing repositories through\nlibgit.a and Perl bindings of course!\n\nBut static variable removal is low on the priority list for this\nproject I think.  Our more important issues are related to some of\nthe other items.\n\n>     o Avoid dying when a function call fails (eg, malloc())\n\nmalloc is a huge problem in the Git code today.  Almost all\nof our malloc calls are actually through the xmalloc wrapper.\nAll xmalloc callers assume xmalloc will *never* fail.  This\nmakes it, uh, interesting.  ;-)\n\nAlthough one could argue that being unable to malloc needed memory\nprobably means you're toast, so die()'ing is good.\n\nBut other areas die when they get given a bad SHA-1 (for example).\nIf the library caller can supply that (possibly bad) SHA-1 to an\nAPI function, that's just mean to die out.  ;-)\n\n>     o Input parameter checking (plus errno setting)\n\nYes, of course.  But most functions (at least those that should be\nmade public) probably already do check their arguments.  Some return\nan error code back to their caller; others die() and abort the\ncurrent process.  And there are probably a few that don't check\ntheir arguments enough.  But I think input parameter checking is\nprobably going to be a relatively small task here.\n\nAlthough sometimes the input checking is done in the program that\ncalls the function, and not the function itself.  So that might\nneed to be refactored in a few spots.\n\n>     o Documentation (eg, doxygen)\n\nYes; very important for the library to be of any use to anyone else.\n\n>     o Unit-tests\n\nOf the public API, yes.  Our current test suite covers some of that\ncode that we want to make public, but does so through programs that\ncall those functions.  We would want unit tests to verify the public\nAPI conforms to the expectations of the unit test's writer.  ;-)\n\n>     o Add prefix (eg, git_*) to public API functions\n\nYes.  But which functions shall we expose?  ;-)\n\nSee below for functionality I'm thinking about; others may have\ndifferent ideas.\n \n      o Build system issues\n\nYou missed this, but I think its an important consideration.\nOur current libgit.a is a static library that has a relatively large\nnumber of symbols its modules are exporting.  These symbol names\naren't namespace-ized (e.g. git_* prefix) so we wouldn't want to\njust offer this library up in its current form.\n\nSome of those symbols would get name changes (as you suggest above),\nbut others might not (e.g. the active_cache that I suggest further\nabove).  These modules might need to be moved out of libgit.a and\nmoved into say a new libgitprivate.a, that our own code can link\nagainst, but that isn't offered to the public as a stable API.\n\n      o Public header definition\n\nWhatever we expose, we will need to draft a public \"git.h\"\n(or somesuch) that callers can rely upon.  It will need to be\nfairly stable, and handle revisions as new features get added.\nE.g. version testing support like the zlib and cURL library have,\nand that we rely upon in Git to do feature checks.  ;-)\n\n>  2. What's the minimum amount of work that need to be done for\n> the SoC project to be considered successful?\n\nI'd like to see enough API support that gitweb.cgi could:\n\n * get the most recent commit date of all refs in all projects\n   (the toplevel project index page);\n * get a shortlog for the main summary page of a project;\n * get the full content of a single commit;\n * get the \"raw\" diff (paths that changed) for two commits;\n\nThere's a thousand other things that gitweb.cgi would still need to\nfully avoid forking Git processes.  But that's a really good start,\nand is probably going to be a decent chunk of work.  Especially to\ncreate high-quality patches that pass our standards review.  ;-)\n\nIn some cases much of the above is already \"internally public\";\nmeaning we already treat parts of that code as a library and invoke\nthem from within processes to get work done.  Much of this project\nis about improving the interfaces and behavior enough to make those\nexisting APIs truely public.\n\nSee refs.h, diff.h, revision.h, commit.h...\n \n>  3. I don't code in Perl, is it a problem? I mean, the project's\n> goal is to have a Perl binding but I think it goes far from\n> that: we could have a python module, a C program, or anything\n> that shows the libgit is useful.\n\nNo, I don't see that as a problem at all.  We have some Perl\nexperts on the mailing list who would like to see Perl bindings.\nSome of the Perl binding is pure C code, and some if it is this\nweird Perl macro language...  so I expect those Perl experts to come\nout of the woodwork and help the community to create a prototype\nset of bindings.  There's also Ruby and Python interests around,\nso we may see bindings for those too.  ;-)\n\n>From a goal perspective of this SoC project, any functioning binding\nthat can support a gitweb type of application would be great.\nIt shows the library works as intended, is useful, and can be\ncontinued to be built upon.  That's a pretty successful project in\nmy mind.\n\n-- \nShawn.\n"},{"id":"37205","messageId":"7vejnpycu1.fsf@assigned-by-dhcp.cox.net","threadId":"7265","inReplyTo":"20070316045928.GB31606@spearce.org","subject":"Re: Libification project (SoC)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-03-16T05:30:46Z","receivedAt":"2007-03-16T05:30:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> On the other hand, many of the variables declared in environment.c\n> are repository specific configuration variables.  These probably\n> should be abstracted into some sort of wrapper, so that multiple\n> repositories can be accessed from within the same process.  Why?\n> a future mod_perl running gitweb.cgi accessing repositories through\n> libgit.a and Perl bindings of course!\n\nI think if you are abstracting them out, into \"struct repo_state\",\nthe index and object store related variables such as packed_git\nshould go there as well, so your recommendation feels very\ninconsistent to me.\n\n>>     o Avoid dying when a function call fails (eg, malloc())\n>\n> malloc is a huge problem in the Git code today.  Almost all\n> of our malloc calls are actually through the xmalloc wrapper.\n> All xmalloc callers assume xmalloc will *never* fail.  This\n> makes it, uh, interesting.  ;-)\n\nActually they do not assume such.  What they assume is worse.\nThey assume that there is nothing else you can do other than\ndying when allocation fails.\n\n> But other areas die when they get given a bad SHA-1 (for example).\n> If the library caller can supply that (possibly bad) SHA-1 to an\n> API function, that's just mean to die out.  ;-)\n\nThat's a real problem, but on the other hand, perl or whatever\nwrapped ones can do the dying (or not dying) before calling into\nlibgit, so it may not be such a big issue.\n\n>>     o Documentation (eg, doxygen)\n>>     o Unit-tests\n>>     o Add prefix (eg, git_*) to public API functions\n>\n> Yes.  But which functions shall we expose?  ;-)\n\nBefore going into that topic, a bigger question is if we are\nhappy with the current internal API and what the goal of\nlibification is.  If the libification is going to say that \"this\nis a published API so we are not going to change it\", I would\nimagine that it would be very hard to accept in the mainline.\nImprovements like the earlier sliding mmap() series need to be\nable to change the interfaces without backward compatibility\nwart.\n\nIn other words, I do not know what idiot ^W ^W who listed the\nlibification stuff on the SoC \"ideas\" page, but I think (1) it\nis premature to promise stable ABI, and (2) if it does not\npromise stable ABI a library is not very useful.\n\n>       o Build system issues\n>\n> You missed this, but I think its an important consideration.\n> Our current libgit.a is a static library that has a relatively large\n> number of symbols its modules are exporting.  These symbol names\n> aren't namespace-ized (e.g. git_* prefix) so we wouldn't want to\n> just offer this library up in its current form.\n\nVery true, in fact, the current libgit.a is _NOT_ a library at\nall.  It is just a way to be terse in our Makefile to make the\nlinker do the work for us, nothing more.\n\nAnd I do not think we would want to rename our \"internally\npublic\" functions such as find_pack_entry_one() and\nsha1_object_info() with git_ prefix only for the purpose of this\nlibification.\n\nIf we can trick the linker to create gitlib.so which defines the\nsymbol git_sha1_object_info() that lets the caller to call our\ninternal sha1_object_info(), without exposing the internal name\nsha1_object_info(), and strip other global names libgit.a and\nplumbing internally use to communicate each other, such as\nfind_pack_entry_one(), from the gitlib.so library, that would be\na good solution.\n\n>>  2. What's the minimum amount of work that need to be done for\n>> the SoC project to be considered successful?\n>\n> I'd like to see enough API support that gitweb.cgi could:\n>\n>  * get the most recent commit date of all refs in all projects\n>    (the toplevel project index page);\n>  * get a shortlog for the main summary page of a project;\n>  * get the full content of a single commit;\n>  * get the \"raw\" diff (paths that changed) for two commits;\n\nI would disagree with tying libification and Perl binding this\nway.  If the goal is to get faster gitweb, then that does not\nnecessarily have to be libified git.  Let one person who does\nthe libification come up with a decent C binding and let others\nworry about Perl bindings.\n\n> In some cases much of the above is already \"internally public\";\n> meaning we already treat parts of that code as a library and invoke\n> them from within processes to get work done.  Much of this project\n> is about improving the interfaces and behavior enough to make those\n> existing APIs truely public.\n\nOne big thing you forgot to mention is that whatever form it\ntakes, the libification should not impact performance of\nexisting plumbing.  These interfaces are \"internally\" public\nexactly because the callers still honor underlying convention\nsuch as not having to clean-up the object flags for the last\ninvocation.  If you libify in a wrong way, you would end up an\nimplementation of the interface that always cleans up (because\nyou would not know if you are part of a long-living process so\nyou will clean-up just in case you will still be called later),\nwhich would be unusable from the plumbing point-of-view.\n"},{"id":"37206","messageId":"20070316060033.GD31606@spearce.org","threadId":"7265","inReplyTo":"7vejnpycu1.fsf@assigned-by-dhcp.cox.net","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-16T06:00:33Z","receivedAt":"2007-03-16T06:00:33Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> > On the other hand, many of the variables declared in environment.c\n> > are repository specific configuration variables.  These probably\n> > should be abstracted into some sort of wrapper, so that multiple\n> > repositories can be accessed from within the same process.  Why?\n> > a future mod_perl running gitweb.cgi accessing repositories through\n> > libgit.a and Perl bindings of course!\n> \n> I think if you are abstracting them out, into \"struct repo_state\",\n> the index and object store related variables such as packed_git\n> should go there as well, so your recommendation feels very\n> inconsistent to me.\n\nI missed packed_git, but you are right, that should definately go\nwith a struct repo_state.  And maybe you are right that the index\nshould go with it... but I'm not sure the index should be tied to the\nrepository at all.  Its strictly convention that the index goes with\nthe repository; GIT_INDEX_FILE lets you say otherwise at the command\nline level, why can't we do otherwise from a library level too?\n \n> >>     o Add prefix (eg, git_*) to public API functions\n> >\n> > Yes.  But which functions shall we expose?  ;-)\n> \n> Before going into that topic, a bigger question is if we are\n> happy with the current internal API and what the goal of\n> libification is.  If the libification is going to say that \"this\n> is a published API so we are not going to change it\", I would\n> imagine that it would be very hard to accept in the mainline.\n\nI'm looking at a middleground between our current \"moving target\"\ninternal API and our \"frozen\" plumbing process based API.  There\nare a number of places where just being able to get data *out*\nof Git easily would be useful, but doing so right now is awkward.\nEither you code against our \"moving target\" internal API by creating\na new builtin (e.g. my builtin-statplog) where its easy to get what\nyou want, or you code against the plumbing based tools, where its\nsometimes not so easy...\n\nMost of the data formats aren't changing; a commit is a commit is\na commit.  It has a tree, parents, author, committer, message.\n\n> Improvements like the earlier sliding mmap() series need to be\n> able to change the interfaces without backward compatibility\n> wart.\n\nI agree.  But I also think the use_mmap() API is just way too low\nlevel for a public library.  That particular change was pretty\nlow level.\n\nThink higher, like \"struct commit\".  That is actually too low still,\nas it doesn't really help you with the author and committer.\n\n> In other words, I do not know what idiot ^W ^W who listed the\n> libification stuff on the SoC \"ideas\" page,\n\nI'm the idiot ^W individual responsible.  ;-)\n\n> I would disagree with tying libification and Perl binding this\n> way.  If the goal is to get faster gitweb, then that does not\n> necessarily have to be libified git.  Let one person who does\n> the libification come up with a decent C binding and let others\n> worry about Perl bindings.\n\nYes.  However Perl bindings are often asked for.  And Marco Costalba\nmight like a working libgit that he could use for revision fetching\nin qgit.  I think that if patches for a library started to appear,\nanother interested party would start to at least play with them.\n \n> One big thing you forgot to mention is that whatever form it\n> takes, the libification should not impact performance of\n> existing plumbing.  These interfaces are \"internally\" public\n> exactly because the callers still honor underlying convention\n> such as not having to clean-up the object flags for the last\n> invocation.  If you libify in a wrong way, you would end up an\n> implementation of the interface that always cleans up (because\n> you would not know if you are part of a long-living process so\n> you will clean-up just in case you will still be called later),\n> which would be unusable from the plumbing point-of-view.\n\nI didn't forget; I just simply did not mention it.  I was considering\nwriting something to that effect, and probably should have.\n\nThis is a really valid point.  Git is insanely fast, partly because\nwe have a lot of \"run once\" types of applications and we have\noptimized for those.  Any sort of \"run many times\" reuse needs to\nnot make the \"run once\" guy pay for something he will not use.\n\nA good example of this is in git-describe, where we use the object\nflags, and only bother to clear them out if there is another commit\nremaining to be described.\n\n-- \nShawn.\n"},{"id":"37210","messageId":"7vps79wueu.fsf@assigned-by-dhcp.cox.net","threadId":"7265","inReplyTo":"20070316060033.GD31606@spearce.org","subject":"Re: Libification project (SoC)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-03-16T06:54:01Z","receivedAt":"2007-03-16T06:54:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Junio C Hamano <junkio@cox.net> wrote:\n>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>> > On the other hand, many of the variables declared in environment.c\n>> > are repository specific configuration variables.  These probably\n>> > should be abstracted into some sort of wrapper, so that multiple\n>> > repositories can be accessed from within the same process.  Why?\n>> > a future mod_perl running gitweb.cgi accessing repositories through\n>> > libgit.a and Perl bindings of course!\n>> \n>> I think if you are abstracting them out, into \"struct repo_state\",\n>> the index and object store related variables such as packed_git\n>> should go there as well, so your recommendation feels very\n>> inconsistent to me.\n>\n> I missed packed_git, but you are right, that should definately go\n> with a struct repo_state.  And maybe you are right that the index\n> should go with it... but I'm not sure the index should be tied to the\n> repository at all.  Its strictly convention that the index goes with\n> the repository; GIT_INDEX_FILE lets you say otherwise at the command\n> line level, why can't we do otherwise from a library level too?\n\nEven within a plumbing, being able to shuffle multiple indices\nat once would be very useful.  For example, if I were to rewrite\nunpack-trees, I would most likely read from the current index\nand trees and populate a new index from emptiness by appending\nto it, thereby avoiding the binary-search and insert costs.\n\nI've thought about the layering when Smurf first brought up the\nlibification (which was a loooong time ago), and concluded three\nlayered approach would be most useful.\n\nThe bottom layer is object store across repositories.  If we\nignore SHA-1 collisions as an issue (and we _will_ ignore it for\nforseeable future), unless you are doing \"read from one\nrepository and write that to another repository\", it is more\nhandy to be able to name an object and get its data without\nknowing which repository's object store it comes from, and it\nwould make \"git log master~A..master~B\" across repositories\n(i.e. 'master' of repository A and 'master' of repository B)\npossible.  An example interface would be like:\n\n(current)\nvoid *read_sha1_file(const unsigned char *sha1,\n\t\t     enum object_type *type,\n\t\t     unsigned long *size);\n\n(libified)\nvoid *git_read_sha1_file(struct gitlib *,\n\t\t\t const unsigned char *sha1,\n\t\t\t enum object_type *type,\n\t\t\t unsigned long *size);\n\nwhere \"struct gitlib\" has a list of \"struct object_store\", and\nwe will have:\n\nint git_add_object_store(struct gitlib *, const char *path);\n\nto add one directory as object store the toplevel gitlib structure\nknows about.  In a sense, \"struct gitlib\" and object store is so\nglobal that we might not even need to have it as a parameter\n(iow, it and \"struct object **obj_hash\" from object.c can stay\nglobal).\n\nThe middle layer is repositories, primarily their refs and\nreflogs.  An example interface would be like:\n\n(current)\nint get_sha1(const char *name, unsigned char *sha1);\n\n(libified)\nint git_get_sha1(struct git_repo *, const char *name, unsigned char *sha1);\n\nwhere \"struct git_repo\" is one repository (and it would have a\npointer to \"struct gitlib *\" so that we can follow objects to\nfollow parents and stuff).\n\nAnd the top layer would have indices, and working trees as\nper-invocation parameter.\n\n(current)\nint cache_name_pos(const char *name, int namelen);\nint unpack_trees(struct object_list *trees, struct unpack_trees_options *o);\n\n(libified)\nint git_cache_name_pos(struct git_cache *, const char *name, int namelen);\nint git_unpack_trees(struct object_list *trees, struct git_unpack_trees_options *o);\n\nwhere \"struct git_cache\" has \"index\" thingies, such as\nactive_cache, active_nr, active_alloc, and active_cache_tree.\nAnd we would have pointer to \"struct git_cache *\" in unpack_trees_options\nstructure.\n"},{"id":"37212","messageId":"45FA501B.FA5B9F30@eudaptics.com","threadId":"7265","inReplyTo":"20070316045928.GB31606@spearce.org","subject":"Re: Libification project (SoC)","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2007-03-16T08:06:51Z","receivedAt":"2007-03-16T08:06:51Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"\"Shawn O. Pearce\" wrote:\n> \"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br> wrote:\n> >     o Avoid dying when a function call fails (eg, malloc())\n> \n> malloc is a huge problem in the Git code today.  Almost all\n> of our malloc calls are actually through the xmalloc wrapper.\n> All xmalloc callers assume xmalloc will *never* fail.  This\n> makes it, uh, interesting.  ;-)\n\nYou could think about longjmp(3)ing out into main(), which would have to\nsetjmp(3). But in order to clean up intermediate frames, you would have\nto have a stack of setjmp/longjmp buffers.\n\nOh, well, how do I *love* them C++ exceptions!\n\n-- Hannes\n"},{"id":"37215","messageId":"vpqveh15zvn.fsf@olympe.imag.fr","threadId":"7265","inReplyTo":"45FA501B.FA5B9F30@eudaptics.com","subject":"Re: Libification project (SoC)","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-03-16T08:58:04Z","receivedAt":"2007-03-16T08:58:04Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johannes Sixt <J.Sixt@eudaptics.com> writes:\n\n> You could think about longjmp(3)ing out into main(), which would have to\n> setjmp(3). But in order to clean up intermediate frames, you would have\n> to have a stack of setjmp/longjmp buffers.\n>\n> Oh, well, how do I *love* them C++ exceptions!\n\nYou can have exceptions in C too.\n\nI've used it a bit while contributing to Baz 1.x (the fork of tla).\nThe library used was cexcept ( http://cexcept.sourceforge.net/ ).\n\nAs you mention, jumping is the easy part, and cleaning up is the hard\none. Baz was using talloc, hacked to somehow work with cexcept. The\nmini-library doesn't seem to be available as a tarball anymore, so I\ndid the checkout+targz in case someone's curious to have a look, and\nlazy enough not to install baz to get it:\n\nhttp://www-verimag.imag.fr/~moy/tmp/talloc-except--2.0.1--patch-2.tar.gz\n\nThis stuff is not supported anymore, but very small anyway.\n\n-- \nMatthieu\n"},{"id":"37228","messageId":"Pine.LNX.4.63.0703161248380.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"vpqveh15zvn.fsf@olympe.imag.fr","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T11:51:13Z","receivedAt":"2007-03-16T11:51:13Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Mar 2007, Matthieu Moy wrote:\n\n> Johannes Sixt <J.Sixt@eudaptics.com> writes:\n> \n> > You could think about longjmp(3)ing out into main(), which would have to\n> > setjmp(3). But in order to clean up intermediate frames, you would have\n> > to have a stack of setjmp/longjmp buffers.\n> >\n> > Oh, well, how do I *love* them C++ exceptions!\n> \n> You can have exceptions in C too.\n> \n> I've used it a bit while contributing to Baz 1.x (the fork of tla).\n> The library used was cexcept ( http://cexcept.sourceforge.net/ ).\n> \n> As you mention, jumping is the easy part, and cleaning up is the hard\n> one. Baz was using talloc, hacked to somehow work with cexcept. The\n> mini-library doesn't seem to be available as a tarball anymore, so I\n> did the checkout+targz in case someone's curious to have a look, and\n> lazy enough not to install baz to get it:\n> \n> http://www-verimag.imag.fr/~moy/tmp/talloc-except--2.0.1--patch-2.tar.gz\n> \n> This stuff is not supported anymore, but very small anyway.\n\nI was thinking about a similar approach some time ago. But that means that \nyou _must not_ have static variables that you rely on being initialised \ncorrectly.\n\nI mean, we have xmalloc(), and it would be easy to enforce xfree(), too \n(which would be good for memory profiling anyway), and we _could_ hack \nthat into tracking which pointers were returned after which checkpoint.\n\nBut we _cannot_ say which static variables should be initialised (and \nhow), after some \"exception\" was thrown at a certain point.\n\nCiao,\nDscho\n"},{"id":"37232","messageId":"Pine.LNX.4.63.0703161251200.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"7vps79wueu.fsf@assigned-by-dhcp.cox.net","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T11:54:52Z","receivedAt":"2007-03-16T11:54:52Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Mar 2007, Junio C Hamano wrote:\n\n> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> \n> > Junio C Hamano <junkio@cox.net> wrote:\n> >> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> >> > On the other hand, many of the variables declared in environment.c\n> >> > are repository specific configuration variables.  These probably\n> >> > should be abstracted into some sort of wrapper, so that multiple\n> >> > repositories can be accessed from within the same process.  Why?\n> >> > a future mod_perl running gitweb.cgi accessing repositories through\n> >> > libgit.a and Perl bindings of course!\n> >> \n> >> I think if you are abstracting them out, into \"struct repo_state\",\n> >> the index and object store related variables such as packed_git\n> >> should go there as well, so your recommendation feels very\n> >> inconsistent to me.\n> >\n> > I missed packed_git, but you are right, that should definately go\n> > with a struct repo_state.  And maybe you are right that the index\n> > should go with it... but I'm not sure the index should be tied to the\n> > repository at all.  Its strictly convention that the index goes with\n> > the repository; GIT_INDEX_FILE lets you say otherwise at the command\n> > line level, why can't we do otherwise from a library level too?\n> \n> Even within a plumbing, being able to shuffle multiple indices\n> at once would be very useful.  For example, if I were to rewrite\n> unpack-trees, I would most likely read from the current index\n> and trees and populate a new index from emptiness by appending\n> to it, thereby avoiding the binary-search and insert costs.\n> \n> I've thought about the layering when Smurf first brought up the\n> libification (which was a loooong time ago), and concluded three\n> layered approach would be most useful.\n> \n> The bottom layer is object store across repositories.  If we\n> ignore SHA-1 collisions as an issue (and we _will_ ignore it for\n> forseeable future), unless you are doing \"read from one\n> repository and write that to another repository\", it is more\n> handy to be able to name an object and get its data without\n> knowing which repository's object store it comes from, and it\n> would make \"git log master~A..master~B\" across repositories\n> (i.e. 'master' of repository A and 'master' of repository B)\n> possible.  An example interface would be like:\n> \n> (current)\n> void *read_sha1_file(const unsigned char *sha1,\n> \t\t     enum object_type *type,\n> \t\t     unsigned long *size);\n> \n> (libified)\n> void *git_read_sha1_file(struct gitlib *,\n> \t\t\t const unsigned char *sha1,\n> \t\t\t enum object_type *type,\n> \t\t\t unsigned long *size);\n> \n> where \"struct gitlib\" has a list of \"struct object_store\", and\n> we will have:\n> \n> int git_add_object_store(struct gitlib *, const char *path);\n> \n> to add one directory as object store the toplevel gitlib structure\n> knows about.  In a sense, \"struct gitlib\" and object store is so\n> global that we might not even need to have it as a parameter\n> (iow, it and \"struct object **obj_hash\" from object.c can stay\n> global).\n> \n> The middle layer is repositories, primarily their refs and\n> reflogs.  An example interface would be like:\n> \n> (current)\n> int get_sha1(const char *name, unsigned char *sha1);\n> \n> (libified)\n> int git_get_sha1(struct git_repo *, const char *name, unsigned char *sha1);\n> \n> where \"struct git_repo\" is one repository (and it would have a\n> pointer to \"struct gitlib *\" so that we can follow objects to\n> follow parents and stuff).\n> \n> And the top layer would have indices, and working trees as\n> per-invocation parameter.\n> \n> (current)\n> int cache_name_pos(const char *name, int namelen);\n> int unpack_trees(struct object_list *trees, struct unpack_trees_options *o);\n> \n> (libified)\n> int git_cache_name_pos(struct git_cache *, const char *name, int namelen);\n> int git_unpack_trees(struct object_list *trees, struct git_unpack_trees_options *o);\n> \n> where \"struct git_cache\" has \"index\" thingies, such as\n> active_cache, active_nr, active_alloc, and active_cache_tree.\n> And we would have pointer to \"struct git_cache *\" in unpack_trees_options\n> structure.\n\nIsn't this an awfully long shot?\n\nI'd be happy if the libification project resulted\n\n- in a (static!) libgit.a which can be linked to qgit or similar (being \n  reentrant, or at least optionally so, and not die()ing all the time), \n  and\n\n- which does not fix the API yet (at least for the most parts).\n\nWe _can_ -- once we agree on a stable API -- expose _some_ functions in a \nlibgit.so, but that does not have to be the goal for the first step!\n\nCiao,\nDscho\n"},{"id":"37236","messageId":"20070316125327.GC4489@pasky.or.cz","threadId":"7265","inReplyTo":"7vejnpycu1.fsf@assigned-by-dhcp.cox.net","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-16T12:53:27Z","receivedAt":"2007-03-16T12:53:27Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, Mar 16, 2007 at 06:30:46AM CET, Junio C Hamano wrote:\n> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> > But other areas die when they get given a bad SHA-1 (for example).\n> > If the library caller can supply that (possibly bad) SHA-1 to an\n> > API function, that's just mean to die out.  ;-)\n> \n> That's a real problem, but on the other hand, perl or whatever\n> wrapped ones can do the dying (or not dying) before calling into\n> libgit, so it may not be such a big issue.\n\nAt least you can catch the die from the library caller using\nset_*_routine(). ;-)\n\n> >>     o Documentation (eg, doxygen)\n> >>     o Unit-tests\n> >>     o Add prefix (eg, git_*) to public API functions\n> >\n> > Yes.  But which functions shall we expose?  ;-)\n> \n> Before going into that topic, a bigger question is if we are\n> happy with the current internal API and what the goal of\n> libification is.  If the libification is going to say that \"this\n> is a published API so we are not going to change it\", I would\n> imagine that it would be very hard to accept in the mainline.\n> Improvements like the earlier sliding mmap() series need to be\n> able to change the interfaces without backward compatibility\n> wart.\n> \n> In other words, I do not know what idiot ^W ^W who listed the\n> libification stuff on the SoC \"ideas\" page, but I think (1) it\n> is premature to promise stable ABI, and (2) if it does not\n> promise stable ABI a library is not very useful.\n\nI disagree, it can live in the \"zero major version\" realm and already be\nvery useful for language bindings (say whatever is bundled with git\nitself) and other nifty stuff.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37237","messageId":"20070316125529.GD4489@pasky.or.cz","threadId":"7265","inReplyTo":"20070316045928.GB31606@spearce.org","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-16T12:55:29Z","receivedAt":"2007-03-16T12:55:29Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, Mar 16, 2007 at 05:59:28AM CET, Shawn O. Pearce wrote:\n> \"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br> wrote:\n> >  3. I don't code in Perl, is it a problem? I mean, the project's\n> > goal is to have a Perl binding but I think it goes far from\n> > that: we could have a python module, a C program, or anything\n> > that shows the libgit is useful.\n> \n> No, I don't see that as a problem at all.  We have some Perl\n> experts on the mailing list who would like to see Perl bindings.\n> Some of the Perl binding is pure C code, and some if it is this\n> weird Perl macro language...  so I expect those Perl experts to come\n> out of the woodwork and help the community to create a prototype\n> set of bindings.  There's also Ruby and Python interests around,\n> so we may see bindings for those too.  ;-)\n\nI'll add perl binding as soon as libgit part is there; the\ninfrastructure is already in place (not now but it's in git history, you\njust have to dig it out), so it should be pretty easy too; so even if I\nwouldn't, someone surely will. ;-) I don't think knowing Perl or\nmoreover the Perl XS horrors should be a prerequisite for this project.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37238","messageId":"20070316130958.GD1783@peter.daprodeges.fqdn.th-h.de","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703161251200.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2007-03-16T13:09:58Z","receivedAt":"2007-03-16T13:09:58Z","isPatch":false,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Johannes Schindelin [07-03-16 12:54:52 +0100] wrote:\n\n[...]\n\n>Isn't this an awfully long shot?\n\n>I'd be happy if the libification project resulted\n\n>- in a (static!) libgit.a which can be linked to qgit or similar (being \n>  reentrant, or at least optionally so, and not die()ing all the time), \n>  and\n\n>- which does not fix the API yet (at least for the most parts).\n\n>We _can_ -- once we agree on a stable API -- expose _some_ functions in a \n>libgit.so, but that does not have to be the goal for the first step!\n\nFirst, I think that would be some cleanup \"only\" since that basically \nwould mean to\n\n   1) make all functions die()ing return some value and handle it and\n   2) wrap all static vars into structures and pass them around\n\nIf you don't choose a design before wrapping things up in structures, \nyou'll probably end up having one structure per source file (at least \ntoo many structures).\n\nPorting things like qgit to it or writting proper perl/python bindings \nis wasted time since you'd have to rewrite all of it once you decided \nwhich functions to expose and which structures to use (calling the \nmain() routines of builtin's doesn't count as real libifaction, it would \nrather be a performance improvement only).\n\nI'd simply try to find a rough consensus on the data structures and the \nlayer model before starting the project, solve 1), afterwards implement \n2) according to it. While 2) happens it would make sense to try to \ndevelop perl, python, C and C++ bindings in parallel to find out early \nenough whether the design details chosen are useful for real consumers \noutside the git-* tools.\n\nYou could put big fat warnings everywhere that parts of the API which \nare exposed are heavily unstable and likely subject to change and that \nprogrammers using them will have to frequently start over. Once it turns \nout that all the git-tools and all \"reference consumers\" work it, you \ncan do some cleanup to get to the final first API version after the \nlibification project is done.\n\n   bye, Rocco\n-- \n:wq!\n"},{"id":"37239","messageId":"20070316104715.483df0d5@localhost","threadId":"7265","inReplyTo":"7vejnpycu1.fsf@assigned-by-dhcp.cox.net","subject":"Re: Libification project (SoC)","fromName":"Luiz Fernando N. Capitulino","fromEmail":"lcapitulino@mandriva.com.br","sentAt":"2007-03-16T13:47:15Z","receivedAt":"2007-03-16T13:47:15Z","isPatch":false,"sender":{"key":"lcapitulino@mandriva.com.br","avatar":null},"body":"Em Thu, 15 Mar 2007 22:30:46 -0700\nJunio C Hamano <junkio@cox.net> escreveu:\n\n| \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n| \n| >>     o Documentation (eg, doxygen)\n| >>     o Unit-tests\n| >>     o Add prefix (eg, git_*) to public API functions\n| >\n| > Yes.  But which functions shall we expose?  ;-)\n| \n| Before going into that topic, a bigger question is if we are\n| happy with the current internal API and what the goal of\n| libification is.  If the libification is going to say that \"this\n| is a published API so we are not going to change it\", I would\n| imagine that it would be very hard to accept in the mainline.\n\n I think you can put this way: do you want/whish to make\ngit more useful than it's today?\n\n If so, such a library is important because it will allow\nusers to write application that use git in a reasonable\nway.\n\n It doesn't need to be the next five-zilion-function-library\nthat will provide the wonders of git in several different\nways.\n\n We could start by fixing the got-an-error-die behaivor and\ndefine a _experimental_ API (just a few functions) just to get\ndata out of git.\n\n This would be enough to write the Perl binding I think?\n\n-- \nLuiz Fernando N. Capitulino\n"},{"id":"37241","messageId":"20070316140855.GE4489@pasky.or.cz","threadId":"7265","inReplyTo":"20070316104715.483df0d5@localhost","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-16T14:08:55Z","receivedAt":"2007-03-16T14:08:55Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, Mar 16, 2007 at 02:47:15PM CET, Luiz Fernando N. Capitulino wrote:\n>  We could start by fixing the got-an-error-die behaivor and\n> define a _experimental_ API (just a few functions) just to get\n> data out of git.\n> \n>  This would be enough to write the Perl binding I think?\n\nActually, well, I've already done this. :-)\n\nThe trouble begins when you want to access multiple repositories from\nthe same process, etc. Without that, writing the Perl binding is\ntrivial; there's already a hook the binding can use to catch dies, I've\nadded it.\n\nSo, the main point of the work is to define a _good_ API and get rid of\nthe static state, I guess.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37243","messageId":"Pine.LNX.4.63.0703161509560.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"20070316130958.GD1783@peter.daprodeges.fqdn.th-h.de","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T15:12:17Z","receivedAt":"2007-03-16T15:12:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[please do not cull the Cc: list]\n\nOn Fri, 16 Mar 2007, Rocco Rutte wrote:\n\n> First, I think that would be some cleanup \"only\" since that basically would\n> mean to\n> \n>   1) make all functions die()ing return some value and handle it and\n>   2) wrap all static vars into structures and pass them around\n> \n> If you don't choose a design before wrapping things up in structures, you'll\n> probably end up having one structure per source file (at least too many\n> structures).\n\nWhy? For some tasks, it should be 1) easier, 2) more elegant, and 3) \nfaster to write a function which re-initialises the static variables.\n\nOf course, if you want to work with multiple repos _at the same time_, \nthis does not help you. But frankly, we don't support that with core-git, \nso why should we in libgit?\n\n> Porting things like qgit to it or writting proper perl/python bindings \n> is wasted time since you'd have to rewrite all of it once you decided \n> which functions to expose and which structures to use (calling the \n> main() routines of builtin's doesn't count as real libifaction, it would \n> rather be a performance improvement only).\n\nNope. It is _not_ a complete rewrite. More likely, it is minimal \nadjustments. It's not like we will replace apples with cars...\n\n> I'd simply try to find a rough consensus on the data structures and the \n> layer model before starting the project, solve 1), afterwards implement \n> 2) according to it.\n\nWe already _have_ the data structures!\n\nAlso, in my experience, defining a complete API, and only after that, \nimplement it, never works. Rather, start with a _small_ part you want to \ndo. Define a clean API _just for that part_. Implement it. Verify that it \nindeed does what it should do (and that means not just _you_ should verify \nit, but it should be stress tested on the list).\n\nWe don't have to create the whole world in one day, you know?\n\nCiao,\nDscho\n"},{"id":"37244","messageId":"Pine.LNX.4.63.0703161612380.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"20070316104715.483df0d5@localhost","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T15:16:25Z","receivedAt":"2007-03-16T15:16:25Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Mar 2007, Luiz Fernando N. Capitulino wrote:\n\n>  It doesn't need to be the next five-zilion-function-library that will \n> provide the wonders of git in several different ways.\n\nYes. Just like we have a really small really stable part of core-git, \nwhich can be used by porcelains, and is expected to work the same in \nfuture versions, we could have eventually with libgit.\n\nThat would mean, for example, that rev_info should always be initialised \nwith malloc() so that future versions can make it bigger, and that new \nmembers be added always at the end.\n\n>  We could start by fixing the got-an-error-die behaivor and define a \n> _experimental_ API (just a few functions) just to get data out of git.\n\nThat sounds very reasonable.\n\nAnd if it does not work out as expected, we don't have to make it part of \n\"official\" Git. It can live on as a fork.\n\nCiao,\nDscho\n"},{"id":"37246","messageId":"alpine.LFD.0.83.0703161145520.5518@xanadu.home","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703161509560.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-16T15:55:52Z","receivedAt":"2007-03-16T15:55:52Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 16 Mar 2007, Johannes Schindelin wrote:\n\n> We already _have_ the data structures!\n\nWell... Shawn and I are contemplating alternate data structures to \nimprove things dramatically.\n\nWith a fixed public API I doubt such improvements could be as effective.\n\nOne thing that was really done right in the Linux kernel is to _not_ \nhave any sort of fixed API at all for drivers.  This is a big upside for \nprogress.  Yet the Linux kernel is regarded as highly useful.\n\nSo... if any API is to be developed, I'd argue that it must be done \n_above_ the existing code with a higher level of abstraction and a much \nnarrower scope.\n\n\nNicolas\n"},{"id":"37248","messageId":"Pine.LNX.4.63.0703161710400.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"alpine.LFD.0.83.0703161145520.5518@xanadu.home","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T16:13:04Z","receivedAt":"2007-03-16T16:13:04Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Mar 2007, Nicolas Pitre wrote:\n\n> On Fri, 16 Mar 2007, Johannes Schindelin wrote:\n> \n> > We already _have_ the data structures!\n> \n> Well... Shawn and I are contemplating alternate data structures to \n> improve things dramatically.\n\nI was alluding to rev_info, not pack_window and friends.\n\n> With a fixed public API I doubt such improvements could be as effective.\n\nJust think of the \"API\" we have for porcelains. It is literally unchanged \nsince the beginning. You can even use the original script git-log.sh \ntoday! _That_ is what I mean by fixed public API: give certain guarantees \nabout what will not go away.\n\n> One thing that was really done right in the Linux kernel is to _not_ \n> have any sort of fixed API at all for drivers.  This is a big upside for \n> progress.  Yet the Linux kernel is regarded as highly useful.\n\nYes. I am a Linux user myself.\n\nCiao,\nDscho\n"},{"id":"37251","messageId":"20070316161752.GA3275@spearce.org","threadId":"7265","inReplyTo":"alpine.LFD.0.83.0703161145520.5518@xanadu.home","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-16T16:17:52Z","receivedAt":"2007-03-16T16:17:52Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 16 Mar 2007, Johannes Schindelin wrote:\n> \n> > We already _have_ the data structures!\n> \n> Well... Shawn and I are contemplating alternate data structures to \n> improve things dramatically.\n\nHang on.  Yes, Nico and I are contemplating alternate disk based\ndata structure, and in some cases, alternate memory based data\nstructures to improve things.\n\nBut these structures are not changing the basic Git data structures\nthat have been with us since way back when. ;-) Commits still\nhave the same fields, with the same data and the same meaning.\nTrees still have the same fields, and same meaning... etc.\n\n> With a fixed public API I doubt such improvements could be as effective.\n\nThey still can be, and without shooting ourselves in the foot in the\nprocess.\n \n> So... if any API is to be developed, I'd argue that it must be done \n> _above_ the existing code with a higher level of abstraction and a much \n> narrower scope.\n\nYes.  Today we have a frozen API for commit walking.  Its called\n`git rev-list --pretty=raw A ^B`.  That output format is pretty\nwell set in stone, and we cannot change it.  Everyone knows what\neach field means, and hopefully knows that additional fields can\nbe added.  ;-)\n\nInstead of formatting out those fields as hex strings, or as decimal\ninteger dates, we can offer them in a struct.  E.g.:\n\n\tstruct git_objid {\n\t\tconst unsigned char *obj_name;\n\t};\n\n\tstruct git_commit {\n\t\tstruct git_objid tree;\n\t\tstruct git_objid *parents;\n\t\tuint32_t nr_parent;\n\t\tconst char *author;\n\t\ttime_t author_date;\n\t\tint author_tz;\n\t\tconst char *committer;\n\t\ttime_t committer_date;\n\t\tint committer_tz;\n\t\tconst char *message;\n\t};\n\nWith the rule that the pointers are to static memory buffers that\nlibgit is loaning out to the caller (the caller should *not* free\nthese buffers).  This lets us play cute tricks down in the lower\ntiers by pointing directly into the packfile dictionary tables\n(saves memcpys); or xstrdup/xmalloc everything we give out if we\nwant to be really paranoid.\n\nJust tossing ideas out - don't think that what I wrote above is my\nfinal suggestion on the matter.  It may change in another day or\ntwo if I think about it more.  ;-)\n\n-- \nShawn.\n"},{"id":"37253","messageId":"alpine.LFD.0.83.0703161218140.5518@xanadu.home","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703161710400.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-16T16:26:25Z","receivedAt":"2007-03-16T16:26:25Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 16 Mar 2007, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Fri, 16 Mar 2007, Nicolas Pitre wrote:\n> \n> > On Fri, 16 Mar 2007, Johannes Schindelin wrote:\n> > \n> > > We already _have_ the data structures!\n> > \n> > Well... Shawn and I are contemplating alternate data structures to \n> > improve things dramatically.\n> \n> I was alluding to rev_info, not pack_window and friends.\n> \n> > With a fixed public API I doubt such improvements could be as effective.\n> \n> Just think of the \"API\" we have for porcelains. It is literally unchanged \n> since the beginning. You can even use the original script git-log.sh \n> today! _That_ is what I mean by fixed public API: give certain guarantees \n> about what will not go away.\n\nSure.  But the output from an executable is a damn good abstraction and \nthe executable itself is an impenetrable boundary.  Anything can change \n(and did change) underneath.\n\nThis is why a public API must be done at a higher level to allow for \nanything to change at the lower level as we wish.\n\n\nNicolas\n"},{"id":"37268","messageId":"e5bfff550703161120o4571769eq18c13ae29ac79957@mail.gmail.com","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703161509560.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-03-16T18:20:26Z","receivedAt":"2007-03-16T18:20:26Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/16/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>\n> > Porting things like qgit to it or writting proper perl/python bindings\n> > is wasted time since you'd have to rewrite all of it once you decided\n> > which functions to expose and which structures to use (calling the\n> > main() routines of builtin's doesn't count as real libifaction, it would\n> > rather be a performance improvement only).\n>\n> Nope. It is _not_ a complete rewrite. More likely, it is minimal\n> adjustments. It's not like we will replace apples with cars...\n>\n\nIMHO probably the truth is in the middle. I wouldn't call it a trivial\nporting, at least for me, but anyway it would be interesting to have\nfun with linking libgit.\n\n*The most important thing for a libgit to be used by qgit is reentrancy*\n\nCurrently an unlimited number of tabs could be open in qgit, I'm not\ntalking about tabs open on different repos, but different views on the\nsame repo: main view, file history of file A, file history of file B,\ntree view, i.e. select some files/directory from directory tree and\nview the revisions that modified that repo subset, and so on. Other\ndifferent views could be added in the future. Because each view has a\ndedicated tab and each tab calls _his_ 'git rev-list' instance (could\nbe called also at the same time) this libgit thing should be able to\nsupport many instance of the libified git-rev-list function running at\nthe same time.\n\nPerhaps currently this need is only for qgit among the GUI browsers,\nbut it would be not too difficult to foreseen a multi view GUI\ninterface as a relative common feature in the future also for the\nremaining crop of git tools.\n\n    Marco\n"},{"id":"37269","messageId":"1174069353.2599.13.camel@mejai","threadId":"7265","inReplyTo":"alpine.LFD.0.83.0703161218140.5518@xanadu.home","subject":"Re: Libification project (SoC)","fromName":"Steve Frécinaux","fromEmail":"nudrema@gmail.com","sentAt":"2007-03-16T18:22:33Z","receivedAt":"2007-03-16T18:22:33Z","isPatch":false,"sender":{"key":"nudrema@gmail.com","avatar":null},"body":"On Fri, 2007-03-16 at 12:26 -0400, Nicolas Pitre wrote:\n\n> Sure.  But the output from an executable is a damn good abstraction and \n> the executable itself is an impenetrable boundary.  Anything can change \n> (and did change) underneath.\n\nStrictly speaking, you can use opaque structures for commits and so on\n(so that the outside world will only ever see a pointer), and use some\ngetter/setters for commonly used stuffs (like datum, title, content).\n\nAlso, I guess what people would expect from a C library is roughly the\nsame as for the current plumbing... just easier to use from another\nprogram. It doesn't need a low-level access to data structure (most\napplications would be to interact with an existing repo or to store data\nfor a third-party software, something that is high-level) and I don't\nthink such an opaque API would be a huge constraint as soon as you keep\nthe Object/Index/Tree/Commit/etc basic opaque structs.\n"},{"id":"37270","messageId":"e5bfff550703161138x5ab1fe3anf7b2aaab81bb77e4@mail.gmail.com","threadId":"7265","inReplyTo":"e5bfff550703161120o4571769eq18c13ae29ac79957@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-03-16T18:38:08Z","receivedAt":"2007-03-16T18:38:08Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/16/07, Marco Costalba <mcostalba@gmail.com> wrote:\n>\n> *The most important thing for a libgit to be used by qgit is reentrancy*\n>\n\nAnother crtitical feature is that this call to git-rev-list-like\nfunction MUST be non-blocking.\n\nReading a big repo could take many seconds, also more then 10 seconds\nin cold cache case for Linux tree, as example. Getting the history of\na file ('git rev-list -- /path/to/file) it's also very slow.\n\nThere is no way that a GUI tool is allowed to *freeze* for that amount\nof time. Currently, because an external process is forked when running\n'git rev-list' all the problem is happly handled by the kernel\nscheduler and the QProcess callback mechanism (based on select()). In\ncase of a libified git-rev-list this could be an issue.\n\n   Marco\n"},{"id":"37271","messageId":"20070316153822.5c842e69@localhost","threadId":"7265","inReplyTo":"20070316140855.GE4489@pasky.or.cz","subject":"Re: Libification project (SoC)","fromName":"Luiz Fernando N. Capitulino","fromEmail":"lcapitulino@mandriva.com.br","sentAt":"2007-03-16T18:38:22Z","receivedAt":"2007-03-16T18:38:22Z","isPatch":false,"sender":{"key":"lcapitulino@mandriva.com.br","avatar":null},"body":"Em Fri, 16 Mar 2007 15:08:55 +0100\nPetr Baudis <pasky@suse.cz> escreveu:\n\n| On Fri, Mar 16, 2007 at 02:47:15PM CET, Luiz Fernando N. Capitulino wrote:\n| >  We could start by fixing the got-an-error-die behaivor and\n| > define a _experimental_ API (just a few functions) just to get\n| > data out of git.\n| > \n| >  This would be enough to write the Perl binding I think?\n| \n| Actually, well, I've already done this. :-)\n\n Not exactly, at least not the way I think it should be done.\n\n| The trouble begins when you want to access multiple repositories from\n| the same process, etc. Without that, writing the Perl binding is\n| trivial; there's already a hook the binding can use to catch dies, I've\n| added it.\n| \n| So, the main point of the work is to define a _good_ API and get rid of\n| the static state, I guess.\n\n Yes, the set_*_routine()s seems a workaround to me, you're only fixing\ndie()'s final effect.\n\n I think the right solution is to get rid of die() from functions that\nare supposed to be an interface, set errno if needed and return -1\nor NULL.\n\n That looks a lot of work BTW, but I'll be pleased to work on it.\n\n Is there more things like the set_*_routine()s added to fix\nother problems?\n\n-- \nLuiz Fernando N. Capitulino\n"},{"id":"37272","messageId":"alpine.LFD.0.83.0703161433300.18328@xanadu.home","threadId":"7265","inReplyTo":"1174069353.2599.13.camel@mejai","subject":"Re: Libification project (SoC)","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-16T18:53:06Z","receivedAt":"2007-03-16T18:53:06Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 16 Mar 2007, Steve Frécinaux wrote:\n\n> Also, I guess what people would expect from a C library is roughly the\n> same as for the current plumbing... just easier to use from another\n> program. It doesn't need a low-level access to data structure (most\n> applications would be to interact with an existing repo or to store data\n> for a third-party software, something that is high-level) and I don't\n> think such an opaque API would be a huge constraint as soon as you keep\n> the Object/Index/Tree/Commit/etc basic opaque structs.\n\nRight.  I like that idea.\n\nA good way to define the lib API needs then might be expressed as \nfollows:\n\n  Each existing plumbing commands must be turned into the minimal \n  implementation required to interact with the libgit public API and\n  display results.\n\n  In other words, the public libgit API should provide the same \n  functionality as existing plumbing commands such that those existing\n  commands will only need the necessary code to bridge the C interface\n  with the existing command line interface.\n\nThen, of course, there is the matter of reentrancy.  But that's still a \nminor API detail even if it is not a trivial issue implementation wise.  \nBut the API must be right as this is what we'll be stuck with even if \nthe implementation may change.  And as far as an API definition is \nneeded I think that it should reflect the current plumbing which is \nactually the real API that grew naturally and has been proven useful.\n\n\nNicolas\n"},{"id":"37273","messageId":"alpine.LFD.0.83.0703161454280.18328@xanadu.home","threadId":"7265","inReplyTo":"e5bfff550703161138x5ab1fe3anf7b2aaab81bb77e4@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-16T18:59:49Z","receivedAt":"2007-03-16T18:59:49Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 16 Mar 2007, Marco Costalba wrote:\n\n> On 3/16/07, Marco Costalba <mcostalba@gmail.com> wrote:\n> > \n> > *The most important thing for a libgit to be used by qgit is reentrancy*\n> > \n> \n> Another crtitical feature is that this call to git-rev-list-like\n> function MUST be non-blocking.\n\nI'm not sure I agree.\n\nThe non-blockingness can be (and probably should be) handled at a higher \nlevel with your own threading facility of choice.  Making GIT \nrestartable has the potential for making the core code much too complex.\n\n\nNicolas\n"},{"id":"37274","messageId":"200703161909.38662.andyparkins@gmail.com","threadId":"7265","inReplyTo":"e5bfff550703161138x5ab1fe3anf7b2aaab81bb77e4@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-03-16T19:09:36Z","receivedAt":"2007-03-16T19:09:36Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2007, March 16, Marco Costalba wrote:\n\n> There is no way that a GUI tool is allowed to *freeze* for that\n> amount of time. Currently, because an external process is forked when\n> running 'git rev-list' all the problem is happly handled by the\n> kernel scheduler and the QProcess callback mechanism (based on\n> select()). In case of a libified git-rev-list this could be an issue.\n\nI don't think that is ever going to be an issue.  At the worst you could \njust fork() and run the libgit command in that.  Threads are fairly \neasy in Qt as well.\n\nIn short, I wouldn't worry about libgit blocking - in fact it's almost a \nguarantee that libgit /will/ block; it would be a nightmare to write an \nasynchronous libgit.\n\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"37277","messageId":"e5bfff550703161407u6afefae9u4a23cf1cb49125ce@mail.gmail.com","threadId":"7265","inReplyTo":"alpine.LFD.0.83.0703161454280.18328@xanadu.home","subject":"Re: Libification project (SoC)","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-03-16T21:07:43Z","receivedAt":"2007-03-16T21:07:43Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/16/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 16 Mar 2007, Marco Costalba wrote:\n>\n> > On 3/16/07, Marco Costalba <mcostalba@gmail.com> wrote:\n> > >\n> > > *The most important thing for a libgit to be used by qgit is reentrancy*\n> > >\n> >\n> > Another crtitical feature is that this call to git-rev-list-like\n> > function MUST be non-blocking.\n>\n> I'm not sure I agree.\n>\n> The non-blockingness can be (and probably should be) handled at a higher\n> level with your own threading facility of choice.  Making GIT\n> restartable has the potential for making the core code much too complex.\n>\n\nThe fact is that the solution is complex anyway, moving the complex\ncode at higher level doesn't simplify the whole issue, it just *moves*\nthe issue somewhere else.\n\nBTW now qgit is single-threaded (as gitk), you suggest that linking\nwith libgit it will involve to go on the multi threading side and I\nthink you are right. But it will be not that easy.\n\nCurrently we have both single threaded GUI tools and blocking git\ncommands and it works nicely not because it's simple but because the\n'complex code' is hidden inside the OS process handling and scheduling\nstuff.\n\nLinking with a synchronous libgit it means, roughly speaking, take the\n'complex code' out from the OS and put somewhere in user space, or in\nlibgit or in the user GUI tool linked with the library.\n\nNow, it happens that Qt has a good multi thread support, but this is\njust incidental and of course cannot be taken as granted by a git\nlibrary that aims to be broadly and possibly easily used.\n\nBecause we are just speaking (well, writing ;-) ) about a possible\nlibrary I think we could take in account what would involve to\nforeseen a callback mechanism in the API, at least for the slowest\nones.\n\n    Marco\n"},{"id":"37285","messageId":"20070316231646.GB4508@spearce.org","threadId":"7265","inReplyTo":"20070316153822.5c842e69@localhost","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-16T23:16:46Z","receivedAt":"2007-03-16T23:16:46Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br> wrote:\n>  I think the right solution is to get rid of die() from functions that\n> are supposed to be an interface, set errno if needed and return -1\n> or NULL.\n\nAnd then make their callers (if they are above the public API layer)\ndie instead.  In some cases this might imply an undesirable change\nin the error message produced, as necessary details that are included\ntoday would be unavailable in the caller.\n \n>  Is there more things like the set_*_routine()s added to fix\n> other problems?\n\nNot that I am aware of.\n\n-- \nShawn.\n"},{"id":"37287","messageId":"Pine.LNX.4.63.0703170014130.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"e5bfff550703161407u6afefae9u4a23cf1cb49125ce@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T23:24:44Z","receivedAt":"2007-03-16T23:24:44Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Mar 2007, Marco Costalba wrote:\n\n> On 3/16/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Fri, 16 Mar 2007, Marco Costalba wrote:\n> > \n> > > On 3/16/07, Marco Costalba <mcostalba@gmail.com> wrote:\n> > > >\n> > > > *The most important thing for a libgit to be used by qgit is \n> > > > reentrancy*\n> > > >\n> > >\n> > > Another crtitical feature is that this call to git-rev-list-like\n> > > function MUST be non-blocking.\n> > \n> > I'm not sure I agree.\n\nI am sure I don't agree.\n\n> > The non-blockingness can be (and probably should be) handled at a \n> > higher level with your own threading facility of choice.  Making GIT \n> > restartable has the potential for making the core code much too \n> > complex.\n> \n> The fact is that the solution is complex anyway, moving the complex code \n> at higher level doesn't simplify the whole issue, it just *moves* the \n> issue somewhere else.\n\nIt not only *moves* the issue somewhere else, but it also cleanly \nseparates the issues.\n\n> BTW now qgit is single-threaded (as gitk), you suggest that linking with \n> libgit it will involve to go on the multi threading side and I think you \n> are right. But it will be not that easy.\n\nWhy?\n\nFirst, it _is_ multi-threaded, since it calls external programs. That is \neven more than a thread. It is a process.\n\nSecond, it _would_ be easy to just use the threads provided by Qt.\n\n> Because we are just speaking (well, writing ;-) ) about a possible \n> library I think we could take in account what would involve to foreseen \n> a callback mechanism in the API, at least for the slowest ones.\n\nWe are talking about libgit. Which should make access to certain common \nfunctions on Git repositories easy. Nothing more than that.\n\nIf you need to do that asynchronously, do _not_ fiddle with libgit. Just \nimagine what this would involve: you'd have to have timeouts (since there \nis _NO_ other way to find out when to return with empty hands, instead of \nblocking), which is _not_ portable. You'd soon be in the same _mess_ we \nare talking about with respect to exceptions.\n\nAlso, you would make _all_ operations expensive, since they _would_ have \nto store state to be restartable.\n\nThe common solution for your problem _is_ to use threads.\n\nAnd you have to admit that _only_ viewers would need asynchronous access \nanyway. I doubt that other tools -- which could take their advantage of a \nlibgit -- would need such an access.\n\nCiao,\nDscho\n"},{"id":"37288","messageId":"Pine.LNX.4.63.0703170025100.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"alpine.LFD.0.83.0703161218140.5518@xanadu.home","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-16T23:26:18Z","receivedAt":"2007-03-16T23:26:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Mar 2007, Nicolas Pitre wrote:\n\n> [...] the output from an executable is a damn good abstraction and the \n> executable itself is an impenetrable boundary.  Anything can change (and \n> did change) underneath.\n> \n> This is why a public API must be done at a higher level to allow for \n> anything to change at the lower level as we wish.\n\nAbsolutely.\n\nCiao,\nDscho\n"},{"id":"37296","messageId":"etfjb1$uof$1@sea.gmane.org","threadId":"7265","inReplyTo":"20070316042406.7e750ed0@home.brethil","subject":"Re: Libification project (SoC)","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-03-17T02:24:09Z","receivedAt":"2007-03-17T02:24:09Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Cc: git@vger.kernel.org]\n\nLuiz Fernando N. Capitulino wrote:\n\n>     o Documentation (eg, doxygen)\n\nI wonder if documenting and finishing documentation of git storage structure\n(format description of: loose objects, packs, pack indices, index, refs and\nsymbolic refs, packed refs) and git protocols (git protocol description,\nlocal/ssh fetch/push pipeline description), perhaps using RFC or RFC-like\nnotation could (and should) be made part of libification effort...\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"37303","messageId":"20070317052258.GB5731@spearce.org","threadId":"7265","inReplyTo":"etfjb1$uof$1@sea.gmane.org","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-17T05:22:58Z","receivedAt":"2007-03-17T05:22:58Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n> [Cc: git@vger.kernel.org]\n> \n> Luiz Fernando N. Capitulino wrote:\n> \n> >     o Documentation (eg, doxygen)\n> \n> I wonder if documenting and finishing documentation of git storage structure\n> (format description of: loose objects, packs, pack indices, index, refs and\n> symbolic refs, packed refs) and git protocols (git protocol description,\n> local/ssh fetch/push pipeline description), perhaps using RFC or RFC-like\n> notation could (and should) be made part of libification effort...\n\nI would consider that out of scope for this project.\n\nIt would be nice if someone did this work, or at least dusted\noff \"A Large Angry SCM\"'s document and made that available in the\nDocumentation/technical folder.  But I don't think it should be part\nof the Libification SoC project, or any of our other current ideas.\n\nUsers of a public API don't need to know the internal formatting\nof an object within a packfile.  They do however need to know that\na commit has a tree, and 0-n parents.  And that's already covered\nin our existing docs.\n\n\nAnd *please* stop breaking the CC chains Jakub.  We've asked you\nto not do that.  I had to go lookup Luiz' email address so I could\nget him back onto it.\n\n-- \nShawn.\n"},{"id":"37304","messageId":"e5bfff550703170004n3ab9075euf66a9e6dd56040d7@mail.gmail.com","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703170014130.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-03-17T07:04:58Z","receivedAt":"2007-03-17T07:04:58Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/17/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>\n> We are talking about libgit. Which should make access to certain common\n> functions on Git repositories easy. Nothing more than that.\n>\n\nFair enough.\n\n> If you need to do that asynchronously, do _not_ fiddle with libgit. Just\n> imagine what this would involve: you'd have to have timeouts (since there\n> is _NO_ other way to find out when to return with empty hands, instead of\n> blocking), which is _not_ portable. You'd soon be in the same _mess_ we\n> are talking about with respect to exceptions.\n>\n> Also, you would make _all_ operations expensive, since they _would_ have\n> to store state to be restartable.\n>\n> The common solution for your problem _is_ to use threads.\n>\n\nI would say, the common solution to have non blocking libgit is to use\nthreads in the tool linked with libgit.\n\nThis is clearly a  design choice and I agree it's an important\nstatement to keep libgit simple and portable (otherwise you'd probably\nneed to use a thread library as pthread in libgit). Thread facility in\nQt is instead already portable and well integrated. Anyway it's a\ndesign choice perhaps worth documenting.\n\n> And you have to admit that _only_ viewers would need asynchronous access\n> anyway. I doubt that other tools -- which could take their advantage of a\n> libgit -- would need such an access.\n>\n\nYes, and you have to admit ;-)  that viewers are the tools that mostly\nwill use libgit.\n\n    Marco\n"},{"id":"37322","messageId":"Pine.LNX.4.63.0703171054580.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"e5bfff550703170004n3ab9075euf66a9e6dd56040d7@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-17T17:29:43Z","receivedAt":"2007-03-17T17:29:43Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 17 Mar 2007, Marco Costalba wrote:\n\n> On 3/17/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> \n> > The common solution for your problem _is_ to use threads.\n> \n> I would say, the common solution to have non blocking libgit is to use \n> threads in the tool linked with libgit.\n\nYes, that's what I tried to say.\n\n> > And you have to admit that _only_ viewers would need asynchronous \n> > access anyway. I doubt that other tools -- which could take their \n> > advantage of a libgit -- would need such an access.\n> \n> Yes, and you have to admit ;-)  that viewers are the tools that mostly \n> will use libgit.\n\nI hope that there are many more users. _And_ not all viewers want to do \nthe display asynchronously. For example, statplot takes the time it takes.\n\nCiao,\nDscho\n"},{"id":"37331","messageId":"20070317195832.2af87c06@home.brethil","threadId":"7265","inReplyTo":"20070316231646.GB4508@spearce.org","subject":"Re: Libification project (SoC)","fromName":"Luiz Fernando N. Capitulino","fromEmail":"lcapitulino@mandriva.com.br","sentAt":"2007-03-17T19:58:32Z","receivedAt":"2007-03-17T19:58:32Z","isPatch":false,"sender":{"key":"lcapitulino@mandriva.com.br","avatar":null},"body":"On Fri, 16 Mar 2007 19:16:46 -0400\n\"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n\n| \"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br> wrote:\n| >  I think the right solution is to get rid of die() from functions that\n| > are supposed to be an interface, set errno if needed and return -1\n| > or NULL.\n| \n| And then make their callers (if they are above the public API layer)\n| die instead.  In some cases this might imply an undesirable change\n| in the error message produced, as necessary details that are included\n| today would be unavailable in the caller.\n\n Exactly!\n\n One simple example of an important error message that would be\nlost can be found in read-cache.c:read_cache_from():\n\n o index file smaller than expected\n\n I've found a possible solution, though.\n\n Take a look at Rusty's solution for the same problem in\nmodule-init-tools:\n\n\"\"\"\n/* We use error numbers in a loose translation... */\nstatic const char *insert_moderror(int err)\n{\n\tswitch (err) {\n\tcase ENOEXEC:\n\t\treturn \"Invalid module format\";\n\tcase ENOENT:\n\t\treturn \"Unknown symbol in module, or unknown parameter (see dmesg)\";\n\tcase ENOSYS:\n\t\treturn \"Kernel does not have module support\";\n\tdefault:\n\t\treturn strerror(err);\n\t}\n}\n\"\"\"\n\n Instead of calling strerror() directly for error generated\nwhen inserting a module, the insmod() function calls insert_moderror()\nwhich provides the desirable mapping.\n\n I think we could have something like that for each git's\nmodule, eg, git_cache_strerror(), git_commit_strerror() and so on.\n\n Does this look reasonable?\n\n-- \nLuiz Fernando N. Capitulino\n"},{"id":"37353","messageId":"20070318052332.GC15885@spearce.org","threadId":"7265","inReplyTo":"20070317195832.2af87c06@home.brethil","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-18T05:23:32Z","receivedAt":"2007-03-18T05:23:32Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br> wrote:\n> On Fri, 16 Mar 2007 19:16:46 -0400\n> \"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n> | And then make their callers (if they are above the public API layer)\n> | die instead.  In some cases this might imply an undesirable change\n> | in the error message produced, as necessary details that are included\n> | today would be unavailable in the caller.\n> \n>  I've found a possible solution, though.\n> \n>  Take a look at Rusty's solution for the same problem in\n> module-init-tools:\n> \n> \"\"\"\n> /* We use error numbers in a loose translation... */\n> static const char *insert_moderror(int err)\n> {\n> \tswitch (err) {\n> \tcase ENOEXEC:\n> \t\treturn \"Invalid module format\";\n> \tcase ENOENT:\n> \t\treturn \"Unknown symbol in module, or unknown parameter (see dmesg)\";\n> \tcase ENOSYS:\n> \t\treturn \"Kernel does not have module support\";\n> \tdefault:\n> \t\treturn strerror(err);\n> \t}\n> }\n> \"\"\"\n\nTake a look at sha1_file.c, open_packed_git_1:\n\n...\n    if (!pack_version_ok(hdr.hdr_version))\n        return error(\"packfile %s is version %u and not supported\"\n            \" (try upgrading GIT to a newer version)\",\n            p->pack_name, ntohl(hdr.hdr_version));\n...\n\nHere we are supplying a lot more than just a simple error code\nthat can be mapped to a static string.\n\nOf course that code is currently feeding it to the error function,\nwhich today calls the error_routine (see usage.c).  We could buffer\nthe strings sent to error()/warn() and let the caller obtain all\nstrings that occurred during the last API call.\n\n-- \nShawn.\n"},{"id":"37355","messageId":"7vzm6bp07f.fsf@assigned-by-dhcp.cox.net","threadId":"7265","inReplyTo":"20070318052332.GC15885@spearce.org","subject":"Re: Libification project (SoC)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-03-18T05:52:52Z","receivedAt":"2007-03-18T05:52:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Take a look at sha1_file.c, open_packed_git_1:\n>\n> ...\n>     if (!pack_version_ok(hdr.hdr_version))\n>         return error(\"packfile %s is version %u and not supported\"\n>             \" (try upgrading GIT to a newer version)\",\n>             p->pack_name, ntohl(hdr.hdr_version));\n> ...\n>\n> Here we are supplying a lot more than just a simple error code\n> that can be mapped to a static string.\n>\n> Of course that code is currently feeding it to the error function,\n> which today calls the error_routine (see usage.c).  We could buffer\n> the strings sent to error()/warn() and let the caller obtain all\n> strings that occurred during the last API call.\n\nActually, since we are talking about the error path,\n\n (1) we do not care performance of what happens there that much, but\n (2) we *do* care about not doing extra allocation.\n\nSo it might make sense to have a preallocated \"error string\"\nbuffer, sprintf the error message in there and return error\ncodes.\n"},{"id":"37373","messageId":"20070318135706.GF4489@pasky.or.cz","threadId":"7265","inReplyTo":"alpine.LFD.0.83.0703161433300.18328@xanadu.home","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-18T13:57:06Z","receivedAt":"2007-03-18T13:57:06Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, Mar 16, 2007 at 07:53:06PM CET, Nicolas Pitre wrote:\n> A good way to define the lib API needs then might be expressed as \n> follows:\n> \n>   Each existing plumbing commands must be turned into the minimal \n>   implementation required to interact with the libgit public API and\n>   display results.\n> \n>   In other words, the public libgit API should provide the same \n>   functionality as existing plumbing commands such that those existing\n>   commands will only need the necessary code to bridge the C interface\n>   with the existing command line interface.\n\nI think this is good definition if interpreted well - that is, git-log\nlibrary equivalent shouldn't spew out textual output but provide\ninterface to retrieve revision information in easy-to-use format.\n\n> Then, of course, there is the matter of reentrancy.  But that's still a \n> minor API detail even if it is not a trivial issue implementation wise.  \n> But the API must be right as this is what we'll be stuck with even if \n> the implementation may change.  And as far as an API definition is \n> needed I think that it should reflect the current plumbing which is \n> actually the real API that grew naturally and has been proven useful.\n\nWell what you said about reentrancy is that \"it's minor API detail but\neven minor API details must be right because we will be stuck with\nthem\". And I don't think it's minor at all either. :-)\n\nAlso, even if the implementation won't be completely re-entrant\ninitially, the question of re-entrancy is something we should decide\nsince it still affects the scope of the librarification work.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37374","messageId":"20070318140816.GG4489@pasky.or.cz","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703161509560.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-18T14:08:16Z","receivedAt":"2007-03-18T14:08:16Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Fri, Mar 16, 2007 at 04:12:17PM CET, Johannes Schindelin wrote:\n> Hi,\n> \n> [please do not cull the Cc: list]\n> \n> On Fri, 16 Mar 2007, Rocco Rutte wrote:\n> \n> > First, I think that would be some cleanup \"only\" since that basically would\n> > mean to\n> > \n> >   1) make all functions die()ing return some value and handle it and\n> >   2) wrap all static vars into structures and pass them around\n> > \n> > If you don't choose a design before wrapping things up in structures, you'll\n> > probably end up having one structure per source file (at least too many\n> > structures).\n> \n> Why? For some tasks, it should be 1) easier, 2) more elegant, and 3) \n> faster to write a function which re-initialises the static variables.\n> \n> Of course, if you want to work with multiple repos _at the same time_, \n> this does not help you. But frankly, we don't support that with core-git, \n> so why should we in libgit?\n\nBecause you don't know who will want to use libgit. Maybe perl bindings\nfrom inside of mod_perl, where single process can multiplex between many\nrepositories based on whichever request just arrived. You talked about\nmemory usage issues, but I think that's just a minor technical issue\nthat can be adjusted, while this is _conceptual_. Maybe someone will\nwant to write repodiff which looks at two repositories and compares them\n(without fetching massive data around). Maybe someone will want to write\nsome other cool hack we didn't think about.\n\nBecause in the other subthread you just suggested the git viewers should\nbe multi-threaded. Of course you can state that \"only a single thread\ncan use libgit at a time\", but then multithreading is just a hack to\nwork around libgit limitations (albeit still legitimate) while it could\nbe used to do so much more cool stuff like fetching old history\ninformation on background while you can already _work_ with the tool and\nlook at the new stuff details (isn't this actually exactly how gitk and\nqgit already work? they couldn't with non-reentrant libgit!).\n\nBecause if you look at the UNIX history, you'll notice that first people\nstarted with non-reentrant stuff because it was \"good enough\" and then\ncame back later and added reentrant versions anyway. Let's learn from\nhistory. It's question of probability but it's very likely this will\nhappen to us as well.\n\nThis is why the _API_ should be designed to be re-entrant. The\nimplementation may not be re-entrant right away, it may take a while to\nget there, but the API really should be.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37383","messageId":"20070318161854.5a6a34e0@home.brethil","threadId":"7265","inReplyTo":"7vzm6bp07f.fsf@assigned-by-dhcp.cox.net","subject":"Re: Libification project (SoC)","fromName":"Luiz Fernando N. Capitulino","fromEmail":"lcapitulino@mandriva.com.br","sentAt":"2007-03-18T16:18:54Z","receivedAt":"2007-03-18T16:18:54Z","isPatch":false,"sender":{"key":"lcapitulino@mandriva.com.br","avatar":null},"body":"On Sat, 17 Mar 2007 22:52:52 -0700\nJunio C Hamano <junkio@cox.net> wrote:\n\n| \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n| \n| > Take a look at sha1_file.c, open_packed_git_1:\n| >\n| > ...\n| >     if (!pack_version_ok(hdr.hdr_version))\n| >         return error(\"packfile %s is version %u and not supported\"\n| >             \" (try upgrading GIT to a newer version)\",\n| >             p->pack_name, ntohl(hdr.hdr_version));\n| > ...\n| >\n| > Here we are supplying a lot more than just a simple error code\n| > that can be mapped to a static string.\n| >\n| > Of course that code is currently feeding it to the error function,\n| > which today calls the error_routine (see usage.c).  We could buffer\n| > the strings sent to error()/warn() and let the caller obtain all\n| > strings that occurred during the last API call.\n| \n| Actually, since we are talking about the error path,\n| \n|  (1) we do not care performance of what happens there that much, but\n|  (2) we *do* care about not doing extra allocation.\n| \n| So it might make sense to have a preallocated \"error string\"\n| buffer, sprintf the error message in there and return error\n| codes.\n\n Other possibility is to let the caller do the job.\n\n I mean, if the information needed to print the error message (packfile\nname and version in this example) is available to the caller, or the\ncaller can get it someway, then the caller could check which error\nhe got and build the message himself.\n\n That seems simpler to me, considering the caller has the needed\ninfo, of course...\n\n\n-- \nLuiz Fernando N. Capitulino\n"},{"id":"37389","messageId":"7v4poimjr2.fsf@assigned-by-dhcp.cox.net","threadId":"7265","inReplyTo":"20070318161854.5a6a34e0@home.brethil","subject":"Re: Libification project (SoC)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-03-18T19:31:13Z","receivedAt":"2007-03-18T19:31:13Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br>\nwrites:\n\n>  I mean, if the information needed to print the error message (packfile\n> name and version in this example) is available to the caller, or the\n> caller can get it someway, then the caller could check which error\n> he got and build the message himself.\n>\n>  That seems simpler to me, considering the caller has the needed\n> info, of course...\n\nIt's a possibility, but that would make it much less nice to\ndiagnose and debug problems, as the caller does not usually have\nnecessary information.\n\nThe caller may ask for object A, and the error is triggered\nbecause a different object C is missing, which is the delta base\nof object B which in turn is the delta base of object A.  The\nbest your \"caller\" can say is \"cannot read object A for some\nreason\", and it cannot say \"cannot read object A because object\nC is missing\".\n"},{"id":"37400","messageId":"alpine.LFD.0.83.0703181709280.18328@xanadu.home","threadId":"7265","inReplyTo":"20070318161854.5a6a34e0@home.brethil","subject":"Re: Libification project (SoC)","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-03-18T21:15:30Z","receivedAt":"2007-03-18T21:15:30Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 18 Mar 2007, Luiz Fernando N. Capitulino wrote:\n\n>  Other possibility is to let the caller do the job.\n> \n>  I mean, if the information needed to print the error message (packfile\n> name and version in this example) is available to the caller, or the\n> caller can get it someway, then the caller could check which error\n> he got and build the message himself.\n\nNah...  The error details should be handled at the failure location.  \nAny error code based mechanism is bound to get out of synch at some \npoint, or people simply won't bother adding new codes for new error \nconditions but simply reuse an existing generic enough code instead.\n\nWe already have this nice error() function.  Right now it simply dumps \nthe message to stderr but it could be made more sophisticated if needed.\n\n\nNicolas\n"},{"id":"37422","messageId":"Pine.LNX.4.63.0703190045520.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"20070318140816.GG4489@pasky.or.cz","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-18T23:48:27Z","receivedAt":"2007-03-18T23:48:27Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 18 Mar 2007, Petr Baudis wrote:\n\n> [...] if you look at the UNIX history, you'll notice that first people \n> started with non-reentrant stuff because it was \"good enough\" and then \n> came back later and added reentrant versions anyway. Let's learn from \n> history. It's question of probability but it's very likely this will \n> happen to us as well.\n\nYes, let's learn from history. Start with a libgit that is good enough. \nAnd when somebody actually needs it to behave a little differently, or \nmore sophisticated, then let that somebody work on it!\n\nCiao,\nDscho\n"},{"id":"37433","messageId":"20070319012111.GS18276@pasky.or.cz","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703190045520.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-19T01:21:11Z","receivedAt":"2007-03-19T01:21:11Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hi,\n\nOn Mon, Mar 19, 2007 at 12:48:27AM CET, Johannes Schindelin wrote:\n> On Sun, 18 Mar 2007, Petr Baudis wrote:\n> \n> > [...] if you look at the UNIX history, you'll notice that first people \n> > started with non-reentrant stuff because it was \"good enough\" and then \n> > came back later and added reentrant versions anyway. Let's learn from \n> > history. It's question of probability but it's very likely this will \n> > happen to us as well.\n> \n> Yes, let's learn from history. Start with a libgit that is good enough. \n> And when somebody actually needs it to behave a little differently, or \n> more sophisticated, then let that somebody work on it!\n\n  I was talking about the API. The API has to be designed to be\nreentrant. And you get pretty much stuck with the API. And requiring\nreentrance isn't that far off once libgit is there, as I tried to point\nout; it's not really any obscure requirement.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37437","messageId":"Pine.LNX.4.63.0703190235330.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"20070319012111.GS18276@pasky.or.cz","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-19T01:43:54Z","receivedAt":"2007-03-19T01:43:54Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 19 Mar 2007, Petr Baudis wrote:\n\n> On Mon, Mar 19, 2007 at 12:48:27AM CET, Johannes Schindelin wrote:\n> > On Sun, 18 Mar 2007, Petr Baudis wrote:\n> > \n> > > [...] if you look at the UNIX history, you'll notice that first \n> > > people started with non-reentrant stuff because it was \"good enough\" \n> > > and then came back later and added reentrant versions anyway. Let's \n> > > learn from history. It's question of probability but it's very \n> > > likely this will happen to us as well.\n> > \n> > Yes, let's learn from history. Start with a libgit that is good \n> > enough. And when somebody actually needs it to behave a little \n> > differently, or more sophisticated, then let that somebody work on it!\n> \n>   I was talking about the API. The API has to be designed to be \n> reentrant. And you get pretty much stuck with the API. And requiring \n> reentrance isn't that far off once libgit is there, as I tried to point \n> out; it's not really any obscure requirement.\n\nI don't see _any_ problem in making an API which works with _one_ repo \nfirst. This has several advantages:\n\n- most users (if any!) will work that way,\n\n- it is easier to implement,\n\n- you are more likely to get that right than the more complex thing you \n  seem to want already in the first version, and\n\n- it is easy enough to extend the API later, _retaining_ the small and \n  beautiful functions.\n\nAs for the memory problems I was pointing out to you on IRC: if you do \nsome operation on one repo, and run out of memory, okay, there is not much \nyou can do about it. Tough luck.\n\nIf you cache different repos in the _same_ process, and run out of memory, \nyou should free the caches of the _other_ repos first, instead of just \nerroring out. This is not entirely trivial, likely to make libgit fragile, \nand quite possibly a performance hit (making libgit unattractive for \nplumbing, which would take away the best test case for libgit).\n\nAlso, when you cache different repos, you want to avoid duplicating \nidentical objects in different caches, which makes the cache handling no \neasier.\n\nBut even if these issues would not exist, isn't it obvious that you should \nstart with something _simple_?\n\nCiao,\nDscho\n"},{"id":"37446","messageId":"20070319025636.GE11371@thunk.org","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703190235330.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-03-19T02:56:36Z","receivedAt":"2007-03-19T02:56:36Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Mar 19, 2007 at 02:43:54AM +0100, Johannes Schindelin wrote:\n> >   I was talking about the API. The API has to be designed to be \n> > reentrant. And you get pretty much stuck with the API. And requiring \n> > reentrance isn't that far off once libgit is there, as I tried to point \n> > out; it's not really any obscure requirement.\n> \n> - it is easy enough to extend the API later, _retaining_ the small and \n>   beautiful functions.\n\nUm, look at what we had to do with gethostbyname() and\ngethostbyname_r().  It wasn't possible to sweep through and fix all of\nthe programs that used gethostbyname(), despite the fact that if a\nprogram called gethostbyname(), then called library function which\nunknowingly to application, could possibly do a DNS or YP lookup (and\nwhose behavior could change depending on some config file like\n/etc/nsswitch.conf), which would blow away the static information.  So\nif the application tryied to use the information returned by _its_\ncall to gethostbyname after calling some other library function, it\ncould get some completely random hostname that wasn't what it\nexpected.\n\nYelch!  And so we have two API's that libc has to support,\ngethostbyname(), and gethostbyname_r(), with the ugly _r() suffix, and\nwhich in a sane world most programs should use since otherwise they\ncan be incredibly fragile unless the _first_ thing they do after\ncalling gethostbyname is to copy the information to someplace stable,\ninstead of relying on the static buffer to remain sane.  (And yet they\ndon't, which means bugs that only show up if optional YP or Hesiod\nlookups are enabled, etc.)\n\nBerkely got it horribly wrong when it tried to start with the \"small\nand beautiful\" functions that were non-reentrant, and we've been\npaying the price ever since.  Do we really want to support two\nversions of the API forever?  Is it really that hard to support a\nreentrant API from the beginning?  I'd submit the answer to these two\nquestions are no, and no, respectively.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"37460","messageId":"20070319035526.GJ20658@spearce.org","threadId":"7265","inReplyTo":"20070319025636.GE11371@thunk.org","subject":"Re: Libification project (SoC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-03-19T03:55:26Z","receivedAt":"2007-03-19T03:55:26Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Theodore Tso <tytso@mit.edu> wrote:\n> Berkely got it horribly wrong when it tried to start with the \"small\n> and beautiful\" functions that were non-reentrant, and we've been\n> paying the price ever since.  Do we really want to support two\n> versions of the API forever?  Is it really that hard to support a\n> reentrant API from the beginning?  I'd submit the answer to these two\n> questions are no, and no, respectively.\n\nI agree entirely, for every reason mentioned by Ted (including\nthose not quoted).  ;-)\n\nI learned about gethostbyname after gethostbyname_r was already\nintroduced, so I have always been asking myself \"uhhhhh, why do we\nhave gethostbyname?\".  ;-)\n\n-- \nShawn.\n"},{"id":"37476","messageId":"e5bfff550703190001k761541c7v2c259ef3f7695b10@mail.gmail.com","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703190235330.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-03-19T07:01:35Z","receivedAt":"2007-03-19T07:01:35Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/19/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>\n> I don't see _any_ problem in making an API which works with _one_ repo\n> first. This has several advantages:\n>\n> - most users (if any!) will work that way,\n>\n\nSometime could be useful to write a list of possible users before\nstarting to code.\n\nPlease which are, in your opinion, the possible tools that could use a\nnon-reentrant, blocking libgit? In case tool is already exsistant\nplease write the name, in case it's a 'would be' one give a brief\ndescription.\n\nI' have tried to do the list myself, but I found only viewers ;-)\namong _currently_ tools I know of, and all the viewers allow loading\nin background _now_ so will not be portable to libgit without main\nsurgery, read multi-thread (BTW none is currently multi-thread).\n\n   Marco\n"},{"id":"37481","messageId":"1174297616.5884.42.camel@mejai","threadId":"7265","inReplyTo":"e5bfff550703190001k761541c7v2c259ef3f7695b10@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Steve Frécinaux","fromEmail":"nudrema@gmail.com","sentAt":"2007-03-19T09:46:56Z","receivedAt":"2007-03-19T09:46:56Z","isPatch":false,"sender":{"key":"nudrema@gmail.com","avatar":null},"body":"On Mon, 2007-03-19 at 08:01 +0100, Marco Costalba wrote:\n\n> I' have tried to do the list myself, but I found only viewers ;-)\n> among _currently_ tools I know of, and all the viewers allow loading\n> in background _now_ so will not be portable to libgit without main\n> surgery, read multi-thread (BTW none is currently multi-thread).\n\nI thought about configuration tools (gconf, kconfig, etc), that could\nthen implement something similar to what the recovery system of WinXP\ndoes: they could store an history of the configuration state, and then\nrecover a previous state if things go wrong. This would be incredibly\nuseful for system administrators.\n\nAlso, more generally, git can be used as a versioned storage system\nwithout direct link to source control. I'm thinking about ikiwiki for\ninstance.\n\nMore SCM-oriented, a cron script that manages a website by checkouting\nseveral repositories (one for the wiki module, another for the blog\nmodule, another for the forum, etc) using, say, the python bindi\n\nThere are probably a zillion other possible uses. The common thing when\nexposing an API is that it ends up being used in a way nobody had\nthought of. So it's dangerous to say \"it's useless\" or \"nobody will do\nit\". You can be sure someone will, it's just a matter of time.\n"},{"id":"37483","messageId":"1174300418.5884.50.camel@mejai","threadId":"7265","inReplyTo":"e5bfff550703190001k761541c7v2c259ef3f7695b10@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Steve Frécinaux","fromEmail":"nudrema@gmail.com","sentAt":"2007-03-19T10:33:38Z","receivedAt":"2007-03-19T10:33:38Z","isPatch":false,"sender":{"key":"nudrema@gmail.com","avatar":null},"body":"On Mon, 2007-03-19 at 08:01 +0100, Marco Costalba wrote:\n\n> Please which are, in your opinion, the possible tools that could use a\n> non-reentrant, blocking libgit? In case tool is already existent\n> please write the name, in case it's a 'would be' one give a brief\n> description.\n\nAnother idea that I just remembered about: two years ago there was a SoC\nproject to make nautilus (the file manager from gnome) able to version\ndirectories. It was using SVN (and failed, but it's another story).\n\nWhile nautilus is heavily multi-threaded, it's a \"single-instance app\",\nso there is at most only one instance of nautilus ever running. Under\nthe hypothesis of a \"versioned directories\" support using libgit (that\nwould be easier to do and support since it doesn't need to set up a\nserver), it's quite obvious that a non-reentrant git would not be\nenough: you are likely to have more than one versioned directories on\nscreen at the same time! OTOH, blocking doesn't look like an issue since\ngnomevfs already deals with quite a number of blocking synchronous libs\nand exposes an async API on top of those (similar to what QT Threading\ndoes, I guess).\n\nBTW, if some Gnome people are reading, if libgit comes into life, such a\nproject is something I'd like to see for real ;-)\n"},{"id":"37489","messageId":"Pine.LNX.4.63.0703191329220.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"e5bfff550703190001k761541c7v2c259ef3f7695b10@mail.gmail.com","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-19T12:37:18Z","receivedAt":"2007-03-19T12:37:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 19 Mar 2007, Marco Costalba wrote:\n\n> On 3/19/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > \n> > I don't see _any_ problem in making an API which works with _one_ repo\n> > first. This has several advantages:\n> > \n> > - most users (if any!) will work that way,\n> > \n> \n> Sometime could be useful to write a list of possible users before\n> starting to code.\n\nFair enough.\n\nI expect the most visible users of libgit to be: the core Git programs! \nBecause if we don't eat our own dog food, why should anybody else?\n\nAnd I am absolutely utterly opposed to make them slower just to support a \nprogram which wants to cache meta data from multiple repositories.\n\nYes, you could write a program which can compare objects from several \nrepos, but that is easy in fact: just set GIT_ALTERNATE_OBJECT_DIRECTORIES \nand you're done. Without changing the core of Git at all!\n\nHaving said that, I never liked the idea of having static variables to \ntalk with config handlers, and would have preferred cb_data like \nfor_each_ref() does. That is a low hanging fruit, which does not affect \nperformance, and is _definitely_ a clean up.\n\nI am not so sure about the impact of changing the index to a non-static \nstructure.\n\nCiao,\nDscho\n"},{"id":"37491","messageId":"20070319125233.GT18276@pasky.or.cz","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703191329220.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-03-19T12:52:33Z","receivedAt":"2007-03-19T12:52:33Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Mon, Mar 19, 2007 at 01:37:18PM CET, Johannes Schindelin wrote:\n> Yes, you could write a program which can compare objects from several \n> repos, but that is easy in fact: just set GIT_ALTERNATE_OBJECT_DIRECTORIES \n> and you're done. Without changing the core of Git at all!\n\nBut you'll also need to access refs.\n\nAnd the key point here is reentrance - handling multiple repositories at\nonce is only part of this, actually probably the much bigger customer\nwould be multi-threaded programs. And easier creation of reusable\ncomponents and other libraries, and so on...\n\nI believe the performance impact will be most likely absolutely\nnegligible. Of course we have no hard data, but I doubt it's this where\nmost of the CPU crunching is.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"37493","messageId":"e5bfff550703190604n6360659cl3880ec5b3a9b5042@mail.gmail.com","threadId":"7265","inReplyTo":"Pine.LNX.4.63.0703191329220.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Libification project (SoC)","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2007-03-19T13:04:47Z","receivedAt":"2007-03-19T13:04:47Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/19/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Mon, 19 Mar 2007, Marco Costalba wrote:\n>\n> > On 3/19/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > >\n> > > I don't see _any_ problem in making an API which works with _one_ repo\n> > > first. This has several advantages:\n> > >\n> > > - most users (if any!) will work that way,\n> > >\n> >\n> > Sometime could be useful to write a list of possible users before\n> > starting to code.\n>\n> Fair enough.\n>\n> I expect the most visible users of libgit to be: the core Git programs!\n> Because if we don't eat our own dog food, why should anybody else?\n>\n\nBut in case you eat your own food, why others should to the same?\n\n\n> And I am absolutely utterly opposed to make them slower just to support a\n> program which wants to cache meta data from multiple repositories.\n>\n\nThe problem, at least with viewers I know, it's not with multiple\nrepositories but with multiple  views of the same repo.\n\n\nAnyway. Just to give my two cent:\n\nThe two possible features we are talking about are:\n\n  - reentrancy (many views open on the same repo)\n\n  - non-blocking behaviour (loading repo in background)\n\nThese two features are _very_ different. I agree an async library it's\nnot a small thing, and probably it involves using an external thread\nlibrary in libgit itself, like pthread, just to not reinventing the\n(difficult) wheel.\n\nRegarding reentrancy I don't know what is involved in avoiding globals\nand the like, but I would think it's really an absolute minimum to get\npeople eating your food ;-)\n\nI completely agree that it's impossible to know how a library will be\nused when you write it, but giving a good look around before to start\nallows you to get a minimum subset of needed features and if you add a\nlittle bit of generalization and you are lucky enough perhaps you will\navoid to rewrite the library in the future.\n\n>From the viewers survey and also from the interesting examples of\nSteve I would say that do not planning for reentarncy would be a big\nno-no\n\n  Marco\n"},{"id":"37498","messageId":"Pine.LNX.4.63.0703191453040.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"20070319125233.GT18276@pasky.or.cz","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-19T13:55:21Z","receivedAt":"2007-03-19T13:55:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 19 Mar 2007, Petr Baudis wrote:\n\n> On Mon, Mar 19, 2007 at 01:37:18PM CET, Johannes Schindelin wrote:\n> > Yes, you could write a program which can compare objects from several \n> > repos, but that is easy in fact: just set GIT_ALTERNATE_OBJECT_DIRECTORIES \n> > and you're done. Without changing the core of Git at all!\n> \n> But you'll also need to access refs.\n\nYes, and you want it to bake some fine pizza, too.\n\n> And the key point here is reentrance - handling multiple repositories at \n> once is only part of this, actually probably the much bigger customer \n> would be multi-threaded programs. And easier creation of reusable \n> components and other libraries, and so on...\n> \n> I believe the performance impact will be most likely absolutely \n> negligible. Of course we have no hard data, but I doubt it's this where \n> most of the CPU crunching is.\n\nMy time is very limited, and I see this thread going nowhere since \neverybody says \"I like this, I like that\", and nobody shows some hard data \n(me included). It almost feels like a Windows user community. Or Slashdot.\n\nAnyway, I refuse to comment on these issues until somebody proves me wrong \nor right in my assumption that the impact on core Git (in terms of time \n_or_ lines of code) would be huge.\n\nCiao,\nDscho\n"},{"id":"37503","messageId":"Pine.LNX.4.63.0703191556210.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"7265","inReplyTo":"20070319025636.GE11371@thunk.org","subject":"Re: Libification project (SoC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-19T14:57:51Z","receivedAt":"2007-03-19T14:57:51Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 18 Mar 2007, Theodore Tso wrote:\n\n> On Mon, Mar 19, 2007 at 02:43:54AM +0100, Johannes Schindelin wrote:\n> > >   I was talking about the API. The API has to be designed to be \n> > > reentrant. And you get pretty much stuck with the API. And requiring \n> > > reentrance isn't that far off once libgit is there, as I tried to point \n> > > out; it's not really any obscure requirement.\n> > \n> > - it is easy enough to extend the API later, _retaining_ the small and \n> >   beautiful functions.\n> \n> Um, look at what we had to do with gethostbyname() and \n> gethostbyname_r().  It wasn't possible to sweep through and fix all of \n> the programs that used gethostbyname(), despite the fact that if a \n> program called gethostbyname(), then called library function which \n> unknowingly to application, could possibly do a DNS or YP lookup (and \n> whose behavior could change depending on some config file like \n> /etc/nsswitch.conf), which would blow away the static information.  So \n> if the application tryied to use the information returned by _its_ call \n> to gethostbyname after calling some other library function, it could get \n> some completely random hostname that wasn't what it expected.\n> \n> Yelch!  And so we have two API's that libc has to support, \n> gethostbyname(), and gethostbyname_r(), with the ugly _r() suffix, and \n> which in a sane world most programs should use since otherwise they can \n> be incredibly fragile unless the _first_ thing they do after calling \n> gethostbyname is to copy the information to someplace stable, instead of \n> relying on the static buffer to remain sane.  (And yet they don't, which \n> means bugs that only show up if optional YP or Hesiod lookups are \n> enabled, etc.)\n> \n> Berkely got it horribly wrong when it tried to start with the \"small and \n> beautiful\" functions that were non-reentrant, and we've been paying the \n> price ever since.  Do we really want to support two versions of the API \n> forever?  Is it really that hard to support a reentrant API from the \n> beginning?  I'd submit the answer to these two questions are no, and no, \n> respectively.\n\nYou make a good case why gethostbyname() was wrong, and should have been \ndefined as gethostbyname_r() to begin with.\n\nHowever, as I wrote in another reply in this thread, I am not prepared to \nsink more time in this discussion, _unless_ somebody who cares about it \nenough shows me some code and/or numbers.\n\nCiao,\nDscho\n"},{"id":"37513","messageId":"20070319130907.23b73273@localhost","threadId":"7265","inReplyTo":"7v4poimjr2.fsf@assigned-by-dhcp.cox.net","subject":"Re: Libification project (SoC)","fromName":"Luiz Fernando N. Capitulino","fromEmail":"lcapitulino@mandriva.com.br","sentAt":"2007-03-19T16:09:07Z","receivedAt":"2007-03-19T16:09:07Z","isPatch":false,"sender":{"key":"lcapitulino@mandriva.com.br","avatar":null},"body":"Em Sun, 18 Mar 2007 12:31:13 -0700\nJunio C Hamano <junkio@cox.net> escreveu:\n\n| \"Luiz Fernando N. Capitulino\" <lcapitulino@mandriva.com.br>\n| writes:\n| \n| >  I mean, if the information needed to print the error message (packfile\n| > name and version in this example) is available to the caller, or the\n| > caller can get it someway, then the caller could check which error\n| > he got and build the message himself.\n| >\n| >  That seems simpler to me, considering the caller has the needed\n| > info, of course...\n| \n| It's a possibility, but that would make it much less nice to\n| diagnose and debug problems, as the caller does not usually have\n| necessary information.\n| \n| The caller may ask for object A, and the error is triggered\n| because a different object C is missing, which is the delta base\n| of object B which in turn is the delta base of object A.  The\n| best your \"caller\" can say is \"cannot read object A for some\n| reason\", and it cannot say \"cannot read object A because object\n| C is missing\".\n\n Okay, you're right. I'm going to let the low-level functions\nfill the error buffer then.\n\n Thanks,\n\n-- \nLuiz Fernando N. Capitulino\n"},{"id":"37516","messageId":"Pine.LNX.4.64.0703190912190.6730@woody.linux-foundation.org","threadId":"7265","inReplyTo":"20070319025636.GE11371@thunk.org","subject":"Re: Libification project (SoC)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-19T16:28:00Z","receivedAt":"2007-03-19T16:28:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 18 Mar 2007, Theodore Tso wrote:\n> \n> Berkely got it horribly wrong when it tried to start with the \"small\n> and beautiful\" functions that were non-reentrant, and we've been\n> paying the price ever since.\n\nI don't think that's a good argument, ESPECIALLY when coming from somebody \nfrom MIT.\n\nBerkeley may have gotten it \"horribly wrong\", but the fact is, BSD kicked \nass and took over the world, in a way that nothing comparable I know of \nfrom MIT ever did. Exactly *because* the BSD people didn't try to make it \nperfect, but made things \"small and easy to *implement*\".\n\n(I would not say \"small and beautiful\". \"Beauty\" had nothing to do with \nit. \"simple\" had. And unlike beauty, simplicity really *is* more than skin \ndeep, and is a fundamentally good design).\n\nI'm a *huge* believer in \"Worse is Better\" (for people who don't know it, \njust google for that phrase, with the quotes around it).\n\nIn fact, I'd argue that the reason git kicks ass is exactly that \"Worse is \nBetter\" design: you need to have a few conceptual (good) ideas to base \nyour design off on, but given those good ideas, it's more important that \nthings _work_well_in_practice_ than some \"wouldn't it be better..\" kind of \nmentality.\n\nThe \"paying the price ever since\" argument is bogus. If you get to that \npoint, you've by definition *already*won*! \n\nHere's the real world according to Linus:\n 1) everybody makes mistakes\n 2) only the winners \"pay the price\" of those mistakes ever since, since \n    the losers will not be around to pay it, and the winners will have \n    made mistakes too (see #1)\n 3) the more complex and subtle you make the interfaces, the more mistakes \n    you'll make, AND the less likely you are to be a winner anyway, since \n    you'll have problems implementing it *and* it will probably be subtle \n    to use too!\n\nSo the motto should always be: \"Just Do It!\", and screw worrying about \npaying the price. You *want* to have to pay the price. It's the best thing \nthat can ever happen to you. And you want to have to start paying the \nprice as early as possible - because that not only means that you won, it \nalso means that you'll now be learning from your mistakes instead of \ntrying to anticipate them, and I will *guarantee* that learnign from \nmistakes is going to be a lot more productive than trying to worry about \nthem up-front.\n\n> Do we really want to support two versions of the API forever?\n\nI'd personally strongly vote for a \"simple library\" interface as a first \ncut.\n\nAnd yes, if that means supporting two versions, I think it's better. You \ncan easily have \"libgit-simple.a\" for trivial non-threaded accesses with \nout-of-memory conditions causing the process to die. That really *is* a \nvery useful schenario, as shown by the fact that *every*single*core*git \nprogram has been happy with it.\n\nClaiming that you need a complicated interface in the face of the *proof* \nthat git itself dosn't need that complicated an interface is to me a bit \ndisingenious.\n\nYes, *some* people will want a thread-safe one. But we're not talking \nsomething like libc here, where the library is so fundamental that it \nneeds to be acceptable for everybody. It's perfectly possible to have a \n\"libgit-simple.a\" that is good for 99% of all uses, and that is simple to \nuse, and less bug-prone simply because is is *simpler* (not just for \nusers, but as an implementation).\n\nAnd then for the small small minority of programs that want something \nfancier, do a \"libgit-complicated.a\" library. IF you ever get it working \nand complete, you can always then implement \"libgit-simple\" in terms of \nthe complicated version.\n\n\n\n  Is it \n> really that hard to support a reentrant API from the beginning?  I'd \n> submit the answer to these two questions are no, and no, respectively.\n> \n> \t\t\t\t\t\t- Ted\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n"},{"id":"37517","messageId":"Pine.LNX.4.64.0703190930060.6730@woody.linux-foundation.org","threadId":"7265","inReplyTo":"Pine.LNX.4.64.0703190912190.6730@woody.linux-foundation.org","subject":"Re: Libification project (SoC)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-19T16:32:13Z","receivedAt":"2007-03-19T16:32:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOops. My fingers are faster than my brain, and that email got sent out \nhalf-completed and without the final editing. But it wasn't reallymissing \nanything else than editing away the parts of the original I didn't respond \nto, and my normal sign-off.\n\nSo I'll just sign this one off twice, instead..\n\n\t\tLinus\n\n\t\tLinus\n"},{"id":"37664","messageId":"46011450.4000200@op5.se","threadId":"7265","inReplyTo":"Pine.LNX.4.64.0703190912190.6730@woody.linux-foundation.org","subject":"Re: Libification project (SoC)","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-03-21T11:17:36Z","receivedAt":"2007-03-21T11:17:36Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> \n> I'm a *huge* believer in \"Worse is Better\" (for people who don't know it, \n> just google for that phrase, with the quotes around it).\n> \n\nI just did, and having read the first page of the document found at \nhttp://www.jwz.org/doc/worse-is-better.html, I must say \"worse-is-better\"\nsounds an awful lot like evolution; \"Start with something that works. When\nsomething else works better, jump train and embrace The New Thing\".\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"37686","messageId":"Pine.LNX.4.64.0703211020070.6730@woody.linux-foundation.org","threadId":"7265","inReplyTo":"46011450.4000200@op5.se","subject":"Re: Libification project (SoC)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-03-21T17:24:07Z","receivedAt":"2007-03-21T17:24:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Mar 2007, Andreas Ericsson wrote:\n\n> Linus Torvalds wrote:\n> > \n> > I'm a *huge* believer in \"Worse is Better\" (for people who don't know it, \n> > just google for that phrase, with the quotes around it).\n> \n> I just did, and having read the first page of the document found at \n> http://www.jwz.org/doc/worse-is-better.html, I must say \"worse-is-better\"\n> sounds an awful lot like evolution; \"Start with something that works. When\n> something else works better, jump train and embrace The New Thing\".\n\nYeah. I'm a huge believer in evolution too (and not just the biological \nkind ;)\n\nThe thing is, most \"designers\" are just totally clueless. Even the \nsmartest people that have done something similar five times before are \nprone to totally mis-design something if they start from scratch and try \nto \"think it through\". You tend to concentrate on the problems of the \nprevious generation, and not even think about everything that worked \nwonderfully well, because that wasn't something you *needed* to think \nabout.\n\nSo \"designing\" stuff is way overrated. You can spend years designing \nsomethign that is total crap, just because you didn't actually try it out \nand _realize_ that it wasn't what the user wanted (it may have been what \nthe user _thought_ and _claimed_ that he wanted, but that was before \nactually tried to use it, and realized that he was wrong).\n\n\t\t\tLinus\n"},{"id":"37732","messageId":"4602519B.7000303@op5.se","threadId":"7265","inReplyTo":"Pine.LNX.4.64.0703211020070.6730@woody.linux-foundation.org","subject":"Re: Libification project (SoC)","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-03-22T09:51:23Z","receivedAt":"2007-03-22T09:51:23Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Linus Torvalds wrote:\n> \n> On Wed, 21 Mar 2007, Andreas Ericsson wrote:\n> \n>> Linus Torvalds wrote:\n>>> I'm a *huge* believer in \"Worse is Better\" (for people who don't know it, \n>>> just google for that phrase, with the quotes around it).\n>> I just did, and having read the first page of the document found at \n>> http://www.jwz.org/doc/worse-is-better.html, I must say \"worse-is-better\"\n>> sounds an awful lot like evolution; \"Start with something that works. When\n>> something else works better, jump train and embrace The New Thing\".\n> \n> Yeah. I'm a huge believer in evolution too (and not just the biological \n> kind ;)\n> \n> So \"designing\" stuff is way overrated. You can spend years designing \n> somethign that is total crap, just because you didn't actually try it out \n> and _realize_ that it wasn't what the user wanted (it may have been what \n> the user _thought_ and _claimed_ that he wanted, but that was before \n> actually tried to use it, and realized that he was wrong).\n> \n\nIndeed. That's probably why Extreme Programming (silly hype-name, but what\nto call it otherwise?) has gained so much popularity from the people that\nreally understand the concept.\n\nTo those that don't wish to google for it, Extreme Programming is about\ntaking small steps that lead to a diffuse goal (\"We shall make a fantasy\nvideo game that millions of people would like to play. Significant lore\nis here, here and here\"). \n\nThe goal and any of the steps might change along the way. Basically, it\nputs \"re-think, re-design, re-factor\" on the table for corporate software\nproduction and promotes rapid implementation over correctness.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"}]}