{"thread":{"id":"16488","subject":"[RFC] Git Perl bindings, and OO interface","startedAt":"2008-11-27T01:58:49Z","lastAt":"2009-07-10T02:08:04Z","messageCount":4,"participants":["Jakub Narebski","nadim khemir","Tom Lanyon"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"96598","messageId":"200811270258.50898.jnareb@gmail.com","threadId":"16488","inReplyTo":null,"subject":"[RFC] Git Perl bindings, and OO interface","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-11-27T01:58:49Z","receivedAt":"2008-11-27T01:58:49Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"There exists many Git bindings for various programming languages, some \nof them using git commands, some of them reimplementing Git, or parts \nof Git.  There is GitPython and PyGit (with some native implementation)\nfor Python, there is (deprecated) Ruby/Git and Grit (with some native\nimplementation) for Ruby, there is #Git for C#, there is ObjectiveGit\nfor Objective C (native), there is JGit (native) and JavaGit for Java,\nthere is Gat for Haskell and eWiki contains something for PHP.\n\nAnd of course there is Git.pm (included with Git) and Git::Repo (part\nof 'gitweb caching' GSoC 2008 project by Lea Wiemann) for Perl.  Now\nGit::Repo didn't get accepted into git.git codebase, but developing it\nsparked a bit of discussion about Perl interface to Git commands, and\nObject Oriented interface to Git.\n\nI'd like to spawn a discussion in this thread about interfaces to Git\nand Object Oriented interface to Git, mainly but not only in Perl.\nI hope that the authors of mentioned (and not mentioned) bindings, \ninterfaces and implementations of Git would contribute to this thread.\n\n\n0. One of points of disagreement between Git.pm and new Git::Repo was\n   using Error module for frontend error handling.  While the \n   explanation in http://www.perl.com/pub/a/2002/11/14/exception.html\n   is compelling, it is not standard Perl technique.  Additionally\n   adding \"cmd_git_try { CODE } ERRORMSG\" syntactic sugar was not very\n   good idea.\n\n   So the first thing I'd like to discuss: to use Error and try/catch,\n   or not in Perl interface (bindings) to Git?  I would really like to\n   hear from Perl experts / Perl hackers here...\n\n1. Git::Cmd\n\n   If I remember correctly Git.pm started as a way to gather in one\n   place safe_cmd and safe_pipe like construct from various git commands\n   implemented in Perl.  The goal here is to provide portable, safe, and\n   working with old Perl interface:\n    * portable: this means trying to work with ActiveState Perl on \n      MS Windows; I don't know how important it is _now_ (if there are\n      common other Perl distributions on MS Windows).\n    * safe: if some of arguments to git commands come from variables,\n      then they have to be safe against shell expansion (whitespace,\n      quoting characters, escape characters, metacharacters, etc.).\n    * compatibile: it should work with as old Perl version as is\n      reasonable; it is possible that you can install git locally, but\n      cannot upgrade Perl.\n\n   Note that some git commands, for example 'git version', 'git\n   ls-remote' and 'git clone' doesn't need git repository to work on.\n\n   We would want to be able to catch git command output to scalar, to\n   list (line by line), and to filehandle. More advanced stuff is bidi\n   pipe (watch for deadlocks!), and redirecting both stdout and stderr\n   of git command to filehandle.\n\n   What instance of Git::Cmd should know is where to find 'git' binary\n   (what is $GIT in gitweb, for example). It could cache/store\n   internally exec_path.\n\n2. Git::Config\n\n   If git command (a piece of code) uses more than one configuration\n   variable, then one would want to get relevant configuration using\n   as few calls to git commands as possible.  Therefore using git-config\n   to read each config variable is usually out of the question (but it\n   is sometimes useful); we would want to read all config in one go,\n   either by using \"git config -l -z\", or by writing config parser in\n   Perl (as some command(s) did).\n\n   The problem with this solution is that we have to implement \"type\n   casting\", i.e. equivalent of --int and --bool options to git-config\n   ourselves. This mean converting to integer with optional size suffix,\n   converting to boolean, and asking for escape codes corresponding to\n   given color. And if we add new type (like proposed --path, expanding\n   for example '~' to HOME, and ~user to home directory of given user)\n   we would have to add it to Perl interface too.\n  \n3. Git::Repo::Bare and Git::Repo::Nonbare\n\n   Git.pm partially implements those, in a kind of mixed way. Git::Repo\n   from Lea Wiemann implements if I remember correctly bare repo only.\n\n   What Git::Repo::Bare (or just Git::Repo) should support is to pass\n   appropriate '--git-dir=<dir>' to Git::Cmd, and support accessing git\n   repository config via Git::Config.  It could have also use\n   long-running pipe to \"git cat-file --batch / --batch-check\"\n   invocation.  For gitweb we only need that part.\n\n   Git::Repo::Nonbare has to additionally pass '--work-tree=<dir>' if\n   needed, ant be able to take care and manipulate where in working\n   directory we are, i.e. what for example \"git rev-parse --show-prefix\"\n   does.\n\n4. Git::Object: Git::Commit, Git::Tag, Git::Blob and Git::Tree\n\n   Here begins \"true\" object-oriented part of Git Perl API.\n\n   The easy part is for Git::Commit and Git::Tag to parse commit and tag\n   objects (perhaps Git::Object should have interface for long-lived\n   \"git cat-file --batch\") into headers and body (commit/tag message).\n   I think we can borrow / be inspired by parse_commit() and other such\n   code in gitweb; we have to remember that there might be in some time\n   some new headers we don't know about but are perfectly valid (see for\n   example \"encoding\" header in commit object format, which was added\n   later, not during initial design).\n\n   The harder part would be to be able to deal with author and committer\n   info, splitting it into parts (author name, author email, date and\n   timezone, etc.), and also generating dates in various formats, like\n   RFC-2822 or ISO-8601.\n\n   The easiest part would be structureless Git::Blob... but there we\n   might want size of blob.\n\n   A bit harder would be Git::Tree object and dealing with elements of\n   a tree (tree entries).  I'm not sure if some kind of iterator access\n   would be useful here.\n\n   Note that for Git::Commit if we are to use plumbing like git-cat-file\n   we would have to take care of fake parents info, namely grafts and\n   shallow info by ourself, in Perl, to have 'effective parents'.\n\n5. Git::Diff::Raw and Git::Diff::Patchset\n\n   Here I am thinking simply about parsing difftree (raw diff output\n   format) and patchset format, as it is used in gitweb.  It is meant\n   to be able to access for example to permissions of a file, or diff\n   status, or diff stats, etc.\n\n   Here we would want to be able to deal also with merge commits and\n   combined diff output format.\n\n6. Git::Log or Git::RevList\n\n   The only difference from list of Git::Commit objects is that \n   depending on parameters like path limiting it might have different\n   effective parents if there is history simplification.\n\n7. Git::Refs\n\n   It is meant to represent references, mainly branches, and be filled\n   using git-for-each-ref... and for example used for ref markers.\n\nThere are probably a few things I have forgot about...\n-- \nJakub Narebski\nPoland\n"},{"id":"96782","messageId":"200811301445.18969.nadim@khemir.net","threadId":"16488","inReplyTo":"200811270258.50898.jnareb@gmail.com","subject":"Re: [RFC] Git Perl bindings, and OO interface","fromName":"nadim khemir","fromEmail":"nadim@khemir.net","sentAt":"2008-11-30T13:45:18Z","receivedAt":"2008-11-30T13:45:18Z","isPatch":false,"sender":{"key":"nadim@khemir.net","avatar":null},"body":"On Thursday 27 November 2008 02.58.49 Jakub Narebski wrote:\n> ...\n>\n> 7. Git::Refs\n>\n>    It is meant to represent references, mainly branches, and be filled\n>    using git-for-each-ref... and for example used for ref markers.\n>\n> There are probably a few things I have forgot about...\n\nThank you for writing the RFC, it's a very good start. I would like to see \nsome strategy for libgit[2] in the RFC. What is your opinion about that?\n\nNadim.\n"},{"id":"96784","messageId":"m37i6ljmjy.fsf@localhost.localdomain","threadId":"16488","inReplyTo":"200811301445.18969.nadim@khemir.net","subject":"Re: [RFC] Git Perl bindings, and OO interface","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-11-30T14:50:38Z","receivedAt":"2008-11-30T14:50:38Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"nadim khemir <nadim@khemir.net> writes:\n> On Thursday 27 November 2008 02.58.49 Jakub Narebski wrote:\n> > ...\n> >\n> > 7. Git::Refs\n> >\n> >    It is meant to represent references, mainly branches, and be filled\n> >    using git-for-each-ref... and for example used for ref markers.\n> >\n> > There are probably a few things I have forgot about...\n> \n> Thank you for writing the RFC, it's a very good start. I would like to see \n> some strategy for libgit[2] in the RFC. What is your opinion about that?\n\nI do not know enought about libgit2 or even git unofficial internal C\nAPI to talk about it.\n\nI did not plan for Perl interface to be actual Perl bindings, using\nlibgit2.  Please remember that earlier effort of using XS (Perl <-> C\ninterface) failed because it relied on GCC support for -fPIC and was\nnot sufficiently portable... if I remember it correctly.  Calling Git\ncommands and massaging output would be enough for me.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"117744","messageId":"E6AB02D1-CF72-4611-91B0-DA524081A2EE@netspot.com.au","threadId":"16488","inReplyTo":"200811270258.50898.jnareb@gmail.com","subject":"Re: [RFC] Git Perl bindings, and OO interface","fromName":"Tom Lanyon","fromEmail":"tom@netspot.com.au","sentAt":"2009-07-10T02:08:04Z","receivedAt":"2009-07-10T02:08:04Z","isPatch":false,"sender":{"key":"tom@netspot.com.au","avatar":null},"body":"On 27/11/2008, at 12:28 PM, Jakub Narebski wrote:\n> 0. One of points of disagreement between Git.pm and new Git::Repo was\n>   using Error module for frontend error handling.  While the\n>   explanation in http://www.perl.com/pub/a/2002/11/14/exception.html\n>   is compelling, it is not standard Perl technique.  Additionally\n>   adding \"cmd_git_try { CODE } ERRORMSG\" syntactic sugar was not very\n>   good idea.\n>\n>   So the first thing I'd like to discuss: to use Error and try/catch,\n>   or not in Perl interface (bindings) to Git?  I would really like to\n>   hear from Perl experts / Perl hackers here...\n\n\nSorry to bring up an old thread - but there was no further discussion  \non this and I've recently run into some grief with Git.pm.\n\nI'm new to Git, but not new to Perl and recently attempted to perform  \nsome simple operations over Git repositories from a Perl application  \n(it needs to clone, push, checkout, merge and that's about it) and  \nfound the Error.pm style handling of errors unintuitive and annoying.  \nIt is currently fairly simple to capture errors into the application  \nby wrapping git_cmd_try { CODE } ERROR into an eval {} block but this  \nreally only provides you with the command's exit status and no  \nmeaningful error messages to display to your users; not to mention  \nit's fairly ugly.\n\nA long standing Perl motto is 'There Is More Than One Way To Do It'  \nand the use of Error.pm here forces developers down a specific path  \nfor error handling - some may like this, some may not, but there's not  \na lot they can do about it. I would suggest that the Perl way for  \nGit.pm to handle errors is for its methods to return the standard 1 or  \n0 for success or failure and perhaps store some meaningful error  \nmessages in an accessor or variable. The module should also not die()  \nif there's an error - leave this up to the users of the module to  \nhandle errors how they prefer - if it dies, we must wrap the methods  \nin eval{} blocks or handle with $SIG{__DIE__}, making for some messy  \nand ugly code.\n\nI would love to be able to:\n\n\tmy $repo = Git->repository( directory => '/some/repo' )\n\t\tor die \"Unable to load git repo /some/repo: $Git::errstr\";\n\n\t$repo->command( 'push', [ 'some-remote' ] )\n\t\tor die \"Unable to push to origin: $Git::errstr\";\n\n... or similar, and have $Git::errstr set to something meaningful like  \nthe \"fatal: 'some-remote': unable to chdir or not a git archive\"  \nreturned by git-push. This also leads into some discussion around git  \ncommands printing to STDERR when there is no error -- example: if  \neverything is fine and up to date, I don't need git-push to tell me  \n\"Everything up-to-date\" in STDERR...\n\nHope this helps.\n\nRegards,\nTom\n"}]}