{"thread":{"id":"49","subject":"using git directory cache code in darcs?","startedAt":"2005-04-16T13:22:36Z","lastAt":"2005-04-18T09:23:03Z","messageCount":14,"participants":["David Roundy","Ingo Molnar","Junio C Hamano","Linus Torvalds","Mike Taht","Nomad Arton","Paul Jackson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"300","messageId":"20050416132231.GJ2551@abridgegame.org","threadId":"49","inReplyTo":null,"subject":"using git directory cache code in darcs?","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-16T13:22:36Z","receivedAt":"2005-04-16T13:22:36Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"Hello Linus and various git developers (ccing darcs developers),\n\nI've been thinking about the possibility of using the git \"current\ndirectory cache\" code in darcs.  Darcs already has an abstraction layer\nover its pristine directory cache, so this shouldn't be too hard--provided\nthe git code is understandable.  The default in darcs is currently to use\nan actual directory (\"_darcs/current/\") as the cache, and we synchronize\nthe file modification times in the cache with those of identical files in\nthe working directory to speed up compares.  We (the darcs developers) have\ntalked for some time about introducing a single-file directory cache, but\nnoone ever got around to it, partly because there wasn't a particularly\ncompelling need.\n\nIt seems that the git directory cache is precisely what we want.  Also, if\nwe switch to (optionally) using the git directory cache, I imagine it'll\nmake interfacing with git a lot easier.  And, of course, it would\nsignificantly speed up a number of darcs commands, which are limited by the\nslowness of the readdir-related code.  We haven't tracked this down why\nthis is, but a recursive directory compare in which we readdir only one of\nthe directories (since we don't care about new files in the other one)\ntakes half the time of a compare in which we readdir both directories.\n\nSo my questions are:\n\n1) Would this actually be a good idea? It seems good to me, but there may\nbe other considerations that I haven't thought of.\n\n2) Will a license be chosen soon for git? Or has one been chosen, and I\nmissed it? I can't really include git code in darcs without a license.  I'd\nprefer GPLv2 or later (since that's how darcs is licensed), but as long as\nit's at least compabible with GPLv2, I'll be all right.\n\n3) Is it likely that git will switch to not using global variables for\nactive_cache, active_nr and active_alloc? I'd be more comfortable passing\nthese things around, since it would make the haskell interface easier and\nsafer.  e.g. I'd like\n\nstruct git_cache {\n        struct cache_entry **cache;\n        unsigned int nr, alloc;\n};\n\ngit_cache *read_cache(char *path_to_index);\n\nor alternatively\n\nint read_cache(char *path_to_index, git_cache *);\n\nWould anyone on the git side be interested in making such changes? If not,\nwould they be likely to be accepted if a darcs person submitted patches?\n\n4) Would there be interest in creating a libgit? I've been imagining taking\ngit source files and including them directly in darcs' code, but in the\nlong run it would be easier if there were a standard git API we could use.\n\nI guess that's about all.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"302","messageId":"20050416133140.GA21552@elte.hu","threadId":"49","inReplyTo":"20050416132231.GJ2551@abridgegame.org","subject":"Re: using git directory cache code in darcs?","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-16T13:31:40Z","receivedAt":"2005-04-16T13:31:40Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* David Roundy <droundy@abridgegame.org> wrote:\n\n> 2) Will a license be chosen soon for git? Or has one been chosen, and \n> I missed it? I can't really include git code in darcs without a \n> license.  I'd prefer GPLv2 or later (since that's how darcs is \n> licensed), but as long as it's at least compabible with GPLv2, I'll be \n> all right.\n\nthere's a license in the latest code, it's GPLv2. Here's the COPYING \nfile:\n\n---------------\n Note that the only valid version of the GPL as far as this project\n is concerned is _this_ particular version of the license (ie v2, not\n v2.2 or v3.x or whatever), unless explicitly otherwise stated.\n\n HOWEVER, in order to allow a migration to GPLv3 if that seems like\n a good idea, I also ask that people involved with the project make\n their preferences known. In particular, if you trust me to make that\n decision, you might note so in your copyright message, ie something\n like\n\n\tThis file is licensed under the GPL v2, or a later version\n\tat the discretion of Linus.\n\n  might avoid issues. But we can also just decide to synchronize and\n  contact all copyright holders on record if/when the occasion arises.\n\n\t\t\tLinus Torvalds\n\n----------------------------------------\n\n\t\t    GNU GENERAL PUBLIC LICENSE\n\t\t       Version 2, June 1991\n\n Copyright (C) 1989, 1991 Free Software Foundation, Inc.\n                       59 Temple Place, Suite 330, Boston, MA  02111-1307  USA\n Everyone is permitted to copy and distribute verbatim copies\n of this license document, but changing it is not allowed.\n\n\t\t\t    Preamble\n\n  The licenses for most software are designed to take away your\nfreedom to share and change it.  By contrast, the GNU General Public\nLicense is intended to guarantee your freedom to share and change free\nsoftware--to make sure the software is free for all its users.  This\nGeneral Public License applies to most of the Free Software\nFoundation's software and to any other program whose authors commit to\nusing it.  (Some other Free Software Foundation software is covered by\nthe GNU Library General Public License instead.)  You can apply it to\nyour programs, too.\n\n  When we speak of free software, we are referring to freedom, not\nprice.  Our General Public Licenses are designed to make sure that you\nhave the freedom to distribute copies of free software (and charge for\nthis service if you wish), that you receive source code or can get it\nif you want it, that you can change the software or use pieces of it\nin new free programs; and that you know you can do these things.\n\n  To protect your rights, we need to make restrictions that forbid\nanyone to deny you these rights or to ask you to surrender the rights.\nThese restrictions translate to certain responsibilities for you if you\ndistribute copies of the software, or if you modify it.\n\n  For example, if you distribute copies of such a program, whether\ngratis or for a fee, you must give the recipients all the rights that\nyou have.  You must make sure that they, too, receive or can get the\nsource code.  And you must show them these terms so they know their\nrights.\n\n  We protect your rights with two steps: (1) copyright the software, and\n(2) offer you this license which gives you legal permission to copy,\ndistribute and/or modify the software.\n\n  Also, for each author's protection and ours, we want to make certain\nthat everyone understands that there is no warranty for this free\nsoftware.  If the software is modified by someone else and passed on, we\nwant its recipients to know that what they have is not the original, so\nthat any problems introduced by others will not reflect on the original\nauthors' reputations.\n\n  Finally, any free program is threatened constantly by software\npatents.  We wish to avoid the danger that redistributors of a free\nprogram will individually obtain patent licenses, in effect making the\nprogram proprietary.  To prevent this, we have made it clear that any\npatent must be licensed for everyone's free use or not licensed at all.\n\n  The precise terms and conditions for copying, distribution and\nmodification follow.\n\f\n\t\t    GNU GENERAL PUBLIC LICENSE\n   TERMS AND CONDITIONS FOR COPYING, DISTRIBUTION AND MODIFICATION\n\n  0. This License applies to any program or other work which contains\na notice placed by the copyright holder saying it may be distributed\nunder the terms of this General Public License.  The \"Program\", below,\nrefers to any such program or work, and a \"work based on the Program\"\nmeans either the Program or any derivative work under copyright law:\nthat is to say, a work containing the Program or a portion of it,\neither verbatim or with modifications and/or translated into another\nlanguage.  (Hereinafter, translation is included without limitation in\nthe term \"modification\".)  Each licensee is addressed as \"you\".\n\nActivities other than copying, distribution and modification are not\ncovered by this License; they are outside its scope.  The act of\nrunning the Program is not restricted, and the output from the Program\nis covered only if its contents constitute a work based on the\nProgram (independent of having been made by running the Program).\nWhether that is true depends on what the Program does.\n\n  1. You may copy and distribute verbatim copies of the Program's\nsource code as you receive it, in any medium, provided that you\nconspicuously and appropriately publish on each copy an appropriate\ncopyright notice and disclaimer of warranty; keep intact all the\nnotices that refer to this License and to the absence of any warranty;\nand give any other recipients of the Program a copy of this License\nalong with the Program.\n\nYou may charge a fee for the physical act of transferring a copy, and\nyou may at your option offer warranty protection in exchange for a fee.\n\n  2. You may modify your copy or copies of the Program or any portion\nof it, thus forming a work based on the Program, and copy and\ndistribute such modifications or work under the terms of Section 1\nabove, provided that you also meet all of these conditions:\n\n    a) You must cause the modified files to carry prominent notices\n    stating that you changed the files and the date of any change.\n\n    b) You must cause any work that you distribute or publish, that in\n    whole or in part contains or is derived from the Program or any\n    part thereof, to be licensed as a whole at no charge to all third\n    parties under the terms of this License.\n\n    c) If the modified program normally reads commands interactively\n    when run, you must cause it, when started running for such\n    interactive use in the most ordinary way, to print or display an\n    announcement including an appropriate copyright notice and a\n    notice that there is no warranty (or else, saying that you provide\n    a warranty) and that users may redistribute the program under\n    these conditions, and telling the user how to view a copy of this\n    License.  (Exception: if the Program itself is interactive but\n    does not normally print such an announcement, your work based on\n    the Program is not required to print an announcement.)\n\f\nThese requirements apply to the modified work as a whole.  If\nidentifiable sections of that work are not derived from the Program,\nand can be reasonably considered independent and separate works in\nthemselves, then this License, and its terms, do not apply to those\nsections when you distribute them as separate works.  But when you\ndistribute the same sections as part of a whole which is a work based\non the Program, the distribution of the whole must be on the terms of\nthis License, whose permissions for other licensees extend to the\nentire whole, and thus to each and every part regardless of who wrote it.\n\nThus, it is not the intent of this section to claim rights or contest\nyour rights to work written entirely by you; rather, the intent is to\nexercise the right to control the distribution of derivative or\ncollective works based on the Program.\n\nIn addition, mere aggregation of another work not based on the Program\nwith the Program (or with a work based on the Program) on a volume of\na storage or distribution medium does not bring the other work under\nthe scope of this License.\n\n  3. You may copy and distribute the Program (or a work based on it,\nunder Section 2) in object code or executable form under the terms of\nSections 1 and 2 above provided that you also do one of the following:\n\n    a) Accompany it with the complete corresponding machine-readable\n    source code, which must be distributed under the terms of Sections\n    1 and 2 above on a medium customarily used for software interchange; or,\n\n    b) Accompany it with a written offer, valid for at least three\n    years, to give any third party, for a charge no more than your\n    cost of physically performing source distribution, a complete\n    machine-readable copy of the corresponding source code, to be\n    distributed under the terms of Sections 1 and 2 above on a medium\n    customarily used for software interchange; or,\n\n    c) Accompany it with the information you received as to the offer\n    to distribute corresponding source code.  (This alternative is\n    allowed only for noncommercial distribution and only if you\n    received the program in object code or executable form with such\n    an offer, in accord with Subsection b above.)\n\nThe source code for a work means the preferred form of the work for\nmaking modifications to it.  For an executable work, complete source\ncode means all the source code for all modules it contains, plus any\nassociated interface definition files, plus the scripts used to\ncontrol compilation and installation of the executable.  However, as a\nspecial exception, the source code distributed need not include\nanything that is normally distributed (in either source or binary\nform) with the major components (compiler, kernel, and so on) of the\noperating system on which the executable runs, unless that component\nitself accompanies the executable.\n\nIf distribution of executable or object code is made by offering\naccess to copy from a designated place, then offering equivalent\naccess to copy the source code from the same place counts as\ndistribution of the source code, even though third parties are not\ncompelled to copy the source along with the object code.\n\f\n  4. You may not copy, modify, sublicense, or distribute the Program\nexcept as expressly provided under this License.  Any attempt\notherwise to copy, modify, sublicense or distribute the Program is\nvoid, and will automatically terminate your rights under this License.\nHowever, parties who have received copies, or rights, from you under\nthis License will not have their licenses terminated so long as such\nparties remain in full compliance.\n\n  5. You are not required to accept this License, since you have not\nsigned it.  However, nothing else grants you permission to modify or\ndistribute the Program or its derivative works.  These actions are\nprohibited by law if you do not accept this License.  Therefore, by\nmodifying or distributing the Program (or any work based on the\nProgram), you indicate your acceptance of this License to do so, and\nall its terms and conditions for copying, distributing or modifying\nthe Program or works based on it.\n\n  6. Each time you redistribute the Program (or any work based on the\nProgram), the recipient automatically receives a license from the\noriginal licensor to copy, distribute or modify the Program subject to\nthese terms and conditions.  You may not impose any further\nrestrictions on the recipients' exercise of the rights granted herein.\nYou are not responsible for enforcing compliance by third parties to\nthis License.\n\n  7. If, as a consequence of a court judgment or allegation of patent\ninfringement or for any other reason (not limited to patent issues),\nconditions are imposed on you (whether by court order, agreement or\notherwise) that contradict the conditions of this License, they do not\nexcuse you from the conditions of this License.  If you cannot\ndistribute so as to satisfy simultaneously your obligations under this\nLicense and any other pertinent obligations, then as a consequence you\nmay not distribute the Program at all.  For example, if a patent\nlicense would not permit royalty-free redistribution of the Program by\nall those who receive copies directly or indirectly through you, then\nthe only way you could satisfy both it and this License would be to\nrefrain entirely from distribution of the Program.\n\nIf any portion of this section is held invalid or unenforceable under\nany particular circumstance, the balance of the section is intended to\napply and the section as a whole is intended to apply in other\ncircumstances.\n\nIt is not the purpose of this section to induce you to infringe any\npatents or other property right claims or to contest validity of any\nsuch claims; this section has the sole purpose of protecting the\nintegrity of the free software distribution system, which is\nimplemented by public license practices.  Many people have made\ngenerous contributions to the wide range of software distributed\nthrough that system in reliance on consistent application of that\nsystem; it is up to the author/donor to decide if he or she is willing\nto distribute software through any other system and a licensee cannot\nimpose that choice.\n\nThis section is intended to make thoroughly clear what is believed to\nbe a consequence of the rest of this License.\n\f\n  8. If the distribution and/or use of the Program is restricted in\ncertain countries either by patents or by copyrighted interfaces, the\noriginal copyright holder who places the Program under this License\nmay add an explicit geographical distribution limitation excluding\nthose countries, so that distribution is permitted only in or among\ncountries not thus excluded.  In such case, this License incorporates\nthe limitation as if written in the body of this License.\n\n  9. The Free Software Foundation may publish revised and/or new versions\nof the General Public License from time to time.  Such new versions will\nbe similar in spirit to the present version, but may differ in detail to\naddress new problems or concerns.\n\nEach version is given a distinguishing version number.  If the Program\nspecifies a version number of this License which applies to it and \"any\nlater version\", you have the option of following the terms and conditions\neither of that version or of any later version published by the Free\nSoftware Foundation.  If the Program does not specify a version number of\nthis License, you may choose any version ever published by the Free Software\nFoundation.\n\n  10. If you wish to incorporate parts of the Program into other free\nprograms whose distribution conditions are different, write to the author\nto ask for permission.  For software which is copyrighted by the Free\nSoftware Foundation, write to the Free Software Foundation; we sometimes\nmake exceptions for this.  Our decision will be guided by the two goals\nof preserving the free status of all derivatives of our free software and\nof promoting the sharing and reuse of software generally.\n\n\t\t\t    NO WARRANTY\n\n  11. BECAUSE THE PROGRAM IS LICENSED FREE OF CHARGE, THERE IS NO WARRANTY\nFOR THE PROGRAM, TO THE EXTENT PERMITTED BY APPLICABLE LAW.  EXCEPT WHEN\nOTHERWISE STATED IN WRITING THE COPYRIGHT HOLDERS AND/OR OTHER PARTIES\nPROVIDE THE PROGRAM \"AS IS\" WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED\nOR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF\nMERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE.  THE ENTIRE RISK AS\nTO THE QUALITY AND PERFORMANCE OF THE PROGRAM IS WITH YOU.  SHOULD THE\nPROGRAM PROVE DEFECTIVE, YOU ASSUME THE COST OF ALL NECESSARY SERVICING,\nREPAIR OR CORRECTION.\n\n  12. IN NO EVENT UNLESS REQUIRED BY APPLICABLE LAW OR AGREED TO IN WRITING\nWILL ANY COPYRIGHT HOLDER, OR ANY OTHER PARTY WHO MAY MODIFY AND/OR\nREDISTRIBUTE THE PROGRAM AS PERMITTED ABOVE, BE LIABLE TO YOU FOR DAMAGES,\nINCLUDING ANY GENERAL, SPECIAL, INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING\nOUT OF THE USE OR INABILITY TO USE THE PROGRAM (INCLUDING BUT NOT LIMITED\nTO LOSS OF DATA OR DATA BEING RENDERED INACCURATE OR LOSSES SUSTAINED BY\nYOU OR THIRD PARTIES OR A FAILURE OF THE PROGRAM TO OPERATE WITH ANY OTHER\nPROGRAMS), EVEN IF SUCH HOLDER OR OTHER PARTY HAS BEEN ADVISED OF THE\nPOSSIBILITY OF SUCH DAMAGES.\n\n\t\t     END OF TERMS AND CONDITIONS\n\f\n\t    How to Apply These Terms to Your New Programs\n\n  If you develop a new program, and you want it to be of the greatest\npossible use to the public, the best way to achieve this is to make it\nfree software which everyone can redistribute and change under these terms.\n\n  To do so, attach the following notices to the program.  It is safest\nto attach them to the start of each source file to most effectively\nconvey the exclusion of warranty; and each file should have at least\nthe \"copyright\" line and a pointer to where the full notice is found.\n\n    <one line to give the program's name and a brief idea of what it does.>\n    Copyright (C) <year>  <name of author>\n\n    This program is free software; you can redistribute it and/or modify\n    it under the terms of the GNU General Public License as published by\n    the Free Software Foundation; either version 2 of the License, or\n    (at your option) any later version.\n\n    This program is distributed in the hope that it will be useful,\n    but WITHOUT ANY WARRANTY; without even the implied warranty of\n    MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n    GNU General Public License for more details.\n\n    You should have received a copy of the GNU General Public License\n    along with this program; if not, write to the Free Software\n    Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA  02111-1307  USA\n\n\nAlso add information on how to contact you by electronic and paper mail.\n\nIf the program is interactive, make it output a short notice like this\nwhen it starts in an interactive mode:\n\n    Gnomovision version 69, Copyright (C) year name of author\n    Gnomovision comes with ABSOLUTELY NO WARRANTY; for details type `show w'.\n    This is free software, and you are welcome to redistribute it\n    under certain conditions; type `show c' for details.\n\nThe hypothetical commands `show w' and `show c' should show the appropriate\nparts of the General Public License.  Of course, the commands you use may\nbe called something other than `show w' and `show c'; they could even be\nmouse-clicks or menu items--whatever suits your program.\n\nYou should also get your employer (if you work as a programmer) or your\nschool, if any, to sign a \"copyright disclaimer\" for the program, if\nnecessary.  Here is a sample; alter the names:\n\n  Yoyodyne, Inc., hereby disclaims all copyright interest in the program\n  `Gnomovision' (which makes passes at compilers) written by James Hacker.\n\n  <signature of Ty Coon>, 1 April 1989\n  Ty Coon, President of Vice\n\nThis General Public License does not permit incorporating your program into\nproprietary programs.  If your program is a subroutine library, you may\nconsider it more useful to permit linking proprietary applications with the\nlibrary.  If this is what you want to do, use the GNU Library General\nPublic License instead of this License.\n"},{"id":"306","messageId":"7vacny94pb.fsf@assigned-by-dhcp.cox.net","threadId":"49","inReplyTo":"20050416132231.GJ2551@abridgegame.org","subject":"Re: using git directory cache code in darcs?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T14:28:00Z","receivedAt":"2005-04-16T14:28:00Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"DR\" == David Roundy <droundy@abridgegame.org> writes:\n\nDR> 1) Would this actually be a good idea?\n\nI think it is sensible, especially if you are doing a lot of\ncomparison between the working area and the pristine.\n\nDR> 3) Is it likely that git will switch to not using global\nDR> variables for active_cache, active_nr and active_alloc?\n\nDR> 4) Would there be interest in creating a libgit?\n\nThese are related.  I have seen some people interested in\nlibifying it, and encapsulating those globals would naturally\nfall out of it.  My impression from the list however is that a\nlot more people are interested in the upper SCM layer than the\ngit layer right now.  And git layer, although solid enough to\nhost itself, is still slushy.  A couple of days ago dircache\nformat was changed from host to network endian.  Last night\nLinus made another change to dircache format, which fortunately\nis upward compatible if you stay within pathnames shorter than\n2^12 bytes ;-).  Another problem I see for somebody to pick up\nand start libifying things right now is that, although there is\none central person on the SCM side (Petr Baudis), git layer is\nstill fractured between Linus and Petr.\n\nPetr syncs with Linus often and he seems to be doing a good job\nat keeping track of public patches, but the git layer Linus\nworks on does not have some patches Petr collected or wrote\nhimself.  In time, a better coordination would emerge, of\ncourse, but the project is still young at this moment.  Stay\ntuned ;-).\n"},{"id":"364","messageId":"Pine.LNX.4.58.0504161531470.7211@ppc970.osdl.org","threadId":"49","inReplyTo":"20050416132231.GJ2551@abridgegame.org","subject":"Re: using git directory cache code in darcs?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T22:43:02Z","receivedAt":"2005-04-16T22:43:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, David Roundy wrote:\n> \n> I've been thinking about the possibility of using the git \"current\n> directory cache\" code in darcs.\n\nGo wild. The license is GPLv2, with the limitation that I really do want\nto see v3 before I re-license anything at all, so if you take it into\ndarcs, you'd need to add that as a per-file comment (I just doing it in\nthe LICENSE file - I hate cluttering up individual files with tons of\ncommentary).\n\n> So my questions are:\n> \n> 1) Would this actually be a good idea? It seems good to me, but there may\n> be other considerations that I haven't thought of.\n\nI really don't know how well the git index file will work with darcs, and \nthe main issue is that the index file names the \"stable copy\" using the \nsha1 hash. If darcs uses something else (and I imagine it does) you'd need \nto do a fair amount of surgery, and I suspect merging changes won't be \nvery easy.\n\nSo it might well make sense to wait a bit, until the git thing has calmed\ndown some more. For example, I made some rather large changes\n(conceptually, if not in layout of the physical file) to the index file\njust yesterday, since git now uses it for merging too.\n\nIn git, the index file isn't just a speedup, it's the \"work\" file _and_\nthe merge entity. It's not just a floor wax, it's a dessert topping too!\n\n> 2) Will a license be chosen soon for git? Or has one been chosen, and I\n> missed it? I can't really include git code in darcs without a license.  I'd\n> prefer GPLv2 or later (since that's how darcs is licensed), but as long as\n> it's at least compabible with GPLv2, I'll be all right.\n\nYup, GPL, with the same \"v2 by default\" that the kernel uses).\n\n> 3) Is it likely that git will switch to not using global variables for\n> active_cache, active_nr and active_alloc?\n\nI wouldn't hate it, although for the intent of git, the global approach \nactually makes sense (dammit, I want the basic plumbing to be so small \nthat trying to abstract it out more is a waste of time). There's simply \nnot a lot of code that should even work at that level.\n\nBut if you wait a while, and bide your time, and then spring a clean patch \non me, I don't see any reason to be difficult about it either.\n\n> 4) Would there be interest in creating a libgit? I've been imagining taking\n> git source files and including them directly in darcs' code, but in the\n> long run it would be easier if there were a standard git API we could use.\n\nI think libgit might make sense, but again, not quite yet. Maybe the new\nmerge model was my last smart thought even on the subject of SCM's (I kind\nof hope so), but maybe it's not.\n\nMy gut _feel_ is that the basic git low-level architecture is done, and\nyou can certainly start looking around and see if it matches darcs at all. \n\n\t\t\tLinus\n"},{"id":"468","messageId":"20050417121712.GA22772@abridgegame.org","threadId":"49","inReplyTo":"Pine.LNX.4.58.0504161531470.7211@ppc970.osdl.org","subject":"Re: using git directory cache code in darcs?","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-17T12:17:16Z","receivedAt":"2005-04-17T12:17:16Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Sat, Apr 16, 2005 at 03:43:02PM -0700, Linus Torvalds wrote:\n> On Sat, 16 Apr 2005, David Roundy wrote:\n> > 1) Would this actually be a good idea? It seems good to me, but there may\n> > be other considerations that I haven't thought of.\n> \n> I really don't know how well the git index file will work with darcs, and\n> the main issue is that the index file names the \"stable copy\" using the\n> sha1 hash. If darcs uses something else (and I imagine it does) you'd\n> need to do a fair amount of surgery, and I suspect merging changes won't\n> be very easy.\n\nOh, I'm starting to see (having just browsed the git code for another half\nhour or so)... I had been under the (false) impression that the index file\nstored the contents of the files themselves, which in retrospect doesn't\nmake any sense.  So when you run update-cache --add, the file data itself\nimmediately goes into its final hashed location, and only the sha1 info\ngoes into the index.\n\nThat's all right.  Darcs would only access the cached data through a\ngit-caching layer, and we've already got an abstraction layer over the\npristine cache.  As long as the git layer can quickly retrieve the contents\nof a given file, we should be fine.\n\nThe sha1 file and tree hashing isn't direcly useful for darcs, but people\nwill want to interoperate with git, and for that it would be nice to be\nable to know what the hash of a given version is.  I imagine something like\n\ndarcs tag --git\n\nwhich would tag the current version with its git hash.  Of course, to\nimplement that we only need to reproduce your algorithm for hashing trees,\nwhich probably would be easier to do ourselves without using any git\ncode... but it would be far faster to recompute with the git backend, since\ngit stores the hashes of all the unmodified files, and since I also imagine\n\ndarcs record --git\n\nwhich would record a change, and then tag the resulting tree with a git\nhash, we might be recomputing the git hashes reasonably often, and we\ncertainly don't want to rehash the entire kernel each time! :)\n\n> So it might well make sense to wait a bit, until the git thing has calmed\n> down some more. For example, I made some rather large changes\n> (conceptually, if not in layout of the physical file) to the index file\n> just yesterday, since git now uses it for merging too.\n> \n> In git, the index file isn't just a speedup, it's the \"work\" file _and_\n> the merge entity. It's not just a floor wax, it's a dessert topping too!\n\nI think that sounds like a pretty reasonable match.  In darcs, there are\ninternally two main datatypes.  One is the Patch (as you might imagine),\nand the other is called a \"Slurpy\", which is basically a tree lazily\n\"slurped\" into memory.\n\nThe pristine cache is then just a way of storing the tree and so we can\n\"slurp\" it again later to retrieve the current state.  So in a sense we'd\nbe using only one side of the index file interface, the \"working directory\"\nside, where you check files out and add files in--treating it as an\nfast filesystem with a few extra-fancy features (like storing inodes of the\nfiles in the working directory).\n\n> I think libgit might make sense, but again, not quite yet. Maybe the new\n> merge model was my last smart thought even on the subject of SCM's (I kind\n> of hope so), but maybe it's not.\n> \n> My gut _feel_ is that the basic git low-level architecture is done, and\n> you can certainly start looking around and see if it matches darcs at all. \n\nSounds good.  That's sort of the feel I had gotten from other people's\nresponses as well.  We'll definitely look into how we can use (and\ninterface with) git.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"493","messageId":"Pine.LNX.4.58.0504170916080.7211@ppc970.osdl.org","threadId":"49","inReplyTo":"20050417121712.GA22772@abridgegame.org","subject":"Re: using git directory cache code in darcs?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-17T16:24:20Z","receivedAt":"2005-04-17T16:24:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 17 Apr 2005, David Roundy wrote:\n> \n> That's all right.  Darcs would only access the cached data through a\n> git-caching layer, and we've already got an abstraction layer over the\n> pristine cache.  As long as the git layer can quickly retrieve the contents\n> of a given file, we should be fine.\n\nYes.\n\nIn fact, one of my hopes was that other SCM's could just use the git\nplumbing. But then I'd really suggest that you use \"git\" itself, not any\n\"libgit\". Ie you take _all_ the plumbing as real programs, and instead of\ntrying to link against individual routines, you'd _script_ it.\n\nIn other words, \"git\" would be an independent cache of the real SCM,\nand/or the \"old history\" (ie an SCM that uses git could decide that the\ngit stuff is fine for archival, and really use git as the base: and then\nthe SCM could entirely concentrate on _only_ the \"interesting\" parts, ie\nthe actual merging etc).\n\nThat was really what I always personally saw \"git\" as, just the plumbing\nbeneath the surface. For example, something like arch, which is based on\n\"patches and tar-balls\" (I think darcs is similar in that respect), could\nuse git as a _hell_ of a better \"history of tar-balls\".\n\nThe thing is, unless you take the git object database approach, using \n_just_ the index part doesn't really mean all that much. Sure, you could \njust keep the \"current objects\" in the object database, but quite \nfrankly, there would probably not be a whole lot of point to that. You'd \nwaste so much time pruning and synchronizing with your \"real\" database \nthat I suspect you'd be better off not using it.\n\n(Or you could prune nightly or something, I guess).\n\n\t\tLinus\n"},{"id":"501","messageId":"4262938E.8010107@timesys.com","threadId":"49","inReplyTo":"Pine.LNX.4.58.0504170916080.7211@ppc970.osdl.org","subject":"Re: using git directory cache code in darcs?","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-04-17T16:49:18Z","receivedAt":"2005-04-17T16:49:18Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Sun, 17 Apr 2005, David Roundy wrote:\n> \n>>That's all right.  Darcs would only access the cached data through a\n>>git-caching layer, and we've already got an abstraction layer over the\n>>pristine cache.  As long as the git layer can quickly retrieve the contents\n>>of a given file, we should be fine.\n> \n> \n> Yes.\n> \n> In fact, one of my hopes was that other SCM's could just use the git\n> plumbing. But then I'd really suggest that you use \"git\" itself, not any\n> \"libgit\". Ie you take _all_ the plumbing as real programs, and instead of\n> trying to link against individual routines, you'd _script_ it.\n\nIf you don't want it, I won't do it. Still makes sense to separate the \nplumbing from the porcelain, though.\n\n-- \n\nMike Taht\n\n\n   \"You can tell how far we have to go, when FORTRAN is the language of\nsupercomputers.\n\t-- Steven Feiner\"\n"},{"id":"560","messageId":"4262E50C.2070006@lazy.shacknet.nu","threadId":"49","inReplyTo":"Pine.LNX.4.58.0504170916080.7211@ppc970.osdl.org","subject":"Re: using git directory cache code in darcs?","fromName":"Nomad Arton","fromEmail":"lkml@lazy.shacknet.nu","sentAt":"2005-04-17T22:37:00Z","receivedAt":"2005-04-17T22:37:00Z","isPatch":false,"sender":{"key":"lkml@lazy.shacknet.nu","avatar":null},"body":"Linus Torvalds schrieb:\n> \n> In fact, one of my hopes was that other SCM's could just use the git\n> plumbing. But then I'd really suggest that you use \"git\" itself, not any\n> \"libgit\". Ie you take _all_ the plumbing as real programs, and instead of\n> trying to link against individual routines, you'd _script_ it.\n\nplease excuse; libgit and scripting to me arent a contradiction. many \nsripting languages are extended by C modules, while still happening to \nhave all the scripting rapidity. its just a matter of how to communicate \nwith the C code, isnt it?\n\nyours,\n\npeter\n"},{"id":"566","messageId":"7vvf6lugw7.fsf@assigned-by-dhcp.cox.net","threadId":"49","inReplyTo":"4262E50C.2070006@lazy.shacknet.nu","subject":"Re: using git directory cache code in darcs?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-17T23:23:36Z","receivedAt":"2005-04-17T23:23:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"NA\" == Nomad Arton <lkml@lazy.shacknet.nu> writes:\n\nNA> Linus Torvalds schrieb:\n>> In fact, one of my hopes was that other SCM's could just use the git\n>> plumbing. But then I'd really suggest that you use \"git\" itself, not any\n>> \"libgit\". Ie you take _all_ the plumbing as real programs, and instead of\n>> trying to link against individual routines, you'd _script_ it.\n\nNA> please excuse; libgit and scripting to me arent a contradiction. many\nNA> sripting languages are extended by C modules, while still happening to\nNA> have all the scripting rapidity. its just a matter of how to\nNA> communicate with the C code, isnt it?\n\nYou are arguing for scripting language binding like what Swig\ncreates.  While that would also be a worthy addition, having\nlanguage binding is not the only way to do _script_.\n\nWhat Linus is saying is that he wants you to talk with git\nplumbing by invoking the executables he have, via system(3),\npopen(3), etc.\n\nThe C-level first has to be libified before you can start\ntalking about host language bindings but that just started to\nhappen and is not ready yet.  However, you can use and benefit\nfrom GIT without waiting for that kind of integration, if you\nuse the \"spawning the executables\" approach.  I agree with him.\n\n"},{"id":"600","messageId":"20050417190001.7e1ae3ac.pj@sgi.com","threadId":"49","inReplyTo":"7vvf6lugw7.fsf@assigned-by-dhcp.cox.net","subject":"Re: using git directory cache code in darcs?","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2005-04-18T02:00:01Z","receivedAt":"2005-04-18T02:00:01Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"Junio wrote:\n> What Linus is saying is that he wants you to talk with git\n> plumbing by invoking the executables he have, via system(3),\n> popen(3), etc.\n\nHopefully, Linus didn't specify system(3) or popen(3) for production\nsoftware.\n\nThey are a rich source of security holes.  Inefficient, too, since they\ninvoke a shell process to interpret the command.\n\nUse execve(2), or exevl(3), execle(3), execv(3).\n\nOr if you really enjoy the path search, use execlp or execvp, but with\nyour own $PATH, not trusting the one passed in via the environment any\nfurther than you can throw it.\n\nHowever, on further consideration, I think Linus is wrong to recommend\nthat the git executables, not a libgit library, be the 'basic user level\non which all else is based.\"\n\nI will reply to a Linus post, expounding on that thought further.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"616","messageId":"20050417195600.6894e576.pj@sgi.com","threadId":"49","inReplyTo":"Pine.LNX.4.58.0504170916080.7211@ppc970.osdl.org","subject":"Re: using git directory cache code in darcs?","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2005-04-18T02:56:00Z","receivedAt":"2005-04-18T02:56:00Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"Linus wrote:\n> But then I'd really suggest that you use \"git\" itself, not any\n> \"libgit\". Ie you take _all_ the plumbing as real programs, and instead of\n> trying to link against individual routines, you'd _script_ it.\n\nI think you've got this upside down, Linus.\n\nTrying to make the executable 'git' commands the bottom layer of the\nuser implementation stack forces inefficiencies on higher layers\nof the stack, and thus encourages stupid workarounds and cheats in\nan effort to speed things up.\n\nI'd encourage you to invite someone to provide a libgit.\n\nSuch work should _start_ with proposing and gaining acceptance on the\nAPI - the calls, the arguments, the types, the rough idea of the\nsemantics. The actual coding is the easy part.  One is not slave to the\nagreed API when coding.  The API will continue to evolve, but if the\noriginally accepted proposal was sound, the evolution will be at a\nmodest rate, with few incompatibilities introduced.\n\nIf several operations should be done as a unit, to preserve the\nintegrity of the .git data or to provide sane results, then libgit need\nonly provide such pre-packaged units, not the incomplete fragments from\nwhich they are composed.  That is, the libgit calls could quite possibly\nbe at roughly the same semantic level as your git commands.  One could\neven code up some of the libgit calls, in early versions of libgit, by\nsimply invoking the corresponding git command.  But, eventually, all the\ngit commands should be recoded on top of the libgit library, and the\nlibgit library become the canonical user interface to git, on which all\nelse is layered.\n\nOne typical way that this choice manifests itself is in the strace\noutput from doing some ordinary git command from a C program that is\nimplementing an SCM system on top of git.  Forcing every operation to be\ndone via a separate git command execution mushrooms the number of kernel\nsystem calls a hundred fold, or two hundred fold if some dang fool uses\nsystem(3S) to invoke the git command.  What might have been a handful of\ncalls to stat/open/read/write/close a file turns into a mini-shell\nsession.  That way lies insanity, or at least painful inefficiency, and\nthe usual parade of bugs, stupid coding tricks and painful user\ninterfaces that follow in the wake.\n\nThe recommended layering of such user facilities is well known, with a C\nlibrary at the bottom.  Granted, the history of source code management\ntools provides few examples of this recommended layering.\n\nOn top of this library go plugin modules for the fancier scripting\nlanguages that accept such.  Swig can be used to aid this construction,\nfor Tcl, Python, Perl, Guile, Java, Ruby, Mzscheme, PHP, Ocaml, Pike,\nC#, Chicken Allegro CL, Modula-3, Javascript and Eiffel.  Though I\npersonally have not worked with Swig enough to gain success with it.\nThe only such modules I've done were handcoded Python modules.\n\nAlso on top of this library one provides a set of command line utilities\nor one multiplexed 'git foo ...' command, for use at shell prompts.  Or\nthe command line utilities can be coded in one of the above higher level\nscripting languages, using in turn the git library plugin.  However many\nof these scripting languages bring runtime requirements that are not\nuniversally satisfied on all target systems, so are a poor choice for\nthis purpose.\n\nIf I am recalling correctly, from the days when I regularly used bk, one\nof the things that Larry did right with bk, which RCS and SCCS did not\ndo right before then, was to provide a low level library to his storage\n- a cleanroom recoded variant of SCCS in his case.\n\nImplementing production source control systems on top of a set of\nexecutable commands is a pain in the arse.  An all too familiar pain.\n\nI'd repeat my encouragement that you invite someone to provide such a\nlibgit, however since I have other commitments for the next month at\nleast, so can't volunteer right away, if ever, it is more appropriate\nthat I shut up now, under the old \"put up code or be quiet\" rule.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"619","messageId":"Pine.LNX.4.58.0504172005450.7211@ppc970.osdl.org","threadId":"49","inReplyTo":"20050417195600.6894e576.pj@sgi.com","subject":"Re: using git directory cache code in darcs?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-18T03:06:59Z","receivedAt":"2005-04-18T03:06:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 17 Apr 2005, Paul Jackson wrote:\n> \n> I'd encourage you to invite someone to provide a libgit.\n\nNot until all the data structures are really really stable.\n\nThat's the thing - we can keep the _program_ interfaces somewhat stable. \nBut internally we may change stuff wildly, and anybody who depends on a \nlibrary interface would be screwed.\n\nErgo: no library interfaces yet. Wait for it to stabilize. Start trying to \njust script the programs.\n\n\t\tLinus\n"},{"id":"623","messageId":"20050417213632.1f099ff9.pj@sgi.com","threadId":"49","inReplyTo":"Pine.LNX.4.58.0504172005450.7211@ppc970.osdl.org","subject":"Re: using git directory cache code in darcs?","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2005-04-18T04:36:32Z","receivedAt":"2005-04-18T04:36:32Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"> Not until all the data structures are really really stable.\n\nFine by me to wait, though perhaps not for the same reason, and perhaps\nnot as long.\n\nA libgit.so can deal with data structure changes just as well as a set\nof command line utilities.  So long as everything funnels through one\nplace, you can change by changing that one place.\n\nHowever I am  willing to agree that its not libgit time yet, for two\nreasons:\n\n 1) everyone who has two clues on the subject is too busy and\n    too productive on more pressing git issues, and\n\n 2) in addition to internal data structures being not yet stable,\n    I suspect that the operations (git commands, options and\n    behaviour) are also not stable.\n\nThe first step of a good libgit is not coding to the internal data\nstructures, but rather designing the interface (the operations,\narguments, data types, and behaviour).\n\nSo, until people have time, and the interface ops are settled down, its\ntoo early to design libgit.  Or at least too early to publish a design\nand seek concensus.  If I had the time the first thing I'd be doing\nright now would be designing libgit on the side, anticipating the day\nwhen it was time to publish a draft and engage the community discussion\nthat leads to an adequate concensus.\n\n===\n\nBy the way, a good libgit design, in my view, would isolate the data\nstructures written to files below .git from the data structures\npresented at the library API, to some extent.  Changes in the file\nstructures must be handled without disrupting the library API.\n\nIf a libgit API didn't isolate the library caller from details of the\nstructures in files below .git, then yes you'd want really really stable\ndata structures, impossibly stable in fact.  That way leads to hacks and\nworkarounds in the future, because the data structures are never\nperfectly stable.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"637","messageId":"42637C77.5070101@timesys.com","threadId":"49","inReplyTo":"20050417213632.1f099ff9.pj@sgi.com","subject":"git options","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-04-18T09:23:03Z","receivedAt":"2005-04-18T09:23:03Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"Would it be useful at this point to make common and centralize some/most \nof the various options that control git? (as well as add some useful \nones). Something like:\n\nstruct _git_opt {\n         int verbose:1;\n         int debug:1;\n\tint dry-run:1;\n         int should_block:1;\n         int remove_lock:1;\n         int allow_add:1;\n         int allow_remove:1;\n         int null_termination:1;\n         int show_cached:1;\n         int show_deleted:1;\n         int show_others:1;\n         int show_ignored:1;\n         int show_stage:1;\n         int show_unmerged:1;\n         int show_edges:1;\n         int show_unreachable:1;\n         int basemask:1;\n         int recursive:1;\n         int force_filename:1;\n         int force:1;\n         int quiet:1;\n         int stage:1;\n};\n\n\n-- \n\nMike Taht\n\n\n   \"Avoid letting temper block progress; keep cool.\n\t-- William Feather\"\n"}]}