{"thread":{"id":"19347","subject":"Narrow clone implementation difficulty estimate","startedAt":"2009-05-14T10:04:30Z","lastAt":"2009-05-16T05:17:01Z","messageCount":3,"participants":["Alexander Gavrilov","Jakub Narebski","Nguyen Thai Ngoc Duy"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"113916","messageId":"200905141404.30695.angavrilov@gmail.com","threadId":"19347","inReplyTo":null,"subject":"Narrow clone implementation difficulty estimate","fromName":"Alexander Gavrilov","fromEmail":"angavrilov@gmail.com","sentAt":"2009-05-14T10:04:30Z","receivedAt":"2009-05-14T10:04:30Z","isPatch":false,"sender":{"key":"angavrilov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/42666?v=4"},"body":"Hello,\n\nWe are considering using Git to manage a large set of mostly binary\nfiles (large images, pdf files, open-office documents, etc). The\namount of data is such that it is infeasible to force every user\nto download all of it, so it is necessary to implement a partial\nretrieval scheme.\n\nIn particular, we need to decide whether it is better to invest\neffort into implementing Narrow Clone, or partitioning and\nreorganizing the data set into submodules (the latter may prove\nto be almost impossible for this data set). We will most likely\ndevelop a new, very simplified GUI for non-technical users,\nso the details of both possible approaches will be hidden\nunder the hood.\n\n\nAfter some looking around, I think that Narrow clone would probably involve:\n\n1. Modifying the revision walk engine used by the pack generator to\nallow filtering blobs using a set of path masks. (Handling the same\ntree object appearing at different paths may be tricky.)\n\n2. Modifying the fetch protocol to allow sending such filter\nexpressions to the server.\n\n3. Adding necessary configuration entries and parameters to commands,\nin order to allow using the new functionality.\n\n4. Resurrecting the sparse checkout series and merging it with the\nnew filtering logic. Narrow clone must imply sparse checkout that\nis a subset of the cloned paths.\n\n5. Fixing all breakage that may be caused by missing blobs.\n\nI feel that the last point involves the most uncertainty, and may also\nprove the most difficult one to implement. However, I cannot judge the\nactual difficulty due to an incomplete understanding of Git internals.\n\n\nI currently see the following additional problems with this approach:\n\n1. Merge conflicts outside the filtered area cannot be handled.\nHowever, in the case of this project they are estimated to be\nextremely unlikely.\n\n2. Changing the filter set is tricky, because extending the watched\narea requires connecting to the server, and requesting missing blobs.\nThis action appears to be mostly identical to initial clone with a\nmore complex filter. On the other hand, shrinking the area would leave\nunnecessary data in the repository, which is difficult to reuse safely\nif the area is extended back. Finally, editing the set without\ndownloading missing data essentially corrupts the repository.\n\n3. One of the goals of using git is building a distributed mirroring\nsystem, similar to gittorrent or mirror-sync proposals. Narrow clone\nsignificantly complicates this because of incomplete data sets.\nA simple solution may be restricting download to peers whose set is\na superset of what's needed, but that may cause the system to degrade\nto a fully centralized one.\n\n\nIn relation to the last point, namely building a mirroring\nnetwork, I also had an idea that perhaps in the current state\nof things bundles are more suited to it, because they can be\ndirectly reused by many peers, and deciding what to put in\nthe bundle is not much of a problem for this particular project.\nI expect that implementation of narrow bundle support should\nnot be much different from narrow clone.\n\n\nCurrently we are evaluating possibilities to approach this\nproblem, and would like to know if this analysis makes sense.\nWe are willing to contribute the results to the Git community\nif/when we implement it.\n\nAlexander\n"},{"id":"113918","messageId":"m38wl0klxt.fsf@localhost.localdomain","threadId":"19347","inReplyTo":"200905141404.30695.angavrilov@gmail.com","subject":"Re: Narrow clone implementation difficulty estimate","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-05-14T10:39:49Z","receivedAt":"2009-05-14T10:39:49Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Alexander Gavrilov <angavrilov@gmail.com> writes:\n\n> We are considering using Git to manage a large set of mostly binary\n> files (large images, pdf files, open-office documents, etc). The\n> amount of data is such that it is infeasible to force every user\n> to download all of it, so it is necessary to implement a partial\n> retrieval scheme.\n> \n> In particular, we need to decide whether it is better to invest\n> effort into implementing Narrow Clone, or partitioning and\n> reorganizing the data set into submodules (the latter may prove\n> to be almost impossible for this data set). We will most likely\n> develop a new, very simplified GUI for non-technical users,\n> so the details of both possible approaches will be hidden\n> under the hood.\n\nFirst, there were quite complete, although as far as I know newer\naccepted into git, work on narrow / sparse / subtree / partial\n*checkout*.  IIRC the general idea about extening or (ab)using\nassume-unchanged mechanism was accepted, but the problem was in the\nuser interface details (I think that porcelain part was quite well\naccepted, except hesitation whether to use/extend existing flag, or\ncreate new for the purpose of narrow checkout).  You can search\narchive for that\n  http://article.gmane.org/gmane.comp.version-control.git/89900\n  http://article.gmane.org/gmane.comp.version-control.git/90016\n  http://article.gmane.org/gmane.comp.version-control.git/77046\n  http://article.gmane.org/gmane.comp.version-control.git/50256\n  ...\nshould give you some idea what to search for. This is of course\nonly part of solution.\n\nSecond, there was an idea to use new \"replace\" mechanism for this\n(currently in 'pu' only, I think, merged as 'cc/replace' branch).\nThis mechanism was created for better bisecting with non-bisectable\ncommits, and is meant to be transferable extension of 'graft'\nmechanism. The \"replace\" mechanism allows to replace also blob objects\n(contents of filename), so you can have two repositories: baseline\nrepository with stub files in place of large binary files, and\nextended repository with replacement in and replacement blobs in\nobject database with 'proper' (and large) contents of those binary\nfiles. But that is just an idea, without implementation.\n\nThird, there was work (a year ago, perhaps?) by Dana How on better\nsupport for large objects. Some of those got accepted, some\ndosn't. You can set maximum size of object in pack, IIRC, and you can\nuse gitattributes to mark (binary) files that are meant to be not\ndeltaified. If all of your repositories are on networked filesystem,\nyou can create separate optimized pack containing only those large\nbinary files, mark it as \"kept\" (using *.keep file, see documentation)\nto avoid repacking those large binary files, and distributed this pack\neither using symlink, or using alternates (keeping only one copy of\nthis pack, and accessing it via networked filesystem when it is\nrequired).\n\nFourth, a long thime ago there was send a patch supposedly adding\nsupport for 'lazy' clone, where you download blob objects from remote\nrepository only as required.  But its was send as a single large\npatch, fairly intrusive.  I don't think it got good review, nevermind\nbeing accepted into git.\n\n\nSome further reading:\n* \"large(25G) repository in git\"\n  http://article.gmane.org/gmane.comp.version-control.git/114351\n* \"Re: Appropriateness of git for digital video production versioning\"\n  http://article.gmane.org/gmane.comp.version-control.git/107696\n* http://git.or.cz/gitwiki/GitTogether08 had some presentation\n  about media files in git, and some thread on git mailing list about\n  that issue was result (which I didn't bookmark).\n\nHTH\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"114074","messageId":"fcaeb9bf0905152217g418c7f38w229f71dd047bb466@mail.gmail.com","threadId":"19347","inReplyTo":"m38wl0klxt.fsf@localhost.localdomain","subject":"Re: Narrow clone implementation difficulty estimate","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2009-05-16T05:17:01Z","receivedAt":"2009-05-16T05:17:01Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, May 14, 2009 at 8:39 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n> Alexander Gavrilov <angavrilov@gmail.com> writes:\n>\n>> We are considering using Git to manage a large set of mostly binary\n>> files (large images, pdf files, open-office documents, etc). The\n>> amount of data is such that it is infeasible to force every user\n>> to download all of it, so it is necessary to implement a partial\n>> retrieval scheme.\n>>\n>> In particular, we need to decide whether it is better to invest\n>> effort into implementing Narrow Clone, or partitioning and\n>> reorganizing the data set into submodules (the latter may prove\n>> to be almost impossible for this data set). We will most likely\n>> develop a new, very simplified GUI for non-technical users,\n>> so the details of both possible approaches will be hidden\n>> under the hood.\n>\n> First, there were quite complete, although as far as I know newer\n> accepted into git, work on narrow / sparse / subtree / partial\n> *checkout*.  IIRC the general idea about extening or (ab)using\n> assume-unchanged mechanism was accepted, but the problem was in the\n> user interface details (I think that porcelain part was quite well\n> accepted, except hesitation whether to use/extend existing flag, or\n> create new for the purpose of narrow checkout).  You can search\n> archive for that\n>  http://article.gmane.org/gmane.comp.version-control.git/89900\n>  http://article.gmane.org/gmane.comp.version-control.git/90016\n>  http://article.gmane.org/gmane.comp.version-control.git/77046\n>  http://article.gmane.org/gmane.comp.version-control.git/50256\n>  ...\n> should give you some idea what to search for. This is of course\n> only part of solution.\n\nFWIW I still maintain the patch series as a merged branch \"tp/sco\"\nunder my branch \"inst\" here\n\nhttp://repo.or.cz/w/git/pclouds.git?a=shortlog;h=refs/heads/inst\n-- \nDuy\n"}]}