{"thread":{"id":"17505","subject":"[wishlist] git-archive -L","startedAt":"2009-02-02T14:34:25Z","lastAt":"2009-02-05T15:04:29Z","messageCount":4,"participants":["Pierre Habouzit","René Scharfe"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"102867","messageId":"20090202143425.GA30667@artemis.corp","threadId":"17505","inReplyTo":null,"subject":"[wishlist] git-archive -L","fromName":"Pierre Habouzit","fromEmail":"madcoder@madism.org","sentAt":"2009-02-02T14:34:25Z","receivedAt":"2009-02-02T14:34:25Z","isPatch":false,"sender":{"key":"madcoder@madism.org","avatar":null},"body":"Hi Rene,\n\nI wanted to do that myself, but I sadly miss the time right now, so I\nwonder if you'd know how to do the following.\n\nWe have in our repository a kind of modular system (for a family of web\nsites) where each web-site uses a (versionned) symlink farm. IOW it\nworks basically that way:\n\n    www/module1\n    www/module2\n    product_A/www/module1 -> ../../www/module1\n    product_A/www/module_A\n    product_B/www/module1 -> ../../www/module1\n    product_B/www/module2 -> ../../www/module2\n    product_B/www/module_B\n\nThough product_A and _B even if they share a fair amount of code, are\nseparate products and when we release, we'd like to be able to perform\nfrom inside:\n\n    git archive --format=tar -L product_$A\n\nwhere -L basically does what it does in cp: dereference symlinks.  To\nmake the thing hairier, we also have symlinks _inside_ www/ (pointing\ninto the same subtree) that we'd like to keep if possible (even if it's\nnot a big deal).\n\nSo I'd suggest something where -L only dereferences the symlink if it\ngoes outside of the list of paths passed to git-archive, and -LL (or -L\n-L) dereferences anything. Of course this would only make sense if the\nsymlinks resolve to something that is tracked :)\n\nFor now we git archive the whole repository, use tar xh; rm what we\ndon't like, reset the symlinks we want to keep, and retar, which is kind\nof counterproductive :)\n\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"102938","messageId":"4987FC03.60607@lsrfire.ath.cx","threadId":"17505","inReplyTo":"20090202143425.GA30667@artemis.corp","subject":"Re: [wishlist] git-archive -L","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2009-02-03T08:10:43Z","receivedAt":"2009-02-03T08:10:43Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Pierre Habouzit schrieb:\n> Hi Rene,\n> \n> I wanted to do that myself, but I sadly miss the time right now, so I\n> wonder if you'd know how to do the following.\n> \n> We have in our repository a kind of modular system (for a family of web\n> sites) where each web-site uses a (versionned) symlink farm. IOW it\n> works basically that way:\n> \n>     www/module1\n>     www/module2\n>     product_A/www/module1 -> ../../www/module1\n>     product_A/www/module_A\n>     product_B/www/module1 -> ../../www/module1\n>     product_B/www/module2 -> ../../www/module2\n>     product_B/www/module_B\n> \n> Though product_A and _B even if they share a fair amount of code, are\n> separate products and when we release, we'd like to be able to perform\n> from inside:\n> \n>     git archive --format=tar -L product_$A\n> \n> where -L basically does what it does in cp: dereference symlinks.  To\n> make the thing hairier, we also have symlinks _inside_ www/ (pointing\n> into the same subtree) that we'd like to keep if possible (even if it's\n> not a big deal).\n> \n> So I'd suggest something where -L only dereferences the symlink if it\n> goes outside of the list of paths passed to git-archive, and -LL (or -L\n> -L) dereferences anything. Of course this would only make sense if the\n> symlinks resolve to something that is tracked :)\n\nLast April, I was working on making archive follow all symlinks pointing\nto internal files.  The goal was a bit different, namely to create\narchives for platforms without symlink support (i.e. it would resolve\nall symlinks pointing to tracked objects).\n\nIIRC the code had some limitations, e.g. it couldn't follow a symlink to\na path containing symlinked directories.  I'll need to rebase it to\nmaster first, though, as the surrounding code has changed a bit in the\nmeantime.\n\nTo follow only symlinks that point outside of the specified paths sounds\nlike a sensible mode of operation, but I'm not sure that it's worth a\none letter option.\n\nGiven your setup you also might want to take a look at submodules and\nthe recent submodule archival support patches by Lars Hjelmi.\n\nAnyway, I'll try to resurrect my old, incomplete symlink following code,\nbut I don't have much time, either. :-/\n\nRené\n"},{"id":"103224","messageId":"498A1E02.6020707@lsrfire.ath.cx","threadId":"17505","inReplyTo":"4987FC03.60607@lsrfire.ath.cx","subject":"Re: [wishlist] git-archive -L","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2009-02-04T23:00:18Z","receivedAt":"2009-02-04T23:00:18Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"René Scharfe schrieb:\n> Anyway, I'll try to resurrect my old, incomplete symlink following code,\n> but I don't have much time, either. :-/\n\nAfter a second and a third look I don't see any salvageable parts in the\nold code any more.  It was a just prototype that taught me something I\nshould have been able to find out by thinking alone: that to follow\nlinks within tracked content we can't simply jump to the target, but we\nhave to walk the whole path step by step.\n\nE.g., consider a repository with these four entries:\n\n\tType\tName\tTarget\n\t-------\t-------\t------\n\tfile\ta/f\n\tsymlink\ta/x\tf\n\tsymlink a/y\t../b/f\n\tsymlink\tb\ta\n\nLet's say our goal is to follow symlinks pointing to tracked content.\n\nWe can easily follow \"a/x\" to get to its target \"f\" by concatenating the\ndirectory part of the symlink's path (\"a/\") with the target (\"f\"), i.e.\nwe only need to do a simple string operation.\n\nIf we do the same for \"a/y\", we'd arrive at \"b/f\", which is not a\ntracked file by itself, though.  We need to look up each path element\none by one and follow symlinks at each step.  That can't be done with\nour existing tree walkers, AFAICS, so we'd need to write a new one.\n\nThe decision to follow a link can be made by the callback and passed to\nread_tree_recursive() as a return value, with, e.g., READ_TREE_FOLLOW\nand READ_TREE_FOLLOW_NON_MATCHES meaning to follow all internal symlinks\nand to follow only those whose target doesn't match the specified paths,\nrespectively.\n\nRené\n"},{"id":"103339","messageId":"20090205150429.GB6434@artemis.corp","threadId":"17505","inReplyTo":"498A1E02.6020707@lsrfire.ath.cx","subject":"Re: [wishlist] git-archive -L","fromName":"Pierre Habouzit","fromEmail":"madcoder@madism.org","sentAt":"2009-02-05T15:04:29Z","receivedAt":"2009-02-05T15:04:29Z","isPatch":false,"sender":{"key":"madcoder@madism.org","avatar":null},"body":"On Wed, Feb 04, 2009 at 11:00:18PM +0000, René Scharfe wrote:\n> René Scharfe schrieb:\n> > Anyway, I'll try to resurrect my old, incomplete symlink following code,\n> > but I don't have much time, either. :-/\n> \n> After a second and a third look I don't see any salvageable parts in the\n> old code any more.  It was a just prototype that taught me something I\n> should have been able to find out by thinking alone: that to follow\n> links within tracked content we can't simply jump to the target, but we\n> have to walk the whole path step by step.\n> \n> E.g., consider a repository with these four entries:\n> \n> \tType\tName\tTarget\n> \t-------\t-------\t------\n> \tfile\ta/f\n> \tsymlink\ta/x\tf\n> \tsymlink a/y\t../b/f\n> \tsymlink\tb\ta\n> \n> Let's say our goal is to follow symlinks pointing to tracked content.\n> \n> We can easily follow \"a/x\" to get to its target \"f\" by concatenating the\n> directory part of the symlink's path (\"a/\") with the target (\"f\"), i.e.\n> we only need to do a simple string operation.\n> \n> If we do the same for \"a/y\", we'd arrive at \"b/f\", which is not a\n> tracked file by itself, though.  We need to look up each path element\n> one by one and follow symlinks at each step.  That can't be done with\n> our existing tree walkers, AFAICS, so we'd need to write a new one.\n\nI mostly stumbled on those issues before I gave up having no time to\nunderstand how tree walkers work :/\n\nBecause of course, our symlinks are exactly symlinks to directories, so\nnot supporting'em is unacceptable to us.\n\n> The decision to follow a link can be made by the callback and passed to\n> read_tree_recursive() as a return value, with, e.g., READ_TREE_FOLLOW\n> and READ_TREE_FOLLOW_NON_MATCHES meaning to follow all internal symlinks\n> and to follow only those whose target doesn't match the specified paths,\n> respectively.\n\nIt has to be more clever. If you consider something like:\n\n    symlink a/b   ..\n\nOr funnier:\n\n    symlink a/b   ../../c\n    symlink c/d   ../../a\n\nIf you don't pay attention, you end up with a nice busy loop, and really\nreally really long path names (a/b/b/b/b/b.... for the first one,\nand a/b/d/b/d/b/d/b/d/b/... for the latter).\n\nThat's why I was thinking of a more straight approach, basicaly doing\nthat:\n  * when meeting a symlink to a blob, see if that blob is tracked or\n    not, and if its \"real\" path in the repository is inside what we're\n    archiving or not. Then match that with what the user asked\n    (following any symlinks -- if we want to, this looks like a pretty\n    big security risk to me, and I see no good reason for that --, only\n    tracked symlinks outside of the archived paths, or only tracked\n    symlinks no matter what), and do it.\n\n    This one is the almost easy bit.\n\n  * when meeting a symlink to a directory, look at the pointee, and like\n    for the file, see if it's \"tracked\" (IOW contains tracked files) and\n    see if the user want symlink replacement or not. If yes, then\n    remember the current <path, pointed directory inside the repository>\n    and put it in a worklist.\n\nWhen finishing the first \"pass\" of archiving, run a new archiving based\non the worklist. Do it a few times. and if you don't converge to a fixed\npoint where the worklist is empty, then you are likely to be in a\nsituation like the ones I depict earlier. *phew*.\n\n\nThough this need quite a reeingeenering of the code, and I had (still\ndon't really have) no time for it. But I think this is the straight\napproach that would work easily (I don't know for zip though, but in tar\nwhere entries are not really sorted, it should work).\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"}]}