{"thread":{"id":"49406","subject":"Import/Export as a fast way to purge files from Git?","startedAt":"2018-09-23T13:05:19Z","lastAt":"2018-11-16T12:29:47Z","messageCount":90,"participants":["Lars Schneider","Eric Sunshine","brian m. carlson","Jeff King","Elijah Newren","Ævar Arnfjörð Bjarmason","Jonathan Nieder","SZEDER Gábor"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"358719","messageId":"F65AF000-7AE0-44C8-81C8-E58D6769FAA3@gmail.com","threadId":"49406","inReplyTo":null,"subject":"Import/Export as a fast way to purge files from Git?","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2018-09-23T13:04:58Z","receivedAt":"2018-09-23T13:05:19Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"Hi,\n\nI recently had to purge files from large Git repos (many files, many commits). \nThe usual recommendation is to use `git filter-branch --index-filter` to purge \nfiles. However, this is *very* slow for large repos (e.g. it takes 45min to\nremove the `builtin` directory from git core). I realized that I can remove\nfiles *way* faster by exporting the repo, removing the file references, \nand then importing the repo (see Perl script below, it takes ~30sec to remove\nthe `builtin` directory from git core). Do you see any problem with this \napproach?\n\nThank you,\nLars\n\n\n\n#!/usr/bin/perl\n#\n# Purge paths from Git repositories.\n#\n# Usage:\n#     git-purge-path [path-regex1] [path-regex2] ...\n#\n# Examples:\n#    Remove the file \"test.bin\" from all directories:\n#    git-purge-path \"/test.bin$\"\n#\n#    Remove all \"*.bin\" files from all directories:\n#    git-purge-path \"\\.bin$\"\n#\n#    Remove all files in the \"/foo\" directory:\n#    git-purge-path \"^/foo/$\"\n#\n# Attention:\n#     You want to run this script on a case sensitive file-system (e.g.\n#     ext4 on Linux). Otherwise the resulting Git repository will not\n#     contain changes that modify the casing of file paths.\n#\n\nuse strict;\nuse warnings;\n\nopen( my $pipe_in, \"git fast-export --progress=100 --no-data HEAD |\" ) or die $!;\nopen( my $pipe_out, \"| git fast-import --force --quiet\" ) or die $!;\n\nLOOP: while ( my $cmd = <$pipe_in> ) {\n    my $data = \"\";\n    if ( $cmd =~ /^data ([0-9]+)$/ ) {\n        # skip data blocks\n        my $skip_bytes = $1;\n        read($pipe_in, $data, $skip_bytes);\n    }\n    elsif ( $cmd =~ /^M [0-9]{6} [0-9a-f]{40} (.+)$/ ) {\n        my $pathname = $1;\n        foreach (@ARGV) {\n            next LOOP if (\"/\" . $pathname) =~ /$_/\n        }\n    }\n    print {$pipe_out} $cmd . $data;\n}\n\n"},{"id":"358724","messageId":"CAPig+cTLjThK4CVzfgV=Uk5OumpjhaQD_YNXmg7pNtkkUFiiyQ@mail.gmail.com","threadId":"49406","inReplyTo":"F65AF000-7AE0-44C8-81C8-E58D6769FAA3@gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2018-09-23T14:55:52Z","receivedAt":"2018-09-23T14:56:07Z","isPatch":false,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Sep 23, 2018 at 9:05 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n> I recently had to purge files from large Git repos (many files, many commits).\n> The usual recommendation is to use `git filter-branch --index-filter` to purge\n> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n> remove the `builtin` directory from git core). I realized that I can remove\n> files *way* faster by exporting the repo, removing the file references,\n> and then importing the repo (see Perl script below, it takes ~30sec to remove\n> the `builtin` directory from git core). Do you see any problem with this\n> approach?\n\nA couple comments:\n\nFor purging files from a history, take a look at BFG[1] which bills\nitself as \"a simpler, faster alternative to git-filter-branch for\ncleansing bad data out of your Git repository history\".\n\nThe approach of exporting to a fast-import stream, modifying the\nstream, and re-importing is quite reasonable. However, rather than\nre-inventing, take a look at reposurgeon[2], which allows you to do\nmajor surgery on fast-import streams. Not only can it purge files from\na repository, but it can slice, dice, puree, and saute pretty much any\nattribute of a repository.\n\n[1]: https://rtyley.github.io/bfg-repo-cleaner/\n[2]: http://www.catb.org/esr/reposurgeon/\n"},{"id":"358725","messageId":"20180923155338.GF432229@genre.crustytoothpaste.net","threadId":"49406","inReplyTo":"F65AF000-7AE0-44C8-81C8-E58D6769FAA3@gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-09-23T15:53:38Z","receivedAt":"2018-09-23T15:54:35Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Sep 23, 2018 at 03:04:58PM +0200, Lars Schneider wrote:\n> Hi,\n> \n> I recently had to purge files from large Git repos (many files, many commits).\n> The usual recommendation is to use `git filter-branch --index-filter` to purge\n> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n> remove the `builtin` directory from git core). I realized that I can remove\n> files *way* faster by exporting the repo, removing the file references,\n> and then importing the repo (see Perl script below, it takes ~30sec to remove\n> the `builtin` directory from git core). Do you see any problem with this\n> approach?\n\nI don't know of any problems with this approach.  I didn't audit your\nspecific Perl script for any issues, though.\n\nI suspect you're gaining speed mostly because you're running three\nprocesses total instead of at least one process (sh) per commit.  So I\ndon't think there's anything that Git can do to make this faster on our\nend without a redesign.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"358726","messageId":"DE9CF60B-7DCB-4B82-9C96-663E145CAD56@gmail.com","threadId":"49406","inReplyTo":"CAPig+cTLjThK4CVzfgV=Uk5OumpjhaQD_YNXmg7pNtkkUFiiyQ@mail.gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2018-09-23T15:58:14Z","receivedAt":"2018-09-23T15:58:20Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n\n> On Sep 23, 2018, at 4:55 PM, Eric Sunshine <sunshine@sunshineco.com> wrote:\n> \n> On Sun, Sep 23, 2018 at 9:05 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n>> I recently had to purge files from large Git repos (many files, many commits).\n>> The usual recommendation is to use `git filter-branch --index-filter` to purge\n>> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n>> remove the `builtin` directory from git core). I realized that I can remove\n>> files *way* faster by exporting the repo, removing the file references,\n>> and then importing the repo (see Perl script below, it takes ~30sec to remove\n>> the `builtin` directory from git core). Do you see any problem with this\n>> approach?\n> \n> A couple comments:\n> \n> For purging files from a history, take a look at BFG[1] which bills\n> itself as \"a simpler, faster alternative to git-filter-branch for\n> cleansing bad data out of your Git repository history\".\n\nYes, BFG is great. Unfortunately, it requires Java which is not available\non every system I have to work with. I required a solution that would work\nin every Git environment. Hence the Perl script :-)\n\n\n> The approach of exporting to a fast-import stream, modifying the\n> stream, and re-importing is quite reasonable.\n\nThanks for the confirmation!\n\n\n> However, rather than\n> re-inventing, take a look at reposurgeon[2], which allows you to do\n> major surgery on fast-import streams. Not only can it purge files from\n> a repository, but it can slice, dice, puree, and saute pretty much any\n> attribute of a repository.\n\nWow. Reposurgeon looks very interesting. Thanks a lot for the pointer!\n\nCheers,\nLars\n\n\n> [1]: https://rtyley.github.io/bfg-repo-cleaner/\n> [2]: http://www.catb.org/esr/reposurgeon/\n\n"},{"id":"358727","messageId":"20180923170404.GA1961@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20180923155338.GF432229@genre.crustytoothpaste.net","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-09-23T17:04:04Z","receivedAt":"2018-09-23T17:04:09Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Sep 23, 2018 at 03:53:38PM +0000, brian m. carlson wrote:\n\n> I suspect you're gaining speed mostly because you're running three\n> processes total instead of at least one process (sh) per commit.  So I\n> don't think there's anything that Git can do to make this faster on our\n> end without a redesign.\n\nIt's not just the process startup overhead that makes it faster. Using\nmultiple processes means they have to communicate somehow. In this case,\ngit-read-tree is writing out the whole index for each commit, which\ngit-rm reads in and modifies, and then git-commit-tree finally converts\nback to a tree. In addition to the raw CPU of that work, there's a bunch\nof latency as each step is performed serially.\n\nWhereas in the proposed pipeline, fast-export is writing out a diff and\nfast-import is turning that directly back into tree objects. And both\nprocesses are proceeding independently, so you benefit from multiple\ncores.\n\nWhich isn't to say I really disagree with \"Git can't really make this\nfaster\". filter-branch has a ton of power to let you replay arbitrary\ncommands (including non-Git commands!), so the speed tradeoff in its\napproach is very intentional. If we could modify the index in-place that\nwould probably make it a little faster, but that probably counts as\n\"redesign\" in your statement. ;)\n\n-Peff\n"},{"id":"358765","messageId":"CABPp-BGL-3_nhZSpt0Bz0EVY-6-mcbgZMmx4YcXEfA_ZrTqFUw@mail.gmail.com","threadId":"49406","inReplyTo":"F65AF000-7AE0-44C8-81C8-E58D6769FAA3@gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-09-24T17:24:23Z","receivedAt":"2018-09-24T17:24:38Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sun, Sep 23, 2018 at 6:08 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n>\n> Hi,\n>\n> I recently had to purge files from large Git repos (many files, many commits).\n> The usual recommendation is to use `git filter-branch --index-filter` to purge\n> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n> remove the `builtin` directory from git core). I realized that I can remove\n> files *way* faster by exporting the repo, removing the file references,\n> and then importing the repo (see Perl script below, it takes ~30sec to remove\n> the `builtin` directory from git core). Do you see any problem with this\n> approach?\n\nIt looks like others have pointed you at other tools, and you're\nalready shifting to that route.  But I think it's a useful question to\nanswer more generally, so for those that are really curious...\n\n\nThe basic approach is fine, though if you try to extend it much you\ncan run into a few possible edge/corner cases (more on that below).\nI've been using this basic approach for years and even created a\nmini-python library[1] designed specifically to allow people to create\n\"fast-filters\", used as\n   git fast-export <options> | your-fast-filter | git fast-import <options>\n\nBut that library didn't really take off; even I have rarely used it,\noften opting for filter-branch despite its horrible performance or a\nsimple fast-export | long-sed-command | fast-import (with some extra\npre-checking to make sure the sed wouldn't unintentionally munge other\ndata).  BFG is great, as long as you're only interested in removing a\nfew big items, but otherwise doesn't seem very useful (to be fair,\nit's very upfront about only wanting to solve that problem).\nRecently, due to continuing questions on filter-branch and folks still\ngetting confused with it, I looked at existing tools, decided I didn't\nthink any quite fit, and started looking into converting\ngit_fast_filter into a filter-branch-like tool instead of just a\nlibary.  Found some bugs and missing features in fast-export along the\nway (and have some patches I still need to send in).  But I kind of\ngot stuck -- if the tool is in python, will that limit adoption too\nmuch?  It'd be kind of nice to have this tool in core git.  But I kind\nof like leaving open the possibility of using it as a tool _or_ as a\nlibrary, the latter for the special cases where case-specific\nprogrammatic filtering is needed.  But a developer-convenience library\nmakes almost no sense unless in a higher level language, such as\npython.  I'm still trying to make up my mind about what I want (and\nwhat others might want), and have been kind of blocking on that.  (If\nothers have opinions, I'm all ears.)\n\n\nAnyway, the edge/corner cases you can watch out for:\n\n  - Signed tags are a problem; you may need to specify\n--signed-tags=strip to fast-export\n\n  - References to other commits in your commit messages will now be\nincorrect.  I think a good tool should either default to rewriting\ncommit ids in commit messages or at least have an option to do so\n(BFG does this; filter-branch doesn't; fast-export format makes it\nreally hard for a filter based on it to do so)\n\n  - If the paths you remove are the only paths modified in a commit,\nthe commit can become empty.  If you're only filtering a few paths\nout, this might be nothing more than a minor inconvenience for you.\nHowever, if you're trying to prune directories (and perhaps several\ntoplevel ones), then it can be extremely annoying to have a new\nhistory with the vast majority of all commits being empty.\n(filter-branch has an option for this; BFG does not; tools based on\nfast-export output can do it with sufficient effort).\n\n  - If you start pruning empty commits, you have to worry about\nrewriting branches and tags to remaining parents.  This _might_ happen\nfor free depending on your history's structure and the fast-export\nstream, but to be correct in general you will have to specify the new\ncommit for some branches or tags.\n\n  - If you start pruning empty commits, you have to decide whether to\nallow pruning of merge commits.  Your first reaction might be to not\nallow it, but if one parent and its entire history are all pruned,\nthen transforming the merge commit to a normal commit and then\nconsidering whether it is empty and allowing it to be pruned is much\nbetter.\n\n  - If you start pruning empty commits, you also have to worry about\nhistory topology changing, beyond the all-ancestors-empty case above.\nFor example, the last non-empty commit in the ancestry of a merge on\nboth sides may be the same commit, making the merge-commit have the\nsame parent twice.  Should the duplicate parent be de-duped,\ntransforming the commit into a normal non-merge commit?  (I'd say yes\n-- this commit is likely to be empty and prunable once you do so, but\nI'm not sure everyone would agree with converting a merge commit to a\nnon-merge.)  Similarly, what if the rewritten parents of a merge have\none parent that is the direct ancestor of another?  Can the extra\nunnecessary parent be removed as a parent?  (And again, such a commit\nis likely to become empty and be prunable itself.)\n\n  - If you try to avoid the extra work involved with pruning empty\ncommits by passing path-specifiers as rev-list-args to fast-export,\nand use the --tag-of-filtered-object=rewrite option if needed, then\ndepending on the topology you can hit any of three bugs: an outright\ndie() (despite the --tag-of-filtered-object=rewrite), a branch being\nreset to a non-existent mark (causing fast-import to die), or find\nthat a ref which you explicitly requested to be part of the export is\nsilently omitted from the stream.  (granted, these aren't fundamental\nissues; they're just bugs in fast-export that I seem to have been the\nfirst to find.)\n\n  -  filter-branch has a nice ability to rewrite only the last few\ncommits using a range specifier like HEAD~10..HEAD.  Trying the same\nwith fast-export will get you a history with only 10 commits, the\nfirst of which squashes all early history together.  Trying to\nduplicate the filter-branch behavior can be done, but it requires\nmultiple exports with different args and usage of --export-marks and\n--import-marks; it's cumbersome and somewhat non-obvious.\n\n  - some filters are difficult; e.g. if you want to mimick\nfilter-branch's --parent-filter, or BFG's --strip-blobs-with-ids, you\nrun into the issue that the fast-export stream doesn't provide the\noriginal sha1sums for commits or blobs and there's no easy way for you\nto associate it with the given mark.\n\n\nThose are the ones I know about.  It's possible that there are others.\n\nHope that helps or is at least interesting.\n\nElijah\n\n\n[1] https://public-inbox.org/git/51419b2c0904072035u1182b507o836a67ac308d32b9@mail.gmail.com/\n"},{"id":"362122","messageId":"91771D9B-166D-403F-BB20-7E574444BB3B@gmail.com","threadId":"49406","inReplyTo":"CABPp-BGL-3_nhZSpt0Bz0EVY-6-mcbgZMmx4YcXEfA_ZrTqFUw@mail.gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2018-10-31T19:15:59Z","receivedAt":"2018-10-31T19:16:05Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n\n> On Sep 24, 2018, at 7:24 PM, Elijah Newren <newren@gmail.com> wrote:\n> \n> On Sun, Sep 23, 2018 at 6:08 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n>> \n>> Hi,\n>> \n>> I recently had to purge files from large Git repos (many files, many commits).\n>> The usual recommendation is to use `git filter-branch --index-filter` to purge\n>> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n>> remove the `builtin` directory from git core). I realized that I can remove\n>> files *way* faster by exporting the repo, removing the file references,\n>> and then importing the repo (see Perl script below, it takes ~30sec to remove\n>> the `builtin` directory from git core). Do you see any problem with this\n>> approach?\n> \n> It looks like others have pointed you at other tools, and you're\n> already shifting to that route.  But I think it's a useful question to\n> answer more generally, so for those that are really curious...\n> \n> \n> The basic approach is fine, though if you try to extend it much you\n> can run into a few possible edge/corner cases (more on that below).\n> I've been using this basic approach for years and even created a\n> mini-python library[1] designed specifically to allow people to create\n> \"fast-filters\", used as\n>   git fast-export <options> | your-fast-filter | git fast-import <options>\n> \n> But that library didn't really take off; even I have rarely used it,\n> often opting for filter-branch despite its horrible performance or a\n> simple fast-export | long-sed-command | fast-import (with some extra\n> pre-checking to make sure the sed wouldn't unintentionally munge other\n> data).  BFG is great, as long as you're only interested in removing a\n> few big items, but otherwise doesn't seem very useful (to be fair,\n> it's very upfront about only wanting to solve that problem).\n> Recently, due to continuing questions on filter-branch and folks still\n> getting confused with it, I looked at existing tools, decided I didn't\n> think any quite fit, and started looking into converting\n> git_fast_filter into a filter-branch-like tool instead of just a\n> libary.  Found some bugs and missing features in fast-export along the\n> way (and have some patches I still need to send in).  But I kind of\n> got stuck -- if the tool is in python, will that limit adoption too\n> much?  It'd be kind of nice to have this tool in core git.  But I kind\n> of like leaving open the possibility of using it as a tool _or_ as a\n> library, the latter for the special cases where case-specific\n> programmatic filtering is needed.  But a developer-convenience library\n> makes almost no sense unless in a higher level language, such as\n> python.  I'm still trying to make up my mind about what I want (and\n> what others might want), and have been kind of blocking on that.  (If\n> others have opinions, I'm all ears.)\n\nThat library sounds like a very interesting idea. Unfortunately, the \nreferenced repo seems not to be available anymore:\n    git://gitorious.org/git_fast_filter/mainline.git\n\nI very much like Python. However, more recently I started to\nwrite Git tools in Perl as they work out of the box on every\nmachine with Git installed ... and I think Perl can be quite\nreadable if no shortcuts are used :-). \n\n\n> Anyway, the edge/corner cases you can watch out for:\n> \n>  - Signed tags are a problem; you may need to specify\n> --signed-tags=strip to fast-export\n> \n>  - References to other commits in your commit messages will now be\n> incorrect.  I think a good tool should either default to rewriting\n> commit ids in commit messages or at least have an option to do so\n> (BFG does this; filter-branch doesn't; fast-export format makes it\n> really hard for a filter based on it to do so)\n> \n>  - If the paths you remove are the only paths modified in a commit,\n> the commit can become empty.  If you're only filtering a few paths\n> out, this might be nothing more than a minor inconvenience for you.\n> However, if you're trying to prune directories (and perhaps several\n> toplevel ones), then it can be extremely annoying to have a new\n> history with the vast majority of all commits being empty.\n> (filter-branch has an option for this; BFG does not; tools based on\n> fast-export output can do it with sufficient effort).\n> \n>  - If you start pruning empty commits, you have to worry about\n> rewriting branches and tags to remaining parents.  This _might_ happen\n> for free depending on your history's structure and the fast-export\n> stream, but to be correct in general you will have to specify the new\n> commit for some branches or tags.\n> \n>  - If you start pruning empty commits, you have to decide whether to\n> allow pruning of merge commits.  Your first reaction might be to not\n> allow it, but if one parent and its entire history are all pruned,\n> then transforming the merge commit to a normal commit and then\n> considering whether it is empty and allowing it to be pruned is much\n> better.\n> \n>  - If you start pruning empty commits, you also have to worry about\n> history topology changing, beyond the all-ancestors-empty case above.\n> For example, the last non-empty commit in the ancestry of a merge on\n> both sides may be the same commit, making the merge-commit have the\n> same parent twice.  Should the duplicate parent be de-duped,\n> transforming the commit into a normal non-merge commit?  (I'd say yes\n> -- this commit is likely to be empty and prunable once you do so, but\n> I'm not sure everyone would agree with converting a merge commit to a\n> non-merge.)  Similarly, what if the rewritten parents of a merge have\n> one parent that is the direct ancestor of another?  Can the extra\n> unnecessary parent be removed as a parent?  (And again, such a commit\n> is likely to become empty and be prunable itself.)\n> \n>  - If you try to avoid the extra work involved with pruning empty\n> commits by passing path-specifiers as rev-list-args to fast-export,\n> and use the --tag-of-filtered-object=rewrite option if needed, then\n> depending on the topology you can hit any of three bugs: an outright\n> die() (despite the --tag-of-filtered-object=rewrite), a branch being\n> reset to a non-existent mark (causing fast-import to die), or find\n> that a ref which you explicitly requested to be part of the export is\n> silently omitted from the stream.  (granted, these aren't fundamental\n> issues; they're just bugs in fast-export that I seem to have been the\n> first to find.)\n> \n>  -  filter-branch has a nice ability to rewrite only the last few\n> commits using a range specifier like HEAD~10..HEAD.  Trying the same\n> with fast-export will get you a history with only 10 commits, the\n> first of which squashes all early history together.  Trying to\n> duplicate the filter-branch behavior can be done, but it requires\n> multiple exports with different args and usage of --export-marks and\n> --import-marks; it's cumbersome and somewhat non-obvious.\n> \n>  - some filters are difficult; e.g. if you want to mimick\n> filter-branch's --parent-filter, or BFG's --strip-blobs-with-ids, you\n> run into the issue that the fast-export stream doesn't provide the\n> original sha1sums for commits or blobs and there's no easy way for you\n> to associate it with the given mark.\n\nThanks a lot for these tips and tricks. I was aware of the empty commits\nbut the signed tags problem was not yet on my radar!\n\nThanks,\nLars"},{"id":"362155","messageId":"CABPp-BEefqYADr8SVvh6uFWkp96PDv7qfKK1c9O1WUnPy3wqrw@mail.gmail.com","threadId":"49406","inReplyTo":"91771D9B-166D-403F-BB20-7E574444BB3B@gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-01T07:12:48Z","receivedAt":"2018-11-01T07:13:02Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Oct 31, 2018 at 12:16 PM Lars Schneider\n<larsxschneider@gmail.com> wrote:\n> > On Sep 24, 2018, at 7:24 PM, Elijah Newren <newren@gmail.com> wrote:\n> > On Sun, Sep 23, 2018 at 6:08 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n> >>\n> >> Hi,\n> >>\n> >> I recently had to purge files from large Git repos (many files, many commits).\n> >> The usual recommendation is to use `git filter-branch --index-filter` to purge\n> >> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n> >> remove the `builtin` directory from git core). I realized that I can remove\n> >> files *way* faster by exporting the repo, removing the file references,\n> >> and then importing the repo (see Perl script below, it takes ~30sec to remove\n> >> the `builtin` directory from git core). Do you see any problem with this\n> >> approach?\n> >\n> > It looks like others have pointed you at other tools, and you're\n> > already shifting to that route.  But I think it's a useful question to\n> > answer more generally, so for those that are really curious...\n> >\n> >\n> > The basic approach is fine, though if you try to extend it much you\n> > can run into a few possible edge/corner cases (more on that below).\n> > I've been using this basic approach for years and even created a\n> > mini-python library[1] designed specifically to allow people to create\n> > \"fast-filters\", used as\n> >   git fast-export <options> | your-fast-filter | git fast-import <options>\n> >\n> > But that library didn't really take off; even I have rarely used it,\n> > often opting for filter-branch despite its horrible performance or a\n> > simple fast-export | long-sed-command | fast-import (with some extra\n> > pre-checking to make sure the sed wouldn't unintentionally munge other\n> > data).  BFG is great, as long as you're only interested in removing a\n> > few big items, but otherwise doesn't seem very useful (to be fair,\n> > it's very upfront about only wanting to solve that problem).\n> > Recently, due to continuing questions on filter-branch and folks still\n> > getting confused with it, I looked at existing tools, decided I didn't\n> > think any quite fit, and started looking into converting\n> > git_fast_filter into a filter-branch-like tool instead of just a\n> > libary.  Found some bugs and missing features in fast-export along the\n> > way (and have some patches I still need to send in).  But I kind of\n> > got stuck -- if the tool is in python, will that limit adoption too\n> > much?  It'd be kind of nice to have this tool in core git.  But I kind\n> > of like leaving open the possibility of using it as a tool _or_ as a\n> > library, the latter for the special cases where case-specific\n> > programmatic filtering is needed.  But a developer-convenience library\n> > makes almost no sense unless in a higher level language, such as\n> > python.  I'm still trying to make up my mind about what I want (and\n> > what others might want), and have been kind of blocking on that.  (If\n> > others have opinions, I'm all ears.)\n>\n> That library sounds like a very interesting idea. Unfortunately, the\n> referenced repo seems not to be available anymore:\n>     git://gitorious.org/git_fast_filter/mainline.git\n\nYeah, gitorious went down at a time when I was busy with enough other\nthings that I never bothered moving my repos to a new hosting site.\nSorry about that.\n\nI've got a copy locally, but I've been editing it heavily, without the\ntesting I should have in place, so I hesitate to point you at it right\nnow.  (Also, the old version failed to handle things like --no-data\noutput, which is important.)  I'll post an updated copy soon; feel\nfree to ping me in a week if you haven't heard anything yet.\n\n> I very much like Python. However, more recently I started to\n> write Git tools in Perl as they work out of the box on every\n> machine with Git installed ... and I think Perl can be quite\n> readable if no shortcuts are used :-).\n\nYeah, when portability matters, perl makes sense.  I thought about\nswitching it over, but I'm not sure I want to rewrite 1-2k lines of\ncode.  Especially since repo-filtering tools are kind of one-shot by\nnature, and only need to be done by one person of a team, on one\nspecific machine, and won't affect daily development thereafter.\n(Also, since I don't depend on any libraries and use only stuff from\nthe default python library, it ought to be relatively portable\nanyway.)\n"},{"id":"362900","messageId":"20181111062312.16342-1-newren@gmail.com","threadId":"49406","inReplyTo":"CABPp-BEefqYADr8SVvh6uFWkp96PDv7qfKK1c9O1WUnPy3wqrw@mail.gmail.com","subject":"[PATCH 00/10] fast export and import fixes and features","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:02Z","receivedAt":"2018-11-11T06:23:24Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"This is a series of ten patches representing two doc corrections, one\npedantic fix, three real bug fixes, one micro code refactor, and three\nnew features.  Each of these ten changes is relatively small in size.\nThese changes predominantly affect fast-export, but there's a couple\nsmall changes for fast-import as well.\n\nI could potentially split these patches up, but I'd just end up\nchaining them sequentially since otherwise there'd be lots of\nconflicts; having 10 different single patch series with lots of\ndependencies sounded like a bigger pain to me, but let me know if you\nwould prefer I split them up and how you suggest doing so.\n\nThese patches were driven by the needs of git-repo-filter[1], but most\nif not all of them should be independently useful.\n\nElijah Newren (10):\n  git-fast-import.txt: fix documentation for --quiet option\n  git-fast-export.txt: clarify misleading documentation about rev-list\n    args\n  fast-export: use value from correct enum\n  fast-export: avoid dying when filtering by paths and old tags exist\n  fast-export: move commit rewriting logic into a function for reuse\n  fast-export: when using paths, avoid corrupt stream with non-existent\n    mark\n  fast-export: ensure we export requested refs\n  fast-export: add --reference-excluded-parents option\n  fast-export: add a --show-original-ids option to show original names\n  fast-export: add --always-show-modify-after-rename\n\n Documentation/git-fast-export.txt |  33 ++++++-\n Documentation/git-fast-import.txt |   7 +-\n builtin/fast-export.c             | 156 +++++++++++++++++++++++-------\n fast-import.c                     |  17 ++++\n t/t9350-fast-export.sh            | 124 +++++++++++++++++++++++-\n 5 files changed, 293 insertions(+), 44 deletions(-)\n\n[1] https://github.com/newren/git-repo-filter if you're really\ncurious, but ***** IT HAS SEVERAL SHARP EDGES *****.  It isn't really\nready for review/testing/usage/announcing/etc; in fact, it's not quite\nWIP/RFC ready.  (Further, it's not clear if it should somehow become\npart of core git, should go into contrib, or just remain separate\nindefinitely.)  Anyway, please do not attempt to use it for anything\nreal yet.  I'll send out an email when I think it's closer to ready.\n\n-- \n2.19.1.866.g82735bcbde\n"},{"id":"362901","messageId":"20181111062312.16342-2-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 01/10] git-fast-import.txt: fix documentation for --quiet option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:03Z","receivedAt":"2018-11-11T06:23:24Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Signed-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-import.txt | 7 ++++---\n 1 file changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\nindex e81117d27f..7ab97745a6 100644\n--- a/Documentation/git-fast-import.txt\n+++ b/Documentation/git-fast-import.txt\n@@ -40,9 +40,10 @@ OPTIONS\n \tnot contain the old commit).\n \n --quiet::\n-\tDisable all non-fatal output, making fast-import silent when it\n-\tis successful.  This option disables the output shown by\n-\t--stats.\n+\tDisable the output shown by --stats, making fast-import usually\n+\tbe silent when it is successful.  However, if the import stream\n+\thas directives intended to show user output (e.g. `progress`\n+\tdirectives), the corresponding messages will still be shown.\n \n --stats::\n \tDisplay some basic statistics about the objects fast-import has\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362902","messageId":"20181111062312.16342-3-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 02/10] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:04Z","receivedAt":"2018-11-11T06:23:25Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Signed-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 3 ++-\n 1 file changed, 2 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex ce954be532..677510b7f7 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -119,7 +119,8 @@ marks the same across runs.\n \t'git rev-list', that specifies the specific objects and references\n \tto export.  For example, `master~10..master` causes the\n \tcurrent master reference to be exported along with all objects\n-\tadded since its 10th ancestor commit.\n+\tadded since its 10th ancestor commit and all files common to\n+\tmaster\\~9 and master~10.\n \n EXAMPLES\n --------\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362903","messageId":"20181111062312.16342-4-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 03/10] fast-export: use value from correct enum","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:05Z","receivedAt":"2018-11-11T06:23:25Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"ABORT and ERROR happen to have the same value, but come from differnt\nenums.  Use the one from the correct enum.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 456797c12a..1a299c2a21 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -752,7 +752,7 @@ static void handle_tag(const char *name, struct tag *tag)\n \ttagged_mark = get_object_mark(tagged);\n \tif (!tagged_mark) {\n \t\tswitch(tag_of_filtered_mode) {\n-\t\tcase ABORT:\n+\t\tcase ERROR:\n \t\t\tdie(\"tag %s tags unexported object; use \"\n \t\t\t    \"--tag-of-filtered-object=<mode> to handle it\",\n \t\t\t    oid_to_hex(&tag->object.oid));\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362904","messageId":"20181111062312.16342-5-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 04/10] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:06Z","receivedAt":"2018-11-11T06:23:28Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If --tag-of-filtered-object=rewrite is specified along with a set of\npaths to limit what is exported, then any tags pointing to old commits\nthat do not contain any of those specified paths cause problems.  Since\nthe old tagged commit is not exported, fast-export attempts to rewrite\nsuch tags to an ancestor commit which was exported.  If no such commit\nexists, then fast-export currently die()s.  Five years after the tag\nrewriting logic was added to fast-export (see commit 2d8ad4691921,\n\"fast-export: Add a --tag-of-filtered-object  option for newly dangling\ntags\", 2009-06-25), fast-import gained the ability to delete refs (see\ncommit 4ee1b225b99f, \"fast-import: add support to delete refs\",\n2014-04-20), so now we do have a valid option to rewrite the tag to.\nDelete these tags instead of dying.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  |  9 ++++++---\n t/t9350-fast-export.sh | 20 ++++++++++++++++++++\n 2 files changed, 26 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 1a299c2a21..89de9d6400 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -774,9 +774,12 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t\tbreak;\n \t\t\t\tif (!(p->object.flags & TREESAME))\n \t\t\t\t\tbreak;\n-\t\t\t\tif (!p->parents)\n-\t\t\t\t\tdie(\"can't find replacement commit for tag %s\",\n-\t\t\t\t\t     oid_to_hex(&tag->object.oid));\n+\t\t\t\tif (!p->parents) {\n+\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\t\tfree(buf);\n+\t\t\t\t\treturn;\n+\t\t\t\t}\n \t\t\t\tp = p->parents->item;\n \t\t\t}\n \t\t\ttagged_mark = get_object_mark(&p->object);\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 6a392e87bc..5bf21b4908 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -325,6 +325,26 @@ test_expect_success 'rewriting tag of filtered out object' '\n )\n '\n \n+test_expect_success 'rewrite tag predating pathspecs to nothing' '\n+\ttest_create_repo rewrite_tag_predating_pathspecs &&\n+\t(\n+\t\tcd rewrite_tag_predating_pathspecs &&\n+\n+\t\ttouch ignored &&\n+\t\tgit add ignored &&\n+\t\ttest_commit initial &&\n+\n+\t\tgit tag -a -m \"Some old tag\" v0.0.0.0.0.0.1 &&\n+\n+\t\techo foo >bar &&\n+\t\tgit add bar &&\n+\t\ttest_commit add-bar &&\n+\n+\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar >output &&\n+\t\tgrep -A 1 refs/tags/v0.0.0.0.0.0.1 output | grep -E ^from.0{40}\n+\t)\n+'\n+\n cat > limit-by-paths/expected << EOF\n blob\n mark :1\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362905","messageId":"20181111062312.16342-6-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 05/10] fast-export: move commit rewriting logic into a function for reuse","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:07Z","receivedAt":"2018-11-11T06:23:28Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Logic to replace a filtered commit with an unfiltered ancestor is useful\nelsewhere; put it into a function we can call.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 37 ++++++++++++++++++++++---------------\n 1 file changed, 22 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 89de9d6400..a3c044b0af 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -187,6 +187,22 @@ static int get_object_mark(struct object *object)\n \treturn ptr_to_mark(decoration);\n }\n \n+static struct commit *rewrite_commit(struct commit *p)\n+{\n+\tfor (;;) {\n+\t\tif (p->parents && p->parents->next)\n+\t\t\tbreak;\n+\t\tif (p->object.flags & UNINTERESTING)\n+\t\t\tbreak;\n+\t\tif (!(p->object.flags & TREESAME))\n+\t\t\tbreak;\n+\t\tif (!p->parents)\n+\t\t\treturn NULL;\n+\t\tp = p->parents->item;\n+\t}\n+\treturn p;\n+}\n+\n static void show_progress(void)\n {\n \tstatic int counter = 0;\n@@ -766,21 +782,12 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t    oid_to_hex(&tag->object.oid),\n \t\t\t\t    type_name(tagged->type));\n \t\t\t}\n-\t\t\tp = (struct commit *)tagged;\n-\t\t\tfor (;;) {\n-\t\t\t\tif (p->parents && p->parents->next)\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (p->object.flags & UNINTERESTING)\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (!(p->object.flags & TREESAME))\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (!p->parents) {\n-\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n-\t\t\t\t\tfree(buf);\n-\t\t\t\t\treturn;\n-\t\t\t\t}\n-\t\t\t\tp = p->parents->item;\n+\t\t\tp = rewrite_commit((struct commit *)tagged);\n+\t\t\tif (!p) {\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tfree(buf);\n+\t\t\t\treturn;\n \t\t\t}\n \t\t\ttagged_mark = get_object_mark(&p->object);\n \t\t}\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362906","messageId":"20181111062312.16342-7-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 06/10] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:08Z","receivedAt":"2018-11-11T06:23:29Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If file paths are specified to fast-export and multiple refs point to a\ncommit that does not touch any of the relevant file paths, then\nfast-export can hit problems.  fast-export has a list of additional refs\nthat it needs to explicitly set after exporting all blobs and commits,\nand when it tries to get_object_mark() on the relevant commit, it can\nget a mark of 0, i.e. \"not found\", because the commit in question did\nnot touch the relevant paths and thus was not exported.  Trying to\nimport a stream with a mark corresponding to an unexported object will\ncause fast-import to crash.\n\nAvoid this problem by taking the commit the ref points to and finding an\nancestor of it that was exported, and make the ref point to that commit\ninstead.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  | 13 ++++++++++++-\n t/t9350-fast-export.sh | 24 ++++++++++++++++++++++++\n 2 files changed, 36 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex a3c044b0af..5648a8ce9c 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -900,7 +900,18 @@ static void handle_tags_and_duplicates(void)\n \t\t\tif (anonymize)\n \t\t\t\tname = anonymize_refname(name);\n \t\t\t/* create refs pointing to already seen commits */\n-\t\t\tcommit = (struct commit *)object;\n+\t\t\tcommit = rewrite_commit((struct commit *)object);\n+\t\t\tif (!commit) {\n+\t\t\t\t/*\n+\t\t\t\t * Neither this object nor any of its\n+\t\t\t\t * ancestors touch any relevant paths, so\n+\t\t\t\t * it has been filtered to nothing.  Delete\n+\t\t\t\t * it.\n+\t\t\t\t */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n \t\t\t       get_object_mark(&commit->object));\n \t\t\tshow_progress();\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 5bf21b4908..dbb560c110 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -386,6 +386,30 @@ test_expect_success 'path limiting with import-marks does not lose unmodified fi\n \tgrep file0 actual\n '\n \n+test_expect_success 'avoid corrupt stream with non-existent mark' '\n+\ttest_create_repo avoid_non_existent_mark &&\n+\t(\n+\t\tcd avoid_non_existent_mark &&\n+\n+\t\ttouch important-path &&\n+\t\tgit add important-path &&\n+\t\ttest_commit initial &&\n+\n+\t\ttouch ignored &&\n+\t\tgit add ignored &&\n+\t\ttest_commit whatever &&\n+\n+\t\tgit branch A &&\n+\t\tgit branch B &&\n+\n+\t\techo foo >>important-path &&\n+\t\tgit add important-path &&\n+\t\ttest_commit more changes &&\n+\n+\t\tgit fast-export --all -- important-path | git fast-import --force\n+\t)\n+'\n+\n test_expect_success 'full-tree re-shows unmodified files'        '\n \tgit checkout -f simple &&\n \tgit fast-export --full-tree simple >actual &&\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362907","messageId":"20181111062312.16342-8-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 07/10] fast-export: ensure we export requested refs","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:09Z","receivedAt":"2018-11-11T06:23:33Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If file paths are specified to fast-export and a ref points to a commit\nthat does not touch any of the relevant paths, then that ref would\nsometimes fail to be exported.  (This depends on whether any ancestors\nof the commit which do touch the relevant paths would be exported with\nthat same ref name or a different ref name.)  To avoid this problem,\nput *all* specified refs into extra_refs to start, and then as we export\neach commit, remove the refname used in the 'commit $REFNAME' directive\nfrom extra_refs.  Then, in handle_tags_and_duplicates() we know which\nrefs actually do need a manual reset directive in order to be included.\n\nThis means that we do need some special handling for excluded refs; e.g.\nif someone runs\n   git fast-export ^master master\nthen they've asked for master to be exported, but they have also asked\nfor the commit which master points to and all of its history to be\nexcluded.  That logically means ref deletion.  Previously, such refs\nwere just silently omitted from being exported despite having been\nexplicitly requested for export.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\nNOTE: I was hoping the strmap API proposal would materialize, but I either\nmissed it or it hasn't shown up.  The usage of string_list in this patch\nwould be better replaced by what Peff suggested.\n\n builtin/fast-export.c  | 48 +++++++++++++++++++++++++++++++-----------\n t/t9350-fast-export.sh | 16 +++++++++++---\n 2 files changed, 49 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 5648a8ce9c..0d0bbd9445 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n+static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n static int anonymize;\n static struct revision_sources revision_sources;\n@@ -611,6 +612,7 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\t\texport_blob(&diff_queued_diff.queue[i]->two->oid);\n \n \trefname = *revision_sources_at(&revision_sources, commit);\n+\tstring_list_remove(&extra_refs, refname, 0);\n \tif (anonymize) {\n \t\trefname = anonymize_refname(refname);\n \t\tanonymize_ident_line(&committer, &committer_end);\n@@ -814,7 +816,7 @@ static struct commit *get_commit(struct rev_cmdline_entry *e, char *full_name)\n \t\t/* handle nested tags */\n \t\twhile (tag && tag->object.type == OBJ_TAG) {\n \t\t\tparse_object(the_repository, &tag->object.oid);\n-\t\t\tstring_list_append(&extra_refs, full_name)->util = tag;\n+\t\t\tstring_list_append(&tag_refs, full_name)->util = tag;\n \t\t\ttag = (struct tag *)tag->tagged;\n \t\t}\n \t\tif (!tag)\n@@ -873,25 +875,30 @@ static void get_tags_and_duplicates(struct rev_cmdline_info *info)\n \t\t}\n \n \t\t/*\n-\t\t * This ref will not be updated through a commit, lets make\n-\t\t * sure it gets properly updated eventually.\n+\t\t * Make sure this ref gets properly updated eventually, whether\n+\t\t * through a commit or manually at the end.\n \t\t */\n-\t\tif (*revision_sources_at(&revision_sources, commit) ||\n-\t\t    commit->object.flags & SHOWN)\n+\t\tif (e->item->type != OBJ_TAG)\n \t\t\tstring_list_append(&extra_refs, full_name)->util = commit;\n+\n \t\tif (!*revision_sources_at(&revision_sources, commit))\n \t\t\t*revision_sources_at(&revision_sources, commit) = full_name;\n \t}\n+\n+\tstring_list_sort(&extra_refs);\n+\tstring_list_remove_duplicates(&extra_refs, 0);\n }\n \n-static void handle_tags_and_duplicates(void)\n+static void handle_tags_and_duplicates(struct string_list *extras)\n {\n \tstruct commit *commit;\n \tint i;\n \n-\tfor (i = extra_refs.nr - 1; i >= 0; i--) {\n-\t\tconst char *name = extra_refs.items[i].string;\n-\t\tstruct object *object = extra_refs.items[i].util;\n+\tfor (i = extras->nr - 1; i >= 0; i--) {\n+\t\tconst char *name = extras->items[i].string;\n+\t\tstruct object *object = extras->items[i].util;\n+\t\tint mark;\n+\n \t\tswitch (object->type) {\n \t\tcase OBJ_TAG:\n \t\t\thandle_tag(name, (struct tag *)object);\n@@ -912,8 +919,24 @@ static void handle_tags_and_duplicates(void)\n \t\t\t\t       name, sha1_to_hex(null_sha1));\n \t\t\t\tcontinue;\n \t\t\t}\n-\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n-\t\t\t       get_object_mark(&commit->object));\n+\n+\t\t\tmark = get_object_mark(&commit->object);\n+\t\t\tif (!mark) {\n+\t\t\t\t/*\n+\t\t\t\t * Getting here means we have a commit which\n+\t\t\t\t * was excluded by a negative refspec (e.g.\n+\t\t\t\t * fast-export ^master master).  If the user\n+\t\t\t\t * wants the branch exported but every commit\n+\t\t\t\t * in its history to be deleted, that sounds\n+\t\t\t\t * like a ref deletion to me.\n+\t\t\t\t */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\n+\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name, mark\n+\t\t\t       );\n \t\t\tshow_progress();\n \t\t\tbreak;\n \t\t}\n@@ -1101,7 +1124,8 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\t}\n \t}\n \n-\thandle_tags_and_duplicates();\n+\thandle_tags_and_duplicates(&extra_refs);\n+\thandle_tags_and_duplicates(&tag_refs);\n \thandle_deletes();\n \n \tif (export_filename && lastimportid != last_idnum)\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex dbb560c110..a0c93f2212 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -552,10 +552,20 @@ test_expect_success 'use refspec' '\n \ttest_cmp expected actual\n '\n \n-test_expect_success 'delete refspec' '\n+test_expect_success 'delete ref because entire history excluded' '\n \tgit branch to-delete &&\n-\tgit fast-export --refspec :refs/heads/to-delete to-delete ^to-delete > actual &&\n-\tcat > expected <<-EOF &&\n+\tgit fast-export to-delete ^to-delete >actual &&\n+\tcat >expected <<-EOF &&\n+\treset refs/heads/to-delete\n+\tfrom 0000000000000000000000000000000000000000\n+\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'delete refspec' '\n+\tgit fast-export --refspec :refs/heads/to-delete >actual &&\n+\tcat >expected <<-EOF &&\n \treset refs/heads/to-delete\n \tfrom 0000000000000000000000000000000000000000\n \n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362908","messageId":"20181111062312.16342-9-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 08/10] fast-export: add --reference-excluded-parents option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:10Z","receivedAt":"2018-11-11T06:23:33Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"git filter-branch has a nifty feature allowing you to rewrite, e.g. just\nthe last 8 commits of a linear history\n  git filter-branch $OPTIONS HEAD~8..HEAD\n\nIf you try the same with git fast-export, you instead get a history of\nonly 8 commits, with HEAD~7 being rewritten into a root commit.  There\nare two alternatives:\n\n  1) Don't use the negative revision specification, and when you're\n     filtering the output to make modifications to the last 8 commits,\n     just be careful to not modify any earlier commits somehow.\n\n  2) First run 'git fast-export --export-marks=somefile HEAD~8', then\n     run 'git fast-export --import-marks=somefile HEAD~8..HEAD'.\n\nBoth are more error prone than I'd like (the first for obvious reasons;\nwith the second option I have sometimes accidentally included too many\nrevisions in the first command and then found that the corresponding\nextra revisions were not exported by the second command and thus were\nnot modified as I expected).  Also, both are poor from a performance\nperspective.\n\nAdd a new --reference-excluded-parents option which will cause\nfast-export to refer to commits outside the specified rev-list-args\nrange by their sha1sum.  Such a stream will only be useful in a\nrepository which already contains the necessary commits (much like the\nrestriction imposed when using --no-data).\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 16 ++++++++++--\n builtin/fast-export.c             | 42 +++++++++++++++++++++++--------\n t/t9350-fast-export.sh            | 11 ++++++++\n 3 files changed, 57 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex 677510b7f7..2916096bdd 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -110,6 +110,17 @@ marks the same across runs.\n \tthe shape of the history and stored tree.  See the section on\n \t`ANONYMIZING` below.\n \n+--reference-excluded-parents::\n+\tBy default, running a command such as `git fast-export\n+\tmaster~5..master` will not include the commit master\\~5 and\n+\twill make master\\~4 no longer have master\\~5 as a parent (though\n+\tboth the old master\\~4 and new master~4 will have all the same\n+\tfiles).  Use --reference-excluded-parents to instead have the\n+\tthe stream refer to commits in the excluded range of history\n+\tby their sha1sum.  Note that the resulting stream can only be\n+\tused by a repository which already contains the necessary\n+\tparent commits.\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\n@@ -119,8 +130,9 @@ marks the same across runs.\n \t'git rev-list', that specifies the specific objects and references\n \tto export.  For example, `master~10..master` causes the\n \tcurrent master reference to be exported along with all objects\n-\tadded since its 10th ancestor commit and all files common to\n-\tmaster\\~9 and master~10.\n+\tadded since its 10th ancestor commit and (unless the\n+\t--reference-excluded-parents option is specified) all files\n+\tcommon to master\\~9 and master~10.\n \n EXAMPLES\n --------\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 0d0bbd9445..ea9c5b1c00 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -37,6 +37,7 @@ static int fake_missing_tagger;\n static int use_done_feature;\n static int no_data;\n static int full_tree;\n+static int reference_excluded_commits;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n@@ -596,7 +597,8 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\tmessage += 2;\n \n \tif (commit->parents &&\n-\t    get_object_mark(&commit->parents->item->object) != 0 &&\n+\t    (get_object_mark(&commit->parents->item->object) != 0 ||\n+\t     reference_excluded_commits) &&\n \t    !full_tree) {\n \t\tparse_commit_or_die(commit->parents->item);\n \t\tdiff_tree_oid(get_commit_tree_oid(commit->parents->item),\n@@ -638,13 +640,21 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \tunuse_commit_buffer(commit, commit_buffer);\n \n \tfor (i = 0, p = commit->parents; p; p = p->next) {\n-\t\tint mark = get_object_mark(&p->item->object);\n-\t\tif (!mark)\n+\t\tstruct object *obj = &p->item->object;\n+\t\tint mark = get_object_mark(obj);\n+\n+\t\tif (!mark && !reference_excluded_commits)\n \t\t\tcontinue;\n \t\tif (i == 0)\n-\t\t\tprintf(\"from :%d\\n\", mark);\n+\t\t\tprintf(\"from \");\n+\t\telse\n+\t\t\tprintf(\"merge \");\n+\t\tif (mark)\n+\t\t\tprintf(\":%d\\n\", mark);\n \t\telse\n-\t\t\tprintf(\"merge :%d\\n\", mark);\n+\t\t\tprintf(\"%s\\n\", sha1_to_hex(anonymize ?\n+\t\t\t\t\t\t   anonymize_sha1(&obj->oid) :\n+\t\t\t\t\t\t   obj->oid.hash));\n \t\ti++;\n \t}\n \n@@ -925,13 +935,22 @@ static void handle_tags_and_duplicates(struct string_list *extras)\n \t\t\t\t/*\n \t\t\t\t * Getting here means we have a commit which\n \t\t\t\t * was excluded by a negative refspec (e.g.\n-\t\t\t\t * fast-export ^master master).  If the user\n+\t\t\t\t * fast-export ^master master).  If we are\n+\t\t\t\t * referencing excluded commits, set the ref\n+\t\t\t\t * to the exact commit.  Otherwise, the user\n \t\t\t\t * wants the branch exported but every commit\n-\t\t\t\t * in its history to be deleted, that sounds\n-\t\t\t\t * like a ref deletion to me.\n+\t\t\t\t * in its history to be deleted, which basically\n+\t\t\t\t * just means deletion of the ref.\n \t\t\t\t */\n-\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tif (!reference_excluded_commits) {\n+\t\t\t\t\t/* delete the ref */\n+\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\t/* set ref to commit using oid, not mark */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\", name,\n+\t\t\t\t       sha1_to_hex(commit->object.oid.hash));\n \t\t\t\tcontinue;\n \t\t\t}\n \n@@ -1068,6 +1087,9 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\tOPT_STRING_LIST(0, \"refspec\", &refspecs_list, N_(\"refspec\"),\n \t\t\t     N_(\"Apply refspec to exported refs\")),\n \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n+\t\tOPT_BOOL(0, \"reference-excluded-parents\",\n+\t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n+\n \t\tOPT_END()\n \t};\n \ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex a0c93f2212..c2f40d6a40 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -66,6 +66,17 @@ test_expect_success 'fast-export master~2..master' '\n \n '\n \n+test_expect_success 'fast-export --reference-excluded-parents master~2..master' '\n+\n+\tgit fast-export --reference-excluded-parents master~2..master >actual &&\n+\tgrep commit.refs/heads/master actual >commit-count &&\n+\ttest_line_count = 2 commit-count &&\n+\tsed \"s/master/rewrite/\" actual |\n+\t\t(cd new &&\n+\t\t git fast-import &&\n+\t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n+'\n+\n test_expect_success 'iso-8859-1' '\n \n \tgit config i18n.commitencoding ISO8859-1 &&\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362909","messageId":"20181111062312.16342-10-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 09/10] fast-export: add a --show-original-ids option to show original names","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:11Z","receivedAt":"2018-11-11T06:23:33Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Knowing the original names (hashes) of commits, blobs, and tags can\nsometimes enable post-filtering that would otherwise be difficult or\nimpossible.  In particular, the desire to rewrite commit messages which\nrefer to other prior commits (on top of whatever other filtering is\nbeing done) is very difficult without knowing the original names of each\ncommit.\n\nThis commit teaches a new --show-original-ids option to fast-export\nwhich will make it add a 'originally <hash>' line to blob, commits, and\ntags.  It also teaches fast-import to parse (and ignore) such lines.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt |  7 +++++++\n builtin/fast-export.c             | 20 +++++++++++++++-----\n fast-import.c                     | 17 +++++++++++++++++\n t/t9350-fast-export.sh            | 17 +++++++++++++++++\n 4 files changed, 56 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex 2916096bdd..4e40f0b99a 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -121,6 +121,13 @@ marks the same across runs.\n \tused by a repository which already contains the necessary\n \tparent commits.\n \n+--show-original-ids::\n+\tAdd an extra directive to the output for commits and blobs,\n+\t`originally <SHA1SUM>`.  While such directives will likely be\n+\tignored by importers such as git-fast-import, it may be useful\n+\tfor intermediary filters (e.g. for rewriting commit messages\n+\twhich refer to older commits, or for stripping blobs by id).\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex ea9c5b1c00..cc01dcc90c 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static int reference_excluded_commits;\n+static int show_original_ids;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n@@ -271,7 +272,10 @@ static void export_blob(const struct object_id *oid)\n \n \tmark_next_object(object);\n \n-\tprintf(\"blob\\nmark :%\"PRIu32\"\\ndata %lu\\n\", last_idnum, size);\n+\tprintf(\"blob\\nmark :%\"PRIu32\"\\n\", last_idnum);\n+\tif (show_original_ids)\n+\t\tprintf(\"originally %s\\n\", oid_to_hex(oid));\n+\tprintf(\"data %lu\\n\", size);\n \tif (size && fwrite(buf, size, 1, stdout) != 1)\n \t\tdie_errno(\"could not write blob '%s'\", oid_to_hex(oid));\n \tprintf(\"\\n\");\n@@ -628,8 +632,10 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\treencoded = reencode_string(message, \"UTF-8\", encoding);\n \tif (!commit->parents)\n \t\tprintf(\"reset %s\\n\", refname);\n-\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n%.*s\\n%.*s\\ndata %u\\n%s\",\n-\t       refname, last_idnum,\n+\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n\", refname, last_idnum);\n+\tif (show_original_ids)\n+\t\tprintf(\"originally %s\\n\", oid_to_hex(&commit->object.oid));\n+\tprintf(\"%.*s\\n%.*s\\ndata %u\\n%s\",\n \t       (int)(author_end - author), author,\n \t       (int)(committer_end - committer), committer,\n \t       (unsigned)(reencoded\n@@ -807,8 +813,10 @@ static void handle_tag(const char *name, struct tag *tag)\n \n \tif (starts_with(name, \"refs/tags/\"))\n \t\tname += 10;\n-\tprintf(\"tag %s\\nfrom :%d\\n%.*s%sdata %d\\n%.*s\\n\",\n-\t       name, tagged_mark,\n+\tprintf(\"tag %s\\nfrom :%d\\n\", name, tagged_mark);\n+\tif (show_original_ids)\n+\t\tprintf(\"originally %s\\n\", oid_to_hex(&tag->object.oid));\n+\tprintf(\"%.*s%sdata %d\\n%.*s\\n\",\n \t       (int)(tagger_end - tagger), tagger,\n \t       tagger == tagger_end ? \"\" : \"\\n\",\n \t       (int)message_size, (int)message_size, message ? message : \"\");\n@@ -1089,6 +1097,8 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n \t\tOPT_BOOL(0, \"reference-excluded-parents\",\n \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n+\t\tOPT_BOOL(0, \"show-original-ids\", &show_original_ids,\n+\t\t\t    N_(\"Show original sha1sums of blobs/commits\")),\n \n \t\tOPT_END()\n \t};\ndiff --git a/fast-import.c b/fast-import.c\nindex 95600c78e0..232b6a8b8d 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -14,11 +14,13 @@ Format of STDIN stream:\n \n   new_blob ::= 'blob' lf\n     mark?\n+    originally?\n     file_content;\n   file_content ::= data;\n \n   new_commit ::= 'commit' sp ref_str lf\n     mark?\n+    originally?\n     ('author' (sp name)? sp '<' email '>' sp when lf)?\n     'committer' (sp name)? sp '<' email '>' sp when lf\n     commit_msg\n@@ -49,6 +51,7 @@ Format of STDIN stream:\n \n   new_tag ::= 'tag' sp tag_str lf\n     'from' sp commit-ish lf\n+    originally?\n     ('tagger' (sp name)? sp '<' email '>' sp when lf)?\n     tag_msg;\n   tag_msg ::= data;\n@@ -73,6 +76,8 @@ Format of STDIN stream:\n   data ::= (delimited_data | exact_data)\n     lf?;\n \n+  originally ::= 'originally' sp not_lf+ lf\n+\n     # note: delim may be any string but must not contain lf.\n     # data_line may contain any data but must not be exactly\n     # delim.\n@@ -1968,6 +1973,13 @@ static void parse_mark(void)\n \t\tnext_mark = 0;\n }\n \n+static void parse_original_identifier(void)\n+{\n+\tconst char *v;\n+\tif (skip_prefix(command_buf.buf, \"originally \", &v))\n+\t\tread_next_command();\n+}\n+\n static int parse_data(struct strbuf *sb, uintmax_t limit, uintmax_t *len_res)\n {\n \tconst char *data;\n@@ -2110,6 +2122,7 @@ static void parse_new_blob(void)\n {\n \tread_next_command();\n \tparse_mark();\n+\tparse_original_identifier();\n \tparse_and_store_blob(&last_blob, NULL, next_mark);\n }\n \n@@ -2733,6 +2746,7 @@ static void parse_new_commit(const char *arg)\n \n \tread_next_command();\n \tparse_mark();\n+\tparse_original_identifier();\n \tif (skip_prefix(command_buf.buf, \"author \", &v)) {\n \t\tauthor = parse_ident(v);\n \t\tread_next_command();\n@@ -2865,6 +2879,9 @@ static void parse_new_tag(const char *arg)\n \t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n \tread_next_command();\n \n+\t/* originally ... */\n+\tparse_original_identifier();\n+\n \t/* tagger ... */\n \tif (skip_prefix(command_buf.buf, \"tagger \", &v)) {\n \t\ttagger = parse_ident(v);\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex c2f40d6a40..5ad6669910 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -77,6 +77,23 @@ test_expect_success 'fast-export --reference-excluded-parents master~2..master'\n \t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n '\n \n+test_expect_success 'fast-export --show-original-ids' '\n+\n+\tgit fast-export --show-original-ids master >output &&\n+\tgrep ^originally output| sed -e s/^originally.// | sort >actual &&\n+\tgit rev-list --objects master muss >objects-and-names &&\n+\tawk \"{print \\$1}\" objects-and-names | sort >commits-trees-blobs &&\n+\tcomm -23 actual commits-trees-blobs >unfound &&\n+\ttest_must_be_empty unfound\n+'\n+\n+test_expect_success 'fast-export --show-original-ids | git fast-import' '\n+\n+\tgit fast-export --show-original-ids master muss | git fast-import --quiet &&\n+\ttest $MASTER = $(git rev-parse --verify refs/heads/master) &&\n+\ttest $MUSS = $(git rev-parse --verify refs/tags/muss)\n+'\n+\n test_expect_success 'iso-8859-1' '\n \n \tgit config i18n.commitencoding ISO8859-1 &&\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362910","messageId":"20181111062312.16342-11-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T06:23:12Z","receivedAt":"2018-11-11T06:23:35Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"fast-export output is traditionally used as an input to a fast-import\nprogram, but it is also useful to help gather statistics about the\nhistory of a repository (particularly when --no-data is also passed).\nFor example, two of the types of information we may want to collect\ncould include:\n  1) general information about renames that have occurred\n  2) what the biggest objects in a repository are and what names\n     they appear under.\n\nThe first bit of information can be gathered by just passing -M to\nfast-export.  The second piece of information can partially be gotten\nfrom running\n    git cat-file --batch-check --batch-all-objects\nHowever, that only shows what the biggest objects in the repository are\nand their sizes, not what names those objects appear as or what commits\nthey were introduced in.  We can get that information from fast-export,\nbut when we only see\n    R oldname newname\ninstead of\n    R oldname newname\n    M 100644 $SHA1 newname\nthen it makes the job more difficult.  Add an option which allows us to\nforce the latter output even when commits have exact renames of files.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 11 ++++++++++\n builtin/fast-export.c             |  7 +++++-\n t/t9350-fast-export.sh            | 36 +++++++++++++++++++++++++++++++\n 3 files changed, 53 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex 4e40f0b99a..946a5aee1f 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -128,6 +128,17 @@ marks the same across runs.\n \tfor intermediary filters (e.g. for rewriting commit messages\n \twhich refer to older commits, or for stripping blobs by id).\n \n+--always-show-modify-after-rename::\n+\tWhen a rename is detected, fast-export normally issues both a\n+\t'R' (rename) and a 'M' (modify) directive.  However, if the\n+\tcontents of the old and new filename match exactly, it will\n+\tonly issue the rename directive.  Use this flag to have it\n+\talways issue the modify directive after the rename, which may\n+\tbe useful for tools which are using the fast-export stream as\n+\ta mechanism for gathering statistics about a repository.  Note\n+\tthat this option only has effect when rename detection is\n+\tactive (see the -M option).\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex cc01dcc90c..db606d1fd0 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static int reference_excluded_commits;\n+static int always_show_modify_after_rename;\n static int show_original_ids;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n@@ -407,7 +408,8 @@ static void show_filemodify(struct diff_queue_struct *q,\n \t\t\t\tputchar('\\n');\n \n \t\t\t\tif (oideq(&ospec->oid, &spec->oid) &&\n-\t\t\t\t    ospec->mode == spec->mode)\n+\t\t\t\t    ospec->mode == spec->mode &&\n+\t\t\t\t    !always_show_modify_after_rename)\n \t\t\t\t\tbreak;\n \t\t\t}\n \t\t\t/* fallthrough */\n@@ -1099,6 +1101,9 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n \t\tOPT_BOOL(0, \"show-original-ids\", &show_original_ids,\n \t\t\t    N_(\"Show original sha1sums of blobs/commits\")),\n+\t\tOPT_BOOL(0, \"always-show-modify-after-rename\",\n+\t\t\t    &always_show_modify_after_rename,\n+\t\t\t N_(\"Always provide 'M' directive after 'R'\")),\n \n \t\tOPT_END()\n \t};\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 5ad6669910..d0c30672ac 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -638,4 +638,40 @@ test_expect_success 'merge commit gets exported with --import-marks' '\n \t)\n '\n \n+test_expect_success 'rename detection and --always-show-modify-after-rename' '\n+\ttest_create_repo renames &&\n+\t(\n+\t\tcd renames &&\n+\t\ttest_seq 0  9  >single_digit &&\n+\t\ttest_seq 10 98 >double_digit &&\n+\t\tgit add . &&\n+\t\tgit commit -m initial &&\n+\n+\t\techo 99 >>double_digit &&\n+\t\tgit mv single_digit single-digit &&\n+\t\tgit mv double_digit double-digit &&\n+\t\tgit add double-digit &&\n+\t\tgit commit -m renames &&\n+\n+\t\t# First, check normal fast-export -M output\n+\t\tgit fast-export -M --no-data master >out &&\n+\n+\t\tgrep double-digit out >out2 &&\n+\t\ttest_line_count = 2 out2 &&\n+\n+\t\tgrep single-digit out >out2 &&\n+\t\ttest_line_count = 1 out2 &&\n+\n+\t\t# Now, test with --always-show-modify-after-rename; should\n+\t\t# have an extra \"M\" directive for \"single-digit\".\n+\t\tgit fast-export -M --no-data --always-show-modify-after-rename master >out &&\n+\n+\t\tgrep double-digit out >out2 &&\n+\t\ttest_line_count = 2 out2 &&\n+\n+\t\tgrep single-digit out >out2 &&\n+\t\ttest_line_count = 2 out2\n+\t)\n+'\n+\n test_done\n-- \n2.19.1.866.g82735bcbde\n\n"},{"id":"362911","messageId":"20181111063359.GA30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-2-newren@gmail.com","subject":"Re: [PATCH 01/10] git-fast-import.txt: fix documentation for --quiet option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T06:33:59Z","receivedAt":"2018-11-11T06:34:06Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:03PM -0800, Elijah Newren wrote:\n\n> Signed-off-by: Elijah Newren <newren@gmail.com>\n> ---\n>  Documentation/git-fast-import.txt | 7 ++++---\n>  1 file changed, 4 insertions(+), 3 deletions(-)\n> \n> diff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\n> index e81117d27f..7ab97745a6 100644\n> --- a/Documentation/git-fast-import.txt\n> +++ b/Documentation/git-fast-import.txt\n> @@ -40,9 +40,10 @@ OPTIONS\n>  \tnot contain the old commit).\n>  \n>  --quiet::\n> -\tDisable all non-fatal output, making fast-import silent when it\n> -\tis successful.  This option disables the output shown by\n> -\t--stats.\n> +\tDisable the output shown by --stats, making fast-import usually\n> +\tbe silent when it is successful.  However, if the import stream\n> +\thas directives intended to show user output (e.g. `progress`\n> +\tdirectives), the corresponding messages will still be shown.\n\nMakes sense. I think one could argue that it should disable those\nmessages, too, but probably the right answer is that the export side\nshould be told to be `--quiet` as well.\n\n-Peff\n"},{"id":"362912","messageId":"20181111063601.GB30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-3-newren@gmail.com","subject":"Re: [PATCH 02/10] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T06:36:01Z","receivedAt":"2018-11-11T06:36:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:04PM -0800, Elijah Newren wrote:\n\n> Signed-off-by: Elijah Newren <newren@gmail.com>\n> ---\n>  Documentation/git-fast-export.txt | 3 ++-\n>  1 file changed, 2 insertions(+), 1 deletion(-)\n> \n> diff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\n> index ce954be532..677510b7f7 100644\n> --- a/Documentation/git-fast-export.txt\n> +++ b/Documentation/git-fast-export.txt\n> @@ -119,7 +119,8 @@ marks the same across runs.\n>  \t'git rev-list', that specifies the specific objects and references\n>  \tto export.  For example, `master~10..master` causes the\n>  \tcurrent master reference to be exported along with all objects\n> -\tadded since its 10th ancestor commit.\n> +\tadded since its 10th ancestor commit and all files common to\n> +\tmaster\\~9 and master~10.\n\nDo you need to backslash the second tilde?  Maybe `master~9` and\n`master~10` instead of escaping?\n\nI'm not sure what this is trying to say. I guess that we'd always show\nall of the blobs necessary to reconstruct the first non-negative commit\n(i.e., `master~9` here)?\n\n-Peff\n"},{"id":"362913","messageId":"20181111063636.GC30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-4-newren@gmail.com","subject":"Re: [PATCH 03/10] fast-export: use value from correct enum","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T06:36:37Z","receivedAt":"2018-11-11T06:36:40Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:05PM -0800, Elijah Newren wrote:\n\n> ABORT and ERROR happen to have the same value, but come from differnt\n> enums.  Use the one from the correct enum.\n\nYikes. :)\n\nThis is a good argument for naming these SIGNED_TAG_ABORT, etc. But this\nis obviously an improvement in the meantime.\n\n-Peff\n"},{"id":"362914","messageId":"20181111064442.GD30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-5-newren@gmail.com","subject":"Re: [PATCH 04/10] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T06:44:43Z","receivedAt":"2018-11-11T06:44:47Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:06PM -0800, Elijah Newren wrote:\n\n> If --tag-of-filtered-object=rewrite is specified along with a set of\n> paths to limit what is exported, then any tags pointing to old commits\n> that do not contain any of those specified paths cause problems.  Since\n> the old tagged commit is not exported, fast-export attempts to rewrite\n> such tags to an ancestor commit which was exported.  If no such commit\n> exists, then fast-export currently die()s.  Five years after the tag\n> rewriting logic was added to fast-export (see commit 2d8ad4691921,\n> \"fast-export: Add a --tag-of-filtered-object  option for newly dangling\n> tags\", 2009-06-25), fast-import gained the ability to delete refs (see\n> commit 4ee1b225b99f, \"fast-import: add support to delete refs\",\n> 2014-04-20), so now we do have a valid option to rewrite the tag to.\n> Delete these tags instead of dying.\n\nHmm. That's the right thing to do if we're considering the export to be\nan independent unit. But what if I'm just rewriting a portion of history\nlike:\n\n  git fast-export HEAD~5..HEAD | some_filter | git fast-import\n\n? If I have a tag pointing to HEAD~10, will this delete that? Ideally I\nthink it would be left alone.\n\n> +test_expect_success 'rewrite tag predating pathspecs to nothing' '\n> +\ttest_create_repo rewrite_tag_predating_pathspecs &&\n> +\t(\n> +\t\tcd rewrite_tag_predating_pathspecs &&\n> +\n> +\t\ttouch ignored &&\n\nWe usually prefer \">ignored\" to create an empty file rather than\n\"touch\".\n\n> +\t\tgit add ignored &&\n> +\t\ttest_commit initial &&\n\nWhat do we need this \"ignored\" for? test_commit should create a file\n\"initial.t\".\n\n> +\t\techo foo >bar &&\n> +\t\tgit add bar &&\n> +\t\ttest_commit add-bar &&\n\nLikewise, \"test_commit bar\" should work by itself (though note the\nfilename is \"bar.t\" in your fast-export command).\n\n> +\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar >output &&\n> +\t\tgrep -A 1 refs/tags/v0.0.0.0.0.0.1 output | grep -E ^from.0{40}\n\nI don't think \"grep -A\" is portable (and we don't seem to otherwise use\nit). You can probably do something similar with sed.\n\nUse $ZERO_OID instead of hard-coding 40, which future-proofs for the\nhash transition (though I suppose the hash is not likely to get\n_shorter_ ;) ).\n\n-Peff\n"},{"id":"362915","messageId":"20181111064751.GE30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-6-newren@gmail.com","subject":"Re: [PATCH 05/10] fast-export: move commit rewriting logic into a function for reuse","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T06:47:51Z","receivedAt":"2018-11-11T06:47:55Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:07PM -0800, Elijah Newren wrote:\n\n> Logic to replace a filtered commit with an unfiltered ancestor is useful\n> elsewhere; put it into a function we can call.\n\nOK. I had to stare at it for a minute to make sure there was not an\nedge case with looking at \"p\" versus \"p->parents\", but I think it is a\nfaithful conversion.\n\n-Peff\n"},{"id":"362916","messageId":"20181111065338.GF30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-7-newren@gmail.com","subject":"Re: [PATCH 06/10] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T06:53:38Z","receivedAt":"2018-11-11T06:53:43Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:08PM -0800, Elijah Newren wrote:\n\n> If file paths are specified to fast-export and multiple refs point to a\n> commit that does not touch any of the relevant file paths, then\n> fast-export can hit problems.  fast-export has a list of additional refs\n> that it needs to explicitly set after exporting all blobs and commits,\n> and when it tries to get_object_mark() on the relevant commit, it can\n> get a mark of 0, i.e. \"not found\", because the commit in question did\n> not touch the relevant paths and thus was not exported.  Trying to\n> import a stream with a mark corresponding to an unexported object will\n> cause fast-import to crash.\n> \n> Avoid this problem by taking the commit the ref points to and finding an\n> ancestor of it that was exported, and make the ref point to that commit\n> instead.\n\nAs with the earlier tag commit, I wonder if this might depend on the\ncontext in which you're using fast-export. I suppose that if you did not\nfeed the ref on the command line that we would not be dealing with it at\nall (and maybe that is the answer to my question about the tag thing,\ntoo).\n\nIt does seem funny that the behavior for the earlier case (bounded\ncommits) and this case (skipping some commits) are different. Would you\never want to keep walking backwards to find an ancestor in the earlier\ncase? Or vice versa, would you ever want to simply delete a tag in a\ncase like this one?\n\nI'm not sure sure, but I suspect you may have thought about it a lot\nharder than I have. :)\n\n> diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n> index a3c044b0af..5648a8ce9c 100644\n> --- a/builtin/fast-export.c\n> +++ b/builtin/fast-export.c\n> @@ -900,7 +900,18 @@ static void handle_tags_and_duplicates(void)\n>  \t\t\tif (anonymize)\n>  \t\t\t\tname = anonymize_refname(name);\n>  \t\t\t/* create refs pointing to already seen commits */\n> -\t\t\tcommit = (struct commit *)object;\n> +\t\t\tcommit = rewrite_commit((struct commit *)object);\n> +\t\t\tif (!commit) {\n> +\t\t\t\t/*\n> +\t\t\t\t * Neither this object nor any of its\n> +\t\t\t\t * ancestors touch any relevant paths, so\n> +\t\t\t\t * it has been filtered to nothing.  Delete\n> +\t\t\t\t * it.\n> +\t\t\t\t */\n> +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n> +\t\t\t\t       name, sha1_to_hex(null_sha1));\n> +\t\t\t\tcontinue;\n> +\t\t\t}\n\nThis hunk makes sense.\n\n> --- a/t/t9350-fast-export.sh\n> +++ b/t/t9350-fast-export.sh\n> @@ -386,6 +386,30 @@ test_expect_success 'path limiting with import-marks does not lose unmodified fi\n>  \tgrep file0 actual\n>  '\n>  \n> +test_expect_success 'avoid corrupt stream with non-existent mark' '\n> +\ttest_create_repo avoid_non_existent_mark &&\n> +\t(\n> +\t\tcd avoid_non_existent_mark &&\n> +\n> +\t\ttouch important-path &&\n> +\t\tgit add important-path &&\n> +\t\ttest_commit initial &&\n> +\n> +\t\ttouch ignored &&\n> +\t\tgit add ignored &&\n> +\t\ttest_commit whatever &&\n> +\n> +\t\tgit branch A &&\n> +\t\tgit branch B &&\n> +\n> +\t\techo foo >>important-path &&\n> +\t\tgit add important-path &&\n> +\t\ttest_commit more changes &&\n> +\n> +\t\tgit fast-export --all -- important-path | git fast-import --force\n> +\t)\n> +'\n\nSimilar comments apply about \"touch\" and \"test_commit\" to what I wrote\nfor the earlier patch.\n\n-Peff\n"},{"id":"362917","messageId":"20181111070240.GG30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-8-newren@gmail.com","subject":"Re: [PATCH 07/10] fast-export: ensure we export requested refs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T07:02:40Z","receivedAt":"2018-11-11T07:02:46Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:09PM -0800, Elijah Newren wrote:\n\n> If file paths are specified to fast-export and a ref points to a commit\n> that does not touch any of the relevant paths, then that ref would\n> sometimes fail to be exported.  (This depends on whether any ancestors\n> of the commit which do touch the relevant paths would be exported with\n> that same ref name or a different ref name.)  To avoid this problem,\n> put *all* specified refs into extra_refs to start, and then as we export\n> each commit, remove the refname used in the 'commit $REFNAME' directive\n> from extra_refs.  Then, in handle_tags_and_duplicates() we know which\n> refs actually do need a manual reset directive in order to be included.\n> \n> This means that we do need some special handling for excluded refs; e.g.\n> if someone runs\n>    git fast-export ^master master\n> then they've asked for master to be exported, but they have also asked\n> for the commit which master points to and all of its history to be\n> excluded.  That logically means ref deletion.  Previously, such refs\n> were just silently omitted from being exported despite having been\n> explicitly requested for export.\n\nHmm. Reading this it makes sense to me, but I remember from discussion\nlong ago that there were a lot of funny corner cases around \"which refs\nto include\" and possibly even some ambiguous cases. Maybe that is all\nsorted these days, with --refspec.\n\n> ---\n> NOTE: I was hoping the strmap API proposal would materialize, but I either\n> missed it or it hasn't shown up.  The usage of string_list in this patch\n> would be better replaced by what Peff suggested.\n\nYou didn't miss it. Junio did some manual conversions using hashmap,\nwhich weren't too bad.  It's not entirely clear to me how often we'd be\nable to use strmap instead of a full-on hashmap, so I haven't really\npursued it.\n\nIt looks like you generate the list here via append, and then sort at\nthe end. That's at least not quadratic. I think the string_list_remove()\nis, though.\n\n-Peff\n"},{"id":"362919","messageId":"20181111071102.GH30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-9-newren@gmail.com","subject":"Re: [PATCH 08/10] fast-export: add --reference-excluded-parents option","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T07:11:03Z","receivedAt":"2018-11-11T07:11:09Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:10PM -0800, Elijah Newren wrote:\n\n> git filter-branch has a nifty feature allowing you to rewrite, e.g. just\n> the last 8 commits of a linear history\n>   git filter-branch $OPTIONS HEAD~8..HEAD\n> \n> If you try the same with git fast-export, you instead get a history of\n> only 8 commits, with HEAD~7 being rewritten into a root commit.  There\n> are two alternatives:\n\nAh, I think this maybe answers some of my earlier questions, too. You\ncannot use fast-import as it stands to do a partial rewrite.\n\n>   1) Don't use the negative revision specification, and when you're\n>      filtering the output to make modifications to the last 8 commits,\n>      just be careful to not modify any earlier commits somehow.\n> \n>   2) First run 'git fast-export --export-marks=somefile HEAD~8', then\n>      run 'git fast-export --import-marks=somefile HEAD~8..HEAD'.\n> \n> Both are more error prone than I'd like (the first for obvious reasons;\n> with the second option I have sometimes accidentally included too many\n> revisions in the first command and then found that the corresponding\n> extra revisions were not exported by the second command and thus were\n> not modified as I expected).  Also, both are poor from a performance\n> perspective.\n\nYeah, this should be O(commits you're touching), and it the current code\ndoes not allow that at all. So I think this feature makes a lot of sense\n(it probably _should_ have been the default, but it's a bit late for\nthat now).\n\n> @@ -638,13 +640,21 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n>  \tunuse_commit_buffer(commit, commit_buffer);\n>  \n>  \tfor (i = 0, p = commit->parents; p; p = p->next) {\n> -\t\tint mark = get_object_mark(&p->item->object);\n> -\t\tif (!mark)\n> +\t\tstruct object *obj = &p->item->object;\n> +\t\tint mark = get_object_mark(obj);\n> +\n> +\t\tif (!mark && !reference_excluded_commits)\n>  \t\t\tcontinue;\n>  \t\tif (i == 0)\n> -\t\t\tprintf(\"from :%d\\n\", mark);\n> +\t\t\tprintf(\"from \");\n> +\t\telse\n> +\t\t\tprintf(\"merge \");\n> +\t\tif (mark)\n> +\t\t\tprintf(\":%d\\n\", mark);\n>  \t\telse\n> -\t\t\tprintf(\"merge :%d\\n\", mark);\n> +\t\t\tprintf(\"%s\\n\", sha1_to_hex(anonymize ?\n> +\t\t\t\t\t\t   anonymize_sha1(&obj->oid) :\n> +\t\t\t\t\t\t   obj->oid.hash));\n>  \t\ti++;\n>  \t}\n\nOK, so this just teaches us to start with the sensible \"from\" directive.\nI think we might be able to do a little more optimization here. If we're\nexporting HEAD^..HEAD and there's an object in HEAD^ which is unchanged\nin HEAD, I think we'd still print it (because it would not be marked\nSHOWN), but we could omit it (by walking the tree of the boundary\ncommits and marking them shown).\n\nI don't think it's a blocker for what you're doing here, but just a\npossible future optimization.\n\n> @@ -925,13 +935,22 @@ static void handle_tags_and_duplicates(struct string_list *extras)\n>  \t\t\t\t/*\n>  \t\t\t\t * Getting here means we have a commit which\n>  \t\t\t\t * was excluded by a negative refspec (e.g.\n> -\t\t\t\t * fast-export ^master master).  If the user\n> +\t\t\t\t * fast-export ^master master).  If we are\n> +\t\t\t\t * referencing excluded commits, set the ref\n> +\t\t\t\t * to the exact commit.  Otherwise, the user\n>  \t\t\t\t * wants the branch exported but every commit\n> -\t\t\t\t * in its history to be deleted, that sounds\n> -\t\t\t\t * like a ref deletion to me.\n> +\t\t\t\t * in its history to be deleted, which basically\n> +\t\t\t\t * just means deletion of the ref.\n>  \t\t\t\t */\n> -\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n> -\t\t\t\t       name, sha1_to_hex(null_sha1));\n> +\t\t\t\tif (!reference_excluded_commits) {\n> +\t\t\t\t\t/* delete the ref */\n> +\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n> +\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n> +\t\t\t\t\tcontinue;\n> +\t\t\t\t}\n> +\t\t\t\t/* set ref to commit using oid, not mark */\n> +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\", name,\n> +\t\t\t\t       sha1_to_hex(commit->object.oid.hash));\n\nOK, and this is basically answering my earlier questions again: yes, you\n_would_ want to keep old tags pointing at their commits. But only in\nthis much more sensible mode.\n\n-Peff\n"},{"id":"362920","messageId":"CABPp-BHwg2U=b+UGK2SufB7uZPmmiPVKXoTpYt+LuHnLwmwuZQ@mail.gmail.com","threadId":"49406","inReplyTo":"20181111063601.GB30850@sigill.intra.peff.net","subject":"Re: [PATCH 02/10] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T07:17:47Z","receivedAt":"2018-11-11T07:18:01Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 10:36 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:04PM -0800, Elijah Newren wrote:\n>\n> > Signed-off-by: Elijah Newren <newren@gmail.com>\n> > ---\n> >  Documentation/git-fast-export.txt | 3 ++-\n> >  1 file changed, 2 insertions(+), 1 deletion(-)\n> >\n> > diff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\n> > index ce954be532..677510b7f7 100644\n> > --- a/Documentation/git-fast-export.txt\n> > +++ b/Documentation/git-fast-export.txt\n> > @@ -119,7 +119,8 @@ marks the same across runs.\n> >       'git rev-list', that specifies the specific objects and references\n> >       to export.  For example, `master~10..master` causes the\n> >       current master reference to be exported along with all objects\n> > -     added since its 10th ancestor commit.\n> > +     added since its 10th ancestor commit and all files common to\n> > +     master\\~9 and master~10.\n>\n> Do you need to backslash the second tilde?  Maybe `master~9` and\n> `master~10` instead of escaping?\n\nOops, yeah, that needs to be consistent.\n\n> I'm not sure what this is trying to say. I guess that we'd always show\n> all of the blobs necessary to reconstruct the first non-negative commit\n> (i.e., `master~9` here)?\n\nFor someone familiar with fast-export or fast-import, sure, you'd\nguess that it'd show all the blobs necessary to reconstruct the first\nnon-negative commit.  But it's not clear to first time users and\nreaders of the docs that the first non-negative commit becomes a root\ncommit; by comparison, filter-branch suggests using a very similar\nconstruction and yet behaves quite differently -- it does not turn the\nfirst non-negative commit into a root but retains the original\nparent(s) of the first non-negative commit without rewriting those\nearlier commits.  The text as previously written, \"along with all\nobjects added since its 10th ancestor commit\", seems to suggest\nbehavior similar to how filter-branch behaves (particularly the\n\"Acked-by example\"), i.e. it implies that files not touched in the\nlast 10 commits are not included.  My wording in this patch was an\nattempt to fix that.  Was my attempt perhaps too clumsy, or was it\njust the case that you had sufficient knowledge of fast-export that\nthe previous text didn't mislead you?\n"},{"id":"362921","messageId":"20181111072007.GI30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-10-newren@gmail.com","subject":"Re: [PATCH 09/10] fast-export: add a --show-original-ids option to show original names","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T07:20:07Z","receivedAt":"2018-11-11T07:20:11Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:11PM -0800, Elijah Newren wrote:\n\n> Knowing the original names (hashes) of commits, blobs, and tags can\n> sometimes enable post-filtering that would otherwise be difficult or\n> impossible.  In particular, the desire to rewrite commit messages which\n> refer to other prior commits (on top of whatever other filtering is\n> being done) is very difficult without knowing the original names of each\n> commit.\n> \n> This commit teaches a new --show-original-ids option to fast-export\n> which will make it add a 'originally <hash>' line to blob, commits, and\n> tags.  It also teaches fast-import to parse (and ignore) such lines.\n\nMakes sense as a feature; I think filter-branch can make its mappings\navailable, too.\n\nDo we need to worry about compatibility with other fast-import programs?\nI think no, because this is not enabled by default (so if sending the\nextra lines to another importer hurts, the answer is \"don't do that\").\n\nI have a vague feeling that there might be some way to combine this with\n--export-marks or --no-data, but I can't really think of a way. They\nseem related, but not quite.\n\n> ---\n>  Documentation/git-fast-export.txt |  7 +++++++\n>  builtin/fast-export.c             | 20 +++++++++++++++-----\n>  fast-import.c                     | 17 +++++++++++++++++\n>  t/t9350-fast-export.sh            | 17 +++++++++++++++++\n>  4 files changed, 56 insertions(+), 5 deletions(-)\n\nThe fast-import format is documented in Documentation/git-fast-import.txt.\nIt might need an update to cover the new format.\n\n> --- a/Documentation/git-fast-export.txt\n> +++ b/Documentation/git-fast-export.txt\n> @@ -121,6 +121,13 @@ marks the same across runs.\n>  \tused by a repository which already contains the necessary\n>  \tparent commits.\n>  \n> +--show-original-ids::\n> +\tAdd an extra directive to the output for commits and blobs,\n> +\t`originally <SHA1SUM>`.  While such directives will likely be\n> +\tignored by importers such as git-fast-import, it may be useful\n> +\tfor intermediary filters (e.g. for rewriting commit messages\n> +\twhich refer to older commits, or for stripping blobs by id).\n\nI'm not quite sure how a blob ends up being rewritten by fast-export (I\nget that commits may change due to dropping parents).\n\nThe name \"originally\" doesn't seem great to me. Probably because I would\ncontinually wonder if it has one \"l\" or two. ;) Perhaps something like\n\"original-oid\" might be better. That's well into bikeshed territory,\nthough.\n\n-Peff\n"},{"id":"362922","messageId":"20181111072356.GJ30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-11-newren@gmail.com","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T07:23:56Z","receivedAt":"2018-11-11T07:24:00Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:12PM -0800, Elijah Newren wrote:\n\n> fast-export output is traditionally used as an input to a fast-import\n> program, but it is also useful to help gather statistics about the\n> history of a repository (particularly when --no-data is also passed).\n> For example, two of the types of information we may want to collect\n> could include:\n>   1) general information about renames that have occurred\n>   2) what the biggest objects in a repository are and what names\n>      they appear under.\n> \n> The first bit of information can be gathered by just passing -M to\n> fast-export.  The second piece of information can partially be gotten\n> from running\n>     git cat-file --batch-check --batch-all-objects\n> However, that only shows what the biggest objects in the repository are\n> and their sizes, not what names those objects appear as or what commits\n> they were introduced in.  We can get that information from fast-export,\n> but when we only see\n>     R oldname newname\n> instead of\n>     R oldname newname\n>     M 100644 $SHA1 newname\n> then it makes the job more difficult.  Add an option which allows us to\n> force the latter output even when commits have exact renames of files.\n\nfast-export seems like a funny tool to look up paths. What about \"git\nlog --find-object=$SHA1\" ?\n\n-Peff\n"},{"id":"362923","messageId":"20181111072716.GK30850@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"Re: [PATCH 00/10] fast export and import fixes and features","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-11T07:27:16Z","receivedAt":"2018-11-11T07:27:19Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 10:23:02PM -0800, Elijah Newren wrote:\n\n> This is a series of ten patches representing two doc corrections, one\n> pedantic fix, three real bug fixes, one micro code refactor, and three\n> new features.  Each of these ten changes is relatively small in size.\n> These changes predominantly affect fast-export, but there's a couple\n> small changes for fast-import as well.\n> \n> I could potentially split these patches up, but I'd just end up\n> chaining them sequentially since otherwise there'd be lots of\n> conflicts; having 10 different single patch series with lots of\n> dependencies sounded like a bigger pain to me, but let me know if you\n> would prefer I split them up and how you suggest doing so.\n\nI think it's fine to put them in sequence when there's a textual\ndependency.  If it turns out that one of them needs more discussion and\nwe don't want it to hold later patches hostage, we can always re-roll at\nthat point.\n\n(I also think it's fine to lump together thematically similar patches\neven when they aren't strictly dependent, even textually. It's less work\nfor the maintainer to consider 1 group of 10 than 10 groups of 1).\n\n> These patches were driven by the needs of git-repo-filter[1], but most\n> if not all of them should be independently useful.\n\nI left lots of comments. Some of the earlier ones may just be showing my\nconfusion about fast-export works (some of which was cleared up by your\nlater patches). But I like the overall direction for sure.\n\n-Peff\n"},{"id":"362925","messageId":"CABPp-BFy1aS3mHGF99Lr=+APruzC3pF5PCEph8SU71uuyOnQ7Q@mail.gmail.com","threadId":"49406","inReplyTo":"20181111064442.GD30850@sigill.intra.peff.net","subject":"Re: [PATCH 04/10] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T07:38:45Z","receivedAt":"2018-11-11T07:40:18Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 10:44 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:06PM -0800, Elijah Newren wrote:\n>\n> > If --tag-of-filtered-object=rewrite is specified along with a set of\n> > paths to limit what is exported, then any tags pointing to old commits\n> > that do not contain any of those specified paths cause problems.  Since\n> > the old tagged commit is not exported, fast-export attempts to rewrite\n> > such tags to an ancestor commit which was exported.  If no such commit\n> > exists, then fast-export currently die()s.  Five years after the tag\n> > rewriting logic was added to fast-export (see commit 2d8ad4691921,\n> > \"fast-export: Add a --tag-of-filtered-object  option for newly dangling\n> > tags\", 2009-06-25), fast-import gained the ability to delete refs (see\n> > commit 4ee1b225b99f, \"fast-import: add support to delete refs\",\n> > 2014-04-20), so now we do have a valid option to rewrite the tag to.\n> > Delete these tags instead of dying.\n>\n> Hmm. That's the right thing to do if we're considering the export to be\n> an independent unit. But what if I'm just rewriting a portion of history\n> like:\n>\n>   git fast-export HEAD~5..HEAD | some_filter | git fast-import\n>\n> ? If I have a tag pointing to HEAD~10, will this delete that? Ideally I\n> think it would be left alone.\n\nA couple things:\n  * This code path only triggers in a very specific case: If a tag is\nrequested for export but points to a commit which is filtered out by\nsomething else (e.g. path limiters and the commit in question didn't\nmodify any of the relevant paths), AND the user explicitly specified\n--tag-of-filtered-object=rewrite (so that the tag in question can be\nrewritten to the nearest non-filtered ancestor).\n  * You didn't specify to export any tags, only HEAD, so this\nsituation isn't relevant (the tag wouldn't be exported or deleted).\n  * You didn't specify --tag-of-filtered-object=rewrite, so this\nsituation isn't relevant (even if you had specified a tag to filter,\nyou'd get an abort instead)\n\nBut let's say you do modify the example some:\n   git fast-export --tag-of-filtered-object=rewrite\n--signed-tags=strip --tags master -- relatively_recent_subdirectory/ |\nsome_filter | git fast-import\n\nThe user asked that all tags and master be exported but only for the\nhistory that touched relatively_recent_subdirectory/, and if any tags\npoint at commits that are pruned by only asking for commits touching\nrelatively_recent_subdirectory/, then rewrite what those tags point to\nso that they instead point to the nearest non-filtered ancestor.  What\nabout a commit like v0.1.0 that likely pre-dated the introduction of\nrelatively_recent_subdirectory/?  It has no nearest ancestor to\nrewrite to.  The previous answer was to abort, which is really bad,\nespecially since the user was clearly asking us to do whatever smart\nrewriting we can (--signed-tags=strip and\n--tag-of-filtered-object=rewrite).\n\nPerhaps there's a different answer that's workable as well, but this\none, in these circumstances, seemed the most reasonable to me.\n\n> > +test_expect_success 'rewrite tag predating pathspecs to nothing' '\n> > +     test_create_repo rewrite_tag_predating_pathspecs &&\n> > +     (\n> > +             cd rewrite_tag_predating_pathspecs &&\n> > +\n> > +             touch ignored &&\n>\n> We usually prefer \">ignored\" to create an empty file rather than\n> \"touch\".\n\nWill fix.\n\n>\n> > +             git add ignored &&\n> > +             test_commit initial &&\n>\n> What do we need this \"ignored\" for? test_commit should create a file\n> \"initial.t\".\n\nI think I original had plain \"git commit\", then switched to\ntest_commit, then didn't recheck.  Thanks, will fix.\n\n> > +             echo foo >bar &&\n> > +             git add bar &&\n> > +             test_commit add-bar &&\n>\n> Likewise, \"test_commit bar\" should work by itself (though note the\n> filename is \"bar.t\" in your fast-export command).\n>\n> > +             git fast-export --tag-of-filtered-object=rewrite --all -- bar >output &&\n> > +             grep -A 1 refs/tags/v0.0.0.0.0.0.1 output | grep -E ^from.0{40}\n>\n> I don't think \"grep -A\" is portable (and we don't seem to otherwise use\n> it). You can probably do something similar with sed.\n>\n> Use $ZERO_OID instead of hard-coding 40, which future-proofs for the\n> hash transition (though I suppose the hash is not likely to get\n> _shorter_ ;) ).\n\nWill fix these up as well...after waiting for more feedback on\npossible alternate suggestions.\n"},{"id":"362926","messageId":"CABPp-BGF8C5vhyVbAwpmXeii452fBgtvL4dPRLWdOPxLiCYR0A@mail.gmail.com","threadId":"49406","inReplyTo":"20181111065338.GF30850@sigill.intra.peff.net","subject":"Re: [PATCH 06/10] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T08:01:43Z","receivedAt":"2018-11-11T08:01:57Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 10:53 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:08PM -0800, Elijah Newren wrote:\n>\n> > If file paths are specified to fast-export and multiple refs point to a\n> > commit that does not touch any of the relevant file paths, then\n> > fast-export can hit problems.  fast-export has a list of additional refs\n> > that it needs to explicitly set after exporting all blobs and commits,\n> > and when it tries to get_object_mark() on the relevant commit, it can\n> > get a mark of 0, i.e. \"not found\", because the commit in question did\n> > not touch the relevant paths and thus was not exported.  Trying to\n> > import a stream with a mark corresponding to an unexported object will\n> > cause fast-import to crash.\n> >\n> > Avoid this problem by taking the commit the ref points to and finding an\n> > ancestor of it that was exported, and make the ref point to that commit\n> > instead.\n>\n> As with the earlier tag commit, I wonder if this might depend on the\n> context in which you're using fast-export. I suppose that if you did not\n> feed the ref on the command line that we would not be dealing with it at\n> all (and maybe that is the answer to my question about the tag thing,\n> too).\n\nRight, if you didn't feed the ref on the command line, we're not\ndealing with the ref at all, so the code here doesn't affect any such\nref.\n\n> It does seem funny that the behavior for the earlier case (bounded\n> commits) and this case (skipping some commits) are different. Would you\n> ever want to keep walking backwards to find an ancestor in the earlier\n> case? Or vice versa, would you ever want to simply delete a tag in a\n> case like this one?\n>\n> I'm not sure sure, but I suspect you may have thought about it a lot\n> harder than I have. :)\n\nI'm not sure why you thought the behavior for the two cases was\ndifferent?  For both patches, my testcases used path limiting; it was\nyou who suggested employing a negative revision to bound the commits.\n\nAnyway, for both patches assuming you haven't bounded the commits, you\ncan attempt to keep walking backwards to find an earlier ancestor, but\nthe fundamental fact is you aren't guaranteed that you can find one\n(i.e. some tag or branch points to a commit that didn't modify any of\nthe specified paths, and nor did any of its ancestors back to any root\ncommits).  I hit that case lots of times.  If the user explicitly\nrequested a tag or branch for export (and requested tag rewriting),\nand limited to certain paths that had never existed in the repository\nas of the time of the tag or branch, then you hit the cases these\npatches worry about.  Patch 4 was about (annotated and signed) tags,\nthis patch is about unannotated tags and branches and other refs.\n\nIf you think about using negative revisions, for both cases, then\nagain you can keep walking back history to try to find a commit that\nyour tag or branch or ref can point to, but if you get back to the\nnegative revisions, then you are in the range the user requested to be\nomitted from the resulting repository.  Sounds like tag/ref deletion\nto me.\n\n>\n> > diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n> > index a3c044b0af..5648a8ce9c 100644\n> > --- a/builtin/fast-export.c\n> > +++ b/builtin/fast-export.c\n> > @@ -900,7 +900,18 @@ static void handle_tags_and_duplicates(void)\n> >                       if (anonymize)\n> >                               name = anonymize_refname(name);\n> >                       /* create refs pointing to already seen commits */\n> > -                     commit = (struct commit *)object;\n> > +                     commit = rewrite_commit((struct commit *)object);\n> > +                     if (!commit) {\n> > +                             /*\n> > +                              * Neither this object nor any of its\n> > +                              * ancestors touch any relevant paths, so\n> > +                              * it has been filtered to nothing.  Delete\n> > +                              * it.\n> > +                              */\n> > +                             printf(\"reset %s\\nfrom %s\\n\\n\",\n> > +                                    name, sha1_to_hex(null_sha1));\n> > +                             continue;\n> > +                     }\n>\n> This hunk makes sense.\n\nCool, this was the entirety of the code...so does this mean that the\ncode makes more sense than my commit message summary did?  ...and\nperhaps that my attempts to answer your questions in this email\nweren't necessary anymore?\n\n> > --- a/t/t9350-fast-export.sh\n> > +++ b/t/t9350-fast-export.sh\n> > @@ -386,6 +386,30 @@ test_expect_success 'path limiting with import-marks does not lose unmodified fi\n> >       grep file0 actual\n> >  '\n> >\n> > +test_expect_success 'avoid corrupt stream with non-existent mark' '\n> > +     test_create_repo avoid_non_existent_mark &&\n> > +     (\n> > +             cd avoid_non_existent_mark &&\n> > +\n> > +             touch important-path &&\n> > +             git add important-path &&\n> > +             test_commit initial &&\n> > +\n> > +             touch ignored &&\n> > +             git add ignored &&\n> > +             test_commit whatever &&\n> > +\n> > +             git branch A &&\n> > +             git branch B &&\n> > +\n> > +             echo foo >>important-path &&\n> > +             git add important-path &&\n> > +             test_commit more changes &&\n> > +\n> > +             git fast-export --all -- important-path | git fast-import --force\n> > +     )\n> > +'\n>\n> Similar comments apply about \"touch\" and \"test_commit\" to what I wrote\n> for the earlier patch.\n\nThanks; will fix.\n"},{"id":"362927","messageId":"CABPp-BGQNsZYKYuaBcY7Umr=u0qzF5gXWFT3yGjLdzAz2ZGs+w@mail.gmail.com","threadId":"49406","inReplyTo":"20181111070240.GG30850@sigill.intra.peff.net","subject":"Re: [PATCH 07/10] fast-export: ensure we export requested refs","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T08:20:16Z","receivedAt":"2018-11-11T08:20:30Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 11:02 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:09PM -0800, Elijah Newren wrote:\n>\n> > If file paths are specified to fast-export and a ref points to a commit\n> > that does not touch any of the relevant paths, then that ref would\n> > sometimes fail to be exported.  (This depends on whether any ancestors\n> > of the commit which do touch the relevant paths would be exported with\n> > that same ref name or a different ref name.)  To avoid this problem,\n> > put *all* specified refs into extra_refs to start, and then as we export\n> > each commit, remove the refname used in the 'commit $REFNAME' directive\n> > from extra_refs.  Then, in handle_tags_and_duplicates() we know which\n> > refs actually do need a manual reset directive in order to be included.\n> >\n> > This means that we do need some special handling for excluded refs; e.g.\n> > if someone runs\n> >    git fast-export ^master master\n> > then they've asked for master to be exported, but they have also asked\n> > for the commit which master points to and all of its history to be\n> > excluded.  That logically means ref deletion.  Previously, such refs\n> > were just silently omitted from being exported despite having been\n> > explicitly requested for export.\n>\n> Hmm. Reading this it makes sense to me, but I remember from discussion\n> long ago that there were a lot of funny corner cases around \"which refs\n> to include\" and possibly even some ambiguous cases. Maybe that is all\n> sorted these days, with --refspec.\n\nOh yeah, there definitely were some funny corner cases around \"which\nrefs to include\" (though I don't think --refspec affects this, either\nbefore or after my patch.)  Before this commit, fast-export would\noften emit unnecessary reset directives at the end, AND fail to export\nsome other refs that had been explicitly requested for export.  It had\nsome simple logic to attempt to cover the cases, but it was just\nwrong.  As far as I can tell, this patch fixes all of those.\n\n...well, almost all.  We still fail on tags of tags of commits (or\nhigher level nestings), but that's a multi-pronged issue that feels\nlike a different beast. (We rewrite tags of tags of commits to just be\ntags of commits, even without any special request from the user\nsomewhat contrary to otherwise requiring --signed-tags and\n--tag-of-filtered-object options.  As far as I can tell, this isn't\ndocumented for fast-export but I saw somewhere in the filter-branch\ndocs where it said it does this kind of thing on purpose.  However, to\nmake it even weirder, if the user requests\n--tag-of-filtered-object=rewrite instead of the default of \"abort\"\nthen we actually abort on tags-of-tags-of-commits instead of\nrewriting.  I don't think it was intentional, but\ntags-of-tags-of-commits inverts the meaning of the\n--tag-of-filtered-object={rewrite vs. abort} flag -- it's very weird).\nI put more time into attempting to fix the nested tags issue than I\nfeel like it was worth.  git.git is the only repo I know of that seems\nto have such tags, so I just gave up on them for now.\n\n> > ---\n> > NOTE: I was hoping the strmap API proposal would materialize, but I either\n> > missed it or it hasn't shown up.  The usage of string_list in this patch\n> > would be better replaced by what Peff suggested.\n>\n> You didn't miss it. Junio did some manual conversions using hashmap,\n> which weren't too bad.  It's not entirely clear to me how often we'd be\n> able to use strmap instead of a full-on hashmap, so I haven't really\n> pursued it.\n>\n> It looks like you generate the list here via append, and then sort at\n> the end. That's at least not quadratic. I think the string_list_remove()\n> is, though.\n\nI think it would have been useful in multiple places in\nmerge-recursive.c, in addition to here.  Maybe that just means I need\nto add strmap to my list of things to do.\n"},{"id":"362928","messageId":"CABPp-BGNt0FcqiT=OqctjOEvY9ewNUJZ-Rs_aVEihjbQt3K8tQ@mail.gmail.com","threadId":"49406","inReplyTo":"20181111072007.GI30850@sigill.intra.peff.net","subject":"Re: [PATCH 09/10] fast-export: add a --show-original-ids option to show original names","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T08:32:22Z","receivedAt":"2018-11-11T08:32:36Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 11:20 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:11PM -0800, Elijah Newren wrote:\n>\n> > Knowing the original names (hashes) of commits, blobs, and tags can\n> > sometimes enable post-filtering that would otherwise be difficult or\n> > impossible.  In particular, the desire to rewrite commit messages which\n> > refer to other prior commits (on top of whatever other filtering is\n> > being done) is very difficult without knowing the original names of each\n> > commit.\n> >\n> > This commit teaches a new --show-original-ids option to fast-export\n> > which will make it add a 'originally <hash>' line to blob, commits, and\n> > tags.  It also teaches fast-import to parse (and ignore) such lines.\n>\n> Makes sense as a feature; I think filter-branch can make its mappings\n> available, too.\n>\n> Do we need to worry about compatibility with other fast-import programs?\n> I think no, because this is not enabled by default (so if sending the\n> extra lines to another importer hurts, the answer is \"don't do that\").\n>\n> I have a vague feeling that there might be some way to combine this with\n> --export-marks or --no-data, but I can't really think of a way. They\n> seem related, but not quite.\n>\n> > ---\n> >  Documentation/git-fast-export.txt |  7 +++++++\n> >  builtin/fast-export.c             | 20 +++++++++++++++-----\n> >  fast-import.c                     | 17 +++++++++++++++++\n> >  t/t9350-fast-export.sh            | 17 +++++++++++++++++\n> >  4 files changed, 56 insertions(+), 5 deletions(-)\n>\n> The fast-import format is documented in Documentation/git-fast-import.txt.\n> It might need an update to cover the new format.\n\nWe document the format in both fast-import.c and\nDocumentation/git-fast-import.txt?  Maybe we should delete the long\ncomments in fast-import.c so this isn't duplicated?\n\n> > --- a/Documentation/git-fast-export.txt\n> > +++ b/Documentation/git-fast-export.txt\n> > @@ -121,6 +121,13 @@ marks the same across runs.\n> >       used by a repository which already contains the necessary\n> >       parent commits.\n> >\n> > +--show-original-ids::\n> > +     Add an extra directive to the output for commits and blobs,\n> > +     `originally <SHA1SUM>`.  While such directives will likely be\n> > +     ignored by importers such as git-fast-import, it may be useful\n> > +     for intermediary filters (e.g. for rewriting commit messages\n> > +     which refer to older commits, or for stripping blobs by id).\n>\n> I'm not quite sure how a blob ends up being rewritten by fast-export (I\n> get that commits may change due to dropping parents).\n\nIt doesn't get rewritten by fast-export; it gets rewritten by other\nintermediary filters, e.g. in something like this:\n\n   git fast-export --show-original-ids --all | intermediary_filter |\ngit fast-import\n\nThe intermediary_filter program may want to strip out blobs by id, or\nremove filemodify and filedelete directives unless they touch certain\npaths, etc.\n\n> The name \"originally\" doesn't seem great to me. Probably because I would\n> continually wonder if it has one \"l\" or two. ;) Perhaps something like\n> \"original-oid\" might be better. That's well into bikeshed territory,\n> though.\n\nI wasn't a huge fan of \"originally\" either, but I just couldn't come\nup with anything else that wasn't really long.  I'd be happy to switch\nto original-oid.\n"},{"id":"362929","messageId":"CABPp-BGREOAvF-6DBymdwsUL2LpyPNqy8dCw0RuUKZf2Da6cJA@mail.gmail.com","threadId":"49406","inReplyTo":"20181111072356.GJ30850@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T08:42:58Z","receivedAt":"2018-11-11T08:43:12Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 11:23 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:12PM -0800, Elijah Newren wrote:\n>\n> > fast-export output is traditionally used as an input to a fast-import\n> > program, but it is also useful to help gather statistics about the\n> > history of a repository (particularly when --no-data is also passed).\n> > For example, two of the types of information we may want to collect\n> > could include:\n> >   1) general information about renames that have occurred\n> >   2) what the biggest objects in a repository are and what names\n> >      they appear under.\n> >\n> > The first bit of information can be gathered by just passing -M to\n> > fast-export.  The second piece of information can partially be gotten\n> > from running\n> >     git cat-file --batch-check --batch-all-objects\n> > However, that only shows what the biggest objects in the repository are\n> > and their sizes, not what names those objects appear as or what commits\n> > they were introduced in.  We can get that information from fast-export,\n> > but when we only see\n> >     R oldname newname\n> > instead of\n> >     R oldname newname\n> >     M 100644 $SHA1 newname\n> > then it makes the job more difficult.  Add an option which allows us to\n> > force the latter output even when commits have exact renames of files.\n>\n> fast-export seems like a funny tool to look up paths. What about \"git\n> log --find-object=$SHA1\" ?\n\nEek, and give me O(N*M) behavior, where N is the number of commits in\nthe repository and M is the number of renames that occur in its\nhistory?  Also, that's the inverse of the lookup I need anyway (I have\nthe commit and filename, but am missing the SHA).\n\nOne of the problems with filter-branch that people often run into is\nthey know what they want at a high-level (e.g. extract the history of\nthis directory for a new repository, or rewrite the history of this\nrepo to appear at a subdirectory so it can be merged into a bigger\nrepo and people passing filenames to log will still get the history of\nthose files, or I want to remove some of the big stuff in my history),\nbut often times that's not quite enough.  They need help finding big\nobjects, or may be unaware that the subset of files they want used to\nbe known by alternative names.\n\nI want a simple --analyze mode that can report on all files that have\nbeen renamed (so users don't just say \"all I care about is these N\nfiles, give me a rewritten history just including those\" -- we can\npoint out to them whether those N files used to be known by other\nnames), as well as reporting on all big files and if they've been\ndeleted, and aggregations of the \"big files\" information across\ndirectories and file extensions.\n"},{"id":"362931","messageId":"CABPp-BGzqpxF_+ubp2cft9dQ-03pgcCxJEP13VOUv5WADHDjnA@mail.gmail.com","threadId":"49406","inReplyTo":"20181111072716.GK30850@sigill.intra.peff.net","subject":"Re: [PATCH 00/10] fast export and import fixes and features","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-11T08:44:47Z","receivedAt":"2018-11-11T08:45:01Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 11:27 PM Jeff King <peff@peff.net> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:23:02PM -0800, Elijah Newren wrote:\n>\n> > This is a series of ten patches representing two doc corrections, one\n> > pedantic fix, three real bug fixes, one micro code refactor, and three\n> > new features.  Each of these ten changes is relatively small in size.\n> > These changes predominantly affect fast-export, but there's a couple\n> > small changes for fast-import as well.\n> >\n> > I could potentially split these patches up, but I'd just end up\n> > chaining them sequentially since otherwise there'd be lots of\n> > conflicts; having 10 different single patch series with lots of\n> > dependencies sounded like a bigger pain to me, but let me know if you\n> > would prefer I split them up and how you suggest doing so.\n>\n> I think it's fine to put them in sequence when there's a textual\n> dependency.  If it turns out that one of them needs more discussion and\n> we don't want it to hold later patches hostage, we can always re-roll at\n> that point.\n>\n> (I also think it's fine to lump together thematically similar patches\n> even when they aren't strictly dependent, even textually. It's less work\n> for the maintainer to consider 1 group of 10 than 10 groups of 1).\n>\n> > These patches were driven by the needs of git-repo-filter[1], but most\n> > if not all of them should be independently useful.\n>\n> I left lots of comments. Some of the earlier ones may just be showing my\n> confusion about fast-export works (some of which was cleared up by your\n> later patches). But I like the overall direction for sure.\n\nThanks for taking the time to read over the series and providing lots\nof feedback!  And, whoops, looks like it's gotten kinda late, so I'll\ncheck any further feedback on Monday.\n"},{"id":"362948","messageId":"87va532x5i.fsf@evledraar.gmail.com","threadId":"49406","inReplyTo":"20181111063636.GC30850@sigill.intra.peff.net","subject":"Re: [PATCH 03/10] fast-export: use value from correct enum","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-11-11T20:10:17Z","receivedAt":"2018-11-11T20:10:23Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Nov 11 2018, Jeff King wrote:\n\n> On Sat, Nov 10, 2018 at 10:23:05PM -0800, Elijah Newren wrote:\n>\n>> ABORT and ERROR happen to have the same value, but come from differnt\n>> enums.  Use the one from the correct enum.\n>\n> Yikes. :)\n>\n> This is a good argument for naming these SIGNED_TAG_ABORT, etc. But this\n> is obviously an improvement in the meantime.\n\nIn C enum values aren't the types of the enum, but I'd thought someone\nwould have added a warning for this:\n\n    #include <stdio.h>\n\n    enum { A, B } foo = A;\n    enum { C, D } bar = C;\n\n    int main(void)\n    {\n        switch (foo) {\n          case C:\n            puts(\"A\");\n            break;\n          case B:\n            puts(\"B\");\n            break;\n        }\n    }\n\nBut none of the 4 C compilers (gcc, clang, suncc & xlc) I have warn\nabout it. Good to know.\n"},{"id":"362976","messageId":"CACBZZX6Ck-M7UPK85UW4ZOYWeodSJo0gp7Kgs__on5SfZDmojA@mail.gmail.com","threadId":"49406","inReplyTo":"87va532x5i.fsf@evledraar.gmail.com","subject":"Re: [PATCH 03/10] fast-export: use value from correct enum","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-11-12T09:12:50Z","receivedAt":"2018-11-12T09:13:04Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Sun, Nov 11, 2018 at 9:10 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Sun, Nov 11 2018, Jeff King wrote:\n>\n> > On Sat, Nov 10, 2018 at 10:23:05PM -0800, Elijah Newren wrote:\n> >\n> >> ABORT and ERROR happen to have the same value, but come from differnt\n> >> enums.  Use the one from the correct enum.\n> >\n> > Yikes. :)\n> >\n> > This is a good argument for naming these SIGNED_TAG_ABORT, etc. But this\n> > is obviously an improvement in the meantime.\n>\n> In C enum values aren't the types of the enum, but I'd thought someone\n> would have added a warning for this:\n>\n>     #include <stdio.h>\n>\n>     enum { A, B } foo = A;\n>     enum { C, D } bar = C;\n>\n>     int main(void)\n>     {\n>         switch (foo) {\n>           case C:\n>             puts(\"A\");\n>             break;\n>           case B:\n>             puts(\"B\");\n>             break;\n>         }\n>     }\n>\n> But none of the 4 C compilers (gcc, clang, suncc & xlc) I have warn\n> about it. Good to know.\n\nAsked GCC to implement it: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=87983\n"},{"id":"362977","messageId":"87r2fq3b9t.fsf@evledraar.gmail.com","threadId":"49406","inReplyTo":"CABPp-BEefqYADr8SVvh6uFWkp96PDv7qfKK1c9O1WUnPy3wqrw@mail.gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-11-12T09:17:34Z","receivedAt":"2018-11-12T09:17:41Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Nov 01 2018, Elijah Newren wrote:\n\n> On Wed, Oct 31, 2018 at 12:16 PM Lars Schneider\n> <larsxschneider@gmail.com> wrote:\n>> > On Sep 24, 2018, at 7:24 PM, Elijah Newren <newren@gmail.com> wrote:\n>> > On Sun, Sep 23, 2018 at 6:08 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n>> >>\n>> >> Hi,\n>> >>\n>> >> I recently had to purge files from large Git repos (many files, many commits).\n>> >> The usual recommendation is to use `git filter-branch --index-filter` to purge\n>> >> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n>> >> remove the `builtin` directory from git core). I realized that I can remove\n>> >> files *way* faster by exporting the repo, removing the file references,\n>> >> and then importing the repo (see Perl script below, it takes ~30sec to remove\n>> >> the `builtin` directory from git core). Do you see any problem with this\n>> >> approach?\n>> >\n>> > It looks like others have pointed you at other tools, and you're\n>> > already shifting to that route.  But I think it's a useful question to\n>> > answer more generally, so for those that are really curious...\n>> >\n>> >\n>> > The basic approach is fine, though if you try to extend it much you\n>> > can run into a few possible edge/corner cases (more on that below).\n>> > I've been using this basic approach for years and even created a\n>> > mini-python library[1] designed specifically to allow people to create\n>> > \"fast-filters\", used as\n>> >   git fast-export <options> | your-fast-filter | git fast-import <options>\n>> >\n>> > But that library didn't really take off; even I have rarely used it,\n>> > often opting for filter-branch despite its horrible performance or a\n>> > simple fast-export | long-sed-command | fast-import (with some extra\n>> > pre-checking to make sure the sed wouldn't unintentionally munge other\n>> > data).  BFG is great, as long as you're only interested in removing a\n>> > few big items, but otherwise doesn't seem very useful (to be fair,\n>> > it's very upfront about only wanting to solve that problem).\n>> > Recently, due to continuing questions on filter-branch and folks still\n>> > getting confused with it, I looked at existing tools, decided I didn't\n>> > think any quite fit, and started looking into converting\n>> > git_fast_filter into a filter-branch-like tool instead of just a\n>> > libary.  Found some bugs and missing features in fast-export along the\n>> > way (and have some patches I still need to send in).  But I kind of\n>> > got stuck -- if the tool is in python, will that limit adoption too\n>> > much?  It'd be kind of nice to have this tool in core git.  But I kind\n>> > of like leaving open the possibility of using it as a tool _or_ as a\n>> > library, the latter for the special cases where case-specific\n>> > programmatic filtering is needed.  But a developer-convenience library\n>> > makes almost no sense unless in a higher level language, such as\n>> > python.  I'm still trying to make up my mind about what I want (and\n>> > what others might want), and have been kind of blocking on that.  (If\n>> > others have opinions, I'm all ears.)\n>>\n>> That library sounds like a very interesting idea. Unfortunately, the\n>> referenced repo seems not to be available anymore:\n>>     git://gitorious.org/git_fast_filter/mainline.git\n>\n> Yeah, gitorious went down at a time when I was busy with enough other\n> things that I never bothered moving my repos to a new hosting site.\n> Sorry about that.\n>\n> I've got a copy locally, but I've been editing it heavily, without the\n> testing I should have in place, so I hesitate to point you at it right\n> now.  (Also, the old version failed to handle things like --no-data\n> output, which is important.)  I'll post an updated copy soon; feel\n> free to ping me in a week if you haven't heard anything yet.\n>\n>> I very much like Python. However, more recently I started to\n>> write Git tools in Perl as they work out of the box on every\n>> machine with Git installed ... and I think Perl can be quite\n>> readable if no shortcuts are used :-).\n>\n> Yeah, when portability matters, perl makes sense.  I thought about\n> switching it over, but I'm not sure I want to rewrite 1-2k lines of\n> code.  Especially since repo-filtering tools are kind of one-shot by\n> nature, and only need to be done by one person of a team, on one\n> specific machine, and won't affect daily development thereafter.\n> (Also, since I don't depend on any libraries and use only stuff from\n> the default python library, it ought to be relatively portable\n> anyway.)\n\nFWIW I'd be very happy to have this tool itself included in git.git\nif/when it's stable / useful enough, and as you point out the language\ndoesn't really matter as much as what features it exposes.\n"},{"id":"362988","messageId":"20181112113123.GA470@sigill.intra.peff.net","threadId":"49406","inReplyTo":"87va532x5i.fsf@evledraar.gmail.com","subject":"Re: [PATCH 03/10] fast-export: use value from correct enum","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T11:31:23Z","receivedAt":"2018-11-12T11:31:27Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2018 at 09:10:17PM +0100, Ævar Arnfjörð Bjarmason wrote:\n\n> > This is a good argument for naming these SIGNED_TAG_ABORT, etc. But this\n> > is obviously an improvement in the meantime.\n> \n> In C enum values aren't the types of the enum, but I'd thought someone\n> would have added a warning for this:\n> \n>     #include <stdio.h>\n> \n>     enum { A, B } foo = A;\n>     enum { C, D } bar = C;\n> \n>     int main(void)\n>     {\n>         switch (foo) {\n>           case C:\n>             puts(\"A\");\n>             break;\n>           case B:\n>             puts(\"B\");\n>             break;\n>         }\n>     }\n> \n> But none of the 4 C compilers (gcc, clang, suncc & xlc) I have warn\n> about it. Good to know.\n\nThere is -Wenum-compare, but it does not seem to catch this (and is\nenabled by -Wall). It (gcc, at least) does catch:\n\n\tenum foo { A, B };\n\tenum bar { C, D };\n\n\tint f(enum foo x)\n\t{\n\t\treturn x == C;\n\t}\n\nbut converting that equality check to:\n\n\tswitch (x) {\n\tcase C:\n\t\treturn 1;\n\tdefault:\n\t\treturn 0;\n\t}\n\nis not (which is essentially the same as your snippet). So I think the\nbug / feature request is to have -Wenum-compare apply to switch\nstatements.\n\nClang has -Wenum-compare-switch, but I cannot seem to get it to complain\nabout even the \"==\" version using -Wenum-compare. Not sure if it's\nbuggy, or if I'm holding it wrong. This patch seems to be what we want:\n\n  https://reviews.llvm.org/D36407\n\n-Peff\n"},{"id":"363001","messageId":"20181112123232.GF3956@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BFy1aS3mHGF99Lr=+APruzC3pF5PCEph8SU71uuyOnQ7Q@mail.gmail.com","subject":"Re: [PATCH 04/10] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T12:32:32Z","receivedAt":"2018-11-12T12:32:36Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Nov 10, 2018 at 11:38:45PM -0800, Elijah Newren wrote:\n\n> > Hmm. That's the right thing to do if we're considering the export to be\n> > an independent unit. But what if I'm just rewriting a portion of history\n> > like:\n> >\n> >   git fast-export HEAD~5..HEAD | some_filter | git fast-import\n> >\n> > ? If I have a tag pointing to HEAD~10, will this delete that? Ideally I\n> > think it would be left alone.\n> \n> A couple things:\n>   * This code path only triggers in a very specific case: If a tag is\n> requested for export but points to a commit which is filtered out by\n> something else (e.g. path limiters and the commit in question didn't\n> modify any of the relevant paths), AND the user explicitly specified\n> --tag-of-filtered-object=rewrite (so that the tag in question can be\n> rewritten to the nearest non-filtered ancestor).\n\nRight, I think this is the bit I was missing: somebody has to have\nexplicitly asked to export the tag. At which point the only sensible\nthing to do is drop it.\n\n-Peff\n"},{"id":"363004","messageId":"20181112124547.GG3956@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BGF8C5vhyVbAwpmXeii452fBgtvL4dPRLWdOPxLiCYR0A@mail.gmail.com","subject":"Re: [PATCH 06/10] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T12:45:47Z","receivedAt":"2018-11-12T12:45:51Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2018 at 12:01:43AM -0800, Elijah Newren wrote:\n\n> > It does seem funny that the behavior for the earlier case (bounded\n> > commits) and this case (skipping some commits) are different. Would you\n> > ever want to keep walking backwards to find an ancestor in the earlier\n> > case? Or vice versa, would you ever want to simply delete a tag in a\n> > case like this one?\n> >\n> > I'm not sure sure, but I suspect you may have thought about it a lot\n> > harder than I have. :)\n> \n> I'm not sure why you thought the behavior for the two cases was\n> different?  For both patches, my testcases used path limiting; it was\n> you who suggested employing a negative revision to bound the commits.\n\nSorry, I think I just got confused. I was thinking about the\ndocumentation fixup you started with, which did regard bounded commits.\nBut that's not relevant here.\n\n> Anyway, for both patches assuming you haven't bounded the commits, you\n> can attempt to keep walking backwards to find an earlier ancestor, but\n> the fundamental fact is you aren't guaranteed that you can find one\n> (i.e. some tag or branch points to a commit that didn't modify any of\n> the specified paths, and nor did any of its ancestors back to any root\n> commits).  I hit that case lots of times.  If the user explicitly\n> requested a tag or branch for export (and requested tag rewriting),\n> and limited to certain paths that had never existed in the repository\n> as of the time of the tag or branch, then you hit the cases these\n> patches worry about.  Patch 4 was about (annotated and signed) tags,\n> this patch is about unannotated tags and branches and other refs.\n\nOK, that makes more sense.\n\nSo I guess my question is: in patch 4, why do we not walk back to find\nan appropriate ancestor pointed to by the signed tag object, as we do\nhere for the unannotated case?\n\nAnd I think the answer is: we already do that. It's just that the\nunannotated case never learned the same trick. So basically it's:\n\n  1. rewriting annotated tags to ancestors is already known on \"master\"\n\n  2. patch 4 further teaches it to drop a tag when that fails\n\n  3. patch 6 teaches both (1) and (2) to the unannotated code path,\n     which knew neither\n\nIs that right?\n\n> > This hunk makes sense.\n> \n> Cool, this was the entirety of the code...so does this mean that the\n> code makes more sense than my commit message summary did?  ...and\n> perhaps that my attempts to answer your questions in this email\n> weren't necessary anymore?\n\nNo, it only made sense that the hunk implemented what you claimed in the\ncommit message. ;)\n\nI think your responses did help me understand that what the commit\nmessage is claiming is a good thing.\n\n-Peff\n"},{"id":"363005","messageId":"20181112125341.GH3956@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BGNt0FcqiT=OqctjOEvY9ewNUJZ-Rs_aVEihjbQt3K8tQ@mail.gmail.com","subject":"Re: [PATCH 09/10] fast-export: add a --show-original-ids option to show original names","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T12:53:41Z","receivedAt":"2018-11-12T12:53:44Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2018 at 12:32:22AM -0800, Elijah Newren wrote:\n\n> > >  Documentation/git-fast-export.txt |  7 +++++++\n> > >  builtin/fast-export.c             | 20 +++++++++++++++-----\n> > >  fast-import.c                     | 17 +++++++++++++++++\n> > >  t/t9350-fast-export.sh            | 17 +++++++++++++++++\n> > >  4 files changed, 56 insertions(+), 5 deletions(-)\n> >\n> > The fast-import format is documented in Documentation/git-fast-import.txt.\n> > It might need an update to cover the new format.\n> \n> We document the format in both fast-import.c and\n> Documentation/git-fast-import.txt?  Maybe we should delete the long\n> comments in fast-import.c so this isn't duplicated?\n\nYes, that is probably worth doing (see the comment at the top of\nfast-import.c). Some information might need to be migrated.\n\nIf we're going to have just one spot, I think it needs to be the\nuser-facing documentation. This is a public interface that other people\nare building compatible implementations for (including your new tool).\n\n> > > +--show-original-ids::\n> > > +     Add an extra directive to the output for commits and blobs,\n> > > +     `originally <SHA1SUM>`.  While such directives will likely be\n> > > +     ignored by importers such as git-fast-import, it may be useful\n> > > +     for intermediary filters (e.g. for rewriting commit messages\n> > > +     which refer to older commits, or for stripping blobs by id).\n> >\n> > I'm not quite sure how a blob ends up being rewritten by fast-export (I\n> > get that commits may change due to dropping parents).\n> \n> It doesn't get rewritten by fast-export; it gets rewritten by other\n> intermediary filters, e.g. in something like this:\n> \n>    git fast-export --show-original-ids --all | intermediary_filter |\n> git fast-import\n> \n> The intermediary_filter program may want to strip out blobs by id, or\n> remove filemodify and filedelete directives unless they touch certain\n> paths, etc.\n\nOK, that matches my understanding. So why does fast-export need to print\nthe blob ids? If the intermediary is rewriting blobs, it can then\nproduce the \"originally\" line itself, can't it?\n\nThe more interesting case I guess is your \"strip out blobs by id\"\nexample. There the intermediary _could_ do so itself, but it would\nrequire recomputing the object id of each blob.\n\nIf you use \"--no-data\", then this just works (we specify tree entries by\nobject id, rather than by mark). But I can see how it would be useful to\nhave the information even without \"--no-data\" (i.e., if you are doing\nmultiple kinds of rewrites on a single stream).\n\nI think the thing that confused me is that this \"originally\" is doing\ntwo things:\n\n  - mentioning blob ids as an optimization / convenience for the reader\n\n  - mentioning rewritten commit (and presumably tag?) ids that were\n    rewritten as part of a partial history export. I suppose even trees\n    could be rewritten that way, too, but fast-import doesn't generally\n    consider trees to be a first-class item.\n\nSo I'm OK with it, but I wonder if there is an easier way to explain it.\n\n-Peff\n"},{"id":"363006","messageId":"20181112125847.GI3956@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BGREOAvF-6DBymdwsUL2LpyPNqy8dCw0RuUKZf2Da6cJA@mail.gmail.com","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T12:58:48Z","receivedAt":"2018-11-12T12:58:51Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2018 at 12:42:58AM -0800, Elijah Newren wrote:\n\n> > > fast-export output is traditionally used as an input to a fast-import\n> > > program, but it is also useful to help gather statistics about the\n> > > history of a repository (particularly when --no-data is also passed).\n> > > For example, two of the types of information we may want to collect\n> > > could include:\n> > >   1) general information about renames that have occurred\n> > >   2) what the biggest objects in a repository are and what names\n> > >      they appear under.\n> > >\n> > > The first bit of information can be gathered by just passing -M to\n> > > fast-export.  The second piece of information can partially be gotten\n> > > from running\n> > >     git cat-file --batch-check --batch-all-objects\n> > > However, that only shows what the biggest objects in the repository are\n> > > and their sizes, not what names those objects appear as or what commits\n> > > they were introduced in.  We can get that information from fast-export,\n> > > but when we only see\n> > >     R oldname newname\n> > > instead of\n> > >     R oldname newname\n> > >     M 100644 $SHA1 newname\n> > > then it makes the job more difficult.  Add an option which allows us to\n> > > force the latter output even when commits have exact renames of files.\n> >\n> > fast-export seems like a funny tool to look up paths. What about \"git\n> > log --find-object=$SHA1\" ?\n> \n> Eek, and give me O(N*M) behavior, where N is the number of commits in\n> the repository and M is the number of renames that occur in its\n> history?  Also, that's the inverse of the lookup I need anyway (I have\n> the commit and filename, but am missing the SHA).\n\nMaybe I don't understand what you're trying to accomplish. I was\nthinking specifically of your \"cat-file can tell you the large objects,\nbut you don't know their names/commits\" from above.\n\nI would do:\n\n   git log --raw $(\n     git cat-file --batch-check='%(objectsize:disk) %(objectname)' --batch-all-objects |\n     sort -rn | head -3 |\n     awk '{print \"--find-object=\" $2 }'\n   )\n\nI'm not sure how renames enter into it at all.\n\n> One of the problems with filter-branch that people often run into is\n> they know what they want at a high-level (e.g. extract the history of\n> this directory for a new repository, or rewrite the history of this\n> repo to appear at a subdirectory so it can be merged into a bigger\n> repo and people passing filenames to log will still get the history of\n> those files, or I want to remove some of the big stuff in my history),\n> but often times that's not quite enough.  They need help finding big\n> objects, or may be unaware that the subset of files they want used to\n> be known by alternative names.\n> \n> I want a simple --analyze mode that can report on all files that have\n> been renamed (so users don't just say \"all I care about is these N\n> files, give me a rewritten history just including those\" -- we can\n> point out to them whether those N files used to be known by other\n> names), as well as reporting on all big files and if they've been\n> deleted, and aggregations of the \"big files\" information across\n> directories and file extensions.\n\nSo this seems like a separate problem than what the commit message talks\nabout.\n\nThere I think you'd want to assemble the list with something like \"git\nlog --follow --name-only paths-of-interest\" except that --follow sucks\ntoo much to handle more than one path at a time.\n\nBut if you wanted to do it manually, then:\n\n  git log --diff-filter=R --name-only\n\nwould be enough to let you track it down, wouldn't it?\n\n-Peff\n"},{"id":"363007","messageId":"20181112130007.GJ3956@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BGzqpxF_+ubp2cft9dQ-03pgcCxJEP13VOUv5WADHDjnA@mail.gmail.com","subject":"Re: [PATCH 00/10] fast export and import fixes and features","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T13:00:07Z","receivedAt":"2018-11-12T13:00:10Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2018 at 12:44:47AM -0800, Elijah Newren wrote:\n\n> > > These patches were driven by the needs of git-repo-filter[1], but most\n> > > if not all of them should be independently useful.\n> >\n> > I left lots of comments. Some of the earlier ones may just be showing my\n> > confusion about fast-export works (some of which was cleared up by your\n> > later patches). But I like the overall direction for sure.\n> \n> Thanks for taking the time to read over the series and providing lots\n> of feedback!  And, whoops, looks like it's gotten kinda late, so I'll\n> check any further feedback on Monday.\n\nThank you for your patience with my sometimes-confused responses. :)\n\nOverall it makes more sense to me now (and everything seems like a good\ndirection), with the exception that I'm still a bit confused about patch\n10.\n\n-Peff\n"},{"id":"363044","messageId":"CABPp-BGVhw6HeCb7wTUubEjqxfW3LopB8PXY1TdHrB9Gfd3_jw@mail.gmail.com","threadId":"49406","inReplyTo":"87r2fq3b9t.fsf@evledraar.gmail.com","subject":"Re: Import/Export as a fast way to purge files from Git?","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-12T15:34:50Z","receivedAt":"2018-11-12T15:35:05Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 12, 2018 at 1:17 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Thu, Nov 01 2018, Elijah Newren wrote:\n>\n> > On Wed, Oct 31, 2018 at 12:16 PM Lars Schneider\n> > <larsxschneider@gmail.com> wrote:\n> >> > On Sep 24, 2018, at 7:24 PM, Elijah Newren <newren@gmail.com> wrote:\n> >> > On Sun, Sep 23, 2018 at 6:08 AM Lars Schneider <larsxschneider@gmail.com> wrote:\n> >> >>\n> >> >> Hi,\n> >> >>\n> >> >> I recently had to purge files from large Git repos (many files, many commits).\n> >> >> The usual recommendation is to use `git filter-branch --index-filter` to purge\n> >> >> files. However, this is *very* slow for large repos (e.g. it takes 45min to\n> >> >> remove the `builtin` directory from git core). I realized that I can remove\n> >> >> files *way* faster by exporting the repo, removing the file references,\n> >> >> and then importing the repo (see Perl script below, it takes ~30sec to remove\n> >> >> the `builtin` directory from git core). Do you see any problem with this\n> >> >> approach?\n> >> >\n> >> > It looks like others have pointed you at other tools, and you're\n> >> > already shifting to that route.  But I think it's a useful question to\n> >> > answer more generally, so for those that are really curious...\n> >> >\n> >> >\n> >> > The basic approach is fine, though if you try to extend it much you\n> >> > can run into a few possible edge/corner cases (more on that below).\n> >> > I've been using this basic approach for years and even created a\n> >> > mini-python library[1] designed specifically to allow people to create\n> >> > \"fast-filters\", used as\n> >> >   git fast-export <options> | your-fast-filter | git fast-import <options>\n> >> >\n> >> > But that library didn't really take off; even I have rarely used it,\n> >> > often opting for filter-branch despite its horrible performance or a\n> >> > simple fast-export | long-sed-command | fast-import (with some extra\n> >> > pre-checking to make sure the sed wouldn't unintentionally munge other\n> >> > data).  BFG is great, as long as you're only interested in removing a\n> >> > few big items, but otherwise doesn't seem very useful (to be fair,\n> >> > it's very upfront about only wanting to solve that problem).\n> >> > Recently, due to continuing questions on filter-branch and folks still\n> >> > getting confused with it, I looked at existing tools, decided I didn't\n> >> > think any quite fit, and started looking into converting\n> >> > git_fast_filter into a filter-branch-like tool instead of just a\n> >> > libary.  Found some bugs and missing features in fast-export along the\n> >> > way (and have some patches I still need to send in).  But I kind of\n> >> > got stuck -- if the tool is in python, will that limit adoption too\n> >> > much?  It'd be kind of nice to have this tool in core git.  But I kind\n> >> > of like leaving open the possibility of using it as a tool _or_ as a\n> >> > library, the latter for the special cases where case-specific\n> >> > programmatic filtering is needed.  But a developer-convenience library\n> >> > makes almost no sense unless in a higher level language, such as\n> >> > python.  I'm still trying to make up my mind about what I want (and\n> >> > what others might want), and have been kind of blocking on that.  (If\n> >> > others have opinions, I'm all ears.)\n> >>\n> >> That library sounds like a very interesting idea. Unfortunately, the\n> >> referenced repo seems not to be available anymore:\n> >>     git://gitorious.org/git_fast_filter/mainline.git\n> >\n> > Yeah, gitorious went down at a time when I was busy with enough other\n> > things that I never bothered moving my repos to a new hosting site.\n> > Sorry about that.\n> >\n> > I've got a copy locally, but I've been editing it heavily, without the\n> > testing I should have in place, so I hesitate to point you at it right\n> > now.  (Also, the old version failed to handle things like --no-data\n> > output, which is important.)  I'll post an updated copy soon; feel\n> > free to ping me in a week if you haven't heard anything yet.\n> >\n> >> I very much like Python. However, more recently I started to\n> >> write Git tools in Perl as they work out of the box on every\n> >> machine with Git installed ... and I think Perl can be quite\n> >> readable if no shortcuts are used :-).\n> >\n> > Yeah, when portability matters, perl makes sense.  I thought about\n> > switching it over, but I'm not sure I want to rewrite 1-2k lines of\n> > code.  Especially since repo-filtering tools are kind of one-shot by\n> > nature, and only need to be done by one person of a team, on one\n> > specific machine, and won't affect daily development thereafter.\n> > (Also, since I don't depend on any libraries and use only stuff from\n> > the default python library, it ought to be relatively portable\n> > anyway.)\n>\n> FWIW I'd be very happy to have this tool itself included in git.git\n> if/when it's stable / useful enough, and as you point out the language\n> doesn't really matter as much as what features it exposes.\n\nWell, I'm happy to propose it for inclusion once it gets to that\npoint.  I'll bring it up on the list to get wider feedback once I've\nremoved at least some of the sharp edges.  I suspect it'll be at least\na few weeks.\n"},{"id":"363046","messageId":"CABPp-BHLEtXe-2OTHNxHe=vypvbd-kFQ3G1FaVGnQ-Gc4+z1uA@mail.gmail.com","threadId":"49406","inReplyTo":"20181112124547.GG3956@sigill.intra.peff.net","subject":"Re: [PATCH 06/10] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-12T15:36:36Z","receivedAt":"2018-11-12T15:36:50Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 12, 2018 at 4:45 AM Jeff King <peff@peff.net> wrote:\n> On Sun, Nov 11, 2018 at 12:01:43AM -0800, Elijah Newren wrote:\n>\n> > > It does seem funny that the behavior for the earlier case (bounded\n> > > commits) and this case (skipping some commits) are different. Would you\n> > > ever want to keep walking backwards to find an ancestor in the earlier\n> > > case? Or vice versa, would you ever want to simply delete a tag in a\n> > > case like this one?\n> > >\n> > > I'm not sure sure, but I suspect you may have thought about it a lot\n> > > harder than I have. :)\n> >\n> > I'm not sure why you thought the behavior for the two cases was\n> > different?  For both patches, my testcases used path limiting; it was\n> > you who suggested employing a negative revision to bound the commits.\n>\n> Sorry, I think I just got confused. I was thinking about the\n> documentation fixup you started with, which did regard bounded commits.\n> But that's not relevant here.\n>\n> > Anyway, for both patches assuming you haven't bounded the commits, you\n> > can attempt to keep walking backwards to find an earlier ancestor, but\n> > the fundamental fact is you aren't guaranteed that you can find one\n> > (i.e. some tag or branch points to a commit that didn't modify any of\n> > the specified paths, and nor did any of its ancestors back to any root\n> > commits).  I hit that case lots of times.  If the user explicitly\n> > requested a tag or branch for export (and requested tag rewriting),\n> > and limited to certain paths that had never existed in the repository\n> > as of the time of the tag or branch, then you hit the cases these\n> > patches worry about.  Patch 4 was about (annotated and signed) tags,\n> > this patch is about unannotated tags and branches and other refs.\n>\n> OK, that makes more sense.\n>\n> So I guess my question is: in patch 4, why do we not walk back to find\n> an appropriate ancestor pointed to by the signed tag object, as we do\n> here for the unannotated case?\n>\n> And I think the answer is: we already do that. It's just that the\n> unannotated case never learned the same trick. So basically it's:\n>\n>   1. rewriting annotated tags to ancestors is already known on \"master\"\n>\n>   2. patch 4 further teaches it to drop a tag when that fails\n>\n>   3. patch 6 teaches both (1) and (2) to the unannotated code path,\n>      which knew neither\n>\n> Is that right?\n\nAh, now I see where the slight disconnect was.  And yes, you are correct.\n\n> > > This hunk makes sense.\n> >\n> > Cool, this was the entirety of the code...so does this mean that the\n> > code makes more sense than my commit message summary did?  ...and\n> > perhaps that my attempts to answer your questions in this email\n> > weren't necessary anymore?\n>\n> No, it only made sense that the hunk implemented what you claimed in the\n> commit message. ;)\n>\n> I think your responses did help me understand that what the commit\n> message is claiming is a good thing.\n"},{"id":"363049","messageId":"CABPp-BG6FJjFm7ZFWpe--n3-vXzAcrQYWXmx4M4hA_kkSPJhkQ@mail.gmail.com","threadId":"49406","inReplyTo":"20181112125341.GH3956@sigill.intra.peff.net","subject":"Re: [PATCH 09/10] fast-export: add a --show-original-ids option to show original names","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-12T15:46:14Z","receivedAt":"2018-11-12T15:46:30Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 12, 2018 at 4:53 AM Jeff King <peff@peff.net> wrote:\n> On Sun, Nov 11, 2018 at 12:32:22AM -0800, Elijah Newren wrote:\n>\n> > > >  Documentation/git-fast-export.txt |  7 +++++++\n> > > >  builtin/fast-export.c             | 20 +++++++++++++++-----\n> > > >  fast-import.c                     | 17 +++++++++++++++++\n> > > >  t/t9350-fast-export.sh            | 17 +++++++++++++++++\n> > > >  4 files changed, 56 insertions(+), 5 deletions(-)\n> > >\n> > > The fast-import format is documented in Documentation/git-fast-import.txt.\n> > > It might need an update to cover the new format.\n> >\n> > We document the format in both fast-import.c and\n> > Documentation/git-fast-import.txt?  Maybe we should delete the long\n> > comments in fast-import.c so this isn't duplicated?\n>\n> Yes, that is probably worth doing (see the comment at the top of\n> fast-import.c). Some information might need to be migrated.\n>\n> If we're going to have just one spot, I think it needs to be the\n> user-facing documentation. This is a public interface that other people\n> are building compatible implementations for (including your new tool).\n\nOkay, I'll work on that.\n\n> OK, that matches my understanding. So why does fast-export need to print\n> the blob ids? If the intermediary is rewriting blobs, it can then\n> produce the \"originally\" line itself, can't it?\n>\n> The more interesting case I guess is your \"strip out blobs by id\"\n> example. There the intermediary _could_ do so itself, but it would\n> require recomputing the object id of each blob.\n>\n> If you use \"--no-data\", then this just works (we specify tree entries by\n> object id, rather than by mark). But I can see how it would be useful to\n> have the information even without \"--no-data\" (i.e., if you are doing\n> multiple kinds of rewrites on a single stream).\n>\n> I think the thing that confused me is that this \"originally\" is doing\n> two things:\n>\n>   - mentioning blob ids as an optimization / convenience for the reader\n>\n>   - mentioning rewritten commit (and presumably tag?) ids that were\n>     rewritten as part of a partial history export. I suppose even trees\n>     could be rewritten that way, too, but fast-import doesn't generally\n>     consider trees to be a first-class item.\n>\n> So I'm OK with it, but I wonder if there is an easier way to explain it.\n\nYeah, I started out just needing to add the original oids for commits.\nOnce I added them there, I wondered whether someone would need them\nfor tags and blobs too (not trees since fast-import doesn't work with\nthose).  For blobs, it made sense as a small performance optimization\n(when running without --no-data), as you pointed out.  I can't think\nof a use for them in tags, but once I've included them in blobs and\ncommits it felt like I might as well include them there for\ncompleteness.  So maybe my commit message should have been something\nmore like:\n\n\"\"\"\nKnowing the original names (hashes) of commits can sometimes enable\npost-filtering that would otherwise be difficult or impossible.  In\nparticular, the desire to rewrite commit messages which refer to other\nprior commits (on top of whatever other filtering is being done) is\nvery difficult without knowing the original names of each commit.\n\nIn addition, knowing the original names (hashes) of blobs can allow\nfiltering by blob-id without requiring re-hashing the content of the\nblob, and is thus useful as a small optimization.\n\nOnce we add original ids for both commits and blobs, we may as well\nadd them for tags too for completeness.  Perhaps someone will have a\nuse for them.\n\nThis commit teaches a new --show-original-ids option to fast-export\nwhich will make it add a 'original-oid <hash>' line to blob, commits,\nand tags.  It also teaches fast-import to parse (and ignore) such\nlines.\n\"\"\"\n\n?\n"},{"id":"363061","messageId":"20181112163124.GA5735@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BG6FJjFm7ZFWpe--n3-vXzAcrQYWXmx4M4hA_kkSPJhkQ@mail.gmail.com","subject":"Re: [PATCH 09/10] fast-export: add a --show-original-ids option to show original names","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-12T16:31:24Z","receivedAt":"2018-11-12T16:31:28Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Nov 12, 2018 at 07:46:14AM -0800, Elijah Newren wrote:\n\n> So maybe my commit message should have been something\n> more like:\n> \n> \"\"\"\n> Knowing the original names (hashes) of commits can sometimes enable\n> post-filtering that would otherwise be difficult or impossible.  In\n> particular, the desire to rewrite commit messages which refer to other\n> prior commits (on top of whatever other filtering is being done) is\n> very difficult without knowing the original names of each commit.\n> \n> In addition, knowing the original names (hashes) of blobs can allow\n> filtering by blob-id without requiring re-hashing the content of the\n> blob, and is thus useful as a small optimization.\n> \n> Once we add original ids for both commits and blobs, we may as well\n> add them for tags too for completeness.  Perhaps someone will have a\n> use for them.\n> \n> This commit teaches a new --show-original-ids option to fast-export\n> which will make it add a 'original-oid <hash>' line to blob, commits,\n> and tags.  It also teaches fast-import to parse (and ignore) such\n> lines.\n> \"\"\"\n> \n> ?\n\nYes, that makes much more sense to me (though of course I've also been\ndiscussing it with you, so just about anything would at this point ;) ).\n\nIt's possible that somebody would want to filter on tree id's, too. A\nfast-import stream just has trees incidentally as part of commit state,\nbut we could say something like \"by the way, this tree is X\". You can\neven do \"fast-export -t\" to see subtrees, though I am not sure if that\nis intentional or just an artifact of being based on the diff code.\n\nI guess that is not all that useful, though. I was mostly thinking about\nit because of your \"we may as well add them for tags too for\ncompleteness\" above. But the issues around trees are sufficiently subtle\nthat we're probably better off not trying to handle them here. There's a\ngood chance we'd get it wrong, making our \"let's just add this for\ncompleteness while we're here\" totally backfire.\n\n-Peff\n"},{"id":"363065","messageId":"CABPp-BHjPq-2JoeXur+FMs+T==arqvMaAW1uLKMSHKBKBS60rA@mail.gmail.com","threadId":"49406","inReplyTo":"20181112125847.GI3956@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-12T18:08:10Z","receivedAt":"2018-11-12T18:08:26Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Mon, Nov 12, 2018 at 4:58 AM Jeff King <peff@peff.net> wrote:\n> On Sun, Nov 11, 2018 at 12:42:58AM -0800, Elijah Newren wrote:\n>\n> Maybe I don't understand what you're trying to accomplish. I was\n> thinking specifically of your \"cat-file can tell you the large objects,\n> but you don't know their names/commits\" from above.\n\nFair enough.  And just to be clear, the first 9 patches were fixes and\nfeatures around trying to rewrite history; patch 10 is orthogonal and\nwas used for a separate run to just gather data.  It is entirely\npossible I could gather that data other ways.\n\n> I would do:\n>\n>    git log --raw $(\n>      git cat-file --batch-check='%(objectsize:disk) %(objectname)' --batch-all-objects |\n>      sort -rn | head -3 |\n>      awk '{print \"--find-object=\" $2 }'\n>    )\n>\n> I'm not sure how renames enter into it at all.\n\nHow did I miss objectsize:disk??  Especially since it is right next to\nobjectsize in the manpage to boot?  That's awesome, thanks for that\npointer.\n\nI do have a separate cat-file --batch-check --batch-all-objects\nprocess already, since I can't get sizes out of either log or\nfast-export.  However, I wouldn't use your 'head -3' since I'm not\nlooking for the N biggest, but reporting on _all_ objects (in reverse\nsize order) and letting the user look over the report and deciding\nwhere to stop reading.  So, this is a big and expensive log command.\nGranted, we will need a big and expensive log command, but let's keep\nin mind that we have this one.\n\n> > One of the problems with filter-branch that people often run into is\n> > they know what they want at a high-level (e.g. extract the history of\n> > this directory for a new repository, or rewrite the history of this\n> > repo to appear at a subdirectory so it can be merged into a bigger\n> > repo and people passing filenames to log will still get the history of\n> > those files, or I want to remove some of the big stuff in my history),\n> > but often times that's not quite enough.  They need help finding big\n> > objects, or may be unaware that the subset of files they want used to\n> > be known by alternative names.\n> >\n> > I want a simple --analyze mode that can report on all files that have\n> > been renamed (so users don't just say \"all I care about is these N\n> > files, give me a rewritten history just including those\" -- we can\n> > point out to them whether those N files used to be known by other\n> > names), as well as reporting on all big files and if they've been\n> > deleted, and aggregations of the \"big files\" information across\n> > directories and file extensions.\n>\n> So this seems like a separate problem than what the commit message talks\n> about.\n>\n> There I think you'd want to assemble the list with something like \"git\n> log --follow --name-only paths-of-interest\" except that --follow sucks\n> too much to handle more than one path at a time.\n>\n> But if you wanted to do it manually, then:\n>\n>   git log --diff-filter=R --name-only\n>\n> would be enough to let you track it down, wouldn't it?\n\nWithout a -M you'd only catch 100% renames, right?  Those aren't the\nonly ones I'd want to catch, so I'd need to add -M.  You are right\nthat we could get basic renames this way, but it doesn't cover\neverything I need.  Let's use this as a starting point, though, and\nbuild up to what I need...\n\nI also want to know when files were deleted.  I've generally found\nthat people are more okay with purging parts of history [corresponding\nto large ojbects] that were deleted longer ago than more recent stuff,\nfor a variety of reasons.  So we could either run yet another log, or\nmodify the command to:\n\n  git log -M --diff-filter=RD --name-status\n\nHowever, I don't just want to know when files were deleted, I'd like\nto know when directories are deleted.  I only knew how to derive that\nfrom knowing what files existed within those directories, so that\nwould take me to:\n\n  git log -M --diff-filter=RAD --name-status\n\n[Edit: I just saw your other email and for the first time learned\nabout the -t rev-list option which might simplify this a little,\nalthough \"need to worry about deleted files being reinstated\" below\nmight require the 'A' anyway.]\n\nAt this point, let's remember that we had another full git-log\ninvocation for mapping object sizes to filenames.  We might as well\ncoalesce the two log commands into one, by extending this latest one\nto:\n\n  git log -M --diff-filter=RAMD --no-abbrev --raw\n\nAlso, I wanted commit date rather than author date, so we need to\nextend the headers a bit.  Also, for reasons I won't bother detailing,\nI think I want to traverse commits in reverse topological order.  So\nour command is:\n\n  git log --pretty=fuller --topo-order --reverse -M --diff-filter=RAMD\n--no-abbrev --raw\n\nBut that still leaves us with four problems, three of which we can\nsolve with further extensions to this command:\n\n1) There are some weird edge cases with deletions and renames.  Lots\nof them in fact.  At a simple level, branching and merging and\nmultiple refs means that \"is-this-deleted\" isn't a binary flag for a\ngiven filename (but rather a binary flag per-ref).  Also, it makes\n\"the set of names associated with a single 'file' as perceived by the\nuser\" possibly rather ill-defined as well.  This can get really hairy,\nbut I'd at least like to handle the very basic cases of (a) \"user\nre-instates filename that used to be deleted\" (i.e. the file isn't\ndeleted anymore) and (b) \"user re-instates a filename that used to\nexist but was renamed to something else\" (in such cases, we can't just\ntreat the two filenames as being different names of the same content).\nHandling the (b) usecase sanely requires some topology information, so\nwe need parents as well.  So our command extends to:\n\n   git log --parents --pretty=fuller --topo-order --reverse -M\n--diff-filter=RAMD --no-abbrev --raw\n\n2) log is not plumbing, so parsing the stuff before the file\nmodifications is not a good idea. This could be fixed by using\n--format:\n\n  git log --format='%H%n%P%n%cd' --date=short --topo-order --reverse\n-M --diff-filter-RAMD --no-abbrev --raw\n\n3) log won't show changes for merge commits by default; we'd need to add -c:\n\n  git log --format='%H%n%P%n%cd' --date=short --topo-order --reverse\n-M --diff-filter-RAMD --no-abbrev --raw -c\n\n4) log is not plumbing, revisited: although at this point I've\nspecified the log output explicitly enough that it ought to be safe to\nparse, there are a few things that make me slightly worried.  I can\ndepend on fast-export to be stable; it only gives 'M' and 'D' unless\nyou explicitly ask for more types (e.g. -M to detect renames will add\n'R').  With log, I'm no so sure; do I need to worry about new types\nappearing in the future?  Also, should I just drop --diff-filter=RAMD\nsince it covers just about everything anyway?  Also, while --raw is\nstable, is the combination of -c and --raw stable?  Is --date=short\nstable (most likely, but still seems more likely to change than\nfast-export would be)?  Is there something else I need to be worried\nabout?  Granted, each of those is only a small worry with log, but\nthey add up and give me pause about whether I should be parsing it\noutput in another tool.\n\n\n\nSo we've come up with an alternate way to get the data I need, though\nwith some worries.\n\nI could potentially switch to using this and drop patch 10/10.  Maybe\nthere's even a good reason to prefer using log.  But at the time I was\nthinking in terms of \"I already have a tool that parses fast-export\noutput and I know it's stable...and it has access to all the\ninformation I need so why not just get the information from it?\"  So I\ndid that, and then realized towards the end that although it had all\nthe needed info, it stripped one piece from me.  Namely, when it had a\n100% rename, I'd only get\n   R oldname newname\nand wouldn't know the sha1sum of newname (for mapping object sizes to\nall their names).  If I cached the information about all file shas for\nall trees I could pull it from that cache (which could be expensive\nmemory-wise for large repos), or I could use the original-oid\ndirective and keep another long running \"git cat-file\n--batch-check='%(objectname)' process and just pass it\n\"$ORIGINAL_OID:$NEWNAME\" lines as I come across them.  However,\nfast-export had the information and did special work to try to avoid\nshowing it when it thought it woudln't be needed, so why not just add\na flag to tell it to just give me the filemodify?\n\nAt this point, if folks don't like this patch, I'm more likely to use\nthe supplementary cat-file process than switching to log, unless\nsomeone can ameliorate my concerns with it and suggest a good reason\nwhy it's actually better.\n\n\n\nAnyway, I hope it makes a little more sense why I created this patch.\nDoes it, or have I just made things even more confusing?\n\n...and if you've read this far, I'm impressed.  Thanks for reading.\n"},{"id":"363095","messageId":"20181112225043.GJ890086@genre.crustytoothpaste.net","threadId":"49406","inReplyTo":"20181111064442.GD30850@sigill.intra.peff.net","subject":"Re: [PATCH 04/10] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-11-12T22:50:43Z","receivedAt":"2018-11-12T22:50:52Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Nov 11, 2018 at 01:44:43AM -0500, Jeff King wrote:\n> > +\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar >output &&\n> > +\t\tgrep -A 1 refs/tags/v0.0.0.0.0.0.1 output | grep -E ^from.0{40}\n> \n> I don't think \"grep -A\" is portable (and we don't seem to otherwise use\n> it). You can probably do something similar with sed.\n> \n> Use $ZERO_OID instead of hard-coding 40, which future-proofs for the\n> hash transition (though I suppose the hash is not likely to get\n> _shorter_ ;) ).\n\nIt would indeed be nice if we used $ZERO_OID.  Also, we prefer to write\n\"egrep\", since some less capable systems don't have a grep with -E.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"363157","messageId":"20181113143822.GA17454@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181112225043.GJ890086@genre.crustytoothpaste.net","subject":"Re: [PATCH 04/10] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-13T14:38:22Z","receivedAt":"2018-11-13T14:38:25Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Nov 12, 2018 at 10:50:43PM +0000, brian m. carlson wrote:\n\n> On Sun, Nov 11, 2018 at 01:44:43AM -0500, Jeff King wrote:\n> > > +\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar >output &&\n> > > +\t\tgrep -A 1 refs/tags/v0.0.0.0.0.0.1 output | grep -E ^from.0{40}\n> > \n> > I don't think \"grep -A\" is portable (and we don't seem to otherwise use\n> > it). You can probably do something similar with sed.\n> > \n> > Use $ZERO_OID instead of hard-coding 40, which future-proofs for the\n> > hash transition (though I suppose the hash is not likely to get\n> > _shorter_ ;) ).\n> \n> It would indeed be nice if we used $ZERO_OID.  Also, we prefer to write\n> \"egrep\", since some less capable systems don't have a grep with -E.\n\nI thought that, too, but it is only \"grep -F\" that has been a problem\nfor us in the past, and we have many \"grep -E\" calls already. c.f.\nhttps://public-inbox.org/git/20180910154453.GA15270@sigill.intra.peff.net/\n\n-Peff\n"},{"id":"363158","messageId":"20181113144554.GB17454@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BHjPq-2JoeXur+FMs+T==arqvMaAW1uLKMSHKBKBS60rA@mail.gmail.com","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-13T14:45:54Z","receivedAt":"2018-11-13T14:45:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Nov 12, 2018 at 10:08:10AM -0800, Elijah Newren wrote:\n\n> > I would do:\n> >\n> >    git log --raw $(\n> >      git cat-file --batch-check='%(objectsize:disk) %(objectname)' --batch-all-objects |\n> >      sort -rn | head -3 |\n> >      awk '{print \"--find-object=\" $2 }'\n> >    )\n> >\n> > I'm not sure how renames enter into it at all.\n> \n> How did I miss objectsize:disk??  Especially since it is right next to\n> objectsize in the manpage to boot?  That's awesome, thanks for that\n> pointer.\n> \n> I do have a separate cat-file --batch-check --batch-all-objects\n> process already, since I can't get sizes out of either log or\n> fast-export.  However, I wouldn't use your 'head -3' since I'm not\n> looking for the N biggest, but reporting on _all_ objects (in reverse\n> size order) and letting the user look over the report and deciding\n> where to stop reading.  So, this is a big and expensive log command.\n> Granted, we will need a big and expensive log command, but let's keep\n> in mind that we have this one.\n\nIt is an expensive log command, but it's the same expense as running\nfast-export, no? And I think maybe that is the disconnect.\n\nI am looking at this problem as \"how do you answer question X in a\nrepository\". And I think you are looking at as \"I am receiving a\nfast-export stream, and I need to answer question X on the fly\".\n\nAnd that would explain why you want to get extra annotations into the\nfast-export stream. Is that right?\n\n> > There I think you'd want to assemble the list with something like \"git\n> > log --follow --name-only paths-of-interest\" except that --follow sucks\n> > too much to handle more than one path at a time.\n> >\n> > But if you wanted to do it manually, then:\n> >\n> >   git log --diff-filter=R --name-only\n> >\n> > would be enough to let you track it down, wouldn't it?\n> \n> Without a -M you'd only catch 100% renames, right?  Those aren't the\n> only ones I'd want to catch, so I'd need to add -M.  You are right\n> that we could get basic renames this way, but it doesn't cover\n> everything I need.  Let's use this as a starting point, though, and\n> build up to what I need...\n\nNo, renames are on by default these days, and that includes inexact\nrenames. That said, if you're scripting you probably ought to be doing:\n\n  git rev-list HEAD | git diff-tree --stdin\n\nand there yes, you'd have to enable \"-M\" yourself (you touched on\nscripting and formatting below; diff-tree can accept the format options\nyou'd want).\n\n> I also want to know when files were deleted.  I've generally found\n> that people are more okay with purging parts of history [corresponding\n> to large ojbects] that were deleted longer ago than more recent stuff,\n> for a variety of reasons.  So we could either run yet another log, or\n> modify the command to:\n> \n>   git log -M --diff-filter=RD --name-status\n> \n> However, I don't just want to know when files were deleted, I'd like\n> to know when directories are deleted.  I only knew how to derive that\n> from knowing what files existed within those directories, so that\n> would take me to:\n> \n>   git log -M --diff-filter=RAD --name-status\n> \n> [Edit: I just saw your other email and for the first time learned\n> about the -t rev-list option which might simplify this a little,\n> although \"need to worry about deleted files being reinstated\" below\n> might require the 'A' anyway.]\n\nYeah, I think \"-t\" would help your tree deletion problem.\n\n> At this point, let's remember that we had another full git-log\n> invocation for mapping object sizes to filenames.  We might as well\n> coalesce the two log commands into one, by extending this latest one\n> to:\n> \n>   git log -M --diff-filter=RAMD --no-abbrev --raw\n\nWhat is there besides RAMD? :)\n\n> I could potentially switch to using this and drop patch 10/10.\n\nSo I'm still not _entirely_ clear on what you're trying to do with\n10/10. I think maybe the \"disconnect\" part I wrote above explains it. If\nthat's correct, then I think framing it in terms of the operations that\nyou'd be able to perform _without running a separate traverse_ would\nmake it more obvious.\n\n> Anyway, I hope it makes a little more sense why I created this patch.\n> Does it, or have I just made things even more confusing?\n\nSome of both, I think.\n\n> ...and if you've read this far, I'm impressed.  Thanks for reading.\n\nI'll admit I skimmed near the end. ;)\n\n-Peff\n"},{"id":"363177","messageId":"CABPp-BFbtusiT30_gU7SgmmMg25NCdgNTSEEJJysoT-1MwSnkA@mail.gmail.com","threadId":"49406","inReplyTo":"20181113144554.GB17454@sigill.intra.peff.net","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-13T17:10:36Z","receivedAt":"2018-11-13T17:10:52Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Tue, Nov 13, 2018 at 6:45 AM Jeff King <peff@peff.net> wrote:\n> It is an expensive log command, but it's the same expense as running\n> fast-export, no? And I think maybe that is the disconnect.\n\nI would expect an expensive log command to generally be the same\nexpense as running fast-export, yes.  But I would expect two expensive\nlog commands to be twice the expense of a single fast-export (and you\nsuggested two log commands: both the --find-object= one and the\n--diff-filter one).\n\n> I am looking at this problem as \"how do you answer question X in a\n> repository\". And I think you are looking at as \"I am receiving a\n> fast-export stream, and I need to answer question X on the fly\".\n>\n> And that would explain why you want to get extra annotations into the\n> fast-export stream. Is that right?\n\nI'm not trying to get information on the fly during a rewrite or\nanything like that.  This is an optional pre-rewrite step (from a\nseparate invocation of the tool) where I have multiple questions I\nwant to answer.  I'd like to answer them all relatively quickly, if\npossible, and I think all of them should be answerable with a single\nhistory traversal (plus a cat-file --batch-all-objects call to get\nobject sizes, since I don't know of another way to get those).  I'd be\nfine with switching from fast-export to log or something else if it\nmet the needs better.\n\nAs far as I can tell, you're trying to split each question apart and\ndo a history traversal for each, and I don't see why that's better.\nSimpler, perhaps, but it seems worse for performance.  Am I missing\nsomething?\n\n> > > There I think you'd want to assemble the list with something like \"git\n> > > log --follow --name-only paths-of-interest\" except that --follow sucks\n> > > too much to handle more than one path at a time.\n> > >\n> > > But if you wanted to do it manually, then:\n> > >\n> > >   git log --diff-filter=R --name-only\n> > >\n> > > would be enough to let you track it down, wouldn't it?\n> >\n> > Without a -M you'd only catch 100% renames, right?  Those aren't the\n> > only ones I'd want to catch, so I'd need to add -M.  You are right\n> > that we could get basic renames this way, but it doesn't cover\n> > everything I need.  Let's use this as a starting point, though, and\n> > build up to what I need...\n>\n> No, renames are on by default these days, and that includes inexact\n> renames. That said, if you're scripting you probably ought to be doing:\n>\n>   git rev-list HEAD | git diff-tree --stdin\n>\n> and there yes, you'd have to enable \"-M\" yourself (you touched on\n> scripting and formatting below; diff-tree can accept the format options\n> you'd want).\n\nAh, I didn't know renames were on by default; I somehow missed that.\nAlso, the rev-list to diff-tree pipe is nice, but I also need parent\nand commit timestamp information.\n\n....\n> Yeah, I think \"-t\" would help your tree deletion problem.\n\nAbsolutely, thanks for the hint.  Much appreciated.  :-)\n\n> > At this point, let's remember that we had another full git-log\n> > invocation for mapping object sizes to filenames.  We might as well\n> > coalesce the two log commands into one, by extending this latest one\n> > to:\n> >\n> >   git log -M --diff-filter=RAMD --no-abbrev --raw\n>\n> What is there besides RAMD? :)\n\nWell, as you pointed out above, log detects renames by default,\nwhereas it didn't used to.\nSo, if someone had written some similar-ish history walking/parsing\ntool years ago that didn't depend need renames and was based on log\noutput, there's a good chance their tool might start failing when\nrename detection was turned on by default, because instead of getting\nboth a 'D' and an 'M' change, they'd get an unexpected 'R'.\n\nFor my case, do I have to worry about similar future changes?  Will\ncopy detection ('C') or break detection ('B') become the default in\nthe future?  Do I have to worry about typechanges ('T\")?  Will new\nchange types be added?  I mean, the fast-export output could maybe\nchange too, but it seems much less likely than with log.\n\n> > I could potentially switch to using this and drop patch 10/10.\n>\n> So I'm still not _entirely_ clear on what you're trying to do with\n> 10/10. I think maybe the \"disconnect\" part I wrote above explains it. If\n> that's correct, then I think framing it in terms of the operations that\n> you'd be able to perform _without running a separate traverse_ would\n> make it more obvious.\n\nLet me try to put it as briefly as I can.  With as few traversals as\npossible, I want to:\n  * Get all blob sizes\n  * Map blob shas to filename(s) they appeared under in the history\n  * Find when files and directories were deleted (and whether they\nwere later reinstated, since that means they aren't actually gone)\n  * Find sets of filenames referring to the same logical 'file'. (e.g.\nfoo->bar in commit A and bar->baz in commit B mean that {foo,bar,baz}\nrefer to the same 'file' so that a user has an easy report to look at\nto find out that if they just want to \"keep baz and its history\" then\nthey need foo & bar & baz.  I need to know about things like another\nfoo or bar being introduced after the rename though, since that breaks\nthe connection between filenames)\n  * Do a few aggregations on the above data as well (e.g. all copies\nof postgres.exe add up to 20M -- why were those checked in anyway?,\n*.webm files in aggregate are .5G, your long-deleted src/video-server/\ndirectory from that aborted experimental project years ago takes up 2G\nof your history, etc.)\n\nRight now, my best solution for this combination of questions is\n'cat-file --batch-all-objects' plus fast-export, if I get patch 10/10\nin place.  I'm totally open to better solutions, including ones that\ndon't use fast-export.\n"},{"id":"363231","messageId":"CABPp-BFo=UvwbqV06R9PVEJ6JyEsvUCr4pe+3eQw8D2W96D96w@mail.gmail.com","threadId":"49406","inReplyTo":"CABPp-BHwg2U=b+UGK2SufB7uZPmmiPVKXoTpYt+LuHnLwmwuZQ@mail.gmail.com","subject":"Re: [PATCH 02/10] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-13T23:25:14Z","receivedAt":"2018-11-13T23:25:30Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Nov 10, 2018 at 11:17 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Sat, Nov 10, 2018 at 10:36 PM Jeff King <peff@peff.net> wrote:\n> >\n> > On Sat, Nov 10, 2018 at 10:23:04PM -0800, Elijah Newren wrote:\n> >\n> > > Signed-off-by: Elijah Newren <newren@gmail.com>\n> > > ---\n> > >  Documentation/git-fast-export.txt | 3 ++-\n> > >  1 file changed, 2 insertions(+), 1 deletion(-)\n> > >\n> > > diff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\n> > > index ce954be532..677510b7f7 100644\n> > > --- a/Documentation/git-fast-export.txt\n> > > +++ b/Documentation/git-fast-export.txt\n> > > @@ -119,7 +119,8 @@ marks the same across runs.\n> > >       'git rev-list', that specifies the specific objects and references\n> > >       to export.  For example, `master~10..master` causes the\n> > >       current master reference to be exported along with all objects\n> > > -     added since its 10th ancestor commit.\n> > > +     added since its 10th ancestor commit and all files common to\n> > > +     master\\~9 and master~10.\n> >\n> > Do you need to backslash the second tilde?  Maybe `master~9` and\n> > `master~10` instead of escaping?\n>\n> Oops, yeah, that needs to be consistent.\n\nActually, no, it actually needs to be inconsistent.\n\nDifferent Input Choices (neither backslashed, both backslashed, then just one):\n  master~9 and master~10\n  master\\~9 and master\\~10\n  master\\~9 and master~10\n\nWhat the outputs look like:\n  master9 and master10\n  master~9 and master\\~10\n  master~9 and master~10\n\nI have no idea why asciidoc behaves this way, but it appears my\nbackslash escaping of just one of the two was necessary.\n"},{"id":"363232","messageId":"20181113233952.GA226088@google.com","threadId":"49406","inReplyTo":"CABPp-BFo=UvwbqV06R9PVEJ6JyEsvUCr4pe+3eQw8D2W96D96w@mail.gmail.com","subject":"Re: [PATCH 02/10] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T23:39:52Z","receivedAt":"2018-11-13T23:39:57Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Elijah Newren wrote:\n\n> Actually, no, it actually needs to be inconsistent.\n>\n> Different Input Choices (neither backslashed, both backslashed, then just one):\n>   master~9 and master~10\n>   master\\~9 and master\\~10\n>   master\\~9 and master~10\n>\n> What the outputs look like:\n>   master9 and master10\n>   master~9 and master\\~10\n>   master~9 and master~10\n>\n> I have no idea why asciidoc behaves this way, but it appears my\n> backslash escaping of just one of the two was necessary.\n\n{tilde} should work consistently.\n\nThanks,\nJonathan\n"},{"id":"363234","messageId":"CABPp-BGz1QuLkpBAEHFZY7krFdYPoZLZHYmufUH+PodLTN4zpw@mail.gmail.com","threadId":"49406","inReplyTo":"20181113233952.GA226088@google.com","subject":"Re: [PATCH 02/10] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:02:03Z","receivedAt":"2018-11-14T00:02:18Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Tue, Nov 13, 2018 at 3:39 PM Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Elijah Newren wrote:\n> > Actually, no, it actually needs to be inconsistent.\n> >\n> > Different Input Choices (neither backslashed, both backslashed, then just one):\n> >   master~9 and master~10\n> >   master\\~9 and master\\~10\n> >   master\\~9 and master~10\n> >\n> > What the outputs look like:\n> >   master9 and master10\n> >   master~9 and master\\~10\n> >   master~9 and master~10\n> >\n> > I have no idea why asciidoc behaves this way, but it appears my\n> > backslash escaping of just one of the two was necessary.\n>\n> {tilde} should work consistently.\n\nIndeed it does (well, outside of `backtick blocks`); thanks for the tip.\n"},{"id":"363261","messageId":"20181114002600.29233-8-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 07/11] fast-export: ensure we export requested refs","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:56Z","receivedAt":"2018-11-14T00:26:17Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If file paths are specified to fast-export and a ref points to a commit\nthat does not touch any of the relevant paths, then that ref would\nsometimes fail to be exported.  (This depends on whether any ancestors\nof the commit which do touch the relevant paths would be exported with\nthat same ref name or a different ref name.)  To avoid this problem,\nput *all* specified refs into extra_refs to start, and then as we export\neach commit, remove the refname used in the 'commit $REFNAME' directive\nfrom extra_refs.  Then, in handle_tags_and_duplicates() we know which\nrefs actually do need a manual reset directive in order to be included.\n\nThis means that we do need some special handling for excluded refs; e.g.\nif someone runs\n   git fast-export ^master master\nthen they've asked for master to be exported, but they have also asked\nfor the commit which master points to and all of its history to be\nexcluded.  That logically means ref deletion.  Previously, such refs\nwere just silently omitted from being exported despite having been\nexplicitly requested for export.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  | 54 ++++++++++++++++++++++++++++++++----------\n t/t9350-fast-export.sh | 16 ++++++++++---\n 2 files changed, 55 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 2eafe351ea..2fef00436b 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n+static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n static int anonymize;\n static struct revision_sources revision_sources;\n@@ -611,6 +612,13 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\t\texport_blob(&diff_queued_diff.queue[i]->two->oid);\n \n \trefname = *revision_sources_at(&revision_sources, commit);\n+\t/*\n+\t * FIXME: string_list_remove() below for each ref is overall\n+\t * O(N^2).  Compared to a history walk and diffing trees, this is\n+\t * just lost in the noise in practice.  However, theoretically a\n+\t * repo may have enough refs for this to become slow.\n+\t */\n+\tstring_list_remove(&extra_refs, refname, 0);\n \tif (anonymize) {\n \t\trefname = anonymize_refname(refname);\n \t\tanonymize_ident_line(&committer, &committer_end);\n@@ -814,7 +822,7 @@ static struct commit *get_commit(struct rev_cmdline_entry *e, char *full_name)\n \t\t/* handle nested tags */\n \t\twhile (tag && tag->object.type == OBJ_TAG) {\n \t\t\tparse_object(the_repository, &tag->object.oid);\n-\t\t\tstring_list_append(&extra_refs, full_name)->util = tag;\n+\t\t\tstring_list_append(&tag_refs, full_name)->util = tag;\n \t\t\ttag = (struct tag *)tag->tagged;\n \t\t}\n \t\tif (!tag)\n@@ -873,25 +881,30 @@ static void get_tags_and_duplicates(struct rev_cmdline_info *info)\n \t\t}\n \n \t\t/*\n-\t\t * This ref will not be updated through a commit, lets make\n-\t\t * sure it gets properly updated eventually.\n+\t\t * Make sure this ref gets properly updated eventually, whether\n+\t\t * through a commit or manually at the end.\n \t\t */\n-\t\tif (*revision_sources_at(&revision_sources, commit) ||\n-\t\t    commit->object.flags & SHOWN)\n+\t\tif (e->item->type != OBJ_TAG)\n \t\t\tstring_list_append(&extra_refs, full_name)->util = commit;\n+\n \t\tif (!*revision_sources_at(&revision_sources, commit))\n \t\t\t*revision_sources_at(&revision_sources, commit) = full_name;\n \t}\n+\n+\tstring_list_sort(&extra_refs);\n+\tstring_list_remove_duplicates(&extra_refs, 0);\n }\n \n-static void handle_tags_and_duplicates(void)\n+static void handle_tags_and_duplicates(struct string_list *extras)\n {\n \tstruct commit *commit;\n \tint i;\n \n-\tfor (i = extra_refs.nr - 1; i >= 0; i--) {\n-\t\tconst char *name = extra_refs.items[i].string;\n-\t\tstruct object *object = extra_refs.items[i].util;\n+\tfor (i = extras->nr - 1; i >= 0; i--) {\n+\t\tconst char *name = extras->items[i].string;\n+\t\tstruct object *object = extras->items[i].util;\n+\t\tint mark;\n+\n \t\tswitch (object->type) {\n \t\tcase OBJ_TAG:\n \t\t\thandle_tag(name, (struct tag *)object);\n@@ -912,8 +925,24 @@ static void handle_tags_and_duplicates(void)\n \t\t\t\t       name, sha1_to_hex(null_sha1));\n \t\t\t\tcontinue;\n \t\t\t}\n-\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n-\t\t\t       get_object_mark(&commit->object));\n+\n+\t\t\tmark = get_object_mark(&commit->object);\n+\t\t\tif (!mark) {\n+\t\t\t\t/*\n+\t\t\t\t * Getting here means we have a commit which\n+\t\t\t\t * was excluded by a negative refspec (e.g.\n+\t\t\t\t * fast-export ^master master).  If the user\n+\t\t\t\t * wants the branch exported but every commit\n+\t\t\t\t * in its history to be deleted, that sounds\n+\t\t\t\t * like a ref deletion to me.\n+\t\t\t\t */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\n+\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name, mark\n+\t\t\t       );\n \t\t\tshow_progress();\n \t\t\tbreak;\n \t\t}\n@@ -1101,7 +1130,8 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\t}\n \t}\n \n-\thandle_tags_and_duplicates();\n+\thandle_tags_and_duplicates(&extra_refs);\n+\thandle_tags_and_duplicates(&tag_refs);\n \thandle_deletes();\n \n \tif (export_filename && lastimportid != last_idnum)\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 299120ba70..50c2fceef4 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -544,10 +544,20 @@ test_expect_success 'use refspec' '\n \ttest_cmp expected actual\n '\n \n-test_expect_success 'delete refspec' '\n+test_expect_success 'delete ref because entire history excluded' '\n \tgit branch to-delete &&\n-\tgit fast-export --refspec :refs/heads/to-delete to-delete ^to-delete > actual &&\n-\tcat > expected <<-EOF &&\n+\tgit fast-export to-delete ^to-delete >actual &&\n+\tcat >expected <<-EOF &&\n+\treset refs/heads/to-delete\n+\tfrom 0000000000000000000000000000000000000000\n+\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'delete refspec' '\n+\tgit fast-export --refspec :refs/heads/to-delete >actual &&\n+\tcat >expected <<-EOF &&\n \treset refs/heads/to-delete\n \tfrom 0000000000000000000000000000000000000000\n \n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363262","messageId":"20181114002600.29233-12-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 11/11] fast-export: add --always-show-modify-after-rename","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:26:00Z","receivedAt":"2018-11-14T00:26:19Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"I wanted a way to gather all the following information efficiently\n(with as few history traversals as possible):\n  * Get all blob sizes\n  * Map blob shas to filename(s) they appeared under in the history\n  * Find when files and directories were deleted (and whether they\n    were later reinstated, since that means they aren't actually gone)\n  * Find sets of filenames referring to the same logical 'file'. (e.g.\n    foo->bar in commit A and bar->baz in commit B mean that\n    {foo,bar,baz} refer to the same 'file', so someone wanting to just\n    \"keep baz and its history\" need all versions of those three\n    filenames).  I need to know about things like another foo or bar\n    being introduced after the rename though, since that breaks the\n    connection between filenames)\nand then I would generate various aggregations on the data and display\nsome type of report for the user.\n\nThe only way I know of to get blob sizes is via\n  cat-file --batch-all-objects --batch-check\n\nThe rest of the data would traditionally be gathered from a log command,\ne.g.\n\n  git log --format='%H%n%P%n%cd' --date=short --topo-order --reverse \\\n      -M --diff-filter=RAMD --no-abbrev --raw -c\n\nhowever, parsing log output seems slightly dangerous given that it is a\nporcelain command.  While we have specified --format and --raw to try\nto avoid the most obvious problems, I'm still slightly concerned about\n--date=short, the combinations of --raw and -c, options that might\ncolorize the output, and also the --diff-filter (there is no current\noption named --no-find-copies or --no-break-rewrites, but what if those\nturn on by default in the future much as we changed the default with\ndetecting renames?).  Each of those is a small worry, but they add up.\n\nA command meant for data serialization, such as fast-export, seems like\na better candidate for this job.  There's just one missing item: in\norder to connect blob sizes to filenames, I need fast-export to tell me\nthe blob sha1sum of any file changes.  It does this for modifies, but\nnot always for renames.  In particular, if a file is a 100% rename, it\nonly prints\n    R oldname newname\ninstead of\n    R oldname newname\n    M 100644 $SHA1 newname\nas occurs when there is a rename+modify.  Add an option which allows us\nto force the latter output even when commits have exact renames of\nfiles.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 11 ++++++++++\n builtin/fast-export.c             |  7 +++++-\n t/t9350-fast-export.sh            | 36 +++++++++++++++++++++++++++++++\n 3 files changed, 53 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex 64c01ba918..b663b6f8af 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -129,6 +129,17 @@ marks the same across runs.\n \tfor intermediary filters (e.g. for rewriting commit messages\n \twhich refer to older commits, or for stripping blobs by id).\n \n+--always-show-modify-after-rename::\n+\tWhen a rename is detected, fast-export normally issues both a\n+\t'R' (rename) and a 'M' (modify) directive.  However, if the\n+\tcontents of the old and new filename match exactly, it will\n+\tonly issue the rename directive.  Use this flag to have it\n+\talways issue the modify directive after the rename, which may\n+\tbe useful for tools which are using the fast-export stream as\n+\ta mechanism for gathering statistics about a repository.  Note\n+\tthat this option only has effect when rename detection is\n+\tactive (see the -M option).\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex e0f794811e..31ad43077a 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static int reference_excluded_commits;\n+static int always_show_modify_after_rename;\n static int show_original_ids;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n@@ -407,7 +408,8 @@ static void show_filemodify(struct diff_queue_struct *q,\n \t\t\t\tputchar('\\n');\n \n \t\t\t\tif (oideq(&ospec->oid, &spec->oid) &&\n-\t\t\t\t    ospec->mode == spec->mode)\n+\t\t\t\t    ospec->mode == spec->mode &&\n+\t\t\t\t    !always_show_modify_after_rename)\n \t\t\t\t\tbreak;\n \t\t\t}\n \t\t\t/* fallthrough */\n@@ -1105,6 +1107,9 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n \t\tOPT_BOOL(0, \"show-original-ids\", &show_original_ids,\n \t\t\t    N_(\"Show original sha1sums of blobs/commits\")),\n+\t\tOPT_BOOL(0, \"always-show-modify-after-rename\",\n+\t\t\t    &always_show_modify_after_rename,\n+\t\t\t N_(\"Always provide 'M' directive after 'R'\")),\n \n \t\tOPT_END()\n \t};\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 5690fe2810..5c20065e39 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -630,4 +630,40 @@ test_expect_success 'merge commit gets exported with --import-marks' '\n \t)\n '\n \n+test_expect_success 'rename detection and --always-show-modify-after-rename' '\n+\ttest_create_repo renames &&\n+\t(\n+\t\tcd renames &&\n+\t\ttest_seq 0  9  >single_digit &&\n+\t\ttest_seq 10 98 >double_digit &&\n+\t\tgit add . &&\n+\t\tgit commit -m initial &&\n+\n+\t\techo 99 >>double_digit &&\n+\t\tgit mv single_digit single-digit &&\n+\t\tgit mv double_digit double-digit &&\n+\t\tgit add double-digit &&\n+\t\tgit commit -m renames &&\n+\n+\t\t# First, check normal fast-export -M output\n+\t\tgit fast-export -M --no-data master >out &&\n+\n+\t\tgrep double-digit out >out2 &&\n+\t\ttest_line_count = 2 out2 &&\n+\n+\t\tgrep single-digit out >out2 &&\n+\t\ttest_line_count = 1 out2 &&\n+\n+\t\t# Now, test with --always-show-modify-after-rename; should\n+\t\t# have an extra \"M\" directive for \"single-digit\".\n+\t\tgit fast-export -M --no-data --always-show-modify-after-rename master >out &&\n+\n+\t\tgrep double-digit out >out2 &&\n+\t\ttest_line_count = 2 out2 &&\n+\n+\t\tgrep single-digit out >out2 &&\n+\t\ttest_line_count = 2 out2\n+\t)\n+'\n+\n test_done\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363263","messageId":"20181114002600.29233-2-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 01/11] git-fast-import.txt: fix documentation for --quiet option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:50Z","receivedAt":"2018-11-14T00:26:20Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Signed-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-import.txt | 7 ++++---\n 1 file changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\nindex e81117d27f..7ab97745a6 100644\n--- a/Documentation/git-fast-import.txt\n+++ b/Documentation/git-fast-import.txt\n@@ -40,9 +40,10 @@ OPTIONS\n \tnot contain the old commit).\n \n --quiet::\n-\tDisable all non-fatal output, making fast-import silent when it\n-\tis successful.  This option disables the output shown by\n-\t--stats.\n+\tDisable the output shown by --stats, making fast-import usually\n+\tbe silent when it is successful.  However, if the import stream\n+\thas directives intended to show user output (e.g. `progress`\n+\tdirectives), the corresponding messages will still be shown.\n \n --stats::\n \tDisplay some basic statistics about the objects fast-import has\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363264","messageId":"20181114002600.29233-11-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 10/11] fast-export: add a --show-original-ids option to show original names","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:59Z","receivedAt":"2018-11-14T00:26:21Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Knowing the original names (hashes) of commits can sometimes enable\npost-filtering that would otherwise be difficult or impossible.  In\nparticular, the desire to rewrite commit messages which refer to other\nprior commits (on top of whatever other filtering is being done) is\nvery difficult without knowing the original names of each commit.\n\nIn addition, knowing the original names (hashes) of blobs can allow\nfiltering by blob-id without requiring re-hashing the content of the\nblob, and is thus useful as a small optimization.\n\nOnce we add original ids for both commits and blobs, we may as well\nadd them for tags too for completeness.  Perhaps someone will have a\nuse for them.\n\nThis commit teaches a new --show-original-ids option to fast-export\nwhich will make it add a 'original-oid <hash>' line to blob, commits,\nand tags.  It also teaches fast-import to parse (and ignore) such\nlines.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt |  7 +++++++\n Documentation/git-fast-import.txt | 16 ++++++++++++++++\n builtin/fast-export.c             | 20 +++++++++++++++-----\n fast-import.c                     | 12 ++++++++++++\n t/t9350-fast-export.sh            | 17 +++++++++++++++++\n 5 files changed, 67 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex f65026662a..64c01ba918 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -122,6 +122,13 @@ marks the same across runs.\n \trepository which already contains the necessary parent\n \tcommits.\n \n+--show-original-ids::\n+\tAdd an extra directive to the output for commits and blobs,\n+\t`original-oid <SHA1SUM>`.  While such directives will likely be\n+\tignored by importers such as git-fast-import, it may be useful\n+\tfor intermediary filters (e.g. for rewriting commit messages\n+\twhich refer to older commits, or for stripping blobs by id).\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\ndiff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\nindex 7ab97745a6..43ab3b1637 100644\n--- a/Documentation/git-fast-import.txt\n+++ b/Documentation/git-fast-import.txt\n@@ -385,6 +385,7 @@ change to the project.\n ....\n \t'commit' SP <ref> LF\n \tmark?\n+\toriginal-oid?\n \t('author' (SP <name>)? SP LT <email> GT SP <when> LF)?\n \t'committer' (SP <name>)? SP LT <email> GT SP <when> LF\n \tdata\n@@ -741,6 +742,19 @@ New marks are created automatically.  Existing marks can be moved\n to another object simply by reusing the same `<idnum>` in another\n `mark` command.\n \n+`original-oid`\n+~~~~~~~~~~~~~~\n+Provides the name of the object in the original source control system.\n+fast-import will simply ignore this directive, but filter processes\n+which operate on and modify the stream before feeding to fast-import\n+may have uses for this information\n+\n+....\n+\t'original-oid' SP <object-identifier> LF\n+....\n+\n+where `<object-identifer>` is any string not containing LF.\n+\n `tag`\n ~~~~~\n Creates an annotated tag referring to a specific commit.  To create\n@@ -749,6 +763,7 @@ lightweight (non-annotated) tags see the `reset` command below.\n ....\n \t'tag' SP <name> LF\n \t'from' SP <commit-ish> LF\n+\toriginal-oid?\n \t'tagger' (SP <name>)? SP LT <email> GT SP <when> LF\n \tdata\n ....\n@@ -823,6 +838,7 @@ assigned mark.\n ....\n \t'blob' LF\n \tmark?\n+\toriginal-oid?\n \tdata\n ....\n \ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 3cc98c31ad..e0f794811e 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static int reference_excluded_commits;\n+static int show_original_ids;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n@@ -271,7 +272,10 @@ static void export_blob(const struct object_id *oid)\n \n \tmark_next_object(object);\n \n-\tprintf(\"blob\\nmark :%\"PRIu32\"\\ndata %lu\\n\", last_idnum, size);\n+\tprintf(\"blob\\nmark :%\"PRIu32\"\\n\", last_idnum);\n+\tif (show_original_ids)\n+\t\tprintf(\"original-oid %s\\n\", oid_to_hex(oid));\n+\tprintf(\"data %lu\\n\", size);\n \tif (size && fwrite(buf, size, 1, stdout) != 1)\n \t\tdie_errno(\"could not write blob '%s'\", oid_to_hex(oid));\n \tprintf(\"\\n\");\n@@ -634,8 +638,10 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\treencoded = reencode_string(message, \"UTF-8\", encoding);\n \tif (!commit->parents)\n \t\tprintf(\"reset %s\\n\", refname);\n-\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n%.*s\\n%.*s\\ndata %u\\n%s\",\n-\t       refname, last_idnum,\n+\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n\", refname, last_idnum);\n+\tif (show_original_ids)\n+\t\tprintf(\"original-oid %s\\n\", oid_to_hex(&commit->object.oid));\n+\tprintf(\"%.*s\\n%.*s\\ndata %u\\n%s\",\n \t       (int)(author_end - author), author,\n \t       (int)(committer_end - committer), committer,\n \t       (unsigned)(reencoded\n@@ -813,8 +819,10 @@ static void handle_tag(const char *name, struct tag *tag)\n \n \tif (starts_with(name, \"refs/tags/\"))\n \t\tname += 10;\n-\tprintf(\"tag %s\\nfrom :%d\\n%.*s%sdata %d\\n%.*s\\n\",\n-\t       name, tagged_mark,\n+\tprintf(\"tag %s\\nfrom :%d\\n\", name, tagged_mark);\n+\tif (show_original_ids)\n+\t\tprintf(\"original-oid %s\\n\", oid_to_hex(&tag->object.oid));\n+\tprintf(\"%.*s%sdata %d\\n%.*s\\n\",\n \t       (int)(tagger_end - tagger), tagger,\n \t       tagger == tagger_end ? \"\" : \"\\n\",\n \t       (int)message_size, (int)message_size, message ? message : \"\");\n@@ -1095,6 +1103,8 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n \t\tOPT_BOOL(0, \"reference-excluded-parents\",\n \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n+\t\tOPT_BOOL(0, \"show-original-ids\", &show_original_ids,\n+\t\t\t    N_(\"Show original sha1sums of blobs/commits\")),\n \n \t\tOPT_END()\n \t};\ndiff --git a/fast-import.c b/fast-import.c\nindex 555d49ad23..71b6cba00f 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1814,6 +1814,13 @@ static void parse_mark(void)\n \t\tnext_mark = 0;\n }\n \n+static void parse_original_identifier(void)\n+{\n+\tconst char *v;\n+\tif (skip_prefix(command_buf.buf, \"original-oid \", &v))\n+\t\tread_next_command();\n+}\n+\n static int parse_data(struct strbuf *sb, uintmax_t limit, uintmax_t *len_res)\n {\n \tconst char *data;\n@@ -1956,6 +1963,7 @@ static void parse_new_blob(void)\n {\n \tread_next_command();\n \tparse_mark();\n+\tparse_original_identifier();\n \tparse_and_store_blob(&last_blob, NULL, next_mark);\n }\n \n@@ -2579,6 +2587,7 @@ static void parse_new_commit(const char *arg)\n \n \tread_next_command();\n \tparse_mark();\n+\tparse_original_identifier();\n \tif (skip_prefix(command_buf.buf, \"author \", &v)) {\n \t\tauthor = parse_ident(v);\n \t\tread_next_command();\n@@ -2711,6 +2720,9 @@ static void parse_new_tag(const char *arg)\n \t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n \tread_next_command();\n \n+\t/* original-oid ... */\n+\tparse_original_identifier();\n+\n \t/* tagger ... */\n \tif (skip_prefix(command_buf.buf, \"tagger \", &v)) {\n \t\ttagger = parse_ident(v);\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex d7d73061d0..5690fe2810 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -77,6 +77,23 @@ test_expect_success 'fast-export --reference-excluded-parents master~2..master'\n \t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n '\n \n+test_expect_success 'fast-export --show-original-ids' '\n+\n+\tgit fast-export --show-original-ids master >output &&\n+\tgrep ^original-oid output| sed -e s/^original-oid.// | sort >actual &&\n+\tgit rev-list --objects master muss >objects-and-names &&\n+\tawk \"{print \\$1}\" objects-and-names | sort >commits-trees-blobs &&\n+\tcomm -23 actual commits-trees-blobs >unfound &&\n+\ttest_must_be_empty unfound\n+'\n+\n+test_expect_success 'fast-export --show-original-ids | git fast-import' '\n+\n+\tgit fast-export --show-original-ids master muss | git fast-import --quiet &&\n+\ttest $MASTER = $(git rev-parse --verify refs/heads/master) &&\n+\ttest $MUSS = $(git rev-parse --verify refs/tags/muss)\n+'\n+\n test_expect_success 'iso-8859-1' '\n \n \tgit config i18n.commitencoding ISO8859-1 &&\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363265","messageId":"20181114002600.29233-6-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 05/11] fast-export: move commit rewriting logic into a function for reuse","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:54Z","receivedAt":"2018-11-14T00:26:21Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Logic to replace a filtered commit with an unfiltered ancestor is useful\nelsewhere; put it into a function we can call.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 37 ++++++++++++++++++++++---------------\n 1 file changed, 22 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex b984a44224..7888fc98b5 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -187,6 +187,22 @@ static int get_object_mark(struct object *object)\n \treturn ptr_to_mark(decoration);\n }\n \n+static struct commit *rewrite_commit(struct commit *p)\n+{\n+\tfor (;;) {\n+\t\tif (p->parents && p->parents->next)\n+\t\t\tbreak;\n+\t\tif (p->object.flags & UNINTERESTING)\n+\t\t\tbreak;\n+\t\tif (!(p->object.flags & TREESAME))\n+\t\t\tbreak;\n+\t\tif (!p->parents)\n+\t\t\treturn NULL;\n+\t\tp = p->parents->item;\n+\t}\n+\treturn p;\n+}\n+\n static void show_progress(void)\n {\n \tstatic int counter = 0;\n@@ -766,21 +782,12 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t    oid_to_hex(&tag->object.oid),\n \t\t\t\t    type_name(tagged->type));\n \t\t\t}\n-\t\t\tp = (struct commit *)tagged;\n-\t\t\tfor (;;) {\n-\t\t\t\tif (p->parents && p->parents->next)\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (p->object.flags & UNINTERESTING)\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (!(p->object.flags & TREESAME))\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (!p->parents) {\n-\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n-\t\t\t\t\tfree(buf);\n-\t\t\t\t\treturn;\n-\t\t\t\t}\n-\t\t\t\tp = p->parents->item;\n+\t\t\tp = rewrite_commit((struct commit *)tagged);\n+\t\t\tif (!p) {\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tfree(buf);\n+\t\t\t\treturn;\n \t\t\t}\n \t\t\ttagged_mark = get_object_mark(&p->object);\n \t\t}\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363266","messageId":"20181114002600.29233-10-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 09/11] fast-import: remove unmaintained duplicate documentation","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:58Z","receivedAt":"2018-11-14T00:26:23Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"fast-import.c has started with a comment for nine and a half years\nre-directing the reader to Documentation/git-fast-import.txt for\nmaintained documentation.  Instead of leaving the unmaintained\ndocumentation in place, just excise it.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n fast-import.c | 154 --------------------------------------------------\n 1 file changed, 154 deletions(-)\n\ndiff --git a/fast-import.c b/fast-import.c\nindex 95600c78e0..555d49ad23 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1,157 +1,3 @@\n-/*\n-(See Documentation/git-fast-import.txt for maintained documentation.)\n-Format of STDIN stream:\n-\n-  stream ::= cmd*;\n-\n-  cmd ::= new_blob\n-        | new_commit\n-        | new_tag\n-        | reset_branch\n-        | checkpoint\n-        | progress\n-        ;\n-\n-  new_blob ::= 'blob' lf\n-    mark?\n-    file_content;\n-  file_content ::= data;\n-\n-  new_commit ::= 'commit' sp ref_str lf\n-    mark?\n-    ('author' (sp name)? sp '<' email '>' sp when lf)?\n-    'committer' (sp name)? sp '<' email '>' sp when lf\n-    commit_msg\n-    ('from' sp commit-ish lf)?\n-    ('merge' sp commit-ish lf)*\n-    (file_change | ls)*\n-    lf?;\n-  commit_msg ::= data;\n-\n-  ls ::= 'ls' sp '\"' quoted(path) '\"' lf;\n-\n-  file_change ::= file_clr\n-    | file_del\n-    | file_rnm\n-    | file_cpy\n-    | file_obm\n-    | file_inm;\n-  file_clr ::= 'deleteall' lf;\n-  file_del ::= 'D' sp path_str lf;\n-  file_rnm ::= 'R' sp path_str sp path_str lf;\n-  file_cpy ::= 'C' sp path_str sp path_str lf;\n-  file_obm ::= 'M' sp mode sp (hexsha1 | idnum) sp path_str lf;\n-  file_inm ::= 'M' sp mode sp 'inline' sp path_str lf\n-    data;\n-  note_obm ::= 'N' sp (hexsha1 | idnum) sp commit-ish lf;\n-  note_inm ::= 'N' sp 'inline' sp commit-ish lf\n-    data;\n-\n-  new_tag ::= 'tag' sp tag_str lf\n-    'from' sp commit-ish lf\n-    ('tagger' (sp name)? sp '<' email '>' sp when lf)?\n-    tag_msg;\n-  tag_msg ::= data;\n-\n-  reset_branch ::= 'reset' sp ref_str lf\n-    ('from' sp commit-ish lf)?\n-    lf?;\n-\n-  checkpoint ::= 'checkpoint' lf\n-    lf?;\n-\n-  progress ::= 'progress' sp not_lf* lf\n-    lf?;\n-\n-     # note: the first idnum in a stream should be 1 and subsequent\n-     # idnums should not have gaps between values as this will cause\n-     # the stream parser to reserve space for the gapped values.  An\n-     # idnum can be updated in the future to a new object by issuing\n-     # a new mark directive with the old idnum.\n-     #\n-  mark ::= 'mark' sp idnum lf;\n-  data ::= (delimited_data | exact_data)\n-    lf?;\n-\n-    # note: delim may be any string but must not contain lf.\n-    # data_line may contain any data but must not be exactly\n-    # delim.\n-  delimited_data ::= 'data' sp '<<' delim lf\n-    (data_line lf)*\n-    delim lf;\n-\n-     # note: declen indicates the length of binary_data in bytes.\n-     # declen does not include the lf preceding the binary data.\n-     #\n-  exact_data ::= 'data' sp declen lf\n-    binary_data;\n-\n-     # note: quoted strings are C-style quoting supporting \\c for\n-     # common escapes of 'c' (e..g \\n, \\t, \\\\, \\\") or \\nnn where nnn\n-     # is the signed byte value in octal.  Note that the only\n-     # characters which must actually be escaped to protect the\n-     # stream formatting is: \\, \" and LF.  Otherwise these values\n-     # are UTF8.\n-     #\n-  commit-ish  ::= (ref_str | hexsha1 | sha1exp_str | idnum);\n-  ref_str     ::= ref;\n-  sha1exp_str ::= sha1exp;\n-  tag_str     ::= tag;\n-  path_str    ::= path    | '\"' quoted(path)    '\"' ;\n-  mode        ::= '100644' | '644'\n-                | '100755' | '755'\n-                | '120000'\n-                ;\n-\n-  declen ::= # unsigned 32 bit value, ascii base10 notation;\n-  bigint ::= # unsigned integer value, ascii base10 notation;\n-  binary_data ::= # file content, not interpreted;\n-\n-  when         ::= raw_when | rfc2822_when;\n-  raw_when     ::= ts sp tz;\n-  rfc2822_when ::= # Valid RFC 2822 date and time;\n-\n-  sp ::= # ASCII space character;\n-  lf ::= # ASCII newline (LF) character;\n-\n-     # note: a colon (':') must precede the numerical value assigned to\n-     # an idnum.  This is to distinguish it from a ref or tag name as\n-     # GIT does not permit ':' in ref or tag strings.\n-     #\n-  idnum   ::= ':' bigint;\n-  path    ::= # GIT style file path, e.g. \"a/b/c\";\n-  ref     ::= # GIT ref name, e.g. \"refs/heads/MOZ_GECKO_EXPERIMENT\";\n-  tag     ::= # GIT tag name, e.g. \"FIREFOX_1_5\";\n-  sha1exp ::= # Any valid GIT SHA1 expression;\n-  hexsha1 ::= # SHA1 in hexadecimal format;\n-\n-     # note: name and email are UTF8 strings, however name must not\n-     # contain '<' or lf and email must not contain any of the\n-     # following: '<', '>', lf.\n-     #\n-  name  ::= # valid GIT author/committer name;\n-  email ::= # valid GIT author/committer email;\n-  ts    ::= # time since the epoch in seconds, ascii base10 notation;\n-  tz    ::= # GIT style timezone;\n-\n-     # note: comments, get-mark, ls-tree, and cat-blob requests may\n-     # appear anywhere in the input, except within a data command. Any\n-     # form of the data command always escapes the related input from\n-     # comment processing.\n-     #\n-     # In case it is not clear, the '#' that starts the comment\n-     # must be the first character on that line (an lf\n-     # preceded it).\n-     #\n-\n-  get_mark ::= 'get-mark' sp idnum lf;\n-  cat_blob ::= 'cat-blob' sp (hexsha1 | idnum) lf;\n-  ls_tree  ::= 'ls' sp (hexsha1 | idnum) sp path_str lf;\n-\n-  comment ::= '#' not_lf* lf;\n-  not_lf  ::= # Any byte that is not ASCII newline (LF);\n-*/\n-\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"repository.h\"\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363267","messageId":"20181114002600.29233-5-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 04/11] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:53Z","receivedAt":"2018-11-14T00:26:24Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If --tag-of-filtered-object=rewrite is specified along with a set of\npaths to limit what is exported, then any tags pointing to old commits\nthat do not contain any of those specified paths cause problems.  Since\nthe old tagged commit is not exported, fast-export attempts to rewrite\nsuch tags to an ancestor commit which was exported.  If no such commit\nexists, then fast-export currently die()s.  Five years after the tag\nrewriting logic was added to fast-export (see commit 2d8ad4691921,\n\"fast-export: Add a --tag-of-filtered-object  option for newly dangling\ntags\", 2009-06-25), fast-import gained the ability to delete refs (see\ncommit 4ee1b225b99f, \"fast-import: add support to delete refs\",\n2014-04-20), so now we do have a valid option to rewrite the tag to.\nDelete these tags instead of dying.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  |  9 ++++++---\n t/t9350-fast-export.sh | 16 ++++++++++++++++\n 2 files changed, 22 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex af724e9937..b984a44224 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -774,9 +774,12 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t\tbreak;\n \t\t\t\tif (!(p->object.flags & TREESAME))\n \t\t\t\t\tbreak;\n-\t\t\t\tif (!p->parents)\n-\t\t\t\t\tdie(\"can't find replacement commit for tag %s\",\n-\t\t\t\t\t     oid_to_hex(&tag->object.oid));\n+\t\t\t\tif (!p->parents) {\n+\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\t\tfree(buf);\n+\t\t\t\t\treturn;\n+\t\t\t\t}\n \t\t\t\tp = p->parents->item;\n \t\t\t}\n \t\t\ttagged_mark = get_object_mark(&p->object);\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 6a392e87bc..3400ebeb51 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -325,6 +325,22 @@ test_expect_success 'rewriting tag of filtered out object' '\n )\n '\n \n+test_expect_success 'rewrite tag predating pathspecs to nothing' '\n+\ttest_create_repo rewrite_tag_predating_pathspecs &&\n+\t(\n+\t\tcd rewrite_tag_predating_pathspecs &&\n+\n+\t\ttest_commit initial &&\n+\n+\t\tgit tag -a -m \"Some old tag\" v0.0.0.0.0.0.1 &&\n+\n+\t\ttest_commit bar &&\n+\n+\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar.t >output &&\n+\t\tgrep from.$ZERO_OID output\n+\t)\n+'\n+\n cat > limit-by-paths/expected << EOF\n blob\n mark :1\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363268","messageId":"20181114002600.29233-1-newren@gmail.com","threadId":"49406","inReplyTo":"20181111062312.16342-1-newren@gmail.com","subject":"[PATCH v2 00/11] fast export and import fixes and features","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:49Z","receivedAt":"2018-11-14T00:26:26Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"This is a series of small fixes and features for fast-export and\nfast-import, mostly on the fast-export side.\n\nChanges since v1 (full range-diff below):\n  - used {tilde} in asciidoc documentation to avoid subscripting and\n    escaping problems\n  - renamed ABORT/ERROR enum values to help avoid further misusage\n  - multiple small testcase cleanups (use $ZERO_OID, remove grep -A, etc.)\n  - add FIXME comment to code about string_list usage\n  - record Peff's idea for a future optimization in patch 8 commit message\n    (is there a better place to put that??)\n  - New patch (9/11): remove the unmaintained copy of fast-import stream\n    format documentation at the beginning of fast-import.c\n  - Rewrite commit message for 10/11 to match the wording Peff liked\n    better, s/originally/original-oid/, and add documentation to\n    git-fast-import.txt\n  - Rewrite commit message for 11/11; the last one didn't make sense to\n    Peff.  I hope this one does.\n\nElijah Newren (11):\n  git-fast-import.txt: fix documentation for --quiet option\n  git-fast-export.txt: clarify misleading documentation about rev-list\n    args\n  fast-export: use value from correct enum\n  fast-export: avoid dying when filtering by paths and old tags exist\n  fast-export: move commit rewriting logic into a function for reuse\n  fast-export: when using paths, avoid corrupt stream with non-existent\n    mark\n  fast-export: ensure we export requested refs\n  fast-export: add --reference-excluded-parents option\n  fast-import: remove unmaintained duplicate documentation\n  fast-export: add a --show-original-ids option to show original names\n  fast-export: add --always-show-modify-after-rename\n\n Documentation/git-fast-export.txt |  34 +++++-\n Documentation/git-fast-import.txt |  23 +++-\n builtin/fast-export.c             | 172 ++++++++++++++++++++++--------\n fast-import.c                     | 166 +++-------------------------\n t/t9350-fast-export.sh            | 116 +++++++++++++++++++-\n 5 files changed, 308 insertions(+), 203 deletions(-)\n\n 1:  0744f65b0d =  1:  8870fb1340 git-fast-import.txt: fix documentation for --quiet option\n 2:  aba1e22fdd !  2:  16d1c3e22d git-fast-export.txt: clarify misleading documentation about rev-list args\n    @@ -13,7 +13,7 @@\n      \tcurrent master reference to be exported along with all objects\n     -\tadded since its 10th ancestor commit.\n     +\tadded since its 10th ancestor commit and all files common to\n    -+\tmaster\\~9 and master~10.\n    ++\tmaster{tilde}9 and master{tilde}10.\n      \n      EXAMPLES\n      --------\n 3:  6983e845b2 <  -:  ---------- fast-export: use value from correct enum\n -:  ---------- >  3:  e19f6b36f9 fast-export: use value from correct enum\n 4:  761ba324d5 !  4:  2b305561d5 fast-export: avoid dying when filtering by paths and old tags exist\n    @@ -49,18 +49,14 @@\n     +\t(\n     +\t\tcd rewrite_tag_predating_pathspecs &&\n     +\n    -+\t\ttouch ignored &&\n    -+\t\tgit add ignored &&\n     +\t\ttest_commit initial &&\n     +\n     +\t\tgit tag -a -m \"Some old tag\" v0.0.0.0.0.0.1 &&\n     +\n    -+\t\techo foo >bar &&\n    -+\t\tgit add bar &&\n    -+\t\ttest_commit add-bar &&\n    ++\t\ttest_commit bar &&\n     +\n    -+\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar >output &&\n    -+\t\tgrep -A 1 refs/tags/v0.0.0.0.0.0.1 output | grep -E ^from.0{40}\n    ++\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar.t >output &&\n    ++\t\tgrep from.$ZERO_OID output\n     +\t)\n     +'\n     +\n 5:  64e9f0d360 =  5:  607b1dc2b2 fast-export: move commit rewriting logic into a function for reuse\n 6:  fd14d9749a !  6:  ec1862e858 fast-export: when using paths, avoid corrupt stream with non-existent mark\n    @@ -54,22 +54,18 @@\n     +\t(\n     +\t\tcd avoid_non_existent_mark &&\n     +\n    -+\t\ttouch important-path &&\n    -+\t\tgit add important-path &&\n    -+\t\ttest_commit initial &&\n    ++\t\ttest_commit important-path &&\n     +\n    -+\t\ttouch ignored &&\n    -+\t\tgit add ignored &&\n    -+\t\ttest_commit whatever &&\n    ++\t\ttest_commit ignored &&\n     +\n     +\t\tgit branch A &&\n     +\t\tgit branch B &&\n     +\n    -+\t\techo foo >>important-path &&\n    -+\t\tgit add important-path &&\n    ++\t\techo foo >>important-path.t &&\n    ++\t\tgit add important-path.t &&\n     +\t\ttest_commit more changes &&\n     +\n    -+\t\tgit fast-export --all -- important-path | git fast-import --force\n    ++\t\tgit fast-export --all -- important-path.t | git fast-import --force\n     +\t)\n     +'\n     +\n 7:  4e67a2bc7f !  7:  9da26e3ccb fast-export: ensure we export requested refs\n    @@ -21,9 +21,6 @@\n         were just silently omitted from being exported despite having been\n         explicitly requested for export.\n     \n    -    NOTE: The usage of string_list should really be replaced with the\n    -    strmap proposal, once it materializes.\n    -\n         Signed-off-by: Elijah Newren <newren@gmail.com>\n     \n      diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n    @@ -41,6 +38,12 @@\n      \t\t\texport_blob(&diff_queued_diff.queue[i]->two->oid);\n      \n      \trefname = *revision_sources_at(&revision_sources, commit);\n    ++\t/*\n    ++\t * FIXME: string_list_remove() below for each ref is overall\n    ++\t * O(N^2).  Compared to a history walk and diffing trees, this is\n    ++\t * just lost in the noise in practice.  However, theoretically a\n    ++\t * repo may have enough refs for this to become slow.\n    ++\t */\n     +\tstring_list_remove(&extra_refs, refname, 0);\n      \tif (anonymize) {\n      \t\trefname = anonymize_refname(refname);\n 8:  be02337f29 !  8:  7e5fe2f02e fast-export: add --reference-excluded-parents option\n    @@ -30,6 +30,15 @@\n         repository which already contains the necessary commits (much like the\n         restriction imposed when using --no-data).\n     \n    +    Note from Peff:\n    +      I think we might be able to do a little more optimization here. If\n    +      we're exporting HEAD^..HEAD and there's an object in HEAD^ which is\n    +      unchanged in HEAD, I think we'd still print it (because it would not\n    +      be marked SHOWN), but we could omit it (by walking the tree of the\n    +      boundary commits and marking them shown).  I don't think it's a\n    +      blocker for what you're doing here, but just a possible future\n    +      optimization.\n    +\n         Signed-off-by: Elijah Newren <newren@gmail.com>\n     \n      diff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\n    @@ -41,14 +50,15 @@\n      \n     +--reference-excluded-parents::\n     +\tBy default, running a command such as `git fast-export\n    -+\tmaster~5..master` will not include the commit master\\~5 and\n    -+\twill make master\\~4 no longer have master\\~5 as a parent (though\n    -+\tboth the old master\\~4 and new master~4 will have all the same\n    -+\tfiles).  Use --reference-excluded-parents to instead have the\n    -+\tthe stream refer to commits in the excluded range of history\n    -+\tby their sha1sum.  Note that the resulting stream can only be\n    -+\tused by a repository which already contains the necessary\n    -+\tparent commits.\n    ++\tmaster~5..master` will not include the commit master{tilde}5\n    ++\tand will make master{tilde}4 no longer have master{tilde}5 as\n    ++\ta parent (though both the old master{tilde}4 and new\n    ++\tmaster{tilde}4 will have all the same files).  Use\n    ++\t--reference-excluded-parents to instead have the the stream\n    ++\trefer to commits in the excluded range of history by their\n    ++\tsha1sum.  Note that the resulting stream can only be used by a\n    ++\trepository which already contains the necessary parent\n    ++\tcommits.\n     +\n      --refspec::\n      \tApply the specified refspec to each ref exported. Multiple of them can\n    @@ -58,10 +68,10 @@\n      \tto export.  For example, `master~10..master` causes the\n      \tcurrent master reference to be exported along with all objects\n     -\tadded since its 10th ancestor commit and all files common to\n    --\tmaster\\~9 and master~10.\n    +-\tmaster{tilde}9 and master{tilde}10.\n     +\tadded since its 10th ancestor commit and (unless the\n     +\t--reference-excluded-parents option is specified) all files\n    -+\tcommon to master\\~9 and master~10.\n    ++\tcommon to master{tilde}9 and master{tilde}10.\n      \n      EXAMPLES\n      --------\n -:  ---------- >  9:  14306a8436 fast-import: remove unmaintained duplicate documentation\n 9:  7ab314849d ! 10:  72487a61e4 fast-export: add a --show-original-ids option to show original names\n    @@ -2,16 +2,24 @@\n     \n         fast-export: add a --show-original-ids option to show original names\n     \n    -    Knowing the original names (hashes) of commits, blobs, and tags can\n    -    sometimes enable post-filtering that would otherwise be difficult or\n    -    impossible.  In particular, the desire to rewrite commit messages which\n    -    refer to other prior commits (on top of whatever other filtering is\n    -    being done) is very difficult without knowing the original names of each\n    -    commit.\n    +    Knowing the original names (hashes) of commits can sometimes enable\n    +    post-filtering that would otherwise be difficult or impossible.  In\n    +    particular, the desire to rewrite commit messages which refer to other\n    +    prior commits (on top of whatever other filtering is being done) is\n    +    very difficult without knowing the original names of each commit.\n    +\n    +    In addition, knowing the original names (hashes) of blobs can allow\n    +    filtering by blob-id without requiring re-hashing the content of the\n    +    blob, and is thus useful as a small optimization.\n    +\n    +    Once we add original ids for both commits and blobs, we may as well\n    +    add them for tags too for completeness.  Perhaps someone will have a\n    +    use for them.\n     \n         This commit teaches a new --show-original-ids option to fast-export\n    -    which will make it add a 'originally <hash>' line to blob, commits, and\n    -    tags.  It also teaches fast-import to parse (and ignore) such lines.\n    +    which will make it add a 'original-oid <hash>' line to blob, commits,\n    +    and tags.  It also teaches fast-import to parse (and ignore) such\n    +    lines.\n     \n         Signed-off-by: Elijah Newren <newren@gmail.com>\n     \n    @@ -19,12 +27,12 @@\n      --- a/Documentation/git-fast-export.txt\n      +++ b/Documentation/git-fast-export.txt\n     @@\n    - \tused by a repository which already contains the necessary\n    - \tparent commits.\n    + \trepository which already contains the necessary parent\n    + \tcommits.\n      \n     +--show-original-ids::\n     +\tAdd an extra directive to the output for commits and blobs,\n    -+\t`originally <SHA1SUM>`.  While such directives will likely be\n    ++\t`original-oid <SHA1SUM>`.  While such directives will likely be\n     +\tignored by importers such as git-fast-import, it may be useful\n     +\tfor intermediary filters (e.g. for rewriting commit messages\n     +\twhich refer to older commits, or for stripping blobs by id).\n    @@ -33,6 +41,54 @@\n      \tApply the specified refspec to each ref exported. Multiple of them can\n      \tbe specified.\n     \n    + diff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\n    + --- a/Documentation/git-fast-import.txt\n    + +++ b/Documentation/git-fast-import.txt\n    +@@\n    + ....\n    + \t'commit' SP <ref> LF\n    + \tmark?\n    ++\toriginal-oid?\n    + \t('author' (SP <name>)? SP LT <email> GT SP <when> LF)?\n    + \t'committer' (SP <name>)? SP LT <email> GT SP <when> LF\n    + \tdata\n    +@@\n    + to another object simply by reusing the same `<idnum>` in another\n    + `mark` command.\n    + \n    ++`original-oid`\n    ++~~~~~~~~~~~~~~\n    ++Provides the name of the object in the original source control system.\n    ++fast-import will simply ignore this directive, but filter processes\n    ++which operate on and modify the stream before feeding to fast-import\n    ++may have uses for this information\n    ++\n    ++....\n    ++\t'original-oid' SP <object-identifier> LF\n    ++....\n    ++\n    ++where `<object-identifer>` is any string not containing LF.\n    ++\n    + `tag`\n    + ~~~~~\n    + Creates an annotated tag referring to a specific commit.  To create\n    +@@\n    + ....\n    + \t'tag' SP <name> LF\n    + \t'from' SP <commit-ish> LF\n    ++\toriginal-oid?\n    + \t'tagger' (SP <name>)? SP LT <email> GT SP <when> LF\n    + \tdata\n    + ....\n    +@@\n    + ....\n    + \t'blob' LF\n    + \tmark?\n    ++\toriginal-oid?\n    + \tdata\n    + ....\n    + \n    +\n      diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n      --- a/builtin/fast-export.c\n      +++ b/builtin/fast-export.c\n    @@ -51,7 +107,7 @@\n     -\tprintf(\"blob\\nmark :%\"PRIu32\"\\ndata %lu\\n\", last_idnum, size);\n     +\tprintf(\"blob\\nmark :%\"PRIu32\"\\n\", last_idnum);\n     +\tif (show_original_ids)\n    -+\t\tprintf(\"originally %s\\n\", oid_to_hex(oid));\n    ++\t\tprintf(\"original-oid %s\\n\", oid_to_hex(oid));\n     +\tprintf(\"data %lu\\n\", size);\n      \tif (size && fwrite(buf, size, 1, stdout) != 1)\n      \t\tdie_errno(\"could not write blob '%s'\", oid_to_hex(oid));\n    @@ -64,7 +120,7 @@\n     -\t       refname, last_idnum,\n     +\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n\", refname, last_idnum);\n     +\tif (show_original_ids)\n    -+\t\tprintf(\"originally %s\\n\", oid_to_hex(&commit->object.oid));\n    ++\t\tprintf(\"original-oid %s\\n\", oid_to_hex(&commit->object.oid));\n     +\tprintf(\"%.*s\\n%.*s\\ndata %u\\n%s\",\n      \t       (int)(author_end - author), author,\n      \t       (int)(committer_end - committer), committer,\n    @@ -77,7 +133,7 @@\n     -\t       name, tagged_mark,\n     +\tprintf(\"tag %s\\nfrom :%d\\n\", name, tagged_mark);\n     +\tif (show_original_ids)\n    -+\t\tprintf(\"originally %s\\n\", oid_to_hex(&tag->object.oid));\n    ++\t\tprintf(\"original-oid %s\\n\", oid_to_hex(&tag->object.oid));\n     +\tprintf(\"%.*s%sdata %d\\n%.*s\\n\",\n      \t       (int)(tagger_end - tagger), tagger,\n      \t       tagger == tagger_end ? \"\" : \"\\n\",\n    @@ -96,44 +152,13 @@\n      --- a/fast-import.c\n      +++ b/fast-import.c\n     @@\n    - \n    -   new_blob ::= 'blob' lf\n    -     mark?\n    -+    originally?\n    -     file_content;\n    -   file_content ::= data;\n    - \n    -   new_commit ::= 'commit' sp ref_str lf\n    -     mark?\n    -+    originally?\n    -     ('author' (sp name)? sp '<' email '>' sp when lf)?\n    -     'committer' (sp name)? sp '<' email '>' sp when lf\n    -     commit_msg\n    -@@\n    - \n    -   new_tag ::= 'tag' sp tag_str lf\n    -     'from' sp commit-ish lf\n    -+    originally?\n    -     ('tagger' (sp name)? sp '<' email '>' sp when lf)?\n    -     tag_msg;\n    -   tag_msg ::= data;\n    -@@\n    -   data ::= (delimited_data | exact_data)\n    -     lf?;\n    - \n    -+  originally ::= 'originally' sp not_lf+ lf\n    -+\n    -     # note: delim may be any string but must not contain lf.\n    -     # data_line may contain any data but must not be exactly\n    -     # delim.\n    -@@\n      \t\tnext_mark = 0;\n      }\n      \n     +static void parse_original_identifier(void)\n     +{\n     +\tconst char *v;\n    -+\tif (skip_prefix(command_buf.buf, \"originally \", &v))\n    ++\tif (skip_prefix(command_buf.buf, \"original-oid \", &v))\n     +\t\tread_next_command();\n     +}\n     +\n    @@ -160,7 +185,7 @@\n      \t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n      \tread_next_command();\n      \n    -+\t/* originally ... */\n    ++\t/* original-oid ... */\n     +\tparse_original_identifier();\n     +\n      \t/* tagger ... */\n    @@ -177,7 +202,7 @@\n     +test_expect_success 'fast-export --show-original-ids' '\n     +\n     +\tgit fast-export --show-original-ids master >output &&\n    -+\tgrep ^originally output| sed -e s/^originally.// | sort >actual &&\n    ++\tgrep ^original-oid output| sed -e s/^original-oid.// | sort >actual &&\n     +\tgit rev-list --objects master muss >objects-and-names &&\n     +\tawk \"{print \\$1}\" objects-and-names | sort >commits-trees-blobs &&\n     +\tcomm -23 actual commits-trees-blobs >unfound &&\n10:  82735bcbde ! 11:  1796373474 fast-export: add --always-show-modify-after-rename\n    @@ -2,29 +2,53 @@\n     \n         fast-export: add --always-show-modify-after-rename\n     \n    -    fast-export output is traditionally used as an input to a fast-import\n    -    program, but it is also useful to help gather statistics about the\n    -    history of a repository (particularly when --no-data is also passed).\n    -    For example, two of the types of information we may want to collect\n    -    could include:\n    -      1) general information about renames that have occurred\n    -      2) what the biggest objects in a repository are and what names\n    -         they appear under.\n    +    I wanted a way to gather all the following information efficiently\n    +    (with as few history traversals as possible):\n    +      * Get all blob sizes\n    +      * Map blob shas to filename(s) they appeared under in the history\n    +      * Find when files and directories were deleted (and whether they\n    +        were later reinstated, since that means they aren't actually gone)\n    +      * Find sets of filenames referring to the same logical 'file'. (e.g.\n    +        foo->bar in commit A and bar->baz in commit B mean that\n    +        {foo,bar,baz} refer to the same 'file', so someone wanting to just\n    +        \"keep baz and its history\" need all versions of those three\n    +        filenames).  I need to know about things like another foo or bar\n    +        being introduced after the rename though, since that breaks the\n    +        connection between filenames)\n    +    and then I would generate various aggregations on the data and display\n    +    some type of report for the user.\n     \n    -    The first bit of information can be gathered by just passing -M to\n    -    fast-export.  The second piece of information can partially be gotten\n    -    from running\n    -        git cat-file --batch-check --batch-all-objects\n    -    However, that only shows what the biggest objects in the repository are\n    -    and their sizes, not what names those objects appear as or what commits\n    -    they were introduced in.  We can get that information from fast-export,\n    -    but when we only see\n    +    The only way I know of to get blob sizes is via\n    +      cat-file --batch-all-objects --batch-check\n    +\n    +    The rest of the data would traditionally be gathered from a log command,\n    +    e.g.\n    +\n    +      git log --format='%H%n%P%n%cd' --date=short --topo-order --reverse \\\n    +          -M --diff-filter=RAMD --no-abbrev --raw -c\n    +\n    +    however, parsing log output seems slightly dangerous given that it is a\n    +    porcelain command.  While we have specified --format and --raw to try\n    +    to avoid the most obvious problems, I'm still slightly concerned about\n    +    --date=short, the combinations of --raw and -c, options that might\n    +    colorize the output, and also the --diff-filter (there is no current\n    +    option named --no-find-copies or --no-break-rewrites, but what if those\n    +    turn on by default in the future much as we changed the default with\n    +    detecting renames?).  Each of those is a small worry, but they add up.\n    +\n    +    A command meant for data serialization, such as fast-export, seems like\n    +    a better candidate for this job.  There's just one missing item: in\n    +    order to connect blob sizes to filenames, I need fast-export to tell me\n    +    the blob sha1sum of any file changes.  It does this for modifies, but\n    +    not always for renames.  In particular, if a file is a 100% rename, it\n    +    only prints\n             R oldname newname\n         instead of\n             R oldname newname\n             M 100644 $SHA1 newname\n    -    then it makes the job more difficult.  Add an option which allows us to\n    -    force the latter output even when commits have exact renames of files.\n    +    as occurs when there is a rename+modify.  Add an option which allows us\n    +    to force the latter output even when commits have exact renames of\n    +    files.\n     \n         Signed-off-by: Elijah Newren <newren@gmail.com>\n\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n"},{"id":"363269","messageId":"20181114002600.29233-7-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 06/11] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:55Z","receivedAt":"2018-11-14T00:26:27Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If file paths are specified to fast-export and multiple refs point to a\ncommit that does not touch any of the relevant file paths, then\nfast-export can hit problems.  fast-export has a list of additional refs\nthat it needs to explicitly set after exporting all blobs and commits,\nand when it tries to get_object_mark() on the relevant commit, it can\nget a mark of 0, i.e. \"not found\", because the commit in question did\nnot touch the relevant paths and thus was not exported.  Trying to\nimport a stream with a mark corresponding to an unexported object will\ncause fast-import to crash.\n\nAvoid this problem by taking the commit the ref points to and finding an\nancestor of it that was exported, and make the ref point to that commit\ninstead.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  | 13 ++++++++++++-\n t/t9350-fast-export.sh | 20 ++++++++++++++++++++\n 2 files changed, 32 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 7888fc98b5..2eafe351ea 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -900,7 +900,18 @@ static void handle_tags_and_duplicates(void)\n \t\t\tif (anonymize)\n \t\t\t\tname = anonymize_refname(name);\n \t\t\t/* create refs pointing to already seen commits */\n-\t\t\tcommit = (struct commit *)object;\n+\t\t\tcommit = rewrite_commit((struct commit *)object);\n+\t\t\tif (!commit) {\n+\t\t\t\t/*\n+\t\t\t\t * Neither this object nor any of its\n+\t\t\t\t * ancestors touch any relevant paths, so\n+\t\t\t\t * it has been filtered to nothing.  Delete\n+\t\t\t\t * it.\n+\t\t\t\t */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n \t\t\t       get_object_mark(&commit->object));\n \t\t\tshow_progress();\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 3400ebeb51..299120ba70 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -382,6 +382,26 @@ test_expect_success 'path limiting with import-marks does not lose unmodified fi\n \tgrep file0 actual\n '\n \n+test_expect_success 'avoid corrupt stream with non-existent mark' '\n+\ttest_create_repo avoid_non_existent_mark &&\n+\t(\n+\t\tcd avoid_non_existent_mark &&\n+\n+\t\ttest_commit important-path &&\n+\n+\t\ttest_commit ignored &&\n+\n+\t\tgit branch A &&\n+\t\tgit branch B &&\n+\n+\t\techo foo >>important-path.t &&\n+\t\tgit add important-path.t &&\n+\t\ttest_commit more changes &&\n+\n+\t\tgit fast-export --all -- important-path.t | git fast-import --force\n+\t)\n+'\n+\n test_expect_success 'full-tree re-shows unmodified files'        '\n \tgit checkout -f simple &&\n \tgit fast-export --full-tree simple >actual &&\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363270","messageId":"20181114002600.29233-9-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 08/11] fast-export: add --reference-excluded-parents option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:57Z","receivedAt":"2018-11-14T00:26:29Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"git filter-branch has a nifty feature allowing you to rewrite, e.g. just\nthe last 8 commits of a linear history\n  git filter-branch $OPTIONS HEAD~8..HEAD\n\nIf you try the same with git fast-export, you instead get a history of\nonly 8 commits, with HEAD~7 being rewritten into a root commit.  There\nare two alternatives:\n\n  1) Don't use the negative revision specification, and when you're\n     filtering the output to make modifications to the last 8 commits,\n     just be careful to not modify any earlier commits somehow.\n\n  2) First run 'git fast-export --export-marks=somefile HEAD~8', then\n     run 'git fast-export --import-marks=somefile HEAD~8..HEAD'.\n\nBoth are more error prone than I'd like (the first for obvious reasons;\nwith the second option I have sometimes accidentally included too many\nrevisions in the first command and then found that the corresponding\nextra revisions were not exported by the second command and thus were\nnot modified as I expected).  Also, both are poor from a performance\nperspective.\n\nAdd a new --reference-excluded-parents option which will cause\nfast-export to refer to commits outside the specified rev-list-args\nrange by their sha1sum.  Such a stream will only be useful in a\nrepository which already contains the necessary commits (much like the\nrestriction imposed when using --no-data).\n\nNote from Peff:\n  I think we might be able to do a little more optimization here. If\n  we're exporting HEAD^..HEAD and there's an object in HEAD^ which is\n  unchanged in HEAD, I think we'd still print it (because it would not\n  be marked SHOWN), but we could omit it (by walking the tree of the\n  boundary commits and marking them shown).  I don't think it's a\n  blocker for what you're doing here, but just a possible future\n  optimization.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 17 +++++++++++--\n builtin/fast-export.c             | 42 +++++++++++++++++++++++--------\n t/t9350-fast-export.sh            | 11 ++++++++\n 3 files changed, 58 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex fda55b3284..f65026662a 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -110,6 +110,18 @@ marks the same across runs.\n \tthe shape of the history and stored tree.  See the section on\n \t`ANONYMIZING` below.\n \n+--reference-excluded-parents::\n+\tBy default, running a command such as `git fast-export\n+\tmaster~5..master` will not include the commit master{tilde}5\n+\tand will make master{tilde}4 no longer have master{tilde}5 as\n+\ta parent (though both the old master{tilde}4 and new\n+\tmaster{tilde}4 will have all the same files).  Use\n+\t--reference-excluded-parents to instead have the the stream\n+\trefer to commits in the excluded range of history by their\n+\tsha1sum.  Note that the resulting stream can only be used by a\n+\trepository which already contains the necessary parent\n+\tcommits.\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\n@@ -119,8 +131,9 @@ marks the same across runs.\n \t'git rev-list', that specifies the specific objects and references\n \tto export.  For example, `master~10..master` causes the\n \tcurrent master reference to be exported along with all objects\n-\tadded since its 10th ancestor commit and all files common to\n-\tmaster{tilde}9 and master{tilde}10.\n+\tadded since its 10th ancestor commit and (unless the\n+\t--reference-excluded-parents option is specified) all files\n+\tcommon to master{tilde}9 and master{tilde}10.\n \n EXAMPLES\n --------\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 2fef00436b..3cc98c31ad 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -37,6 +37,7 @@ static int fake_missing_tagger;\n static int use_done_feature;\n static int no_data;\n static int full_tree;\n+static int reference_excluded_commits;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n@@ -596,7 +597,8 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\tmessage += 2;\n \n \tif (commit->parents &&\n-\t    get_object_mark(&commit->parents->item->object) != 0 &&\n+\t    (get_object_mark(&commit->parents->item->object) != 0 ||\n+\t     reference_excluded_commits) &&\n \t    !full_tree) {\n \t\tparse_commit_or_die(commit->parents->item);\n \t\tdiff_tree_oid(get_commit_tree_oid(commit->parents->item),\n@@ -644,13 +646,21 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \tunuse_commit_buffer(commit, commit_buffer);\n \n \tfor (i = 0, p = commit->parents; p; p = p->next) {\n-\t\tint mark = get_object_mark(&p->item->object);\n-\t\tif (!mark)\n+\t\tstruct object *obj = &p->item->object;\n+\t\tint mark = get_object_mark(obj);\n+\n+\t\tif (!mark && !reference_excluded_commits)\n \t\t\tcontinue;\n \t\tif (i == 0)\n-\t\t\tprintf(\"from :%d\\n\", mark);\n+\t\t\tprintf(\"from \");\n+\t\telse\n+\t\t\tprintf(\"merge \");\n+\t\tif (mark)\n+\t\t\tprintf(\":%d\\n\", mark);\n \t\telse\n-\t\t\tprintf(\"merge :%d\\n\", mark);\n+\t\t\tprintf(\"%s\\n\", sha1_to_hex(anonymize ?\n+\t\t\t\t\t\t   anonymize_sha1(&obj->oid) :\n+\t\t\t\t\t\t   obj->oid.hash));\n \t\ti++;\n \t}\n \n@@ -931,13 +941,22 @@ static void handle_tags_and_duplicates(struct string_list *extras)\n \t\t\t\t/*\n \t\t\t\t * Getting here means we have a commit which\n \t\t\t\t * was excluded by a negative refspec (e.g.\n-\t\t\t\t * fast-export ^master master).  If the user\n+\t\t\t\t * fast-export ^master master).  If we are\n+\t\t\t\t * referencing excluded commits, set the ref\n+\t\t\t\t * to the exact commit.  Otherwise, the user\n \t\t\t\t * wants the branch exported but every commit\n-\t\t\t\t * in its history to be deleted, that sounds\n-\t\t\t\t * like a ref deletion to me.\n+\t\t\t\t * in its history to be deleted, which basically\n+\t\t\t\t * just means deletion of the ref.\n \t\t\t\t */\n-\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\tif (!reference_excluded_commits) {\n+\t\t\t\t\t/* delete the ref */\n+\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\t/* set ref to commit using oid, not mark */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\", name,\n+\t\t\t\t       sha1_to_hex(commit->object.oid.hash));\n \t\t\t\tcontinue;\n \t\t\t}\n \n@@ -1074,6 +1093,9 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\tOPT_STRING_LIST(0, \"refspec\", &refspecs_list, N_(\"refspec\"),\n \t\t\t     N_(\"Apply refspec to exported refs\")),\n \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n+\t\tOPT_BOOL(0, \"reference-excluded-parents\",\n+\t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n+\n \t\tOPT_END()\n \t};\n \ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 50c2fceef4..d7d73061d0 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -66,6 +66,17 @@ test_expect_success 'fast-export master~2..master' '\n \n '\n \n+test_expect_success 'fast-export --reference-excluded-parents master~2..master' '\n+\n+\tgit fast-export --reference-excluded-parents master~2..master >actual &&\n+\tgrep commit.refs/heads/master actual >commit-count &&\n+\ttest_line_count = 2 commit-count &&\n+\tsed \"s/master/rewrite/\" actual |\n+\t\t(cd new &&\n+\t\t git fast-import &&\n+\t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n+'\n+\n test_expect_success 'iso-8859-1' '\n \n \tgit config i18n.commitencoding ISO8859-1 &&\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363272","messageId":"20181114002600.29233-4-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 03/11] fast-export: use value from correct enum","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:52Z","receivedAt":"2018-11-14T00:26:31Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"ABORT and ERROR happen to have the same value, but come from differnt\nenums.  Use the one from the correct enum, and while at it, rename the\nvalues to avoid such problems.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 456797c12a..af724e9937 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -31,8 +31,8 @@ static const char *fast_export_usage[] = {\n };\n \n static int progress;\n-static enum { ABORT, VERBATIM, WARN, WARN_STRIP, STRIP } signed_tag_mode = ABORT;\n-static enum { ERROR, DROP, REWRITE } tag_of_filtered_mode = ERROR;\n+static enum { SIGNED_TAG_ABORT, VERBATIM, WARN, WARN_STRIP, STRIP } signed_tag_mode = SIGNED_TAG_ABORT;\n+static enum { TAG_FILTERING_ABORT, DROP, REWRITE } tag_of_filtered_mode = TAG_FILTERING_ABORT;\n static int fake_missing_tagger;\n static int use_done_feature;\n static int no_data;\n@@ -46,7 +46,7 @@ static int parse_opt_signed_tag_mode(const struct option *opt,\n \t\t\t\t     const char *arg, int unset)\n {\n \tif (unset || !strcmp(arg, \"abort\"))\n-\t\tsigned_tag_mode = ABORT;\n+\t\tsigned_tag_mode = SIGNED_TAG_ABORT;\n \telse if (!strcmp(arg, \"verbatim\") || !strcmp(arg, \"ignore\"))\n \t\tsigned_tag_mode = VERBATIM;\n \telse if (!strcmp(arg, \"warn\"))\n@@ -64,7 +64,7 @@ static int parse_opt_tag_of_filtered_mode(const struct option *opt,\n \t\t\t\t\t  const char *arg, int unset)\n {\n \tif (unset || !strcmp(arg, \"abort\"))\n-\t\ttag_of_filtered_mode = ERROR;\n+\t\ttag_of_filtered_mode = TAG_FILTERING_ABORT;\n \telse if (!strcmp(arg, \"drop\"))\n \t\ttag_of_filtered_mode = DROP;\n \telse if (!strcmp(arg, \"rewrite\"))\n@@ -727,7 +727,7 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t\t       \"\\n-----BEGIN PGP SIGNATURE-----\\n\");\n \t\tif (signature)\n \t\t\tswitch(signed_tag_mode) {\n-\t\t\tcase ABORT:\n+\t\t\tcase SIGNED_TAG_ABORT:\n \t\t\t\tdie(\"encountered signed tag %s; use \"\n \t\t\t\t    \"--signed-tags=<mode> to handle it\",\n \t\t\t\t    oid_to_hex(&tag->object.oid));\n@@ -752,7 +752,7 @@ static void handle_tag(const char *name, struct tag *tag)\n \ttagged_mark = get_object_mark(tagged);\n \tif (!tagged_mark) {\n \t\tswitch(tag_of_filtered_mode) {\n-\t\tcase ABORT:\n+\t\tcase TAG_FILTERING_ABORT:\n \t\t\tdie(\"tag %s tags unexported object; use \"\n \t\t\t    \"--tag-of-filtered-object=<mode> to handle it\",\n \t\t\t    oid_to_hex(&tag->object.oid));\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363271","messageId":"20181114002600.29233-3-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v2 02/11] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T00:25:51Z","receivedAt":"2018-11-14T00:26:32Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Signed-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 3 ++-\n 1 file changed, 2 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex ce954be532..fda55b3284 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -119,7 +119,8 @@ marks the same across runs.\n \t'git rev-list', that specifies the specific objects and references\n \tto export.  For example, `master~10..master` causes the\n \tcurrent master reference to be exported along with all objects\n-\tadded since its 10th ancestor commit.\n+\tadded since its 10th ancestor commit and all files common to\n+\tmaster{tilde}9 and master{tilde}10.\n \n EXAMPLES\n --------\n-- \n2.19.1.1063.g2b8e4a4f82.dirty\n\n"},{"id":"363326","messageId":"20181114071454.GB19904@sigill.intra.peff.net","threadId":"49406","inReplyTo":"CABPp-BFbtusiT30_gU7SgmmMg25NCdgNTSEEJJysoT-1MwSnkA@mail.gmail.com","subject":"Re: [PATCH 10/10] fast-export: add --always-show-modify-after-rename","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-14T07:14:55Z","receivedAt":"2018-11-14T07:14:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Nov 13, 2018 at 09:10:36AM -0800, Elijah Newren wrote:\n\n> > I am looking at this problem as \"how do you answer question X in a\n> > repository\". And I think you are looking at as \"I am receiving a\n> > fast-export stream, and I need to answer question X on the fly\".\n> >\n> > And that would explain why you want to get extra annotations into the\n> > fast-export stream. Is that right?\n> \n> I'm not trying to get information on the fly during a rewrite or\n> anything like that.  This is an optional pre-rewrite step (from a\n> separate invocation of the tool) where I have multiple questions I\n> want to answer.  I'd like to answer them all relatively quickly, if\n> possible, and I think all of them should be answerable with a single\n> history traversal (plus a cat-file --batch-all-objects call to get\n> object sizes, since I don't know of another way to get those).  I'd be\n> fine with switching from fast-export to log or something else if it\n> met the needs better.\n\nAh, OK. Yes, if we're just trying to query, then I think you should be\nable to do what you want with the existing traversal and diff tools. And\nif not, we should think about a new feature there, and not try to\nshoe-horn it into fast-export.\n\n> As far as I can tell, you're trying to split each question apart and\n> do a history traversal for each, and I don't see why that's better.\n> Simpler, perhaps, but it seems worse for performance.  Am I missing\n> something?\n\nI was only trying to address each possible query individually. I agree\nthat if you are querying both things, you should be able to do it in a\nsingle traversal (and that is strictly better). It may require a little\nmore parsing of the output (e.g., `--find-object` is easy to implement\nyourself looking at --raw output).\n\n> Ah, I didn't know renames were on by default; I somehow missed that.\n> Also, the rev-list to diff-tree pipe is nice, but I also need parent\n> and commit timestamp information.\n\ndiff-tree will format the commit info as well (before git-log was a C\nbuiltin, it was just a rev-list/diff-tree pipeline in a shell script).\nSo you can do:\n\n  git rev-list ... |\n  git diff-tree --stdin --format='%h %ct %p' --raw -r -M\n\nand get dump very similar to what fast-export would give you.\n\n> > >   git log -M --diff-filter=RAMD --no-abbrev --raw\n> >\n> > What is there besides RAMD? :)\n> \n> Well, as you pointed out above, log detects renames by default,\n> whereas it didn't used to.\n> So, if someone had written some similar-ish history walking/parsing\n> tool years ago that didn't depend need renames and was based on log\n> output, there's a good chance their tool might start failing when\n> rename detection was turned on by default, because instead of getting\n> both a 'D' and an 'M' change, they'd get an unexpected 'R'.\n\nMostly I just meant: your diff-filter includes basically everything, so\nwhy bother filtering? You're going to have to parse the result anyway,\nand you can throw away uninteresting bits there.\n\n> For my case, do I have to worry about similar future changes?  Will\n> copy detection ('C') or break detection ('B') become the default in\n> the future?  Do I have to worry about typechanges ('T\")?  Will new\n> change types be added?  I mean, the fast-export output could maybe\n> change too, but it seems much less likely than with log.\n\nIf you use diff-tree, then it won't ever enable copy or break detection\nwithout you explicitly asking for it.\n\n> Let me try to put it as briefly as I can.  With as few traversals as\n> possible, I want to:\n>   * Get all blob sizes\n>   * Map blob shas to filename(s) they appeared under in the history\n>   * Find when files and directories were deleted (and whether they\n> were later reinstated, since that means they aren't actually gone)\n>   * Find sets of filenames referring to the same logical 'file'. (e.g.\n> foo->bar in commit A and bar->baz in commit B mean that {foo,bar,baz}\n> refer to the same 'file' so that a user has an easy report to look at\n> to find out that if they just want to \"keep baz and its history\" then\n> they need foo & bar & baz.  I need to know about things like another\n> foo or bar being introduced after the rename though, since that breaks\n> the connection between filenames)\n>   * Do a few aggregations on the above data as well (e.g. all copies\n> of postgres.exe add up to 20M -- why were those checked in anyway?,\n> *.webm files in aggregate are .5G, your long-deleted src/video-server/\n> directory from that aborted experimental project years ago takes up 2G\n> of your history, etc.)\n> \n> Right now, my best solution for this combination of questions is\n> 'cat-file --batch-all-objects' plus fast-export, if I get patch 10/10\n> in place.  I'm totally open to better solutions, including ones that\n> don't use fast-export.\n\nOK, I think I understand your problem better now. I don't think there's\nanything fast-export can show that log/diff-tree could not, aside from\nactual blob contents. But I don't think you want them (and if you did,\nyou can use \"cat-file --batch\" to selectively request them).\n\nI think there's a general problem with any serialized output (log or\nfast-export) that things like rename tracking depend on the topology. If\nI rename \"foo\" to \"bar\" on one branch, and \"bar\" to \"baz\" on another\nbranch, without reconstructing the parent graph you don't realize that\nthose two things were on parallel branches, and not a sequence.  But\nwith the parent ids, you can delve as deep as you like in your analysis\nscript.\n\n-Peff\n"},{"id":"363327","messageId":"20181114072505.GC19904@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"Re: [PATCH v2 00/11] fast export and import fixes and features","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-14T07:25:06Z","receivedAt":"2018-11-14T07:25:20Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Nov 13, 2018 at 04:25:49PM -0800, Elijah Newren wrote:\n\n> This is a series of small fixes and features for fast-export and\n> fast-import, mostly on the fast-export side.\n\nI looked over this, and I think you've addressed all of my questions.\n\nA few quick comments:\n\n> Changes since v1 (full range-diff below):\n>   - used {tilde} in asciidoc documentation to avoid subscripting and\n>     escaping problems\n\nI think just using backticks would make the source more readable, as\nwell as make the output prettier. But that's pretty minor.\n\n>   - renamed ABORT/ERROR enum values to help avoid further misusage\n\nThis is an improvement, I think. It's a little funny that we still have\nbare names for the non-ABORT bits, though (there's less semantic\noverlap, but if it's a good practice to use qualified enum names, we\nshould probably just do so consistently).\n\n>   - multiple small testcase cleanups (use $ZERO_OID, remove grep -A, etc.)\n\nLooks good.\n\n>   - add FIXME comment to code about string_list usage\n\nMakes sense.\n\n>   - record Peff's idea for a future optimization in patch 8 commit message\n>     (is there a better place to put that??)\n\nSeems like a reasonable place (though you are welcome to restate it if\nyou like).\n\n>   - New patch (9/11): remove the unmaintained copy of fast-import stream\n>     format documentation at the beginning of fast-import.c\n\nLooks good. I wondered if there might be bits that need migrated, but\ngiven the length of time that comment has been there, it's unlikely. And\nin the worst case, if somebody finds some information missing from\ngit-fast-import.txt, they can still consult the history.\n\n>   - Rewrite commit message for 10/11 to match the wording Peff liked\n>     better, s/originally/original-oid/, and add documentation to\n>     git-fast-import.txt\n\nLooks good.\n\n>   - Rewrite commit message for 11/11; the last one didn't make sense to\n>     Peff.  I hope this one does.\n\nThanks for your patience in getting me to understand what you're trying\nto do. At this point I still think that using rev-list and diff-tree is\nprobably the right solution for your use case.\n\n-Peff\n"},{"id":"363390","messageId":"20181114191741.GJ30222@szeder.dev","threadId":"49406","inReplyTo":"20181114002600.29233-5-newren@gmail.com","subject":"Re: [PATCH v2 04/11] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-11-14T19:17:41Z","receivedAt":"2018-11-14T19:17:47Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Nov 13, 2018 at 04:25:53PM -0800, Elijah Newren wrote:\n> diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n> index af724e9937..b984a44224 100644\n> --- a/builtin/fast-export.c\n> +++ b/builtin/fast-export.c\n> @@ -774,9 +774,12 @@ static void handle_tag(const char *name, struct tag *tag)\n>  \t\t\t\t\tbreak;\n>  \t\t\t\tif (!(p->object.flags & TREESAME))\n>  \t\t\t\t\tbreak;\n> -\t\t\t\tif (!p->parents)\n> -\t\t\t\t\tdie(\"can't find replacement commit for tag %s\",\n> -\t\t\t\t\t     oid_to_hex(&tag->object.oid));\n> +\t\t\t\tif (!p->parents) {\n> +\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n> +\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n\nPlease use oid_to_hex(&null_oid) instead.\n\n> +\t\t\t\t\tfree(buf);\n> +\t\t\t\t\treturn;\n> +\t\t\t\t}\n>  \t\t\t\tp = p->parents->item;\n>  \t\t\t}\n>  \t\t\ttagged_mark = get_object_mark(&p->object);\n"},{"id":"363391","messageId":"20181114192758.GK30222@szeder.dev","threadId":"49406","inReplyTo":"20181114002600.29233-9-newren@gmail.com","subject":"Re: [PATCH v2 08/11] fast-export: add --reference-excluded-parents option","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-11-14T19:27:58Z","receivedAt":"2018-11-14T19:28:05Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Nov 13, 2018 at 04:25:57PM -0800, Elijah Newren wrote:\n> diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n> index 2fef00436b..3cc98c31ad 100644\n> --- a/builtin/fast-export.c\n> +++ b/builtin/fast-export.c\n> @@ -37,6 +37,7 @@ static int fake_missing_tagger;\n>  static int use_done_feature;\n>  static int no_data;\n>  static int full_tree;\n> +static int reference_excluded_commits;\n>  static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n>  static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n>  static struct refspec refspecs = REFSPEC_INIT_FETCH;\n> @@ -596,7 +597,8 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n>  \t\tmessage += 2;\n>  \n>  \tif (commit->parents &&\n> -\t    get_object_mark(&commit->parents->item->object) != 0 &&\n> +\t    (get_object_mark(&commit->parents->item->object) != 0 ||\n> +\t     reference_excluded_commits) &&\n>  \t    !full_tree) {\n>  \t\tparse_commit_or_die(commit->parents->item);\n>  \t\tdiff_tree_oid(get_commit_tree_oid(commit->parents->item),\n> @@ -644,13 +646,21 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n>  \tunuse_commit_buffer(commit, commit_buffer);\n>  \n>  \tfor (i = 0, p = commit->parents; p; p = p->next) {\n> -\t\tint mark = get_object_mark(&p->item->object);\n> -\t\tif (!mark)\n> +\t\tstruct object *obj = &p->item->object;\n> +\t\tint mark = get_object_mark(obj);\n> +\n> +\t\tif (!mark && !reference_excluded_commits)\n>  \t\t\tcontinue;\n>  \t\tif (i == 0)\n> -\t\t\tprintf(\"from :%d\\n\", mark);\n> +\t\t\tprintf(\"from \");\n> +\t\telse\n> +\t\t\tprintf(\"merge \");\n> +\t\tif (mark)\n> +\t\t\tprintf(\":%d\\n\", mark);\n>  \t\telse\n> -\t\t\tprintf(\"merge :%d\\n\", mark);\n> +\t\t\tprintf(\"%s\\n\", sha1_to_hex(anonymize ?\n> +\t\t\t\t\t\t   anonymize_sha1(&obj->oid) :\n> +\t\t\t\t\t\t   obj->oid.hash));\n\nSince we intend to move away from SHA-1, would this be a good time to\nadd an anonymize_oid() function, \"while at it\"?\n\n>  \t\ti++;\n>  \t}\n>  \n> @@ -931,13 +941,22 @@ static void handle_tags_and_duplicates(struct string_list *extras)\n>  \t\t\t\t/*\n>  \t\t\t\t * Getting here means we have a commit which\n>  \t\t\t\t * was excluded by a negative refspec (e.g.\n> -\t\t\t\t * fast-export ^master master).  If the user\n> +\t\t\t\t * fast-export ^master master).  If we are\n> +\t\t\t\t * referencing excluded commits, set the ref\n> +\t\t\t\t * to the exact commit.  Otherwise, the user\n>  \t\t\t\t * wants the branch exported but every commit\n> -\t\t\t\t * in its history to be deleted, that sounds\n> -\t\t\t\t * like a ref deletion to me.\n> +\t\t\t\t * in its history to be deleted, which basically\n> +\t\t\t\t * just means deletion of the ref.\n>  \t\t\t\t */\n> -\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n> -\t\t\t\t       name, sha1_to_hex(null_sha1));\n> +\t\t\t\tif (!reference_excluded_commits) {\n> +\t\t\t\t\t/* delete the ref */\n> +\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n> +\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n> +\t\t\t\t\tcontinue;\n> +\t\t\t\t}\n> +\t\t\t\t/* set ref to commit using oid, not mark */\n> +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\", name,\n> +\t\t\t\t       sha1_to_hex(commit->object.oid.hash));\n\nPlease use oid_to_hex(&commit->object.oid) instead.\n\n>  \t\t\t\tcontinue;\n>  \t\t\t}\n>  \n"},{"id":"363405","messageId":"CABPp-BFGL_rbEydXASVSW=L+gZLcCOmfATdi=bmp0Kk4av90eg@mail.gmail.com","threadId":"49406","inReplyTo":"20181114191741.GJ30222@szeder.dev","subject":"Re: [PATCH v2 04/11] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T23:13:04Z","receivedAt":"2018-11-14T23:13:18Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Nov 14, 2018 at 11:17 AM SZEDER Gábor <szeder.dev@gmail.com> wrote:\n> On Tue, Nov 13, 2018 at 04:25:53PM -0800, Elijah Newren wrote:\n> > diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n> > index af724e9937..b984a44224 100644\n> > --- a/builtin/fast-export.c\n> > +++ b/builtin/fast-export.c\n> > @@ -774,9 +774,12 @@ static void handle_tag(const char *name, struct tag *tag)\n> >                                       break;\n> >                               if (!(p->object.flags & TREESAME))\n> >                                       break;\n> > -                             if (!p->parents)\n> > -                                     die(\"can't find replacement commit for tag %s\",\n> > -                                          oid_to_hex(&tag->object.oid));\n> > +                             if (!p->parents) {\n> > +                                     printf(\"reset %s\\nfrom %s\\n\\n\",\n> > +                                            name, sha1_to_hex(null_sha1));\n>\n> Please use oid_to_hex(&null_oid) instead.\n\nWill do.  Looks like origin/master:builtin/fast-export.c already had\ntwo sha1_to_hex() calls, so I'll add a cleanup patch fixing those too.\n"},{"id":"363406","messageId":"CABPp-BGqS4KgSaJ85Kku1xFaFw4Cq0UsVDkk389EC+VLtU__3Q@mail.gmail.com","threadId":"49406","inReplyTo":"20181114192758.GK30222@szeder.dev","subject":"Re: [PATCH v2 08/11] fast-export: add --reference-excluded-parents option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-14T23:16:25Z","receivedAt":"2018-11-14T23:16:39Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Wed, Nov 14, 2018 at 11:28 AM SZEDER Gábor <szeder.dev@gmail.com> wrote:\n>\n> On Tue, Nov 13, 2018 at 04:25:57PM -0800, Elijah Newren wrote:\n> > diff --git a/builtin/fast-export.c b/builtin/fast-export.c\n> > index 2fef00436b..3cc98c31ad 100644\n> > --- a/builtin/fast-export.c\n> > +++ b/builtin/fast-export.c\n> > @@ -37,6 +37,7 @@ static int fake_missing_tagger;\n> >  static int use_done_feature;\n> >  static int no_data;\n> >  static int full_tree;\n> > +static int reference_excluded_commits;\n> >  static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n> >  static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n> >  static struct refspec refspecs = REFSPEC_INIT_FETCH;\n> > @@ -596,7 +597,8 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n> >               message += 2;\n> >\n> >       if (commit->parents &&\n> > -         get_object_mark(&commit->parents->item->object) != 0 &&\n> > +         (get_object_mark(&commit->parents->item->object) != 0 ||\n> > +          reference_excluded_commits) &&\n> >           !full_tree) {\n> >               parse_commit_or_die(commit->parents->item);\n> >               diff_tree_oid(get_commit_tree_oid(commit->parents->item),\n> > @@ -644,13 +646,21 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n> >       unuse_commit_buffer(commit, commit_buffer);\n> >\n> >       for (i = 0, p = commit->parents; p; p = p->next) {\n> > -             int mark = get_object_mark(&p->item->object);\n> > -             if (!mark)\n> > +             struct object *obj = &p->item->object;\n> > +             int mark = get_object_mark(obj);\n> > +\n> > +             if (!mark && !reference_excluded_commits)\n> >                       continue;\n> >               if (i == 0)\n> > -                     printf(\"from :%d\\n\", mark);\n> > +                     printf(\"from \");\n> > +             else\n> > +                     printf(\"merge \");\n> > +             if (mark)\n> > +                     printf(\":%d\\n\", mark);\n> >               else\n> > -                     printf(\"merge :%d\\n\", mark);\n> > +                     printf(\"%s\\n\", sha1_to_hex(anonymize ?\n> > +                                                anonymize_sha1(&obj->oid) :\n> > +                                                obj->oid.hash));\n>\n> Since we intend to move away from SHA-1, would this be a good time to\n> add an anonymize_oid() function, \"while at it\"?\n\nSince I already need to add a cleanup commit to remove the\npre-existing sha1_to_hex() calls, I'll just\ns/anonymize_sha1/anonymize_oid/ while at it in the same commit; it's\nnot called from any other file.\n\n> >               i++;\n> >       }\n> >\n> > @@ -931,13 +941,22 @@ static void handle_tags_and_duplicates(struct string_list *extras)\n> >                               /*\n> >                                * Getting here means we have a commit which\n> >                                * was excluded by a negative refspec (e.g.\n> > -                              * fast-export ^master master).  If the user\n> > +                              * fast-export ^master master).  If we are\n> > +                              * referencing excluded commits, set the ref\n> > +                              * to the exact commit.  Otherwise, the user\n> >                                * wants the branch exported but every commit\n> > -                              * in its history to be deleted, that sounds\n> > -                              * like a ref deletion to me.\n> > +                              * in its history to be deleted, which basically\n> > +                              * just means deletion of the ref.\n> >                                */\n> > -                             printf(\"reset %s\\nfrom %s\\n\\n\",\n> > -                                    name, sha1_to_hex(null_sha1));\n> > +                             if (!reference_excluded_commits) {\n> > +                                     /* delete the ref */\n> > +                                     printf(\"reset %s\\nfrom %s\\n\\n\",\n> > +                                            name, sha1_to_hex(null_sha1));\n> > +                                     continue;\n> > +                             }\n> > +                             /* set ref to commit using oid, not mark */\n> > +                             printf(\"reset %s\\nfrom %s\\n\\n\", name,\n> > +                                    sha1_to_hex(commit->object.oid.hash));\n>\n> Please use oid_to_hex(&commit->object.oid) instead.\n\nYeah, there were a couple others I introduced too.  I'll fix them all up.\n"},{"id":"363470","messageId":"20181116075956.27047-4-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 03/11] git-fast-export.txt: clarify misleading documentation about rev-list args","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:48Z","receivedAt":"2018-11-16T08:00:17Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Signed-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 3 ++-\n 1 file changed, 2 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex ce954be532..fda55b3284 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -119,7 +119,8 @@ marks the same across runs.\n \t'git rev-list', that specifies the specific objects and references\n \tto export.  For example, `master~10..master` causes the\n \tcurrent master reference to be exported along with all objects\n-\tadded since its 10th ancestor commit.\n+\tadded since its 10th ancestor commit and all files common to\n+\tmaster{tilde}9 and master{tilde}10.\n \n EXAMPLES\n --------\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363471","messageId":"20181116075956.27047-8-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 07/11] fast-export: when using paths, avoid corrupt stream with non-existent mark","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:52Z","receivedAt":"2018-11-16T08:00:17Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If file paths are specified to fast-export and multiple refs point to a\ncommit that does not touch any of the relevant file paths, then\nfast-export can hit problems.  fast-export has a list of additional refs\nthat it needs to explicitly set after exporting all blobs and commits,\nand when it tries to get_object_mark() on the relevant commit, it can\nget a mark of 0, i.e. \"not found\", because the commit in question did\nnot touch the relevant paths and thus was not exported.  Trying to\nimport a stream with a mark corresponding to an unexported object will\ncause fast-import to crash.\n\nAvoid this problem by taking the commit the ref points to and finding an\nancestor of it that was exported, and make the ref point to that commit\ninstead.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  | 13 ++++++++++++-\n t/t9350-fast-export.sh | 20 ++++++++++++++++++++\n 2 files changed, 32 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 43e98a38a8..227488ae84 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -901,7 +901,18 @@ static void handle_tags_and_duplicates(void)\n \t\t\tif (anonymize)\n \t\t\t\tname = anonymize_refname(name);\n \t\t\t/* create refs pointing to already seen commits */\n-\t\t\tcommit = (struct commit *)object;\n+\t\t\tcommit = rewrite_commit((struct commit *)object);\n+\t\t\tif (!commit) {\n+\t\t\t\t/*\n+\t\t\t\t * Neither this object nor any of its\n+\t\t\t\t * ancestors touch any relevant paths, so\n+\t\t\t\t * it has been filtered to nothing.  Delete\n+\t\t\t\t * it.\n+\t\t\t\t */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, oid_to_hex(&null_oid));\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n \t\t\t       get_object_mark(&commit->object));\n \t\t\tshow_progress();\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 3400ebeb51..299120ba70 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -382,6 +382,26 @@ test_expect_success 'path limiting with import-marks does not lose unmodified fi\n \tgrep file0 actual\n '\n \n+test_expect_success 'avoid corrupt stream with non-existent mark' '\n+\ttest_create_repo avoid_non_existent_mark &&\n+\t(\n+\t\tcd avoid_non_existent_mark &&\n+\n+\t\ttest_commit important-path &&\n+\n+\t\ttest_commit ignored &&\n+\n+\t\tgit branch A &&\n+\t\tgit branch B &&\n+\n+\t\techo foo >>important-path.t &&\n+\t\tgit add important-path.t &&\n+\t\ttest_commit more changes &&\n+\n+\t\tgit fast-export --all -- important-path.t | git fast-import --force\n+\t)\n+'\n+\n test_expect_success 'full-tree re-shows unmodified files'        '\n \tgit checkout -f simple &&\n \tgit fast-export --full-tree simple >actual &&\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363472","messageId":"20181116075956.27047-2-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 01/11] fast-export: convert sha1 to oid","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:46Z","receivedAt":"2018-11-16T08:00:18Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Rename anonymize_sha1() to anonymize_oid(() and change its signature,\nand switch from sha1_to_hex() to oid_to_hex() and from GIT_SHA1_RAWSZ to\nthe_hash_algo->rawsz.  Also change a comment and a die string to mention\noid instead of sha1.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 25 +++++++++++++------------\n 1 file changed, 13 insertions(+), 12 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 456797c12a..f5166ac71e 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -243,7 +243,7 @@ static void export_blob(const struct object_id *oid)\n \t\tif (!buf)\n \t\t\tdie(\"could not read blob %s\", oid_to_hex(oid));\n \t\tif (check_object_signature(oid, buf, size, type_name(type)) < 0)\n-\t\t\tdie(\"sha1 mismatch in blob %s\", oid_to_hex(oid));\n+\t\t\tdie(\"oid mismatch in blob %s\", oid_to_hex(oid));\n \t\tobject = parse_object_buffer(the_repository, oid, type,\n \t\t\t\t\t     size, buf, &eaten);\n \t}\n@@ -330,17 +330,18 @@ static void print_path(const char *path)\n \n static void *generate_fake_oid(const void *old, size_t *len)\n {\n-\tstatic uint32_t counter = 1; /* avoid null sha1 */\n-\tunsigned char *out = xcalloc(GIT_SHA1_RAWSZ, 1);\n-\tput_be32(out + GIT_SHA1_RAWSZ - 4, counter++);\n+\tstatic uint32_t counter = 1; /* avoid null oid */\n+\tconst unsigned hashsz = the_hash_algo->rawsz;\n+\tunsigned char *out = xcalloc(hashsz, 1);\n+\tput_be32(out + hashsz - 4, counter++);\n \treturn out;\n }\n \n-static const unsigned char *anonymize_sha1(const struct object_id *oid)\n+static const struct object_id *anonymize_oid(const struct object_id *oid)\n {\n-\tstatic struct hashmap sha1s;\n-\tsize_t len = GIT_SHA1_RAWSZ;\n-\treturn anonymize_mem(&sha1s, generate_fake_oid, oid, &len);\n+\tstatic struct hashmap objs;\n+\tsize_t len = the_hash_algo->rawsz;\n+\treturn anonymize_mem(&objs, generate_fake_oid, oid, &len);\n }\n \n static void show_filemodify(struct diff_queue_struct *q,\n@@ -399,9 +400,9 @@ static void show_filemodify(struct diff_queue_struct *q,\n \t\t\t */\n \t\t\tif (no_data || S_ISGITLINK(spec->mode))\n \t\t\t\tprintf(\"M %06o %s \", spec->mode,\n-\t\t\t\t       sha1_to_hex(anonymize ?\n-\t\t\t\t\t\t   anonymize_sha1(&spec->oid) :\n-\t\t\t\t\t\t   spec->oid.hash));\n+\t\t\t\t       oid_to_hex(anonymize ?\n+\t\t\t\t\t\t  anonymize_oid(&spec->oid) :\n+\t\t\t\t\t\t  &spec->oid));\n \t\t\telse {\n \t\t\t\tstruct object *object = lookup_object(the_repository,\n \t\t\t\t\t\t\t\t      spec->oid.hash);\n@@ -988,7 +989,7 @@ static void handle_deletes(void)\n \t\t\tcontinue;\n \n \t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\trefspec->dst, sha1_to_hex(null_sha1));\n+\t\t\t\trefspec->dst, oid_to_hex(&null_oid));\n \t}\n }\n \n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363473","messageId":"20181116075956.27047-3-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 02/11] git-fast-import.txt: fix documentation for --quiet option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:47Z","receivedAt":"2018-11-16T08:00:19Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Signed-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-import.txt | 7 ++++---\n 1 file changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\nindex e81117d27f..7ab97745a6 100644\n--- a/Documentation/git-fast-import.txt\n+++ b/Documentation/git-fast-import.txt\n@@ -40,9 +40,10 @@ OPTIONS\n \tnot contain the old commit).\n \n --quiet::\n-\tDisable all non-fatal output, making fast-import silent when it\n-\tis successful.  This option disables the output shown by\n-\t--stats.\n+\tDisable the output shown by --stats, making fast-import usually\n+\tbe silent when it is successful.  However, if the import stream\n+\thas directives intended to show user output (e.g. `progress`\n+\tdirectives), the corresponding messages will still be shown.\n \n --stats::\n \tDisplay some basic statistics about the objects fast-import has\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363474","messageId":"20181116075956.27047-9-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 08/11] fast-export: ensure we export requested refs","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:53Z","receivedAt":"2018-11-16T08:00:21Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If file paths are specified to fast-export and a ref points to a commit\nthat does not touch any of the relevant paths, then that ref would\nsometimes fail to be exported.  (This depends on whether any ancestors\nof the commit which do touch the relevant paths would be exported with\nthat same ref name or a different ref name.)  To avoid this problem,\nput *all* specified refs into extra_refs to start, and then as we export\neach commit, remove the refname used in the 'commit $REFNAME' directive\nfrom extra_refs.  Then, in handle_tags_and_duplicates() we know which\nrefs actually do need a manual reset directive in order to be included.\n\nThis means that we do need some special handling for excluded refs; e.g.\nif someone runs\n   git fast-export ^master master\nthen they've asked for master to be exported, but they have also asked\nfor the commit which master points to and all of its history to be\nexcluded.  That logically means ref deletion.  Previously, such refs\nwere just silently omitted from being exported despite having been\nexplicitly requested for export.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  | 54 ++++++++++++++++++++++++++++++++----------\n t/t9350-fast-export.sh | 16 ++++++++++---\n 2 files changed, 55 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 227488ae84..d71e0333d4 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n+static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n static int anonymize;\n static struct revision_sources revision_sources;\n@@ -612,6 +613,13 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\t\texport_blob(&diff_queued_diff.queue[i]->two->oid);\n \n \trefname = *revision_sources_at(&revision_sources, commit);\n+\t/*\n+\t * FIXME: string_list_remove() below for each ref is overall\n+\t * O(N^2).  Compared to a history walk and diffing trees, this is\n+\t * just lost in the noise in practice.  However, theoretically a\n+\t * repo may have enough refs for this to become slow.\n+\t */\n+\tstring_list_remove(&extra_refs, refname, 0);\n \tif (anonymize) {\n \t\trefname = anonymize_refname(refname);\n \t\tanonymize_ident_line(&committer, &committer_end);\n@@ -815,7 +823,7 @@ static struct commit *get_commit(struct rev_cmdline_entry *e, char *full_name)\n \t\t/* handle nested tags */\n \t\twhile (tag && tag->object.type == OBJ_TAG) {\n \t\t\tparse_object(the_repository, &tag->object.oid);\n-\t\t\tstring_list_append(&extra_refs, full_name)->util = tag;\n+\t\t\tstring_list_append(&tag_refs, full_name)->util = tag;\n \t\t\ttag = (struct tag *)tag->tagged;\n \t\t}\n \t\tif (!tag)\n@@ -874,25 +882,30 @@ static void get_tags_and_duplicates(struct rev_cmdline_info *info)\n \t\t}\n \n \t\t/*\n-\t\t * This ref will not be updated through a commit, lets make\n-\t\t * sure it gets properly updated eventually.\n+\t\t * Make sure this ref gets properly updated eventually, whether\n+\t\t * through a commit or manually at the end.\n \t\t */\n-\t\tif (*revision_sources_at(&revision_sources, commit) ||\n-\t\t    commit->object.flags & SHOWN)\n+\t\tif (e->item->type != OBJ_TAG)\n \t\t\tstring_list_append(&extra_refs, full_name)->util = commit;\n+\n \t\tif (!*revision_sources_at(&revision_sources, commit))\n \t\t\t*revision_sources_at(&revision_sources, commit) = full_name;\n \t}\n+\n+\tstring_list_sort(&extra_refs);\n+\tstring_list_remove_duplicates(&extra_refs, 0);\n }\n \n-static void handle_tags_and_duplicates(void)\n+static void handle_tags_and_duplicates(struct string_list *extras)\n {\n \tstruct commit *commit;\n \tint i;\n \n-\tfor (i = extra_refs.nr - 1; i >= 0; i--) {\n-\t\tconst char *name = extra_refs.items[i].string;\n-\t\tstruct object *object = extra_refs.items[i].util;\n+\tfor (i = extras->nr - 1; i >= 0; i--) {\n+\t\tconst char *name = extras->items[i].string;\n+\t\tstruct object *object = extras->items[i].util;\n+\t\tint mark;\n+\n \t\tswitch (object->type) {\n \t\tcase OBJ_TAG:\n \t\t\thandle_tag(name, (struct tag *)object);\n@@ -913,8 +926,24 @@ static void handle_tags_and_duplicates(void)\n \t\t\t\t       name, oid_to_hex(&null_oid));\n \t\t\t\tcontinue;\n \t\t\t}\n-\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n-\t\t\t       get_object_mark(&commit->object));\n+\n+\t\t\tmark = get_object_mark(&commit->object);\n+\t\t\tif (!mark) {\n+\t\t\t\t/*\n+\t\t\t\t * Getting here means we have a commit which\n+\t\t\t\t * was excluded by a negative refspec (e.g.\n+\t\t\t\t * fast-export ^master master).  If the user\n+\t\t\t\t * wants the branch exported but every commit\n+\t\t\t\t * in its history to be deleted, that sounds\n+\t\t\t\t * like a ref deletion to me.\n+\t\t\t\t */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, oid_to_hex(&null_oid));\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\n+\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name, mark\n+\t\t\t       );\n \t\t\tshow_progress();\n \t\t\tbreak;\n \t\t}\n@@ -1102,7 +1131,8 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\t}\n \t}\n \n-\thandle_tags_and_duplicates();\n+\thandle_tags_and_duplicates(&extra_refs);\n+\thandle_tags_and_duplicates(&tag_refs);\n \thandle_deletes();\n \n \tif (export_filename && lastimportid != last_idnum)\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 299120ba70..50c2fceef4 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -544,10 +544,20 @@ test_expect_success 'use refspec' '\n \ttest_cmp expected actual\n '\n \n-test_expect_success 'delete refspec' '\n+test_expect_success 'delete ref because entire history excluded' '\n \tgit branch to-delete &&\n-\tgit fast-export --refspec :refs/heads/to-delete to-delete ^to-delete > actual &&\n-\tcat > expected <<-EOF &&\n+\tgit fast-export to-delete ^to-delete >actual &&\n+\tcat >expected <<-EOF &&\n+\treset refs/heads/to-delete\n+\tfrom 0000000000000000000000000000000000000000\n+\n+\tEOF\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success 'delete refspec' '\n+\tgit fast-export --refspec :refs/heads/to-delete >actual &&\n+\tcat >expected <<-EOF &&\n \treset refs/heads/to-delete\n \tfrom 0000000000000000000000000000000000000000\n \n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363475","messageId":"20181116075956.27047-10-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 09/11] fast-export: add --reference-excluded-parents option","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:54Z","receivedAt":"2018-11-16T08:00:22Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"git filter-branch has a nifty feature allowing you to rewrite, e.g. just\nthe last 8 commits of a linear history\n  git filter-branch $OPTIONS HEAD~8..HEAD\n\nIf you try the same with git fast-export, you instead get a history of\nonly 8 commits, with HEAD~7 being rewritten into a root commit.  There\nare two alternatives:\n\n  1) Don't use the negative revision specification, and when you're\n     filtering the output to make modifications to the last 8 commits,\n     just be careful to not modify any earlier commits somehow.\n\n  2) First run 'git fast-export --export-marks=somefile HEAD~8', then\n     run 'git fast-export --import-marks=somefile HEAD~8..HEAD'.\n\nBoth are more error prone than I'd like (the first for obvious reasons;\nwith the second option I have sometimes accidentally included too many\nrevisions in the first command and then found that the corresponding\nextra revisions were not exported by the second command and thus were\nnot modified as I expected).  Also, both are poor from a performance\nperspective.\n\nAdd a new --reference-excluded-parents option which will cause\nfast-export to refer to commits outside the specified rev-list-args\nrange by their sha1sum.  Such a stream will only be useful in a\nrepository which already contains the necessary commits (much like the\nrestriction imposed when using --no-data).\n\nNote from Peff:\n  I think we might be able to do a little more optimization here. If\n  we're exporting HEAD^..HEAD and there's an object in HEAD^ which is\n  unchanged in HEAD, I think we'd still print it (because it would not\n  be marked SHOWN), but we could omit it (by walking the tree of the\n  boundary commits and marking them shown).  I don't think it's a\n  blocker for what you're doing here, but just a possible future\n  optimization.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt | 17 +++++++++++--\n builtin/fast-export.c             | 42 +++++++++++++++++++++++--------\n t/t9350-fast-export.sh            | 11 ++++++++\n 3 files changed, 58 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex fda55b3284..f65026662a 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -110,6 +110,18 @@ marks the same across runs.\n \tthe shape of the history and stored tree.  See the section on\n \t`ANONYMIZING` below.\n \n+--reference-excluded-parents::\n+\tBy default, running a command such as `git fast-export\n+\tmaster~5..master` will not include the commit master{tilde}5\n+\tand will make master{tilde}4 no longer have master{tilde}5 as\n+\ta parent (though both the old master{tilde}4 and new\n+\tmaster{tilde}4 will have all the same files).  Use\n+\t--reference-excluded-parents to instead have the the stream\n+\trefer to commits in the excluded range of history by their\n+\tsha1sum.  Note that the resulting stream can only be used by a\n+\trepository which already contains the necessary parent\n+\tcommits.\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\n@@ -119,8 +131,9 @@ marks the same across runs.\n \t'git rev-list', that specifies the specific objects and references\n \tto export.  For example, `master~10..master` causes the\n \tcurrent master reference to be exported along with all objects\n-\tadded since its 10th ancestor commit and all files common to\n-\tmaster{tilde}9 and master{tilde}10.\n+\tadded since its 10th ancestor commit and (unless the\n+\t--reference-excluded-parents option is specified) all files\n+\tcommon to master{tilde}9 and master{tilde}10.\n \n EXAMPLES\n --------\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex d71e0333d4..78fc67b03a 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -37,6 +37,7 @@ static int fake_missing_tagger;\n static int use_done_feature;\n static int no_data;\n static int full_tree;\n+static int reference_excluded_commits;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n@@ -597,7 +598,8 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\tmessage += 2;\n \n \tif (commit->parents &&\n-\t    get_object_mark(&commit->parents->item->object) != 0 &&\n+\t    (get_object_mark(&commit->parents->item->object) != 0 ||\n+\t     reference_excluded_commits) &&\n \t    !full_tree) {\n \t\tparse_commit_or_die(commit->parents->item);\n \t\tdiff_tree_oid(get_commit_tree_oid(commit->parents->item),\n@@ -645,13 +647,21 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \tunuse_commit_buffer(commit, commit_buffer);\n \n \tfor (i = 0, p = commit->parents; p; p = p->next) {\n-\t\tint mark = get_object_mark(&p->item->object);\n-\t\tif (!mark)\n+\t\tstruct object *obj = &p->item->object;\n+\t\tint mark = get_object_mark(obj);\n+\n+\t\tif (!mark && !reference_excluded_commits)\n \t\t\tcontinue;\n \t\tif (i == 0)\n-\t\t\tprintf(\"from :%d\\n\", mark);\n+\t\t\tprintf(\"from \");\n+\t\telse\n+\t\t\tprintf(\"merge \");\n+\t\tif (mark)\n+\t\t\tprintf(\":%d\\n\", mark);\n \t\telse\n-\t\t\tprintf(\"merge :%d\\n\", mark);\n+\t\t\tprintf(\"%s\\n\", oid_to_hex(anonymize ?\n+\t\t\t\t\t\t  anonymize_oid(&obj->oid) :\n+\t\t\t\t\t\t  &obj->oid));\n \t\ti++;\n \t}\n \n@@ -932,13 +942,22 @@ static void handle_tags_and_duplicates(struct string_list *extras)\n \t\t\t\t/*\n \t\t\t\t * Getting here means we have a commit which\n \t\t\t\t * was excluded by a negative refspec (e.g.\n-\t\t\t\t * fast-export ^master master).  If the user\n+\t\t\t\t * fast-export ^master master).  If we are\n+\t\t\t\t * referencing excluded commits, set the ref\n+\t\t\t\t * to the exact commit.  Otherwise, the user\n \t\t\t\t * wants the branch exported but every commit\n-\t\t\t\t * in its history to be deleted, that sounds\n-\t\t\t\t * like a ref deletion to me.\n+\t\t\t\t * in its history to be deleted, which basically\n+\t\t\t\t * just means deletion of the ref.\n \t\t\t\t */\n-\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\t       name, oid_to_hex(&null_oid));\n+\t\t\t\tif (!reference_excluded_commits) {\n+\t\t\t\t\t/* delete the ref */\n+\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t\t       name, oid_to_hex(&null_oid));\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\t/* set ref to commit using oid, not mark */\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\", name,\n+\t\t\t\t       oid_to_hex(&commit->object.oid));\n \t\t\t\tcontinue;\n \t\t\t}\n \n@@ -1075,6 +1094,9 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\tOPT_STRING_LIST(0, \"refspec\", &refspecs_list, N_(\"refspec\"),\n \t\t\t     N_(\"Apply refspec to exported refs\")),\n \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n+\t\tOPT_BOOL(0, \"reference-excluded-parents\",\n+\t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by object id\")),\n+\n \t\tOPT_END()\n \t};\n \ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 50c2fceef4..d7d73061d0 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -66,6 +66,17 @@ test_expect_success 'fast-export master~2..master' '\n \n '\n \n+test_expect_success 'fast-export --reference-excluded-parents master~2..master' '\n+\n+\tgit fast-export --reference-excluded-parents master~2..master >actual &&\n+\tgrep commit.refs/heads/master actual >commit-count &&\n+\ttest_line_count = 2 commit-count &&\n+\tsed \"s/master/rewrite/\" actual |\n+\t\t(cd new &&\n+\t\t git fast-import &&\n+\t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n+'\n+\n test_expect_success 'iso-8859-1' '\n \n \tgit config i18n.commitencoding ISO8859-1 &&\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363476","messageId":"20181116075956.27047-11-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 10/11] fast-import: remove unmaintained duplicate documentation","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:55Z","receivedAt":"2018-11-16T08:00:23Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"fast-import.c has started with a comment for nine and a half years\nre-directing the reader to Documentation/git-fast-import.txt for\nmaintained documentation.  Instead of leaving the unmaintained\ndocumentation in place, just excise it.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n fast-import.c | 154 --------------------------------------------------\n 1 file changed, 154 deletions(-)\n\ndiff --git a/fast-import.c b/fast-import.c\nindex 95600c78e0..555d49ad23 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1,157 +1,3 @@\n-/*\n-(See Documentation/git-fast-import.txt for maintained documentation.)\n-Format of STDIN stream:\n-\n-  stream ::= cmd*;\n-\n-  cmd ::= new_blob\n-        | new_commit\n-        | new_tag\n-        | reset_branch\n-        | checkpoint\n-        | progress\n-        ;\n-\n-  new_blob ::= 'blob' lf\n-    mark?\n-    file_content;\n-  file_content ::= data;\n-\n-  new_commit ::= 'commit' sp ref_str lf\n-    mark?\n-    ('author' (sp name)? sp '<' email '>' sp when lf)?\n-    'committer' (sp name)? sp '<' email '>' sp when lf\n-    commit_msg\n-    ('from' sp commit-ish lf)?\n-    ('merge' sp commit-ish lf)*\n-    (file_change | ls)*\n-    lf?;\n-  commit_msg ::= data;\n-\n-  ls ::= 'ls' sp '\"' quoted(path) '\"' lf;\n-\n-  file_change ::= file_clr\n-    | file_del\n-    | file_rnm\n-    | file_cpy\n-    | file_obm\n-    | file_inm;\n-  file_clr ::= 'deleteall' lf;\n-  file_del ::= 'D' sp path_str lf;\n-  file_rnm ::= 'R' sp path_str sp path_str lf;\n-  file_cpy ::= 'C' sp path_str sp path_str lf;\n-  file_obm ::= 'M' sp mode sp (hexsha1 | idnum) sp path_str lf;\n-  file_inm ::= 'M' sp mode sp 'inline' sp path_str lf\n-    data;\n-  note_obm ::= 'N' sp (hexsha1 | idnum) sp commit-ish lf;\n-  note_inm ::= 'N' sp 'inline' sp commit-ish lf\n-    data;\n-\n-  new_tag ::= 'tag' sp tag_str lf\n-    'from' sp commit-ish lf\n-    ('tagger' (sp name)? sp '<' email '>' sp when lf)?\n-    tag_msg;\n-  tag_msg ::= data;\n-\n-  reset_branch ::= 'reset' sp ref_str lf\n-    ('from' sp commit-ish lf)?\n-    lf?;\n-\n-  checkpoint ::= 'checkpoint' lf\n-    lf?;\n-\n-  progress ::= 'progress' sp not_lf* lf\n-    lf?;\n-\n-     # note: the first idnum in a stream should be 1 and subsequent\n-     # idnums should not have gaps between values as this will cause\n-     # the stream parser to reserve space for the gapped values.  An\n-     # idnum can be updated in the future to a new object by issuing\n-     # a new mark directive with the old idnum.\n-     #\n-  mark ::= 'mark' sp idnum lf;\n-  data ::= (delimited_data | exact_data)\n-    lf?;\n-\n-    # note: delim may be any string but must not contain lf.\n-    # data_line may contain any data but must not be exactly\n-    # delim.\n-  delimited_data ::= 'data' sp '<<' delim lf\n-    (data_line lf)*\n-    delim lf;\n-\n-     # note: declen indicates the length of binary_data in bytes.\n-     # declen does not include the lf preceding the binary data.\n-     #\n-  exact_data ::= 'data' sp declen lf\n-    binary_data;\n-\n-     # note: quoted strings are C-style quoting supporting \\c for\n-     # common escapes of 'c' (e..g \\n, \\t, \\\\, \\\") or \\nnn where nnn\n-     # is the signed byte value in octal.  Note that the only\n-     # characters which must actually be escaped to protect the\n-     # stream formatting is: \\, \" and LF.  Otherwise these values\n-     # are UTF8.\n-     #\n-  commit-ish  ::= (ref_str | hexsha1 | sha1exp_str | idnum);\n-  ref_str     ::= ref;\n-  sha1exp_str ::= sha1exp;\n-  tag_str     ::= tag;\n-  path_str    ::= path    | '\"' quoted(path)    '\"' ;\n-  mode        ::= '100644' | '644'\n-                | '100755' | '755'\n-                | '120000'\n-                ;\n-\n-  declen ::= # unsigned 32 bit value, ascii base10 notation;\n-  bigint ::= # unsigned integer value, ascii base10 notation;\n-  binary_data ::= # file content, not interpreted;\n-\n-  when         ::= raw_when | rfc2822_when;\n-  raw_when     ::= ts sp tz;\n-  rfc2822_when ::= # Valid RFC 2822 date and time;\n-\n-  sp ::= # ASCII space character;\n-  lf ::= # ASCII newline (LF) character;\n-\n-     # note: a colon (':') must precede the numerical value assigned to\n-     # an idnum.  This is to distinguish it from a ref or tag name as\n-     # GIT does not permit ':' in ref or tag strings.\n-     #\n-  idnum   ::= ':' bigint;\n-  path    ::= # GIT style file path, e.g. \"a/b/c\";\n-  ref     ::= # GIT ref name, e.g. \"refs/heads/MOZ_GECKO_EXPERIMENT\";\n-  tag     ::= # GIT tag name, e.g. \"FIREFOX_1_5\";\n-  sha1exp ::= # Any valid GIT SHA1 expression;\n-  hexsha1 ::= # SHA1 in hexadecimal format;\n-\n-     # note: name and email are UTF8 strings, however name must not\n-     # contain '<' or lf and email must not contain any of the\n-     # following: '<', '>', lf.\n-     #\n-  name  ::= # valid GIT author/committer name;\n-  email ::= # valid GIT author/committer email;\n-  ts    ::= # time since the epoch in seconds, ascii base10 notation;\n-  tz    ::= # GIT style timezone;\n-\n-     # note: comments, get-mark, ls-tree, and cat-blob requests may\n-     # appear anywhere in the input, except within a data command. Any\n-     # form of the data command always escapes the related input from\n-     # comment processing.\n-     #\n-     # In case it is not clear, the '#' that starts the comment\n-     # must be the first character on that line (an lf\n-     # preceded it).\n-     #\n-\n-  get_mark ::= 'get-mark' sp idnum lf;\n-  cat_blob ::= 'cat-blob' sp (hexsha1 | idnum) lf;\n-  ls_tree  ::= 'ls' sp (hexsha1 | idnum) sp path_str lf;\n-\n-  comment ::= '#' not_lf* lf;\n-  not_lf  ::= # Any byte that is not ASCII newline (LF);\n-*/\n-\n #include \"builtin.h\"\n #include \"cache.h\"\n #include \"repository.h\"\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363477","messageId":"20181116075956.27047-7-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 06/11] fast-export: move commit rewriting logic into a function for reuse","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:51Z","receivedAt":"2018-11-16T08:00:23Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Logic to replace a filtered commit with an unfiltered ancestor is useful\nelsewhere; put it into a function we can call.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 37 ++++++++++++++++++++++---------------\n 1 file changed, 22 insertions(+), 15 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 7d50f5414e..43e98a38a8 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -187,6 +187,22 @@ static int get_object_mark(struct object *object)\n \treturn ptr_to_mark(decoration);\n }\n \n+static struct commit *rewrite_commit(struct commit *p)\n+{\n+\tfor (;;) {\n+\t\tif (p->parents && p->parents->next)\n+\t\t\tbreak;\n+\t\tif (p->object.flags & UNINTERESTING)\n+\t\t\tbreak;\n+\t\tif (!(p->object.flags & TREESAME))\n+\t\t\tbreak;\n+\t\tif (!p->parents)\n+\t\t\treturn NULL;\n+\t\tp = p->parents->item;\n+\t}\n+\treturn p;\n+}\n+\n static void show_progress(void)\n {\n \tstatic int counter = 0;\n@@ -767,21 +783,12 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t    oid_to_hex(&tag->object.oid),\n \t\t\t\t    type_name(tagged->type));\n \t\t\t}\n-\t\t\tp = (struct commit *)tagged;\n-\t\t\tfor (;;) {\n-\t\t\t\tif (p->parents && p->parents->next)\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (p->object.flags & UNINTERESTING)\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (!(p->object.flags & TREESAME))\n-\t\t\t\t\tbreak;\n-\t\t\t\tif (!p->parents) {\n-\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n-\t\t\t\t\t       name, oid_to_hex(&null_oid));\n-\t\t\t\t\tfree(buf);\n-\t\t\t\t\treturn;\n-\t\t\t\t}\n-\t\t\t\tp = p->parents->item;\n+\t\t\tp = rewrite_commit((struct commit *)tagged);\n+\t\t\tif (!p) {\n+\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t       name, oid_to_hex(&null_oid));\n+\t\t\t\tfree(buf);\n+\t\t\t\treturn;\n \t\t\t}\n \t\t\ttagged_mark = get_object_mark(&p->object);\n \t\t}\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363478","messageId":"20181116075956.27047-1-newren@gmail.com","threadId":"49406","inReplyTo":"20181114002600.29233-1-newren@gmail.com","subject":"[PATCH v3 00/11] fast export and import fixes and features","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:45Z","receivedAt":"2018-11-16T08:00:25Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"This is a series of small fixes and features for fast-export and\nfast-import, mostly on the fast-export side.\n\nChanges since v2 (full range-diff below):\n  * Dropped the final patch; going to try to use Peff's suggestion of\n    rev-list and diff-tree to get what I need instead\n  * Inserted a new patch at the beginning to convert pre-existing sha1\n    stuff to oid (rename sha1_to_hex() -> oid_to_hex(), rename\n    anonymize_sha1() to anonymize_oid(), etc.)\n  * Modified other patches in the series to add calls to oid_to_hex() rather\n    than sha1_to_hex()\n\nElijah Newren (11):\n  fast-export: convert sha1 to oid\n  git-fast-import.txt: fix documentation for --quiet option\n  git-fast-export.txt: clarify misleading documentation about rev-list\n    args\n  fast-export: use value from correct enum\n  fast-export: avoid dying when filtering by paths and old tags exist\n  fast-export: move commit rewriting logic into a function for reuse\n  fast-export: when using paths, avoid corrupt stream with non-existent\n    mark\n  fast-export: ensure we export requested refs\n  fast-export: add --reference-excluded-parents option\n  fast-import: remove unmaintained duplicate documentation\n  fast-export: add a --show-original-ids option to show original names\n\n Documentation/git-fast-export.txt |  23 +++-\n Documentation/git-fast-import.txt |  23 +++-\n builtin/fast-export.c             | 190 +++++++++++++++++++++---------\n fast-import.c                     | 166 ++------------------------\n t/t9350-fast-export.sh            |  80 ++++++++++++-\n 5 files changed, 268 insertions(+), 214 deletions(-)\n\n -:  ---------- >  1:  4c3370c85f fast-export: convert sha1 to oid\n 1:  8870fb1340 =  2:  6ffa30e3c7 git-fast-import.txt: fix documentation for --quiet option\n 2:  16d1c3e22d =  3:  1e278f009a git-fast-export.txt: clarify misleading documentation about rev-list args\n 3:  e19f6b36f9 =  4:  9d7b2aef49 fast-export: use value from correct enum\n 4:  2b305561d5 !  5:  b65a591d4d fast-export: avoid dying when filtering by paths and old tags exist\n    @@ -29,7 +29,7 @@\n     -\t\t\t\t\t     oid_to_hex(&tag->object.oid));\n     +\t\t\t\tif (!p->parents) {\n     +\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    -+\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n    ++\t\t\t\t\t       name, oid_to_hex(&null_oid));\n     +\t\t\t\t\tfree(buf);\n     +\t\t\t\t\treturn;\n     +\t\t\t\t}\n 5:  607b1dc2b2 !  6:  dde52c9cb6 fast-export: move commit rewriting logic into a function for reuse\n    @@ -47,7 +47,7 @@\n     -\t\t\t\t\tbreak;\n     -\t\t\t\tif (!p->parents) {\n     -\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    --\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n    +-\t\t\t\t\t       name, oid_to_hex(&null_oid));\n     -\t\t\t\t\tfree(buf);\n     -\t\t\t\t\treturn;\n     -\t\t\t\t}\n    @@ -55,7 +55,7 @@\n     +\t\t\tp = rewrite_commit((struct commit *)tagged);\n     +\t\t\tif (!p) {\n     +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    -+\t\t\t\t       name, sha1_to_hex(null_sha1));\n    ++\t\t\t\t       name, oid_to_hex(&null_oid));\n     +\t\t\t\tfree(buf);\n     +\t\t\t\treturn;\n      \t\t\t}\n 6:  ec1862e858 !  7:  d9b2e326f0 fast-export: when using paths, avoid corrupt stream with non-existent mark\n    @@ -35,7 +35,7 @@\n     +\t\t\t\t * it.\n     +\t\t\t\t */\n     +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    -+\t\t\t\t       name, sha1_to_hex(null_sha1));\n    ++\t\t\t\t       name, oid_to_hex(&null_oid));\n     +\t\t\t\tcontinue;\n     +\t\t\t}\n      \t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n 7:  9da26e3ccb !  8:  9ddb155a70 fast-export: ensure we export requested refs\n    @@ -97,7 +97,7 @@\n      \t\tcase OBJ_TAG:\n      \t\t\thandle_tag(name, (struct tag *)object);\n     @@\n    - \t\t\t\t       name, sha1_to_hex(null_sha1));\n    + \t\t\t\t       name, oid_to_hex(&null_oid));\n      \t\t\t\tcontinue;\n      \t\t\t}\n     -\t\t\tprintf(\"reset %s\\nfrom :%d\\n\\n\", name,\n    @@ -114,7 +114,7 @@\n     +\t\t\t\t * like a ref deletion to me.\n     +\t\t\t\t */\n     +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    -+\t\t\t\t       name, sha1_to_hex(null_sha1));\n    ++\t\t\t\t       name, oid_to_hex(&null_oid));\n     +\t\t\t\tcontinue;\n     +\t\t\t}\n     +\n 8:  7e5fe2f02e !  9:  595d2e5d30 fast-export: add --reference-excluded-parents option\n    @@ -117,9 +117,9 @@\n     +\t\t\tprintf(\":%d\\n\", mark);\n      \t\telse\n     -\t\t\tprintf(\"merge :%d\\n\", mark);\n    -+\t\t\tprintf(\"%s\\n\", sha1_to_hex(anonymize ?\n    -+\t\t\t\t\t\t   anonymize_sha1(&obj->oid) :\n    -+\t\t\t\t\t\t   obj->oid.hash));\n    ++\t\t\tprintf(\"%s\\n\", oid_to_hex(anonymize ?\n    ++\t\t\t\t\t\t  anonymize_oid(&obj->oid) :\n    ++\t\t\t\t\t\t  &obj->oid));\n      \t\ti++;\n      \t}\n      \n    @@ -138,16 +138,16 @@\n     +\t\t\t\t * just means deletion of the ref.\n      \t\t\t\t */\n     -\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    --\t\t\t\t       name, sha1_to_hex(null_sha1));\n    +-\t\t\t\t       name, oid_to_hex(&null_oid));\n     +\t\t\t\tif (!reference_excluded_commits) {\n     +\t\t\t\t\t/* delete the ref */\n     +\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n    -+\t\t\t\t\t       name, sha1_to_hex(null_sha1));\n    ++\t\t\t\t\t       name, oid_to_hex(&null_oid));\n     +\t\t\t\t\tcontinue;\n     +\t\t\t\t}\n     +\t\t\t\t/* set ref to commit using oid, not mark */\n     +\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\", name,\n    -+\t\t\t\t       sha1_to_hex(commit->object.oid.hash));\n    ++\t\t\t\t       oid_to_hex(&commit->object.oid));\n      \t\t\t\tcontinue;\n      \t\t\t}\n      \n    @@ -156,7 +156,7 @@\n      \t\t\t     N_(\"Apply refspec to exported refs\")),\n      \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n     +\t\tOPT_BOOL(0, \"reference-excluded-parents\",\n    -+\t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n    ++\t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by object id\")),\n     +\n      \t\tOPT_END()\n      \t};\n 9:  14306a8436 = 10:  2686246a89 fast-import: remove unmaintained duplicate documentation\n10:  72487a61e4 ! 11:  b78d548e7d fast-export: add a --show-original-ids option to show original names\n    @@ -141,9 +141,9 @@\n     @@\n      \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n      \t\tOPT_BOOL(0, \"reference-excluded-parents\",\n    - \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by sha1sum\")),\n    + \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by object id\")),\n     +\t\tOPT_BOOL(0, \"show-original-ids\", &show_original_ids,\n    -+\t\t\t    N_(\"Show original sha1sums of blobs/commits\")),\n    ++\t\t\t    N_(\"Show original object ids of blobs/commits\")),\n      \n      \t\tOPT_END()\n      \t};\n\n-- \n2.19.1.1063.g1796373474.dirty\n"},{"id":"363479","messageId":"20181116075956.27047-12-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 11/11] fast-export: add a --show-original-ids option to show original names","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:56Z","receivedAt":"2018-11-16T08:00:26Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Knowing the original names (hashes) of commits can sometimes enable\npost-filtering that would otherwise be difficult or impossible.  In\nparticular, the desire to rewrite commit messages which refer to other\nprior commits (on top of whatever other filtering is being done) is\nvery difficult without knowing the original names of each commit.\n\nIn addition, knowing the original names (hashes) of blobs can allow\nfiltering by blob-id without requiring re-hashing the content of the\nblob, and is thus useful as a small optimization.\n\nOnce we add original ids for both commits and blobs, we may as well\nadd them for tags too for completeness.  Perhaps someone will have a\nuse for them.\n\nThis commit teaches a new --show-original-ids option to fast-export\nwhich will make it add a 'original-oid <hash>' line to blob, commits,\nand tags.  It also teaches fast-import to parse (and ignore) such\nlines.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n Documentation/git-fast-export.txt |  7 +++++++\n Documentation/git-fast-import.txt | 16 ++++++++++++++++\n builtin/fast-export.c             | 20 +++++++++++++++-----\n fast-import.c                     | 12 ++++++++++++\n t/t9350-fast-export.sh            | 17 +++++++++++++++++\n 5 files changed, 67 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/git-fast-export.txt b/Documentation/git-fast-export.txt\nindex f65026662a..64c01ba918 100644\n--- a/Documentation/git-fast-export.txt\n+++ b/Documentation/git-fast-export.txt\n@@ -122,6 +122,13 @@ marks the same across runs.\n \trepository which already contains the necessary parent\n \tcommits.\n \n+--show-original-ids::\n+\tAdd an extra directive to the output for commits and blobs,\n+\t`original-oid <SHA1SUM>`.  While such directives will likely be\n+\tignored by importers such as git-fast-import, it may be useful\n+\tfor intermediary filters (e.g. for rewriting commit messages\n+\twhich refer to older commits, or for stripping blobs by id).\n+\n --refspec::\n \tApply the specified refspec to each ref exported. Multiple of them can\n \tbe specified.\ndiff --git a/Documentation/git-fast-import.txt b/Documentation/git-fast-import.txt\nindex 7ab97745a6..43ab3b1637 100644\n--- a/Documentation/git-fast-import.txt\n+++ b/Documentation/git-fast-import.txt\n@@ -385,6 +385,7 @@ change to the project.\n ....\n \t'commit' SP <ref> LF\n \tmark?\n+\toriginal-oid?\n \t('author' (SP <name>)? SP LT <email> GT SP <when> LF)?\n \t'committer' (SP <name>)? SP LT <email> GT SP <when> LF\n \tdata\n@@ -741,6 +742,19 @@ New marks are created automatically.  Existing marks can be moved\n to another object simply by reusing the same `<idnum>` in another\n `mark` command.\n \n+`original-oid`\n+~~~~~~~~~~~~~~\n+Provides the name of the object in the original source control system.\n+fast-import will simply ignore this directive, but filter processes\n+which operate on and modify the stream before feeding to fast-import\n+may have uses for this information\n+\n+....\n+\t'original-oid' SP <object-identifier> LF\n+....\n+\n+where `<object-identifer>` is any string not containing LF.\n+\n `tag`\n ~~~~~\n Creates an annotated tag referring to a specific commit.  To create\n@@ -749,6 +763,7 @@ lightweight (non-annotated) tags see the `reset` command below.\n ....\n \t'tag' SP <name> LF\n \t'from' SP <commit-ish> LF\n+\toriginal-oid?\n \t'tagger' (SP <name>)? SP LT <email> GT SP <when> LF\n \tdata\n ....\n@@ -823,6 +838,7 @@ assigned mark.\n ....\n \t'blob' LF\n \tmark?\n+\toriginal-oid?\n \tdata\n ....\n \ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex 78fc67b03a..36c2575de5 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -38,6 +38,7 @@ static int use_done_feature;\n static int no_data;\n static int full_tree;\n static int reference_excluded_commits;\n+static int show_original_ids;\n static struct string_list extra_refs = STRING_LIST_INIT_NODUP;\n static struct string_list tag_refs = STRING_LIST_INIT_NODUP;\n static struct refspec refspecs = REFSPEC_INIT_FETCH;\n@@ -271,7 +272,10 @@ static void export_blob(const struct object_id *oid)\n \n \tmark_next_object(object);\n \n-\tprintf(\"blob\\nmark :%\"PRIu32\"\\ndata %lu\\n\", last_idnum, size);\n+\tprintf(\"blob\\nmark :%\"PRIu32\"\\n\", last_idnum);\n+\tif (show_original_ids)\n+\t\tprintf(\"original-oid %s\\n\", oid_to_hex(oid));\n+\tprintf(\"data %lu\\n\", size);\n \tif (size && fwrite(buf, size, 1, stdout) != 1)\n \t\tdie_errno(\"could not write blob '%s'\", oid_to_hex(oid));\n \tprintf(\"\\n\");\n@@ -635,8 +639,10 @@ static void handle_commit(struct commit *commit, struct rev_info *rev,\n \t\treencoded = reencode_string(message, \"UTF-8\", encoding);\n \tif (!commit->parents)\n \t\tprintf(\"reset %s\\n\", refname);\n-\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n%.*s\\n%.*s\\ndata %u\\n%s\",\n-\t       refname, last_idnum,\n+\tprintf(\"commit %s\\nmark :%\"PRIu32\"\\n\", refname, last_idnum);\n+\tif (show_original_ids)\n+\t\tprintf(\"original-oid %s\\n\", oid_to_hex(&commit->object.oid));\n+\tprintf(\"%.*s\\n%.*s\\ndata %u\\n%s\",\n \t       (int)(author_end - author), author,\n \t       (int)(committer_end - committer), committer,\n \t       (unsigned)(reencoded\n@@ -814,8 +820,10 @@ static void handle_tag(const char *name, struct tag *tag)\n \n \tif (starts_with(name, \"refs/tags/\"))\n \t\tname += 10;\n-\tprintf(\"tag %s\\nfrom :%d\\n%.*s%sdata %d\\n%.*s\\n\",\n-\t       name, tagged_mark,\n+\tprintf(\"tag %s\\nfrom :%d\\n\", name, tagged_mark);\n+\tif (show_original_ids)\n+\t\tprintf(\"original-oid %s\\n\", oid_to_hex(&tag->object.oid));\n+\tprintf(\"%.*s%sdata %d\\n%.*s\\n\",\n \t       (int)(tagger_end - tagger), tagger,\n \t       tagger == tagger_end ? \"\" : \"\\n\",\n \t       (int)message_size, (int)message_size, message ? message : \"\");\n@@ -1096,6 +1104,8 @@ int cmd_fast_export(int argc, const char **argv, const char *prefix)\n \t\tOPT_BOOL(0, \"anonymize\", &anonymize, N_(\"anonymize output\")),\n \t\tOPT_BOOL(0, \"reference-excluded-parents\",\n \t\t\t &reference_excluded_commits, N_(\"Reference parents which are not in fast-export stream by object id\")),\n+\t\tOPT_BOOL(0, \"show-original-ids\", &show_original_ids,\n+\t\t\t    N_(\"Show original object ids of blobs/commits\")),\n \n \t\tOPT_END()\n \t};\ndiff --git a/fast-import.c b/fast-import.c\nindex 555d49ad23..71b6cba00f 100644\n--- a/fast-import.c\n+++ b/fast-import.c\n@@ -1814,6 +1814,13 @@ static void parse_mark(void)\n \t\tnext_mark = 0;\n }\n \n+static void parse_original_identifier(void)\n+{\n+\tconst char *v;\n+\tif (skip_prefix(command_buf.buf, \"original-oid \", &v))\n+\t\tread_next_command();\n+}\n+\n static int parse_data(struct strbuf *sb, uintmax_t limit, uintmax_t *len_res)\n {\n \tconst char *data;\n@@ -1956,6 +1963,7 @@ static void parse_new_blob(void)\n {\n \tread_next_command();\n \tparse_mark();\n+\tparse_original_identifier();\n \tparse_and_store_blob(&last_blob, NULL, next_mark);\n }\n \n@@ -2579,6 +2587,7 @@ static void parse_new_commit(const char *arg)\n \n \tread_next_command();\n \tparse_mark();\n+\tparse_original_identifier();\n \tif (skip_prefix(command_buf.buf, \"author \", &v)) {\n \t\tauthor = parse_ident(v);\n \t\tread_next_command();\n@@ -2711,6 +2720,9 @@ static void parse_new_tag(const char *arg)\n \t\tdie(\"Invalid ref name or SHA1 expression: %s\", from);\n \tread_next_command();\n \n+\t/* original-oid ... */\n+\tparse_original_identifier();\n+\n \t/* tagger ... */\n \tif (skip_prefix(command_buf.buf, \"tagger \", &v)) {\n \t\ttagger = parse_ident(v);\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex d7d73061d0..5690fe2810 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -77,6 +77,23 @@ test_expect_success 'fast-export --reference-excluded-parents master~2..master'\n \t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n '\n \n+test_expect_success 'fast-export --show-original-ids' '\n+\n+\tgit fast-export --show-original-ids master >output &&\n+\tgrep ^original-oid output| sed -e s/^original-oid.// | sort >actual &&\n+\tgit rev-list --objects master muss >objects-and-names &&\n+\tawk \"{print \\$1}\" objects-and-names | sort >commits-trees-blobs &&\n+\tcomm -23 actual commits-trees-blobs >unfound &&\n+\ttest_must_be_empty unfound\n+'\n+\n+test_expect_success 'fast-export --show-original-ids | git fast-import' '\n+\n+\tgit fast-export --show-original-ids master muss | git fast-import --quiet &&\n+\ttest $MASTER = $(git rev-parse --verify refs/heads/master) &&\n+\ttest $MUSS = $(git rev-parse --verify refs/tags/muss)\n+'\n+\n test_expect_success 'iso-8859-1' '\n \n \tgit config i18n.commitencoding ISO8859-1 &&\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363481","messageId":"20181116075956.27047-6-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 05/11] fast-export: avoid dying when filtering by paths and old tags exist","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:50Z","receivedAt":"2018-11-16T08:00:28Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"If --tag-of-filtered-object=rewrite is specified along with a set of\npaths to limit what is exported, then any tags pointing to old commits\nthat do not contain any of those specified paths cause problems.  Since\nthe old tagged commit is not exported, fast-export attempts to rewrite\nsuch tags to an ancestor commit which was exported.  If no such commit\nexists, then fast-export currently die()s.  Five years after the tag\nrewriting logic was added to fast-export (see commit 2d8ad4691921,\n\"fast-export: Add a --tag-of-filtered-object  option for newly dangling\ntags\", 2009-06-25), fast-import gained the ability to delete refs (see\ncommit 4ee1b225b99f, \"fast-import: add support to delete refs\",\n2014-04-20), so now we do have a valid option to rewrite the tag to.\nDelete these tags instead of dying.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c  |  9 ++++++---\n t/t9350-fast-export.sh | 16 ++++++++++++++++\n 2 files changed, 22 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex e2be35f41e..7d50f5414e 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -775,9 +775,12 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t\tbreak;\n \t\t\t\tif (!(p->object.flags & TREESAME))\n \t\t\t\t\tbreak;\n-\t\t\t\tif (!p->parents)\n-\t\t\t\t\tdie(\"can't find replacement commit for tag %s\",\n-\t\t\t\t\t     oid_to_hex(&tag->object.oid));\n+\t\t\t\tif (!p->parents) {\n+\t\t\t\t\tprintf(\"reset %s\\nfrom %s\\n\\n\",\n+\t\t\t\t\t       name, oid_to_hex(&null_oid));\n+\t\t\t\t\tfree(buf);\n+\t\t\t\t\treturn;\n+\t\t\t\t}\n \t\t\t\tp = p->parents->item;\n \t\t\t}\n \t\t\ttagged_mark = get_object_mark(&p->object);\ndiff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\nindex 6a392e87bc..3400ebeb51 100755\n--- a/t/t9350-fast-export.sh\n+++ b/t/t9350-fast-export.sh\n@@ -325,6 +325,22 @@ test_expect_success 'rewriting tag of filtered out object' '\n )\n '\n \n+test_expect_success 'rewrite tag predating pathspecs to nothing' '\n+\ttest_create_repo rewrite_tag_predating_pathspecs &&\n+\t(\n+\t\tcd rewrite_tag_predating_pathspecs &&\n+\n+\t\ttest_commit initial &&\n+\n+\t\tgit tag -a -m \"Some old tag\" v0.0.0.0.0.0.1 &&\n+\n+\t\ttest_commit bar &&\n+\n+\t\tgit fast-export --tag-of-filtered-object=rewrite --all -- bar.t >output &&\n+\t\tgrep from.$ZERO_OID output\n+\t)\n+'\n+\n cat > limit-by-paths/expected << EOF\n blob\n mark :1\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363480","messageId":"20181116075956.27047-5-newren@gmail.com","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"[PATCH v3 04/11] fast-export: use value from correct enum","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-11-16T07:59:49Z","receivedAt":"2018-11-16T08:00:29Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"ABORT and ERROR happen to have the same value, but come from differnt\nenums.  Use the one from the correct enum, and while at it, rename the\nvalues to avoid such problems.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n builtin/fast-export.c | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/fast-export.c b/builtin/fast-export.c\nindex f5166ac71e..e2be35f41e 100644\n--- a/builtin/fast-export.c\n+++ b/builtin/fast-export.c\n@@ -31,8 +31,8 @@ static const char *fast_export_usage[] = {\n };\n \n static int progress;\n-static enum { ABORT, VERBATIM, WARN, WARN_STRIP, STRIP } signed_tag_mode = ABORT;\n-static enum { ERROR, DROP, REWRITE } tag_of_filtered_mode = ERROR;\n+static enum { SIGNED_TAG_ABORT, VERBATIM, WARN, WARN_STRIP, STRIP } signed_tag_mode = SIGNED_TAG_ABORT;\n+static enum { TAG_FILTERING_ABORT, DROP, REWRITE } tag_of_filtered_mode = TAG_FILTERING_ABORT;\n static int fake_missing_tagger;\n static int use_done_feature;\n static int no_data;\n@@ -46,7 +46,7 @@ static int parse_opt_signed_tag_mode(const struct option *opt,\n \t\t\t\t     const char *arg, int unset)\n {\n \tif (unset || !strcmp(arg, \"abort\"))\n-\t\tsigned_tag_mode = ABORT;\n+\t\tsigned_tag_mode = SIGNED_TAG_ABORT;\n \telse if (!strcmp(arg, \"verbatim\") || !strcmp(arg, \"ignore\"))\n \t\tsigned_tag_mode = VERBATIM;\n \telse if (!strcmp(arg, \"warn\"))\n@@ -64,7 +64,7 @@ static int parse_opt_tag_of_filtered_mode(const struct option *opt,\n \t\t\t\t\t  const char *arg, int unset)\n {\n \tif (unset || !strcmp(arg, \"abort\"))\n-\t\ttag_of_filtered_mode = ERROR;\n+\t\ttag_of_filtered_mode = TAG_FILTERING_ABORT;\n \telse if (!strcmp(arg, \"drop\"))\n \t\ttag_of_filtered_mode = DROP;\n \telse if (!strcmp(arg, \"rewrite\"))\n@@ -728,7 +728,7 @@ static void handle_tag(const char *name, struct tag *tag)\n \t\t\t\t\t       \"\\n-----BEGIN PGP SIGNATURE-----\\n\");\n \t\tif (signature)\n \t\t\tswitch(signed_tag_mode) {\n-\t\t\tcase ABORT:\n+\t\t\tcase SIGNED_TAG_ABORT:\n \t\t\t\tdie(\"encountered signed tag %s; use \"\n \t\t\t\t    \"--signed-tags=<mode> to handle it\",\n \t\t\t\t    oid_to_hex(&tag->object.oid));\n@@ -753,7 +753,7 @@ static void handle_tag(const char *name, struct tag *tag)\n \ttagged_mark = get_object_mark(tagged);\n \tif (!tagged_mark) {\n \t\tswitch(tag_of_filtered_mode) {\n-\t\tcase ABORT:\n+\t\tcase TAG_FILTERING_ABORT:\n \t\t\tdie(\"tag %s tags unexported object; use \"\n \t\t\t    \"--tag-of-filtered-object=<mode> to handle it\",\n \t\t\t    oid_to_hex(&tag->object.oid));\n-- \n2.19.1.1063.g1796373474.dirty\n\n"},{"id":"363487","messageId":"20181116085016.GA20828@sigill.intra.peff.net","threadId":"49406","inReplyTo":"20181116075956.27047-1-newren@gmail.com","subject":"Re: [PATCH v3 00/11] fast export and import fixes and features","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-16T08:50:17Z","receivedAt":"2018-11-16T08:50:20Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Nov 15, 2018 at 11:59:45PM -0800, Elijah Newren wrote:\n\n> This is a series of small fixes and features for fast-export and\n> fast-import, mostly on the fast-export side.\n> \n> Changes since v2 (full range-diff below):\n>   * Dropped the final patch; going to try to use Peff's suggestion of\n>     rev-list and diff-tree to get what I need instead\n>   * Inserted a new patch at the beginning to convert pre-existing sha1\n>     stuff to oid (rename sha1_to_hex() -> oid_to_hex(), rename\n>     anonymize_sha1() to anonymize_oid(), etc.)\n>   * Modified other patches in the series to add calls to oid_to_hex() rather\n>     than sha1_to_hex()\n\nThanks, these changes all look good to me. I have no more nits to pick.\n:)\n\n-Peff\n"},{"id":"363509","messageId":"20181116122940.GL30222@szeder.dev","threadId":"49406","inReplyTo":"20181116075956.27047-12-newren@gmail.com","subject":"Re: [PATCH v3 11/11] fast-export: add a --show-original-ids option to show original names","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-11-16T12:29:40Z","receivedAt":"2018-11-16T12:29:47Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Nov 15, 2018 at 11:59:56PM -0800, Elijah Newren wrote:\n\n> diff --git a/t/t9350-fast-export.sh b/t/t9350-fast-export.sh\n> index d7d73061d0..5690fe2810 100755\n> --- a/t/t9350-fast-export.sh\n> +++ b/t/t9350-fast-export.sh\n> @@ -77,6 +77,23 @@ test_expect_success 'fast-export --reference-excluded-parents master~2..master'\n>  \t\t test $MASTER = $(git rev-parse --verify refs/heads/rewrite))\n>  '\n>  \n> +test_expect_success 'fast-export --show-original-ids' '\n> +\n> +\tgit fast-export --show-original-ids master >output &&\n> +\tgrep ^original-oid output| sed -e s/^original-oid.// | sort >actual &&\n\nNit: 'sed' can do what this 'grep' does:\n\n  sed -n -e s/^original-oid.//p output | sort >actual &&\n\nthus sparing a process.\n\n> +\tgit rev-list --objects master muss >objects-and-names &&\n> +\tawk \"{print \\$1}\" objects-and-names | sort >commits-trees-blobs &&\n> +\tcomm -23 actual commits-trees-blobs >unfound &&\n> +\ttest_must_be_empty unfound\n> +'\n> +\n> +test_expect_success 'fast-export --show-original-ids | git fast-import' '\n> +\n> +\tgit fast-export --show-original-ids master muss | git fast-import --quiet &&\n> +\ttest $MASTER = $(git rev-parse --verify refs/heads/master) &&\n> +\ttest $MUSS = $(git rev-parse --verify refs/tags/muss)\n> +'\n> +\n>  test_expect_success 'iso-8859-1' '\n>  \n>  \tgit config i18n.commitencoding ISO8859-1 &&\n> -- \n> 2.19.1.1063.g1796373474.dirty\n> \n"}]}