{"thread":{"id":"31717","subject":"What's cooking in git.git (Oct 2012, #01; Tue, 2)","startedAt":"2012-10-02T23:20:22Z","lastAt":"2012-10-30T12:15:18Z","messageCount":28,"participants":["Junio C Hamano","Nguyen Thai Ngoc Duy","Nguyễn Thái Ngọc Duy","Michael Haggerty","David Michael Barr","Johannes Sixt","Andreas Schwab","Matthieu Moy","Florian Achleitner"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"200374","messageId":"7vmx045umh.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":null,"subject":"What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-02T23:20:22Z","receivedAt":"2012-10-02T23:20:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Here are the topics that have been cooking.  Commits prefixed with\n'-' are only in 'pu' (proposed updates) while commits prefixed with\n'+' are in 'next'.\n\nThe tip of 'master' is a bit past v1.8.0-rc0; we are more or less\nfeature complete, except for a few topics marked as \"Will merge to\n'master' in the list below.\n\nI do not expect the topics in the stalled category will be ready for\nthe upcoming release; they may appear in the release after v1.8.0 if\nthey are rerolled.  Also some of the topics that have been cooking\nin 'next' have to be deferred to the next cycle.\n\nI'm planning to keep this cycle reasonably short and aim for tagging\nthe result as 1.8.0 at the end of 9th week, on October 21st, after\nwhich I'd disappear for a few weeks.  http://tinyurl.com/gitCal is\nwhere you can always find my rough tagging schedule at.\n\nYou can find the changes described here in the integration branches of the\nrepositories listed at\n\n    http://git-blame.blogspot.com/p/git-public-repositories.html\n\n--------------------------------------------------\n[New Topics]\n\n* bw/config-lift-variable-name-length-limit (2012-10-01) 1 commit\n - Remove the hard coded length limit on variable names in config files\n\n The configuration parser had an unnecessary hardcoded limit on\n variable names that was not checked consistently. Lift the limit.\n\n Will merge to 'next'.\n\n\n* jc/maint-t1450-fsck-order-fix (2012-10-02) 1 commit\n - t1450: the order the objects are checked is undefined\n\n The fsck test assumed too much on what kind of error it will\n detect. The only important thing is the inconsistency is detected\n as an error.\n\n Will merge to 'next'.\n\n--------------------------------------------------\n[Graduated to \"master\"]\n\n* da/mergetool-custom (2012-09-25) 1 commit\n  (merged to 'next' on 2012-09-27 at f910851)\n + mergetool--lib: Allow custom commands to override built-ins\n\n The actual external command to run for mergetool backend can be\n specified with difftool/mergetool.$name.cmd configuration\n variables, but this mechanism was ignored for the backends we\n natively support.\n\n\n* ep/malloc-check-perturb (2012-09-26) 1 commit\n  (merged to 'next' on 2012-09-27 at f115b8b)\n + MALLOC_CHECK: enable it, unless disabled explicitly\n\n Fixes a brown-paper bag bug.\n\n\n* jc/blame-follows-renames (2012-09-21) 1 commit\n  (merged to 'next' on 2012-09-27 at 12634c1)\n + git blame: document that it always follows origin across whole-file renames\n (this branch is used by jc/blame-no-follow.)\n\n Clarify the \"blame\" documentation to tell the users that there is\n no need to ask for \"--follow\".\n\n\n* jk/completion-tests (2012-09-27) 2 commits\n  (merged to 'next' on 2012-09-27 at 5cd6968)\n + t9902: add completion tests for \"odd\" filenames\n + t9902: add a few basic completion tests\n\n\n* jk/receive-pack-unpack-error-to-pusher (2012-09-21) 3 commits\n  (merged to 'next' on 2012-09-27 at 90f0c6f)\n + receive-pack: drop \"n/a\" on unpacker errors\n + receive-pack: send pack-processing stderr over sideband\n + receive-pack: redirect unpack-objects stdout to /dev/null\n\n Send errors from \"unpack-objects\" and \"index-pack\" back to the \"git\n push\" over the git and smart-http protocols, just like it is done\n for a push over the ssh protocol.\n\n\n* os/commit-submodule-ignore (2012-09-24) 1 commit\n  (merged to 'next' on 2012-09-27 at 9cd2cfd)\n + commit: pay attention to submodule.$name.ignore in .gitmodules\n\n \"git status\" honored the ignore=dirty settings in .gitmodules but\n \"git commit\" didn't.\n\n\n* rt/maint-clone-single (2012-09-20) 1 commit\n  (merged to 'next' on 2012-09-27 at a47d54d)\n + clone --single: limit the fetch refspec to fetched branch\n\n Running \"git fetch\" in a repository made with \"git clone --single\"\n slurps all the branches, defeating the point of \"--single\".\n\n--------------------------------------------------\n[Stalled]\n\n* rc/maint-complete-git-p4 (2012-09-24) 1 commit\n  (merged to 'next' on 2012-09-25 at 116e58f)\n + Teach git-completion about git p4\n\n Comment from Pete will need to be addressed in a follow-up patch.\n\n\n* as/check-ignore (2012-09-27) 17 commits\n - [SQUASH-FIX] 283d072 (Add git-check-ignore sub-command, 2012-09-20)\n - [SQUASH-FIX] 283d072 (Add git-check-ignore sub-command, 2012-09-20)\n - [REROLL NEEDED] minimum compilation fix\n - Add git-check-ignore sub-command\n - dir.c: provide free_directory() for reclaiming dir_struct memory\n - pathspec.c: move reusable code from builtin/add.c\n - dir.c: refactor treat_gitlinks()\n - dir.c: keep track of where patterns came from\n - dir.c: refactor is_path_excluded()\n - dir.c: refactor is_excluded()\n - dir.c: refactor is_excluded_from_list()\n - dir.c: rename excluded() to is_excluded()\n - dir.c: rename excluded_from_list() to is_excluded_from_list()\n - dir.c: rename path_excluded() to is_path_excluded()\n - dir.c: rename cryptic 'which' variable to more consistent name\n - Improve documentation and comments regarding directory traversal API\n - Update directory listing API doc to match code\n\n Expecting a reroll.\n\n\n* as/test-tweaks (2012-09-20) 7 commits\n - tests: paint unexpectedly fixed known breakages in bold red\n - tests: test the test framework more thoroughly\n - [SQUASH] t/t0000-basic.sh: quoting of TEST_DIRECTORY is screwed up\n - tests: refactor mechanics of testing in a sub test-lib\n - tests: paint skipped tests in bold blue\n - tests: test number comes first in 'not ok $count - $message'\n - tests: paint known breakages in bold yellow\n\n Various minor tweaks to the test framework to paint its output\n lines in colors that match what they mean better.\n\n Has the \"is this really blue?\" issue Peff raised resolved???\n\n\n* fa/remote-svn (2012-09-19) 16 commits\n - Add a test script for remote-svn\n - remote-svn: add marks-file regeneration\n - Add a svnrdump-simulator replaying a dump file for testing\n - remote-svn: add incremental import\n - remote-svn: Activate import/export-marks for fast-import\n - Create a note for every imported commit containing svn metadata\n - vcs-svn: add fast_export_note to create notes\n - Allow reading svn dumps from files via file:// urls\n - remote-svn, vcs-svn: Enable fetching to private refs\n - When debug==1, start fast-import with \"--stats\" instead of \"--quiet\"\n - Add documentation for the 'bidi-import' capability of remote-helpers\n - Connect fast-import to the remote-helper via pipe, adding 'bidi-import' capability\n - Add argv_array_detach and argv_array_free_detached\n - Add svndump_init_fd to allow reading dumps from arbitrary FDs\n - Add git-remote-testsvn to Makefile\n - Implement a remote helper for svn in C\n (this branch is used by fa/vcs-svn.)\n\n A GSoC project.\n Waiting for comments from mentors and stakeholders.\n\n\n* fa/vcs-svn (2012-09-19) 4 commits\n - vcs-svn: remove repo_tree\n - vcs-svn/svndump: rewrite handle_node(), begin|end_revision()\n - vcs-svn/svndump: restructure node_ctx, rev_ctx handling\n - svndump: move struct definitions to .h\n (this branch uses fa/remote-svn.)\n\n A GSoC project.\n Waiting for comments from mentors and stakeholders.\n\n\n* jc/maint-name-rev (2012-09-17) 7 commits\n - describe --contains: use \"name-rev --algorithm=weight\"\n - name-rev --algorithm=weight: tests and documentation\n - name-rev --algorithm=weight: cache the computed weight in notes\n - name-rev --algorithm=weight: trivial optimization\n - name-rev: --algorithm option\n - name_rev: clarify the logic to assign a new tip-name to a commit\n - name-rev: lose unnecessary typedef\n\n \"git name-rev\" names the given revision based on a ref that can be\n reached in the smallest number of steps from the rev, but that is\n not useful when the caller wants to know which tag is the oldest one\n that contains the rev.  This teaches a new mode to the command that\n uses the oldest ref among those which contain the rev.\n\n I am not sure if this is worth it; for one thing, even with the help\n from notes-cache, it seems to make the \"describe --contains\" even\n slower. Also the command will be unusably slow for a user who does\n not have a write access (hence unable to create or update the\n notes-cache).\n\n Stalled mostly due to lack of responses.\n\n\n* jc/xprm-generation (2012-09-14) 1 commit\n - test-generation: compute generation numbers and clock skews\n\n A toy to analyze how bad the clock skews are in histories of real\n world projects.\n\n Stalled mostly due to lack of responses.\n\n\n* jc/blame-no-follow (2012-09-21) 2 commits\n - blame: pay attention to --no-follow\n - diff: accept --no-follow option\n\n Teaches \"--no-follow\" option to \"git blame\" to disable its\n whole-file rename detection.\n\n Stalled mostly due to lack of responses.\n\n\n* mk/maint-graph-infinity-loop (2012-09-25) 1 commit\n - graph.c: infinite loop in git whatchanged --graph -m\n\n The --graph code fell into infinite loop when asked to do what the\n code did not expect ;-)\n\n Anybody who worked on \"--graph\" wants to comment?\n Stalled mostly due to lack of responses.\n\n\n* ph/credential-refactor (2012-09-02) 5 commits\n - wincred: port to generic credential helper\n - Merge branch 'ef/win32-cred-helper' into ph/credential-refactor\n - osxkeychain: port to generic credential helper implementation\n - gnome-keyring: port to generic helper implementation\n - contrib: add generic credential helper\n\n Attempts to refactor to share code among OSX keychain, Gnome keyring\n and Win32 credential helpers.\n\n\n* ms/contrib-thunderbird-updates (2012-08-31) 2 commits\n - [SQUASH] minimum fixup\n - Thunderbird: fix appp.sh format problems\n\n Update helper to send out format-patch output using Thunderbird.\n Seems to have design regression for silent users.\n\n\n* jx/test-real-path (2012-08-27) 1 commit\n - test: set the realpath of CWD as TRASH_DIRECTORY\n\n Running tests with the \"trash\" directory elsewhere with the \"--root\"\n option did not work well if the directory was specified by a symbolic\n link pointing at it.\n\n Seems broken as it makes $(pwd) and TRASH_DIRECTORY inconsistent.\n Will discard.\n\n\n* jc/maint-push-refs-all (2012-08-27) 2 commits\n - get_fetch_map(): tighten checks on dest refs\n - [BROKEN] fetch/push: allow refs/*:refs/*\n\n Allows pushing and fetching everything including refs/stash.\n This is broken (see the log message there).\n\n Not ready.\n\n\n* jc/add-delete-default (2012-08-13) 1 commit\n - git add: notice removal of tracked paths by default\n\n \"git add dir/\" updated modified files and added new files, but does\n not notice removed files, which may be \"Huh?\" to some users.  They\n can of course use \"git add -A dir/\", but why should they?\n\n Resurrected from graveyard, as I thought it was a worthwhile thing\n to do in the longer term.\n\n Waiting for comments.\n\n\n* tx/relative-in-the-future (2012-08-16) 2 commits\n - date: show relative dates in the future\n - date: refactor the relative date logic from presentation\n\n Not my itch; rewritten an earlier submission by Tom Xue into\n somewhat more maintainable form, though it breaks existing i18n.\n\n Waiting for a voluteer to fix it up.\n Otherwise may discard.\n\n\n* mb/remote-default-nn-origin (2012-07-11) 6 commits\n - Teach get_default_remote to respect remote.default.\n - Test that plain \"git fetch\" uses remote.default when on a detached HEAD.\n - Teach clone to set remote.default.\n - Teach \"git remote\" about remote.default.\n - Teach remote.c about the remote.default configuration setting.\n - Rename remote.c's default_remote_name static variables.\n\n When the user does not specify what remote to interact with, we\n often attempt to use 'origin'.  This can now be customized via a\n configuration variable.\n\n Expecting a reroll.\n\n \"The first remote becomes the default\" bit is better done as a\n separate step.\n\n\n* jc/split-blob (2012-04-03) 6 commits\n - chunked-object: streaming checkout\n - chunked-object: fallback checkout codepaths\n - bulk-checkin: support chunked-object encoding\n - bulk-checkin: allow the same data to be multiply hashed\n - new representation types in the packstream\n - packfile: use varint functions\n\n Not ready.\n\n I finished the streaming checkout codepath, but as explained in\n 127b177 (bulk-checkin: support chunked-object encoding, 2011-11-30),\n these are still early steps of a long and painful journey. At least\n pack-objects and fsck need to learn the new encoding for the series\n to be usable locally, and then index-pack/unpack-objects needs to\n learn it to be used remotely.\n\n Given that I heard a lot of noise that people want large files, and\n that I was asked by somebody at GitTogether'11 privately for an\n advice on how to pay developers (not me) to help adding necessary\n support, I am somewhat dissapointed that the original patch series\n that was sent long time ago still remains here without much comments\n and updates from the developer community. I even made the interface\n to the logic that decides where to split chunks easily replaceable,\n and I deliberately made the logic in the original patch extremely\n stupid to entice others, especially the \"bup\" fanbois, to come up\n with a better logic, thinking that giving people an easy target to\n shoot for, they may be encouraged to help out. The plan is not\n working :-<.\n\n--------------------------------------------------\n[Cooking]\n\n* jl/submodule-add-by-name (2012-09-30) 2 commits\n - submodule add: Fail when .git/modules/<name> already exists unless forced\n - Teach \"git submodule add\" the --name option\n\n If you remove a submodule, in order to keep the repository so that\n \"git checkout\" to an older commit in the superproject history can\n resurrect the submodule, the real repository will stay in $GIT_DIR\n of the superproject.  A later \"git submodule add $path\" to add a\n different submodule at the same path will fail.  Diagnose this case\n a bit better, and if the user really wants to add an unrelated\n submodule at the same path, give the \"--name\" option to give it a\n place in $GIT_DIR of the superproject that does not conflict with\n the original submodule.\n\n Will merge to 'next'.\n\n\n* nd/grep-reflog (2012-09-29) 4 commits\n  (merged to 'next' on 2012-10-01 at 57773a6)\n + revision: make --grep search in notes too if shown\n + log --grep-reflog: reject the option without -g\n + revision: add --grep-reflog to filter commits by reflog messages\n + grep: prepare for new header field filter\n\n Teach the commands from the \"log\" family the \"--grep-reflog\" option\n to limit output by string that appears in the reflog entry when the\n \"--walk-reflogs\" option is in effect.\n\n Will merge to 'master'.\n\n\n* lt/mailinfo-handle-attachment-more-sanely (2012-09-30) 1 commit\n  (merged to 'next' on 2012-10-01 at 2a1cecc)\n + mailinfo: don't require \"text\" mime type for attachments\n\n A patch attached as application/octet-stream (e.g. not text/*) were\n mishandled, not correctly honoring Content-Transfer-Encoding\n (e.g. base64).\n\n Will merge to 'master'.\n\n\n* tu/gc-auto-quiet (2012-09-27) 1 commit\n  (merged to 'next' on 2012-10-01 at ad8b91b)\n + silence git gc --auto --quiet output\n\n \"gc --auto\" notified the user that auto-packing has triggered even\n under the \"--quiet\" option.\n\n Will merge to 'master'.\n\n\n* jm/diff-context-config (2012-10-02) 2 commits\n  (merged to 'next' on 2012-10-02 at e57700a)\n + t4055: avoid use of sed 'a' command\n  (merged to 'next' on 2012-10-01 at 509a558)\n + diff: diff.context configuration gives default to -U\n\n Teaches a new configuration variable to \"git diff\" Porcelain and\n its friends.\n\n Will defer to the next cycle.\n\n\n* mh/ceiling (2012-09-29) 9 commits\n - t1504: stop resolving symlinks in GIT_CEILING_DIRECTORIES\n - longest_ancestor_length(): resolve symlinks before comparing paths\n - longest_ancestor_length(): use string_list_longest_prefix()\n - longest_ancestor_length(): always add a slash to the end of prefixes\n - longest_ancestor_length(): explicitly filter list before loop\n - longest_ancestor_length(): use string_list_split()\n - Introduce new function real_path_if_valid()\n - real_path_internal(): add comment explaining use of cwd\n - Introduce new static function real_path_internal()\n\n Elements of GIT_CEILING_DIRECTORIES list may not match the real\n pathname we obtain from getcwd(), leading the GIT_DIR discovery\n logic to escape the ceilings the user thought to have specified.\n\n The solution felt a bit unnecessarily convoluted to me.\n Expecting a reroll.\n\n\n* jl/submodule-rm (2012-09-29) 1 commit\n  (merged to 'next' on 2012-10-01 at 4e5c4fc)\n + submodule: teach rm to remove submodules unless they contain a git directory\n\n \"git rm submodule\" cannot blindly remove a submodule directory as\n its working tree may have local changes, and worse yet, it may even\n have its repository embedded in it.  Teach it some special cases\n where it is safe to remove a submodule, specifically, when there is\n no local changes in the submodule working tree, and its repository\n is not embedded in its working tree but is elsewhere and uses the\n gitfile mechanism to point at it.\n\n Will defer to the next cycle.\n\n\n* nd/wildmatch (2012-09-27) 5 commits\n - Support \"**\" in .gitignore and .gitattributes patterns using wildmatch()\n - Integrate wildmatch to git\n - compat/wildmatch: fix case-insensitive matching\n - compat/wildmatch: remove static variable force_lower_case\n - Import wildmatch from rsync\n\n Allows pathname patterns in .gitignore and .gitattributes files\n with double-asterisks \"foo/**/bar\" to match any number of directory\n hiearchies.\n\n It was pointed out that some symbols that do not have to be global\n are left global. I think this reroll fixed most of them.\n\n Will merge to 'next'.\n\n\n* nd/pretty-placeholder-with-color-option (2012-09-30) 9 commits\n - pretty: support %>> that steal trailing spaces\n - pretty: support truncating in %>, %< and %><\n - pretty: support padding placeholders, %< %> and %><\n - pretty: two phase conversion for non utf-8 commits\n - utf8.c: add utf8_strnwidth() with the ability to skip ansi sequences\n - utf8.c: move display_mode_esc_sequence_len() for use by other functions\n - pretty: support %C(auto[,N]) to turn on coloring on next placeholder(s)\n - pretty: split parsing %C into a separate function\n - pretty: share code between format_decoration and show_decorations\n\n\n* jk/no-more-pre-exec-callback (2012-06-05) 1 commit\n - pager: drop \"wait for output to run less\" hack\n\n (Originally merged to 'next' on 2012-07-23)\n\n Will defer to the next cycle.\n"},{"id":"200428","messageId":"CACsJy8BGuoW6K_9vEgGrb2XC2bNtR=0jNRU3JQhsv7_diGQpbA@mail.gmail.com","threadId":"31717","inReplyTo":"7vmx045umh.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-03T15:23:04Z","receivedAt":"2012-10-03T15:23:04Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Oct 3, 2012 at 6:20 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> * nd/wildmatch (2012-09-27) 5 commits\n>  - Support \"**\" in .gitignore and .gitattributes patterns using wildmatch()\n>  - Integrate wildmatch to git\n>  - compat/wildmatch: fix case-insensitive matching\n>  - compat/wildmatch: remove static variable force_lower_case\n>  - Import wildmatch from rsync\n>\n>  Allows pathname patterns in .gitignore and .gitattributes files\n>  with double-asterisks \"foo/**/bar\" to match any number of directory\n>  hiearchies.\n>\n>  It was pointed out that some symbols that do not have to be global\n>  are left global. I think this reroll fixed most of them.\n>\n>  Will merge to 'next'.\n\nJust a bit of finding lately, in case you want to postpone the merge.\n\nThere's an interesting case: \"**foo\". According to our rules, that\npattern does not contain slashes therefore is basename match. But some\nmight find that confusing because \"**\" can match slashes, as opposed\nto ordinary wildcards which cannot. So we could either go with our\nrules and consider \"**\" just like \"*\" in this case (do we need\ndocument clarification?), or redefine it that the presence of \"**\"\nimplies FNM_PATHNAME.\n\nI think the latter makes more sense. When users put \"**\" they expect\nto match some slashes. But that may call for a refactoring in\npath_matches() in attr.c. Putting strstr(pattern, \"**\") in that\nmatching function may increase overhead unnecessarily.\n\nThe third option is just die() and let users decide either \"*foo\",\n\"**/foo\" or \"/**foo\", never \"**foo\".\n-- \nDuy\n"},{"id":"200439","messageId":"7vbogj5sji.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"CACsJy8BGuoW6K_9vEgGrb2XC2bNtR=0jNRU3JQhsv7_diGQpbA@mail.gmail.com","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-03T18:17:37Z","receivedAt":"2012-10-03T18:17:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n\n> There's an interesting case: \"**foo\". According to our rules, that\n> pattern does not contain slashes therefore is basename match. But some\n> might find that confusing because \"**\" can match slashes,...\n\nBy \"our rules\", if you mean \"if a pattern has slash, it is anchored\",\nthat obviously need to be updated with this series, if \"**\" is meant\nto match multiple hierarchies.\n> I think the latter makes more sense. When users put \"**\" they expect\n> to match some slashes. But that may call for a refactoring in\n> path_matches() in attr.c. Putting strstr(pattern, \"**\") in that\n> matching function may increase overhead unnecessarily.\n>\n> The third option is just die() and let users decide either \"*foo\",\n> \"**/foo\" or \"/**foo\", never \"**foo\".\n\nFor the double-star at the beginning, you should just turn it into \"**/\"\nif it is not followed by a slash internally, I think.\n\nWhat is the semantics of ** in the first place?  Is it described to\na reasonable level of detail in the documentation updates?  For\nexample does \"**foo\" match \"afoo\", \"a/b/foo\", \"a/bfoo\", \"a/foo/b\",\n\"a/bfoo/c\"?  Does \"x**y\" match \"xy\", \"xay\", \"xa/by\", \"x/a/y\"?\n\nI am guessing that the only sensible definition is that \"**\"\nrequires anything that comes before it (if exists) is at a proper\nhierarchy boundary, and anything matches it is also at a proper\nhierarchy boundary, so \"x**y\" matches \"x/a/y\" and not \"xy\", \"xay\",\nnor \"xa/by\" in the above example.  If \"x**y\" can match \"xy\" or \"xay\"\n(or \"**foo\" can match \"afoo\"), it would be unreasonable to say it\nimplies the pattern is anchored at any level, no?\n"},{"id":"200464","messageId":"CACsJy8CAGaEzGZBJq7pOW_2SDpRDLiPqJK0t2WjpuNqLU+yewQ@mail.gmail.com","threadId":"31717","inReplyTo":"7vbogj5sji.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T01:56:36Z","receivedAt":"2012-10-04T01:56:36Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Oct 4, 2012 at 1:17 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> For the double-star at the beginning, you should just turn it into \"**/\"\n> if it is not followed by a slash internally, I think.\n>\n> What is the semantics of ** in the first place?  Is it described to\n> a reasonable level of detail in the documentation updates?  For\n> example does \"**foo\" match \"afoo\", \"a/b/foo\", \"a/bfoo\", \"a/foo/b\",\n> \"a/bfoo/c\"?  Does \"x**y\" match \"xy\", \"xay\", \"xa/by\", \"x/a/y\"?\n\nIt's basically what rsync describes: use ’**’ to match anything,\nincluding slashes.\n\nReading rsync's man page again, I notice I missed two other rules related to **:\n\n - If the pattern contains a / (not counting a trailing /) or a \"**\",\nthen it is matched against the full pathname, including any leading\ndirectories.  If  the  pattern  doesn't contain  a / or a \"**\", then\nit is matched only against the final component of the filename.\n(Remember that the algorithm is applied recursively so \"full filename\"\ncan actually be any portion of a path from the starting directory on\ndown.)\n\n - A trailing \"dir_name/***\" will match both the directory (as if\n\"dir_name/\" had been specified) and everything in the directory (as if\n\"dir_name/**\" had been specified).  This behavior was added in version\n2.6.7.\n\nFrom what you wrote, I think we'll go with the first rule. The second\nrule looks irrelevant to what git's doing.\n\n> I am guessing that the only sensible definition is that \"**\"\n> requires anything that comes before it (if exists) is at a proper\n> hierarchy boundary, and anything matches it is also at a proper\n> hierarchy boundary, so \"x**y\" matches \"x/a/y\"\n\nand \"x/y\" too? (As opposed to \"x/**/y\" which does not)\n\n> and not \"xy\", \"xay\",\n> nor \"xa/by\" in the above example.  If \"x**y\" can match \"xy\" or \"xay\"\n> (or \"**foo\" can match \"afoo\"), it would be unreasonable to say it\n> implies the pattern is anchored at any level, no?\n\nYeah. That makes things easier to reason, though not exactly what we're having.\n-- \nDuy\n"},{"id":"200477","messageId":"7v626q3hen.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"CACsJy8CAGaEzGZBJq7pOW_2SDpRDLiPqJK0t2WjpuNqLU+yewQ@mail.gmail.com","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-04T06:01:04Z","receivedAt":"2012-10-04T06:01:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n\n>> I am guessing that the only sensible definition is that \"**\"\n>> requires anything that comes before it (if exists) is at a proper\n>> hierarchy boundary, and anything matches it is also at a proper\n>> hierarchy boundary, so \"x**y\" matches \"x/a/y\"\n>\n> and \"x/y\" too? (As opposed to \"x/**/y\" which does not)\n\nYeah, x**y would match x/y under that \"sensible\" semantics.\n\n>> and not \"xy\", \"xay\",\n>> nor \"xa/by\" in the above example.  If \"x**y\" can match \"xy\" or \"xay\"\n>> (or \"**foo\" can match \"afoo\"), it would be unreasonable to say it\n>> implies the pattern is anchored at any level, no?\n>\n> Yeah. That makes things easier to reason, though not exactly what we're having.\n\nIt sounds like that \"x**y\" with the code you imported would match\n\"xy\" and \"xa/b/cy\", and I do not think of a concise and good way to\ndescribe what it does to the end users.\n\n\"matches anything including '/'\" is not a useful description for the\npurpose of allowing the user to intuitively understand why \"x**y\" is\nanchored at the level (or is not anchored and can appear anywhere).\n\nPerhaps the wildmatch code may not be what we want X-<.\n"},{"id":"200513","messageId":"1349336392-1772-1-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"7v626q3hen.fsf@alter.siamese.dyndns.org","subject":"[PATCH 0/6] wildmatch part 2","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:46Z","receivedAt":"2012-10-04T07:39:46Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Oct 4, 2012 at 1:01 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Perhaps the wildmatch code may not be what we want X-<.\n\nWhen I imported wildmatch I was hoping to make minimum changes to it.\nBut wildmatch is probably the only practical way to support \"**\" even\nif we later need to change it the way we want. Other options are base\nour work on top of compat/fnmatch.c, which is an #ifdef spaghetti\nmess, or write a new fnmatch()-compatible function. Both unattractive\nto me.\n\nAnyway, this is on top of nd/wildmatch, which makes \"ab**cd\" match\nfull pathname.\n\nattr patches port .gitignore optimizations over. In long term, we\nshould probably have a shared matching implementation instead. I tried\nthat road once and failed so I won't attempt again any time soon. If\nwe drop wildmatch, I can split these attr patches out as a separate\nseries. It's a good thing to do anyway.\n\nThe last patch just reflects that current \"**\" is not exactly what we\nwant. I'm not sure if I could look into wildmatch.c and change it.\nAnybody is welcome to step up, of course.\n\nNguyễn Thái Ngọc Duy (6):\n  attr: remove the union in struct match_attr\n  attr: avoid strlen() on every match\n  attr: avoid searching for basename on every match\n  attr: more matching optimizations from .gitignore\n  gitignore: do not do basename match with patterns that have '**'\n  t3001: note about expected \"**\" behavior\n\n Documentation/gitignore.txt        |  10 ++--\n attr.c                             | 101 +++++++++++++++++++++++++++----------\n dir.c                              |   6 +--\n dir.h                              |   2 +\n t/t0003-attributes.sh              |  16 ++++++\n t/t3001-ls-files-others-exclude.sh |  18 +++++++\n 6 files changed, 118 insertions(+), 35 deletions(-)\n\n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200527","messageId":"1349336392-1772-2-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 1/6] attr: remove the union in struct match_attr","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:47Z","receivedAt":"2012-10-04T07:39:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We're going to add more attributes to u.pattern so it'll become bigger\nin size than a pointer. There's no point in sharing the same room with\nu.attr.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n attr.c | 25 ++++++++++++-------------\n 1 file changed, 12 insertions(+), 13 deletions(-)\n\ndiff --git a/attr.c b/attr.c\nindex 15ebaa1..48df800 100644\n--- a/attr.c\n+++ b/attr.c\n@@ -119,10 +119,10 @@ struct attr_state {\n /*\n  * One rule, as from a .gitattributes file.\n  *\n- * If is_macro is true, then u.attr is a pointer to the git_attr being\n+ * If is_macro is true, then attr is a pointer to the git_attr being\n  * defined.\n  *\n- * If is_macro is false, then u.pattern points at the filename pattern\n+ * If is_macro is false, then pattern points at the filename pattern\n  * to which the rule applies.  (The memory pointed to is part of the\n  * memory block allocated for the match_attr instance.)\n  *\n@@ -131,10 +131,8 @@ struct attr_state {\n  * listed as they appear in the file (macros unexpanded).\n  */\n struct match_attr {\n-\tunion {\n-\t\tchar *pattern;\n-\t\tstruct git_attr *attr;\n-\t} u;\n+\tconst char *pattern;\n+\tstruct git_attr *attr;\n \tchar is_macro;\n \tunsigned num_attr;\n \tstruct attr_state state[FLEX_ARRAY];\n@@ -240,11 +238,12 @@ static struct match_attr *parse_attr_line(const char *line, const char *src,\n \t\t      sizeof(struct attr_state) * num_attr +\n \t\t      (is_macro ? 0 : namelen + 1));\n \tif (is_macro)\n-\t\tres->u.attr = git_attr_internal(name, namelen);\n+\t\tres->attr = git_attr_internal(name, namelen);\n \telse {\n-\t\tres->u.pattern = (char *)&(res->state[num_attr]);\n-\t\tmemcpy(res->u.pattern, name, namelen);\n-\t\tres->u.pattern[namelen] = 0;\n+\t\tchar *p = (char *)&(res->state[num_attr]);\n+\t\tmemcpy(p, name, namelen);\n+\t\tp[namelen] = 0;\n+\t\tres->pattern = p;\n \t}\n \tres->is_macro = is_macro;\n \tres->num_attr = num_attr;\n@@ -682,7 +681,7 @@ static int fill_one(const char *what, struct match_attr *a, int rem)\n \n \t\tif (*n == ATTR__UNKNOWN) {\n \t\t\tdebug_set(what,\n-\t\t\t\t  a->is_macro ? a->u.attr->name : a->u.pattern,\n+\t\t\t\t  a->is_macro ? a->attr->name : a->pattern,\n \t\t\t\t  attr, v);\n \t\t\t*n = v;\n \t\t\trem--;\n@@ -702,7 +701,7 @@ static int fill(const char *path, int pathlen, struct attr_stack *stk, int rem)\n \t\tif (a->is_macro)\n \t\t\tcontinue;\n \t\tif (path_matches(path, pathlen,\n-\t\t\t\t a->u.pattern, base, strlen(base)))\n+\t\t\t\t a->pattern, base, strlen(base)))\n \t\t\trem = fill_one(\"fill\", a, rem);\n \t}\n \treturn rem;\n@@ -722,7 +721,7 @@ static int macroexpand_one(int attr_nr, int rem)\n \t\t\tstruct match_attr *ma = stk->attrs[i];\n \t\t\tif (!ma->is_macro)\n \t\t\t\tcontinue;\n-\t\t\tif (ma->u.attr->attr_nr == attr_nr)\n+\t\t\tif (ma->attr->attr_nr == attr_nr)\n \t\t\t\ta = ma;\n \t\t}\n \n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200510","messageId":"1349336392-1772-3-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 2/6] attr: avoid strlen() on every match","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:48Z","receivedAt":"2012-10-04T07:39:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n attr.c | 6 ++++--\n 1 file changed, 4 insertions(+), 2 deletions(-)\n\ndiff --git a/attr.c b/attr.c\nindex 48df800..66b96d9 100644\n--- a/attr.c\n+++ b/attr.c\n@@ -277,6 +277,7 @@ static struct match_attr *parse_attr_line(const char *line, const char *src,\n static struct attr_stack {\n \tstruct attr_stack *prev;\n \tchar *origin;\n+\tsize_t originlen;\n \tunsigned num_matches;\n \tunsigned alloc;\n \tstruct match_attr **attrs;\n@@ -532,6 +533,7 @@ static void bootstrap_attr_stack(void)\n \tif (!is_bare_repository() || direction == GIT_ATTR_INDEX) {\n \t\telem = read_attr(GITATTRIBUTES_FILE, 1);\n \t\telem->origin = xstrdup(\"\");\n+\t\telem->originlen = 0;\n \t\telem->prev = attr_stack;\n \t\tattr_stack = elem;\n \t\tdebug_push(elem);\n@@ -625,7 +627,7 @@ static void prepare_attr_stack(const char *path)\n \t\t\tstrbuf_addstr(&pathbuf, GITATTRIBUTES_FILE);\n \t\t\telem = read_attr(pathbuf.buf, 0);\n \t\t\tstrbuf_setlen(&pathbuf, cp - path);\n-\t\t\telem->origin = strbuf_detach(&pathbuf, NULL);\n+\t\t\telem->origin = strbuf_detach(&pathbuf, &elem->originlen);\n \t\t\telem->prev = attr_stack;\n \t\t\tattr_stack = elem;\n \t\t\tdebug_push(elem);\n@@ -701,7 +703,7 @@ static int fill(const char *path, int pathlen, struct attr_stack *stk, int rem)\n \t\tif (a->is_macro)\n \t\t\tcontinue;\n \t\tif (path_matches(path, pathlen,\n-\t\t\t\t a->pattern, base, strlen(base)))\n+\t\t\t\t a->pattern, base, stk->originlen))\n \t\t\trem = fill_one(\"fill\", a, rem);\n \t}\n \treturn rem;\n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200491","messageId":"1349336392-1772-4-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 3/6] attr: avoid searching for basename on every match","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:49Z","receivedAt":"2012-10-04T07:39:49Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n attr.c | 15 +++++++++------\n 1 file changed, 9 insertions(+), 6 deletions(-)\n\ndiff --git a/attr.c b/attr.c\nindex 66b96d9..eb576ac 100644\n--- a/attr.c\n+++ b/attr.c\n@@ -644,13 +644,11 @@ static void prepare_attr_stack(const char *path)\n }\n \n static int path_matches(const char *pathname, int pathlen,\n+\t\t\tconst char *basename,\n \t\t\tconst char *pattern,\n \t\t\tconst char *base, int baselen)\n {\n \tif (!strchr(pattern, '/')) {\n-\t\t/* match basename */\n-\t\tconst char *basename = strrchr(pathname, '/');\n-\t\tbasename = basename ? basename + 1 : pathname;\n \t\treturn (fnmatch_icase(pattern, basename, 0) == 0);\n \t}\n \t/*\n@@ -693,7 +691,8 @@ static int fill_one(const char *what, struct match_attr *a, int rem)\n \treturn rem;\n }\n \n-static int fill(const char *path, int pathlen, struct attr_stack *stk, int rem)\n+static int fill(const char *path, int pathlen, const char *basename,\n+\t\tstruct attr_stack *stk, int rem)\n {\n \tint i;\n \tconst char *base = stk->origin ? stk->origin : \"\";\n@@ -702,7 +701,7 @@ static int fill(const char *path, int pathlen, struct attr_stack *stk, int rem)\n \t\tstruct match_attr *a = stk->attrs[i];\n \t\tif (a->is_macro)\n \t\t\tcontinue;\n-\t\tif (path_matches(path, pathlen,\n+\t\tif (path_matches(path, pathlen, basename,\n \t\t\t\t a->pattern, base, stk->originlen))\n \t\t\trem = fill_one(\"fill\", a, rem);\n \t}\n@@ -741,15 +740,19 @@ static void collect_all_attrs(const char *path)\n {\n \tstruct attr_stack *stk;\n \tint i, pathlen, rem;\n+\tconst char *basename;\n \n \tprepare_attr_stack(path);\n \tfor (i = 0; i < attr_nr; i++)\n \t\tcheck_all_attr[i].value = ATTR__UNKNOWN;\n \n+\tbasename = strrchr(path, '/');\n+\tbasename = basename ? basename + 1 : path;\n+\n \tpathlen = strlen(path);\n \trem = attr_nr;\n \tfor (stk = attr_stack; 0 < rem && stk; stk = stk->prev)\n-\t\trem = fill(path, pathlen, stk, rem);\n+\t\trem = fill(path, pathlen, basename, stk, rem);\n }\n \n int git_check_attr(const char *path, int num, struct git_attr_check *check)\n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200496","messageId":"1349336392-1772-5-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 4/6] attr: more matching optimizations from .gitignore","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:50Z","receivedAt":"2012-10-04T07:39:50Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":".gitattributes and .gitignore share the same pattern syntax but has\nseparate matching implementation. Over the years, ignore's\nimplementation accumulates more optimizations while attr's stays the\nsame.\n\nThis patch adds those optimizations to .gitattributes. Basically it\ntries to avoid fnmatch/wildmatch in favor of strncmp as much as\npossible.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n attr.c | 63 +++++++++++++++++++++++++++++++++++++++++++++++++++++----------\n dir.c  |  4 ++--\n dir.h  |  2 ++\n 3 files changed, 57 insertions(+), 12 deletions(-)\n\ndiff --git a/attr.c b/attr.c\nindex eb576ac..3fde9fa 100644\n--- a/attr.c\n+++ b/attr.c\n@@ -116,6 +116,13 @@ struct attr_state {\n \tconst char *setto;\n };\n \n+struct pattern {\n+\tconst char *pattern;\n+\tint patternlen;\n+\tint nowildcardlen;\n+\tint flags;\t\t/* EXC_FLAG_* */\n+};\n+\n /*\n  * One rule, as from a .gitattributes file.\n  *\n@@ -131,7 +138,7 @@ struct attr_state {\n  * listed as they appear in the file (macros unexpanded).\n  */\n struct match_attr {\n-\tconst char *pattern;\n+\tstruct pattern pat;\n \tstruct git_attr *attr;\n \tchar is_macro;\n \tunsigned num_attr;\n@@ -243,7 +250,13 @@ static struct match_attr *parse_attr_line(const char *line, const char *src,\n \t\tchar *p = (char *)&(res->state[num_attr]);\n \t\tmemcpy(p, name, namelen);\n \t\tp[namelen] = 0;\n-\t\tres->pattern = p;\n+\t\tres->pat.pattern = p;\n+\t\tres->pat.patternlen = strlen(p);\n+\t\tres->pat.nowildcardlen = simple_length(p);\n+\t\tif (!strchr(p, '/'))\n+\t\t\tres->pat.flags |= EXC_FLAG_NODIR;\n+\t\tif (*p == '*' && no_wildcard(p+1))\n+\t\t\tres->pat.flags |= EXC_FLAG_ENDSWITH;\n \t}\n \tres->is_macro = is_macro;\n \tres->num_attr = num_attr;\n@@ -645,26 +658,56 @@ static void prepare_attr_stack(const char *path)\n \n static int path_matches(const char *pathname, int pathlen,\n \t\t\tconst char *basename,\n-\t\t\tconst char *pattern,\n+\t\t\tconst struct pattern *pat,\n \t\t\tconst char *base, int baselen)\n {\n-\tif (!strchr(pattern, '/')) {\n+\tconst char *pattern = pat->pattern;\n+\tint prefix = pat->nowildcardlen;\n+\tconst char *name;\n+\tint namelen;\n+\n+\tif (pat->flags & EXC_FLAG_NODIR) {\n+\t\tif (prefix == pat->patternlen &&\n+\t\t    !strcmp_icase(pattern, basename))\n+\t\t\treturn 1;\n+\n+\t\tif (pat->flags & EXC_FLAG_ENDSWITH &&\n+\t\t    pat->patternlen - 1 <= pathlen &&\n+\t\t    !strcmp_icase(pattern + 1, pathname +\n+\t\t\t\t  pathlen - pat->patternlen + 1))\n+\t\t\treturn 1;\n+\n \t\treturn (fnmatch_icase(pattern, basename, 0) == 0);\n \t}\n \t/*\n \t * match with FNM_PATHNAME; the pattern has base implicitly\n \t * in front of it.\n \t */\n-\tif (*pattern == '/')\n+\tif (*pattern == '/') {\n \t\tpattern++;\n+\t\tprefix--;\n+\t}\n+\n+\t/*\n+\t * note: unlike excluded_from_list, baselen here does not\n+\t * contain the trailing slash\n+\t */\n+\n \tif (pathlen < baselen ||\n \t    (baselen && pathname[baselen] != '/') ||\n \t    strncmp(pathname, base, baselen))\n \t\treturn 0;\n-\tif (baselen != 0)\n-\t\tbaselen++;\n-\treturn (ignore_case && iwildmatch(pattern, pathname + baselen)) ||\n-\t\t(!ignore_case && wildmatch(pattern, pathname + baselen));\n+\n+\tnamelen = baselen ? pathlen - baselen - 1 : pathlen;\n+\tname = pathname + pathlen - namelen;\n+\n+\t/* if the non-wildcard part is longer than the remaining\n+\t   pathname, surely it cannot match */\n+\tif (!namelen || prefix > namelen)\n+\t\treturn 0;\n+\n+\treturn (ignore_case && iwildmatch(pattern, name)) ||\n+\t\t(!ignore_case && wildmatch(pattern, name));\n }\n \n static int macroexpand_one(int attr_nr, int rem);\n@@ -702,7 +745,7 @@ static int fill(const char *path, int pathlen, const char *basename,\n \t\tif (a->is_macro)\n \t\t\tcontinue;\n \t\tif (path_matches(path, pathlen, basename,\n-\t\t\t\t a->pattern, base, stk->originlen))\n+\t\t\t\t &a->pat, base, stk->originlen))\n \t\t\trem = fill_one(\"fill\", a, rem);\n \t}\n \treturn rem;\ndiff --git a/dir.c b/dir.c\nindex 92cda82..fd49336 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -292,7 +292,7 @@ int match_pathspec_depth(const struct pathspec *ps,\n /*\n  * Return the length of the \"simple\" part of a path match limiter.\n  */\n-static int simple_length(const char *match)\n+int simple_length(const char *match)\n {\n \tint len = -1;\n \n@@ -304,7 +304,7 @@ static int simple_length(const char *match)\n \t}\n }\n \n-static int no_wildcard(const char *string)\n+int no_wildcard(const char *string)\n {\n \treturn string[simple_length(string)] == '\\0';\n }\ndiff --git a/dir.h b/dir.h\nindex 893465a..7ea8678 100644\n--- a/dir.h\n+++ b/dir.h\n@@ -101,6 +101,8 @@ extern void add_exclude(const char *string, const char *base,\n \t\t\tint baselen, struct exclude_list *which);\n extern void free_excludes(struct exclude_list *el);\n extern int file_exists(const char *);\n+extern int simple_length(const char *match);\n+extern int no_wildcard(const char *string);\n \n extern int is_inside_dir(const char *dir);\n extern int dir_inside_of(const char *subdir, const char *dir);\n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200505","messageId":"1349336392-1772-6-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 5/6] gitignore: do not do basename match with patterns that have '**'","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:51Z","receivedAt":"2012-10-04T07:39:51Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\"**\" can match slashes, not like \"*\". \"ab**ef\" should be able to match\n\"ab/cd/ef\", or \"ab/c/d/ef\" and so on. Turn off the EXC_FLAG_NODIR in\nthis case otherwise the pattern is only checked against the base\nname. This behavior is in sync with rsync.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/gitignore.txt        | 10 +++++-----\n attr.c                             |  2 +-\n dir.c                              |  2 +-\n t/t0003-attributes.sh              | 16 ++++++++++++++++\n t/t3001-ls-files-others-exclude.sh | 10 ++++++++++\n 5 files changed, 33 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/gitignore.txt b/Documentation/gitignore.txt\nindex eb81d31..4dfe8bd 100644\n--- a/Documentation/gitignore.txt\n+++ b/Documentation/gitignore.txt\n@@ -81,11 +81,11 @@ PATTERN FORMAT\n    regular file or a symbolic link `foo` (this is consistent\n    with the way how pathspec works in general in git).\n \n- - If the pattern does not contain a slash '/', git treats it as\n-   a shell glob pattern and checks for a match against the\n-   pathname relative to the location of the `.gitignore` file\n-   (relative to the toplevel of the work tree if not from a\n-   `.gitignore` file).\n+ - If the pattern does not contain a slash '/' nor '**', git\n+   treats it as a shell glob pattern and checks for a match\n+   against the pathname relative to the location of the\n+   `.gitignore` file (relative to the toplevel of the work tree\n+   if not from a `.gitignore` file).\n \n  - Otherwise, git treats the pattern as a shell glob suitable\n    for consumption by fnmatch(3) with the FNM_PATHNAME flag:\ndiff --git a/attr.c b/attr.c\nindex 3fde9fa..634b39c 100644\n--- a/attr.c\n+++ b/attr.c\n@@ -253,7 +253,7 @@ static struct match_attr *parse_attr_line(const char *line, const char *src,\n \t\tres->pat.pattern = p;\n \t\tres->pat.patternlen = strlen(p);\n \t\tres->pat.nowildcardlen = simple_length(p);\n-\t\tif (!strchr(p, '/'))\n+\t\tif (!strchr(p, '/') && !strstr(p, \"**\"))\n \t\t\tres->pat.flags |= EXC_FLAG_NODIR;\n \t\tif (*p == '*' && no_wildcard(p+1))\n \t\t\tres->pat.flags |= EXC_FLAG_ENDSWITH;\ndiff --git a/dir.c b/dir.c\nindex fd49336..6a5de98 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -340,7 +340,7 @@ void add_exclude(const char *string, const char *base,\n \tx->base = base;\n \tx->baselen = baselen;\n \tx->flags = flags;\n-\tif (!strchr(string, '/'))\n+\tif (!strchr(string, '/') && !strstr(string, \"**\"))\n \t\tx->flags |= EXC_FLAG_NODIR;\n \tx->nowildcardlen = simple_length(string);\n \tif (*string == '*' && no_wildcard(string+1))\ndiff --git a/t/t0003-attributes.sh b/t/t0003-attributes.sh\nindex 6c3c554..9b534a0 100755\n--- a/t/t0003-attributes.sh\n+++ b/t/t0003-attributes.sh\n@@ -249,4 +249,20 @@ EOF\n \ttest_line_count = 0 err\n '\n \n+test_expect_success '\"**\" with no slashes test' '\n+\techo \"a**f foo=bar\" >.gitattributes &&\n+\tcat <<\\EOF >expect &&\n+f: foo: unspecified\n+a/f: foo: bar\n+a/b/f: foo: bar\n+a/b/c/f: foo: bar\n+EOF\n+\tgit check-attr foo -- \"f\" >actual 2>err &&\n+\tgit check-attr foo -- \"a/f\" >>actual 2>>err &&\n+\tgit check-attr foo -- \"a/b/f\" >>actual 2>>err &&\n+\tgit check-attr foo -- \"a/b/c/f\" >>actual 2>>err &&\n+\ttest_cmp expect actual &&\n+\ttest_line_count = 0 err\n+'\n+\n test_done\ndiff --git a/t/t3001-ls-files-others-exclude.sh b/t/t3001-ls-files-others-exclude.sh\nindex 67c8bcf..6a5a4ab 100755\n--- a/t/t3001-ls-files-others-exclude.sh\n+++ b/t/t3001-ls-files-others-exclude.sh\n@@ -225,4 +225,14 @@ EOF\n \ttest_cmp expect actual\n '\n \n+\n+test_expect_success 'ls-files with \"**\" patterns and no slashes' '\n+\tcat <<\\EOF >expect &&\n+one/a.1\n+one/two/a.1\n+EOF\n+\tgit ls-files -o -i --exclude \"one**a.1\" >actual\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200509","messageId":"1349336392-1772-7-git-send-email-pclouds@gmail.com","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"[PATCH 6/6] t3001: note about expected \"**\" behavior","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T07:39:52Z","receivedAt":"2012-10-04T07:39:52Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\"**\" currently matches any characters including slashes. It's probably\ntoo powerful. A more sensible definition may be match any characters\nthat the but the whole match must be wrapped by slashes. So \"**\" can\nmatch none, \"/\", \"/aaa/\", \"/aa/bb/\" and so on but not \"aa/bb\".\n\nNote it in the test suite.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t3001-ls-files-others-exclude.sh | 8 ++++++++\n 1 file changed, 8 insertions(+)\n\ndiff --git a/t/t3001-ls-files-others-exclude.sh b/t/t3001-ls-files-others-exclude.sh\nindex 6a5a4ab..99b5f5c 100755\n--- a/t/t3001-ls-files-others-exclude.sh\n+++ b/t/t3001-ls-files-others-exclude.sh\n@@ -235,4 +235,12 @@ EOF\n \ttest_cmp expect actual\n '\n \n+# We might want ** to match at directory boundary, e.g. a**b matches\n+# a/b, a/x/b, a/x/x/b... but not ax/xb.\n+test_expect_failure 'ls-files with \"**\" patterns and no slashes' '\n+\t: >expect &&\n+\tgit ls-files -o -i --exclude \"o**a.1\" >actual\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n1.7.12.1.405.gb727dc9\n"},{"id":"200517","messageId":"A4A111D1488E49FFA4D71D85DD6B87A4@rr-dav.id.au","threadId":"31717","inReplyTo":"7vmx045umh.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"David Michael Barr","fromEmail":"b@rr-dav.id.au","sentAt":"2012-10-04T08:17:52Z","receivedAt":"2012-10-04T08:17:52Z","isPatch":false,"sender":{"key":"b@rr-dav.id.au","avatar":"https://gravatar.com/avatar/1c0f0df262aa1749c478ee3586cef5da6d58382c06cb220882b7ef9b93cbec6f?d=mp&s=160"},"body":"\nOn Wednesday, 3 October 2012 at 9:20 AM, Junio C Hamano wrote: \n> \n> * fa/remote-svn (2012-09-19) 16 commits\n> - Add a test script for remote-svn\n> - remote-svn: add marks-file regeneration\n> - Add a svnrdump-simulator replaying a dump file for testing\n> - remote-svn: add incremental import\n> - remote-svn: Activate import/export-marks for fast-import\n> - Create a note for every imported commit containing svn metadata\n> - vcs-svn: add fast_export_note to create notes\n> - Allow reading svn dumps from files via file:// urls\n> - remote-svn, vcs-svn: Enable fetching to private refs\n> - When debug==1, start fast-import with \"--stats\" instead of \"--quiet\"\n> - Add documentation for the 'bidi-import' capability of remote-helpers\n> - Connect fast-import to the remote-helper via pipe, adding 'bidi-import' capability\n> - Add argv_array_detach and argv_array_free_detached\n> - Add svndump_init_fd to allow reading dumps from arbitrary FDs\n> - Add git-remote-testsvn to Makefile\n> - Implement a remote helper for svn in C\n> (this branch is used by fa/vcs-svn.)\n> \n> A GSoC project.\n> Waiting for comments from mentors and stakeholders.\n\nI have reviewed this topic and am happy with the design and implementation.\nI support this topic for inclusion.\n\nAcked-by: David Michael Barr <b@rr-dav.id.au>\n> \n> * fa/vcs-svn (2012-09-19) 4 commits\n> - vcs-svn: remove repo_tree\n> - vcs-svn/svndump: rewrite handle_node(), begin|end_revision()\n> - vcs-svn/svndump: restructure node_ctx, rev_ctx handling\n> - svndump: move struct definitions to .h\n> (this branch uses fa/remote-svn.)\n> \n> A GSoC project.\n> Waiting for comments from mentors and stakeholders.\n\nThis follow-on topic I'm not so sure on, some of the design decisions make me uncomfortable and I need some convincing before I can get behind this topic. \n\n--\nDavid Michael Barr\n"},{"id":"200499","messageId":"506D5837.6020708@alum.mit.edu","threadId":"31717","inReplyTo":"7vbogj5sji.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2012-10-04T09:34:47Z","receivedAt":"2012-10-04T09:34:47Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 10/03/2012 08:17 PM, Junio C Hamano wrote:\n> Nguyen Thai Ngoc Duy <pclouds@gmail.com> writes:\n> \n>> There's an interesting case: \"**foo\". According to our rules, that\n>> pattern does not contain slashes therefore is basename match. But some\n>> might find that confusing because \"**\" can match slashes,...\n> \n> By \"our rules\", if you mean \"if a pattern has slash, it is anchored\",\n> that obviously need to be updated with this series, if \"**\" is meant\n> to match multiple hierarchies.\n>> I think the latter makes more sense. When users put \"**\" they expect\n>> to match some slashes. But that may call for a refactoring in\n>> path_matches() in attr.c. Putting strstr(pattern, \"**\") in that\n>> matching function may increase overhead unnecessarily.\n>>\n>> The third option is just die() and let users decide either \"*foo\",\n>> \"**/foo\" or \"/**foo\", never \"**foo\".\n> \n> For the double-star at the beginning, you should just turn it into \"**/\"\n> if it is not followed by a slash internally, I think.\n> \n> What is the semantics of ** in the first place?  Is it described to\n> a reasonable level of detail in the documentation updates?  For\n> example does \"**foo\" match \"afoo\", \"a/b/foo\", \"a/bfoo\", \"a/foo/b\",\n> \"a/bfoo/c\"?  Does \"x**y\" match \"xy\", \"xay\", \"xa/by\", \"x/a/y\"?\n> \n> I am guessing that the only sensible definition is that \"**\"\n> requires anything that comes before it (if exists) is at a proper\n> hierarchy boundary, and anything matches it is also at a proper\n> hierarchy boundary, so \"x**y\" matches \"x/a/y\" and not \"xy\", \"xay\",\n> nor \"xa/by\" in the above example.  If \"x**y\" can match \"xy\" or \"xay\"\n> (or \"**foo\" can match \"afoo\"), it would be unreasonable to say it\n> implies the pattern is anchored at any level, no?\n\nGiven that there is no obvious interpretation for what a construct like\n\"x**y\" would mean, and many plausible guesses (most of which sound\nrather useless), I suggest that we forbid it.  This will make the\nfeature easier to explain and make .gitignore files that use it easier\nto understand.\n\nI think that 98% of the usefulness of \"**\" would be in constructs where\nit replaces a proper part of the pathname, like \"**/SOMETHING\" or\n\"SOMETHING/**/SOMETHING\"; in other words, where its use matches the\nregexp \"(^|/)\\*\\*/\".  In these constructs the only ambiguity is whether\n\"**/\" matches regexp\n\n    \"([^/]+/)+\"\n\nor\n\n    \"([^/]+/)*\"\n\n(e.g., whether \"foo/**/bar\" matches \"foo/bar\").  I personally prefer the\nsecond, because the first behavior can be had using the second\ninterpretation by using \"SOMETHING/*/**/SOMETHING\", whereas the second\nbehavior cannot be implemented in terms of the first in a single line of\nthe .gitignore file.\n\nOptionally, one might also like to support \"SOMETHING/**\" or \"**\" alone\nin the obvious ways.\n\nAs for the implementation, it is quite easy to textually convert a glob\npattern, including \"**\" parts, into a regexp.  I happen to have written\nsome Python code that does this for another project (see below).  An\nobvious optimization would be to read any literal parts of the path off\nthe beginning of the glob pattern and only use regexps for the tail\npart.  Would a regexp-based implementation be too slow?\n\nMichael\n\n_filename_char_pattern = r'[^/]'\n_glob_patterns = [\n    ('?', _filename_char_pattern),\n    ('/**', r'(/.+)?'),\n    ('**/', r'(.+/)?'),\n    ('*', _filename_char_pattern + r'*'),\n    ]\n\n\ndef glob_to_regexp(pattern):\n    pattern = os.path.normpath(pattern) # remove trivial redundancies\n\n    if pattern == '**':\n        # This case has to be handled separately because it doesn't\n        # involve a '/' character adjacent to the '**' pattern.  (Such\n        # slashes otherwise have to be considered part of the pattern\n        # to handle the matching of zero path components.)\n        return re.compile(\n            r'^' + _filename_char_pattern + r'(.+' +\n_filename_char_pattern + r')?$'\n            )\n\n    regexp = [r'^']\n    i = 0\n    while i < len(pattern):\n        for (s, r) in _glob_patterns:\n            if pattern.startswith(s, i):\n                regexp.append(r)\n                i += len(s)\n                break\n        else:\n            # AFAIK it's a normal character.  Escape it and add it to\n            # pattern.\n            regexp.append(re.escape(pattern[i]))\n            i += 1\n\n    regexp.append(r'$')\n\n    return re.compile(''.join(regexp))\n\n\n\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"200487","messageId":"CACsJy8DUmjwrkDTePr_8zAU_gcm1kh11J4NVWANMXKsqA6Pb1A@mail.gmail.com","threadId":"31717","inReplyTo":"506D5837.6020708@alum.mit.edu","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-04T11:46:07Z","receivedAt":"2012-10-04T11:46:07Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Oct 4, 2012 at 4:34 PM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\nOn Thu, Oct 4, 2012 at 4:34 PM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> Given that there is no obvious interpretation for what a construct like\n> \"x**y\" would mean, and many plausible guesses (most of which sound\n> rather useless), I suggest that we forbid it.  This will make the\n> feature easier to explain and make .gitignore files that use it easier\n> to understand.\n\nYep, sounds like a good short term plan.\n\n> As for the implementation, it is quite easy to textually convert a glob\n> pattern, including \"**\" parts, into a regexp.\n\nOr we could introduce regexp syntax as an alternative and let users\nchoose (and pay associated price). Patterns starting with // are never\nmatched (we don't normalize paths in .gitignore). Any patterns started\nwith \"//regex:\" is followed by regex. Reject all other // patterns for\nfuture use.\n\n> _filename_char_pattern = r'[^/]'\n> _glob_patterns = [\n>     ('?', _filename_char_pattern),\n>     ('/**', r'(/.+)?'),\n>     ('**/', r'(.+/)?'),\n>     ('*', _filename_char_pattern + r'*'),\n>     ]\n\nI don't fully understand the rest (never been a big fan of python) but\nwhat about bracket expressions like [!abc] and [:alnum:]?\n-- \nDuy\n"},{"id":"200514","messageId":"506DA8A4.5080105@alum.mit.edu","threadId":"31717","inReplyTo":"CACsJy8DUmjwrkDTePr_8zAU_gcm1kh11J4NVWANMXKsqA6Pb1A@mail.gmail.com","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2012-10-04T15:17:56Z","receivedAt":"2012-10-04T15:17:56Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 10/04/2012 01:46 PM, Nguyen Thai Ngoc Duy wrote:\n> On Thu, Oct 4, 2012 at 4:34 PM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> As for the implementation, it is quite easy to textually convert a glob\n>> pattern, including \"**\" parts, into a regexp.\n> \n> Or we could introduce regexp syntax as an alternative and let users\n> choose (and pay associated price).\n\nIt seems like overkill to me.  For filenames, globs are usually adequate.\n\n>> _filename_char_pattern = r'[^/]'\n>> _glob_patterns = [\n>>     ('?', _filename_char_pattern),\n>>     ('/**', r'(/.+)?'),\n>>     ('**/', r'(.+/)?'),\n>>     ('*', _filename_char_pattern + r'*'),\n>>     ]\n> \n> I don't fully understand the rest (never been a big fan of python) but\n> what about bracket expressions like [!abc] and [:alnum:]?\n\nYou're right; I forgot that the code that I posted doesn't support brackets.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"200503","messageId":"7vwqz619uu.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"A4A111D1488E49FFA4D71D85DD6B87A4@rr-dav.id.au","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-04T16:27:05Z","receivedAt":"2012-10-04T16:27:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Michael Barr <b@rr-dav.id.au> writes:\n\n> On Wednesday, 3 October 2012 at 9:20 AM, Junio C Hamano wrote: \n>> \n>> * fa/remote-svn (2012-09-19) 16 commits\n>> ...\n>> \n>> A GSoC project.\n>> Waiting for comments from mentors and stakeholders.\n>\n> I have reviewed this topic and am happy with the design and implementation.\n> I support this topic for inclusion.\n>\n> Acked-by: David Michael Barr <b@rr-dav.id.au>\n>> \n>> * fa/vcs-svn (2012-09-19) 4 commits\n>> ...\n>\n> This follow-on topic I'm not so sure on, some of the design\n> decisions make me uncomfortable and I need some convincing before\n> I can get behind this topic.\n\nThanks for a feedback.\n"},{"id":"200532","messageId":"7vobki19ax.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"506D5837.6020708@alum.mit.edu","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-04T16:39:02Z","receivedAt":"2012-10-04T16:39:02Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu> writes:\n\n> On 10/03/2012 08:17 PM, Junio C Hamano wrote:\n>> \n>> What is the semantics of ** in the first place?  Is it described to\n>> a reasonable level of detail in the documentation updates?  For\n>> example does \"**foo\" match \"afoo\", \"a/b/foo\", \"a/bfoo\", \"a/foo/b\",\n>> \"a/bfoo/c\"?  Does \"x**y\" match \"xy\", \"xay\", \"xa/by\", \"x/a/y\"?\n>> \n>> I am guessing that the only sensible definition is that \"**\"\n>> requires anything that comes before it (if exists) is at a proper\n>> hierarchy boundary, and anything matches it is also at a proper\n>> hierarchy boundary, so \"x**y\" matches \"x/a/y\" and not \"xy\", \"xay\",\n>> nor \"xa/by\" in the above example.  If \"x**y\" can match \"xy\" or \"xay\"\n>> (or \"**foo\" can match \"afoo\"), it would be unreasonable to say it\n>> implies the pattern is anchored at any level, no?\n>\n> Given that there is no obvious interpretation for what a construct like\n> \"x**y\" would mean, and many plausible guesses (most of which sound\n> rather useless), I suggest that we forbid it.  This will make the\n> feature easier to explain and make .gitignore files that use it easier\n> to understand.\n>\n> I think that 98% of the usefulness of \"**\" would be in constructs where\n> it replaces a proper part of the pathname, like \"**/SOMETHING\" or\n> \"SOMETHING/**/SOMETHING\"...\n\nI think it is a good way to go in the longer term, if we all agree\nthat \"**\" matching anything does not give us a useful semantics\n[*1*].\n\nIs it something we can easily get by simple patch into the wildmatch\ncode?  I'd hate to see us parsing the input and validating it before\npassing it to the library, as we will surely botch the quoting or\nsomething while doing so.\n\nWhen we require \"x/**/y\", I think we still want it to match \"x/y\".\nDo people agree, or are there good reasons to require at least one\nlevel between x and y for such a pattern?  Assuming that we do want\nto match \"x/y\" with \"x/**/y\", I suspect that \"'**' matches anything\nincluding a slash\" would not give us that semantics. Is it something\nwe can easily fix in the wildmatch code?\n\n\n[Footnote]\n\n*1* The message you are responding to was written in a somewhat\nprovocative way on purpose so that people who like the way rsync\nmatches \"**\" can vocally object. I would like to see arguments from\nthe both sides to see if it makes sense.\n"},{"id":"200538","messageId":"7vwqz6yvyj.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"1349336392-1772-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 0/6] wildmatch part 2","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-04T17:43:16Z","receivedAt":"2012-10-04T17:43:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy <pclouds@gmail.com> writes:\n\n> On Thu, Oct 4, 2012 at 1:01 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>> Perhaps the wildmatch code may not be what we want X-<.\n>\n> When I imported wildmatch I was hoping to make minimum changes to it.\n> But wildmatch is probably the only practical way to support \"**\" even\n> if we later need to change it the way we want. Other options are base\n> our work on top of compat/fnmatch.c, which is an #ifdef spaghetti\n> mess, or write a new fnmatch()-compatible function. Both unattractive\n> to me.\n>\n> Anyway, this is on top of nd/wildmatch, which makes \"ab**cd\" match\n> full pathname.\n\nI do not think we are in a hurry to push \"**\" support in before we\nknow what semantics we want to get out of it.  Pushing half-baked\n\"this is good enough at least to me for now\" topics before they are\nready will cost the users in the longer term.\n\nOn the other hand, the three patches (2/3/4) in this series look\nlike a good improvement regardless of what kind of matching engine\nwe use.  I would have preferred to see them _before_ nd/wildmatch.\n\nI do not agree with the reasoning behind [1/6] that changes\n\n\tunion {\n\t\tchar *pattern;\n\t\tstrict git_attr *attr;\n\t} u;\n\tchar is_macro;\n\nto\n\n\tchar *pattern;\n\tstrict git_attr *attr;\n\tchar is_macro;\n\nby the way.\n\nThe union is much less about space saving but is more about the\nnature of the usage; we use pattern or attr but not both at the same\ntime.  Even if the pattern becomes richer with later patches, that\ndoes not change the fundamental premise that depending on the value\nof \"is_macro\", either \"attr\" is used or \"pattern\" is used but they\nwon't be in effect at the same time.  The evolution of this should\ngo more like this, I think:\n\n\tunion {\n\t\tstruct {\n\t\t\tconst char *string;\n\t\t\tint pattern_length;\n\t\t\tint prefix_literal_length;\n\t\t\tint flags;\n\t\t} pattern;\n\t\tstruct git_attr *attr;\n\t} u;\n\tchar is_macro;\n\nThanks.\n"},{"id":"200553","messageId":"7vsj9uyv6y.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"1349336392-1772-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 5/6] gitignore: do not do basename match with patterns that have '**'","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-04T17:59:49Z","receivedAt":"2012-10-04T17:59:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy <pclouds@gmail.com> writes:\n\n> -\t\tif (!strchr(p, '/'))\n> +\t\tif (!strchr(p, '/') && !strstr(p, \"**\"))\n\nDoesn't wildmatch allow these to be quoted, similar to the way usual\nglob works, e.g.\n\n\t$ >ff\n        $ >\\?f\n        $ echo ??\n        ?f ff\n        $ echo \\?f\n        ?f\n\nEven If wildmatch out-of-the-box doesn't, I would assume that we\nwould want to fix it so that it does.  And if that is the case,\nwe would want to be careful about \"two/asterisks\\**in path\" to\navoid triggering this logic, no?\n"},{"id":"200567","messageId":"7vobkiyuzt.fsf@alter.siamese.dyndns.org","threadId":"31717","inReplyTo":"1349336392-1772-7-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 6/6] t3001: note about expected \"**\" behavior","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-04T18:04:06Z","receivedAt":"2012-10-04T18:04:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy <pclouds@gmail.com> writes:\n\n> \"**\" currently matches any characters including slashes. It's probably\n> too powerful. A more sensible definition may be match any characters\n> that the but the whole match must be wrapped by slashes. So \"**\" can\n> match none, \"/\", \"/aaa/\", \"/aa/bb/\" and so on but not \"aa/bb\".\n\nI do not think this is something we want to retroactively change\nafter releasing it to the public, especially when we _know_ it is a\nproblem from the get-go (unlike the case we did not notice it had a\nproblem, release it to the public and then realize it and have to\nscramble to devise a fix to bring more sanity in a backward\ncompatible way).\n\nWe must either declare that \"**\" that matches any characters is the\nsane semantics and promise we will never change it, or have \"**\"\nthat matches \\(/[^/]*\\)*/ (sorry for a line noise^W^W^Wregexp) from\nthe beginning.\n"},{"id":"200597","messageId":"506E85BF.8010302@viscovery.net","threadId":"31717","inReplyTo":"1349336392-1772-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH 5/6] gitignore: do not do basename match with patterns that have '**'","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2012-10-05T07:01:19Z","receivedAt":"2012-10-05T07:01:19Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 10/4/2012 9:39, schrieb Nguyễn Thái Ngọc Duy:\n> - - If the pattern does not contain a slash '/', git treats it as\n> -   a shell glob pattern and checks for a match against the\n> -   pathname relative to the location of the `.gitignore` file\n> -   (relative to the toplevel of the work tree if not from a\n> -   `.gitignore` file).\n> + - If the pattern does not contain a slash '/' nor '**', git\n> +   treats it as a shell glob pattern and checks for a match\n> +   against the pathname relative to the location of the\n> +   `.gitignore` file (relative to the toplevel of the work tree\n> +   if not from a `.gitignore` file).\n\n> +test_expect_success '\"**\" with no slashes test' '\n> +\techo \"a**f foo=bar\" >.gitattributes &&\n> +\tcat <<\\EOF >expect &&\n> +f: foo: unspecified\n> +a/f: foo: bar\n> +a/b/f: foo: bar\n> +a/b/c/f: foo: bar\n> +EOF\n\nShould the above .gitattributes match nested paths, such as b/a/c/f?\n\nI think it should, because the user can easily say \"/a**f\" that nested\npaths should not be matched.\n\nBut if it does not match, as your documentation update implies, which\noptions does the user have to match nested paths? Only to add more\npatterns for each nested directory, such as \"b/a**f\".\n\n-- Hannes\n"},{"id":"200619","messageId":"CACsJy8BTWEeWvdwCDGBdoLndh8hXpgCgXrQxhWeaO1m9Qrqvgw@mail.gmail.com","threadId":"31717","inReplyTo":"506E85BF.8010302@viscovery.net","subject":"Re: [PATCH 5/6] gitignore: do not do basename match with patterns that have '**'","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-05T11:23:23Z","receivedAt":"2012-10-05T11:23:23Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Oct 5, 2012 at 2:01 PM, Johannes Sixt <j.sixt@viscovery.net> wrote:\n> Am 10/4/2012 9:39, schrieb Nguyễn Thái Ngọc Duy:\n>> - - If the pattern does not contain a slash '/', git treats it as\n>> -   a shell glob pattern and checks for a match against the\n>> -   pathname relative to the location of the `.gitignore` file\n>> -   (relative to the toplevel of the work tree if not from a\n>> -   `.gitignore` file).\n>> + - If the pattern does not contain a slash '/' nor '**', git\n>> +   treats it as a shell glob pattern and checks for a match\n>> +   against the pathname relative to the location of the\n>> +   `.gitignore` file (relative to the toplevel of the work tree\n>> +   if not from a `.gitignore` file).\n\nI think in the latest round, we forbid this case (i.e. a/**, **/b and\na/**/b are ok, but a**b is not), exactly because it's hard to define\nhow it should do. Thanks for another example.\n\n>> +test_expect_success '\"**\" with no slashes test' '\n>> +     echo \"a**f foo=bar\" >.gitattributes &&\n>> +     cat <<\\EOF >expect &&\n>> +f: foo: unspecified\n>> +a/f: foo: bar\n>> +a/b/f: foo: bar\n>> +a/b/c/f: foo: bar\n>> +EOF\n>\n> Should the above .gitattributes match nested paths, such as b/a/c/f?\n>\n> I think it should, because the user can easily say \"/a**f\" that nested\n> paths should not be matched.\n\nThe user can also say **/a/**f to match b/a/c/f.\n-- \nDuy\n"},{"id":"200620","messageId":"m2391t1589.fsf@igel.home","threadId":"31717","inReplyTo":"7vobki19ax.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2012-10-05T12:19:18Z","receivedAt":"2012-10-05T12:19:18Z","isPatch":false,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> When we require \"x/**/y\", I think we still want it to match \"x/y\".\n\nFWIW, in bash (+extglob), ksh and zsh it doesn't.\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5\n\"And now for something completely different.\"\n"},{"id":"200621","messageId":"vpqpq4x14ox.fsf@grenoble-inp.fr","threadId":"31717","inReplyTo":"m2391t1589.fsf@igel.home","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@grenoble-inp.fr","sentAt":"2012-10-05T12:30:54Z","receivedAt":"2012-10-05T12:30:54Z","isPatch":false,"sender":{"key":"matthieu.moy@grenoble-inp.fr","avatar":"https://gravatar.com/avatar/72c8a2705971a25dfaff23cece15130d405685845d911aedd5667ace277f3fc5?d=mp&s=160"},"body":"Andreas Schwab <schwab@linux-m68k.org> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> When we require \"x/**/y\", I think we still want it to match \"x/y\".\n>\n> FWIW, in bash (+extglob), ksh and zsh it doesn't.\n\nYou're right about bash, but I see the opposite for zsh and ksh:\n\nzsh$ echo x/**/y\nx/y x/z/y\n\nksh$ echo x/**/y\nx/y x/z/y\n\n(didn't check the doc so see whether this was configurable, but I've set\nHOME=/ when launching the shell to disable my own configuration)\n\n-- \nMatthieu Moy\nhttp://www-verimag.imag.fr/~moy/\n"},{"id":"200622","messageId":"20121005132132.GA13591@do","threadId":"31717","inReplyTo":"7vobki19ax.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-05T13:21:32Z","receivedAt":"2012-10-05T13:21:32Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Oct 04, 2012 at 09:39:02AM -0700, Junio C Hamano wrote:\n> Assuming that we do want to match \"x/y\" with \"x/**/y\", I suspect\n> that \"'**' matches anything including a slash\" would not give us\n> that semantics. Is it something we can easily fix in the wildmatch\n> code?\n\nSomething like this may suffice. Lightly tested with \"git add -n\".\nReading the code, I think we can even distinguish \"match zero or more\ndirectories\" and \"match one or more directories\" with \"/**/\" and maybe\n\"/***/\". Right now **, ***, ****... are the same. So are /**/, /***/,\n/****/...\n\n-- 8< --\ndiff --git a/wildmatch.c b/wildmatch.c\nindex f153f8a..81eadc8 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -98,8 +98,12 @@ static int dowild(const uchar *p, const uchar *text,\n \t    continue;\n \t  case '*':\n \t    if (*++p == '*') {\n+\t\tint slashstarstar = p[-2] == '/';\n \t\twhile (*++p == '*') {}\n \t\tspecial = TRUE;\n+\t\tif (slashstarstar && *p == '/' &&\n+\t\t    dowild(p + 1, text, a, force_lower_case) == TRUE)\n+\t\t    return TRUE;\n \t    } else\n \t\tspecial = FALSE;\n \t    if (*p == '\\0') {\n-- 8< --\n-- \nDuy\n"},{"id":"200626","messageId":"m2y5jlyph0.fsf@igel.home","threadId":"31717","inReplyTo":"vpqpq4x14ox.fsf@grenoble-inp.fr","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2012-10-05T14:15:39Z","receivedAt":"2012-10-05T14:15:39Z","isPatch":false,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"Matthieu Moy <Matthieu.Moy@grenoble-inp.fr> writes:\n\n> Andreas Schwab <schwab@linux-m68k.org> writes:\n>\n>> Junio C Hamano <gitster@pobox.com> writes:\n>>\n>>> When we require \"x/**/y\", I think we still want it to match \"x/y\".\n>>\n>> FWIW, in bash (+extglob), ksh and zsh it doesn't.\n>\n> You're right about bash, but I see the opposite for zsh and ksh:\n>\n> zsh$ echo x/**/y\n> x/y x/z/y\n>\n> ksh$ echo x/**/y\n> x/y x/z/y\n\nLooks like this is different between filename expansion and case pattern\nmatching (I only tested the latter).\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5\n\"And now for something completely different.\"\n"},{"id":"202201","messageId":"1674207.s6eW8JjC7x@flomedio","threadId":"31717","inReplyTo":"7vmx045umh.fsf@alter.siamese.dyndns.org","subject":"Re: What's cooking in git.git (Oct 2012, #01; Tue, 2)","fromName":"Florian Achleitner","fromEmail":"florian.achleitner2.6.31@gmail.com","sentAt":"2012-10-30T12:15:18Z","receivedAt":"2012-10-30T12:15:18Z","isPatch":false,"sender":{"key":"florian.achleitner.2.6.31@gmail.com","avatar":"https://avatars.githubusercontent.com/u/880777?v=4"},"body":"Sorry for reacting so late, I didn't read the list carefully in the last weeks \nand my gmail filter somehow didn't trigger on that.\n\nOn Tuesday 02 October 2012 16:20:22 Junio C Hamano wrote:\n> * fa/remote-svn (2012-09-19) 16 commits\n>  - Add a test script for remote-svn\n>  - remote-svn: add marks-file regeneration\n>  - Add a svnrdump-simulator replaying a dump file for testing\n>  - remote-svn: add incremental import\n>  - remote-svn: Activate import/export-marks for fast-import\n>  - Create a note for every imported commit containing svn metadata\n>  - vcs-svn: add fast_export_note to create notes\n>  - Allow reading svn dumps from files via file:// urls\n>  - remote-svn, vcs-svn: Enable fetching to private refs\n>  - When debug==1, start fast-import with \"--stats\" instead of \"--quiet\"\n>  - Add documentation for the 'bidi-import' capability of remote-helpers\n>  - Connect fast-import to the remote-helper via pipe, adding 'bidi-import'\n> capability - Add argv_array_detach and argv_array_free_detached\n>  - Add svndump_init_fd to allow reading dumps from arbitrary FDs\n>  - Add git-remote-testsvn to Makefile\n>  - Implement a remote helper for svn in C\n>  (this branch is used by fa/vcs-svn.)\n> \n>  A GSoC project.\n>  Waiting for comments from mentors and stakeholders.\n\n>From my point of view, this is rather complete. It got eight review cycles on \nthe list.\nNote that the remote helper can only fetch, pushing is not possible at all.\n\n> \n> \n> * fa/vcs-svn (2012-09-19) 4 commits\n>  - vcs-svn: remove repo_tree\n>  - vcs-svn/svndump: rewrite handle_node(), begin|end_revision()\n>  - vcs-svn/svndump: restructure node_ctx, rev_ctx handling\n>  - svndump: move struct definitions to .h\n>  (this branch uses fa/remote-svn.)\n> \n>  A GSoC project.\n>  Waiting for comments from mentors and stakeholders.\n\nThis is the result of what I did when I wanted to start implementing branch \ndetection. I found that the existing code is not suitable and restructured it.\n\nThe main goal is to seperate svn revision parsing from git commit creation. \nBecause for creating commits, you need to know on which branch to create the \ncommit.\nWhile for finding out which branch is the right one, you need to read the \ncomplete svn revision first to see what dirs are changed and how.\n\nIt is rather invasive and it doesn't make sense without using it later on.\nSo I'm not surprised that you may not like it.\nAnyways it passes all existing tests (that doesn't mean it's good of course \n;))\n\nFlorian\n"}]}