{"thread":{"id":"64525","subject":"Filter smudge for secret restoration: no disk access?","startedAt":"2025-11-24T07:39:24Z","lastAt":"2025-11-25T08:55:26Z","messageCount":7,"participants":["Kache Hit","Johannes Sixt","Chris Torek","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"531205","messageId":"DEGR5XSM0EVG.27IMOKOK1O98Y@gmail.com","threadId":"64525","inReplyTo":null,"subject":"Filter smudge for secret restoration: no disk access?","fromName":"Kache Hit","fromEmail":"kache.hit@gmail.com","sentAt":"2025-11-24T07:39:22Z","receivedAt":"2025-11-24T07:39:24Z","isPatch":false,"sender":{"key":"kache.hit@gmail.com","avatar":null},"body":"I was working on a git redaction script that restores working copy\nsecrets when applied via `.gitattributes` clean/smudge filters, but\nencountered `smudge` not having access to the \"working file\" on disk.\n\nI see it's documented as intended in\nhttps://git-scm.com/docs/gitattributes:\n\n> Note that \"%f\" is the name of the path that is being worked on.\n> Depending on the version that is being filtered, the corresponding\n> file on disk may not exist, or may have different contents. So, smudge\n> and clean commands should not try to access the file on disk, but only\n> act as filters on the content provided to them on standard input.\n\nAny chance there's a way around this or some alternative? Python\nimplementation below for reference.\n\nAnd also for my understanding, why _shouldn't_ smudge access disk?\n\n```py\n#!/usr/bin/env python3\n\"\"\"\nGit clean/smudge filter for redactions that retains working secrets\n\nIf the following is in the repo as `bar/foo_secrets.yml`:\n```\n    foo_token: ##REDACTED##\n    other: \"not secret\"\n```\n\nThe local token won't be overwritten on checkout/restore:\n```\n    foo_token: secret_value\n    other: \"not secret\"\n```\n\nSetup & example usage:\n\nSave this file in repo root as `git_redact_filter.py`\n\n`.gitattributes`:\n```\nbar/foo_secrets.yml filter=foo_token\n```\n\n`.gitconfig`:\n```\n[filter \"foo_token\"]\n  clean = ./git_redact_filter.py --prefix foo_token:\n  smudge = ./git_redact_filter.py --prefix foo_token: --smudge %f\n```\n\"\"\"\nimport inspect\nimport re\nimport sys\nfrom argparse import ArgumentParser\nfrom pathlib import Path\nfrom typing import TextIO\n\nREDACTED = '##REDACTED##'\n\n\ndef clean(workfile: TextIO, prefixes: list[str], out=None):\n    pat = prefix_secret_rgx(prefixes)\n\n    for line in workfile.readlines():\n        if match := pat.match(line):\n            print(match['prefix'] + REDACTED, file=out)\n        else:\n            print(line, end='', file=out)\n\n\ndef smudge(repofile: TextIO, prefixes: list[str], path: Path, out=None):\n    pat = prefix_secret_rgx(prefixes)\n\n    with path.open() as workfile:  # fails: FileNotFoundError\n        secrets = {\n            str(match['prefix']): match\n            for match in map(pat.match, workfile.readlines())\n            if match\n        }\n\n    for line in repofile.readlines():\n        match = pat.match(line)\n        secret = match and secrets.get(match['prefix'])\n\n        if match and secret and match['secret'] == REDACTED:\n            print(match['prefix'] + secret['secret'], file=out)\n        else:\n            print(line, end='', file=out)\n\n\ndef prefix_secret_rgx(prefixes_unsafe: list[str]):\n    keys = '|'.join(map(re.escape, prefixes_unsafe))\n    pat = rf\"(?P<prefix>\\s*({keys})\\s*)(?P<secret>.*)\"\n    return re.compile(pat if keys else r'$^')\n\n\ndef heredoc(s: str):\n    return inspect.cleandoc(s) + '\\n'\n\n\ndef main():\n    desc = \"Git clean/smudge filter for redactions\"\n    list_arg = {'action': 'append', 'default': []}\n    parser = ArgumentParser(description=desc)\n    parser.add_argument('-p', '--prefix', **list_arg, metavar='PREFIX')\n    parser.add_argument('--smudge', type=Path, metavar='PATH')\n    args = parser.parse_args()\n\n    if args.smudge:\n        return smudge(sys.stdin, args.prefix, args.smudge)\n    else:\n        return clean(sys.stdin, args.prefix)\n\n\nif __name__ == '__main__':\n    sys.exit(main())\n\n\nimport io\nfrom unittest.mock import Mock\n\nimport pytest\nfrom pytest import CaptureFixture\n\n\nwork_file = io.StringIO(heredoc(\"\"\"\n    foo_token: secret_value\n    other: \"not secret\"\n\"\"\"))\nclean_file = io.StringIO(heredoc(\"\"\"\n    foo_token: ##REDACTED##\n    other: \"not secret\"\n\"\"\"))\n\nempty_file = io.StringIO()\nwork_file_secret_removed = io.StringIO(heredoc(\"\"\"\n    other: \"not secret\"\n\"\"\"))\nwork_file_lines_added = io.StringIO(heredoc(\"\"\"\n    new_other: 123\n    foo_token: secret_value\n    other: \"not secret\"\n\"\"\"))\n\n\ndef test_clean(capsys: CaptureFixture):\n    clean(work_file, ['foo_token:'])\n    captured = capsys.readouterr()\n    assert captured.out == clean_file.getvalue(), \"should be redacted\"\n\n\ndef test_clean_idempotent():\n    out, out2 = io.StringIO(), io.StringIO()\n    clean(work_file, ['foo_token:'], out)\n    clean(io.StringIO(out.getvalue()), ['foo_token:'], out2)\n    assert out2.getvalue() == clean_file.getvalue()\n\n\n@pytest.mark.parametrize(['workfile', 'expected', 'msg'], [\n    (work_file,                work_file,  \"secrets should be kept\"),\n    (work_file_lines_added,    work_file,  \"should retain secret\"),\n    (work_file_secret_removed, clean_file, \"should restore redacted\"),\n])\ndef test_smudge_goal(capsys: CaptureFixture, workfile, expected, msg):\n    path = Mock()\n    path.open.side_effect = lambda: io.StringIO(workfile.getvalue())\n\n    smudge(clean_file, ['foo_token:'], path)\n    captured = capsys.readouterr()\n    assert captured.out == expected.getvalue(), msg\n\n\ndef test_smudge_idempotent():\n    path = Mock()\n    path.open.side_effect = lambda: io.StringIO(work_file.getvalue())\n    cleaned, cleaned2 = io.StringIO(), io.StringIO()\n\n    smudge(clean_file, ['foo_token:'], path, cleaned)\n    cleaned.seek(0)\n    smudge(cleaned, ['foo_token:'], path, cleaned2)\n    assert cleaned.getvalue() == cleaned2.getvalue()\n\n\ngit_doc_url = \"https://git-scm.com/docs/gitattributes\"\n@pytest.mark.xfail(reason=f\"should access file on disk: {git_doc_url}\")\n@pytest.mark.parametrize(['workfile', 'expected', 'msg'], [\n    (work_file,                work_file,  \"secrets should be kept\"),\n    (work_file_lines_added,    work_file,  \"should retain secret\"),\n    (work_file_secret_removed, clean_file, \"should restore redacted\"),\n])\ndef test_smudge_actual(capsys: CaptureFixture, workfile, expected, msg):\n    msg = \"[Errno 2] No such file or directory: 'bar/foo_secrets.yml'\"\n    err = FileNotFoundError(msg)\n    mock_workfile_path = Mock()\n    mock_workfile_path.open.side_effect = err\n\n    smudge(clean_file, ['foo_token:'], mock_workfile_path)\n    captured = capsys.readouterr()\n    assert captured.out == expected.getvalue(), msg\n\n\n@pytest.fixture(autouse=True)\ndef reset_files():\n    for file in [work_file, clean_file]:\n        file.seek(0)\n```\n\nThanks,\n\nKache\n"},{"id":"531207","messageId":"9aa7cfdb-fc50-4ceb-936c-2ed441c462a3@kdbg.org","threadId":"64525","inReplyTo":"DEGR5XSM0EVG.27IMOKOK1O98Y@gmail.com","subject":"Re: Filter smudge for secret restoration: no disk access?","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2025-11-24T09:01:19Z","receivedAt":"2025-11-24T09:01:36Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 24.11.25 um 08:39 schrieb Kache Hit:\n> I was working on a git redaction script that restores working copy\n> secrets when applied via `.gitattributes` clean/smudge filters, but\n> encountered `smudge` not having access to the \"working file\" on disk.\n> \n> I see it's documented as intended in\n> https://git-scm.com/docs/gitattributes:\n> \n>> Note that \"%f\" is the name of the path that is being worked on.\n>> Depending on the version that is being filtered, the corresponding\n>> file on disk may not exist, or may have different contents. So, smudge\n>> and clean commands should not try to access the file on disk, but only\n>> act as filters on the content provided to them on standard input.\n> \n> Any chance there's a way around this or some alternative? Python\n> implementation below for reference.\n> \n> And also for my understanding, why _shouldn't_ smudge access disk?\n\nA smudge filter must read its stdin and write the result to stdout. The\npresence of %f in the configuration does not change this.\n\nThe filter can inspect the file name it receives via the %f token (note:\nthe *name* of the file, not the file itself) to draw additional hints\nhow to process the data, but it still has to read stdin and write to stdout.\n\n-- Hannes\n\n"},{"id":"531208","messageId":"CAPx1GvcXkXMpWgOyMWdfHXGEDJQY4wJrJV0p7LHBMeQFPMDHnQ@mail.gmail.com","threadId":"64525","inReplyTo":"9aa7cfdb-fc50-4ceb-936c-2ed441c462a3@kdbg.org","subject":"Re: Filter smudge for secret restoration: no disk access?","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2025-11-24T09:49:58Z","receivedAt":"2025-11-24T09:50:12Z","isPatch":false,"sender":{"key":"chris.torek@gmail.com","avatar":"https://avatars.githubusercontent.com/u/16826774?v=4"},"body":"On Mon, Nov 24, 2025 at 1:01 AM Johannes Sixt <j6t@kdbg.org> wrote:\n> The filter can inspect the file name it receives via the %f token (note:\n> the *name* of the file, not the file itself) to draw additional hints\n> how to process the data, but it still has to read stdin and write to stdout.\n\nIt can, of course, also read and/or write anything else on disk.\n\nWhen and how this is actually useful is another matter entirely.\n\nFor sanity purposes, if no other reasons, it might be wise to store a\n\"file with secrets\" under a file with a name such that it is **never**\ncontrolled by Git (i.e., always listed in a .gitignore or equivalent,\nor outside the working tree entirely), and to store instead, in Git, a\n\"template file with secrets that are replaced\". That way, the secrets\neither exist on disk (and are secret because Git is blind to them), or\ndo not exist at all (and are therefore secret to Git). The template\nfile controls the template and nothing else; the secret-data file has\nboth secrets and, perhaps, data that are extracted from the\nGit-controlled file as well.\n\nIn this manner, a \"to-be-smudged\" file named foo.template might\ncontrol some external-to-Git manipulation of an invisible-to-Gt file\nnamed foo.secret, and no clean filter would be required at all, though\none could inspect and strip secrets accidentally copied into a\nfoo.template.\n\nChris\n"},{"id":"531234","messageId":"DEH58DEF5MGO.2CFIKCM2CAQY2@gmail.com","threadId":"64525","inReplyTo":"CAPx1GvcXkXMpWgOyMWdfHXGEDJQY4wJrJV0p7LHBMeQFPMDHnQ@mail.gmail.com","subject":"Re: Filter smudge for secret restoration: no disk access?","fromName":"Kache Hit","fromEmail":"kache.hit@gmail.com","sentAt":"2025-11-24T18:40:49Z","receivedAt":"2025-11-24T18:40:50Z","isPatch":false,"sender":{"key":"kache.hit@gmail.com","avatar":null},"body":"On Mon Nov 24, 2025 at 1:01 AM PST, Johannes Sixt wrote:\n> A smudge filter must read its stdin and write the result to stdout. The\n> presence of %f in the configuration does not change this.\n>\n> The filter can inspect the file name it receives via the %f token (note:\n> the *name* of the file, not the file itself) to draw additional hints\n> how to process the data, but it still has to read stdin and write to stdout.\n\nYes, I underststand. I'm asking why it's necessary that smudge not read\nfrom disk, even as it properly satisfies that stdin/stdout operation, as\nin my Python implementation of `smudge()`\n\nOn Mon Nov 24, 2025 at 1:49 AM PST, Chris Torek wrote:\n> For sanity purposes, if no other reasons, it might be wise to store a\n> \"file with secrets\" under a file with a name such that it is **never**\n> controlled by Git (i.e., always listed in a .gitignore or equivalent,\n> or outside the working tree entirely), and to store instead, in Git, a\n> \"template file with secrets that are replaced\". That way, the secrets\n> either exist on disk (and are secret because Git is blind to them), or\n> do not exist at all (and are therefore secret to Git). The template\n> file controls the template and nothing else; the secret-data file has\n> both secrets and, perhaps, data that are extracted from the\n> Git-controlled file as well.\n>\n> In this manner, a \"to-be-smudged\" file named foo.template might\n> control some external-to-Git manipulation of an invisible-to-Gt file\n> named foo.secret, and no clean filter would be required at all, though\n> one could inspect and strip secrets accidentally copied into a\n> foo.template.\n\nI'm familiar with this practice, e.g. committing an `.env.template`\nwhich is used to create an `.env` file with secrets within.\n\nHowever, this is my dotfiles repo that includes `~/.config`. There are\nconfig files that store credentials right next to configuration, managed\nby software that I don't control.\n\nAlthough I could still apply that pattern by ignoring `foo.yml` and\ncommitting a redacted `foo.template.yml`, I'd have to manually upstream\nchanges back to the template as the config file changes.\n\nAnother use case is to ignore changes to a specific line without losing\nthe working copy. Some software saves a volatile \"last_updated_at\" or\n\"last_opened\" field into config that doesn't need to be committed. This\ncould also be useful for https://stackoverflow.com/questions/16244969\nand https://stackoverflow.com/questions/61091219\n\n- Kache\n"},{"id":"531235","messageId":"xmqqms4bw7f7.fsf@gitster.g","threadId":"64525","inReplyTo":"DEH58DEF5MGO.2CFIKCM2CAQY2@gmail.com","subject":"Re: Filter smudge for secret restoration: no disk access?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-24T19:35:08Z","receivedAt":"2025-11-24T19:35:11Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Kache Hit\" <kache.hit@gmail.com> writes:\n\n> On Mon Nov 24, 2025 at 1:01 AM PST, Johannes Sixt wrote:\n>> A smudge filter must read its stdin and write the result to stdout. The\n>> presence of %f in the configuration does not change this.\n>>\n>> The filter can inspect the file name it receives via the %f token (note:\n>> the *name* of the file, not the file itself) to draw additional hints\n>> how to process the data, but it still has to read stdin and write to stdout.\n>\n> Yes, I underststand. I'm asking why it's necessary that smudge not read\n> from disk, even as it properly satisfies that stdin/stdout operation, as\n> in my Python implementation of `smudge()`\n\nI do not think it is a total dogmatic prohibition, but is a\npractical piece of advice to be prepared in a situation where the\nfile %f does not exist on the disk in the working tree.  Also even\nwhen the file %f does exist, its contents would not match (because\nit was smudged when it was checked out, and the user may have\nfurther modified it) what in the tree of the commit you are\nswitching out of.\n\nSuppose you added a path F and G with a SAME smudge/clean filter\npair to the history at commit X.  You check out a commit before that\nhappened:\n\n\t$ git checkout -b practice X~1\n\nand then try to come back to commit after X:\n\n\t$ git checkout X\n\nGit would read the cleaned contents of blobs X:F and X:G, invokes\nyour smudge filter once for each of these blobs, and feeds the blob\ncontents to it.  Your smudge filter learns in its one of the two\ninvocations that it is being handed the clean contents and it is\nexpected to smudge it for path F via %f, and then the other\ninvocation of the same smudge filter is told that it is now being\nasked to smudge for path G.\n\nIf F or G exists on the disk, surely, the smudge filter can read it,\nbut in this situation, because you are coming from X~1 before F and\nG appeared in the history, these files are not on disk in your\nworking tree.\n\nThe smudge filter needs to be careful about a similar situation\nwhere commit Y that is a descendant of X modifies F and/or G.  When\nY is checked out and you want to switch to X, working tree may have\nsmudged versions of F and G from Y when your smudge filter is\ncalled.  Or it may happen during a checkout of F or G, and one of\nthe things the checkout needs to do may be to remove the existing\nfile from the working tree, and then create a file anew (probably in\na temporary file) and move it to the final place, in which case,\nyour smudge filter may be called during \"create a file anew\" phase,\nwhere the old file F or G may be missing from the working tree.\nEven if F and G are there, it may be from commit Y and their\ncontents may have nothing to do with the version of the files your\nsmudge filter is trying to turn the clean blob data taken from\ncommit X.\n\nThe note from the \"git help attributes\" you cited summarizes the\nadvice concisely.\n\n    Note that \"%f\" is the name of the path that is being worked on. Depending\n    on the version that is being filtered, the corresponding file on disk may\n    not exist, or may have different contents. So, smudge and clean commands\n    should not try to access the file on disk, but only act as filters on the\n    content provided to them on standard input.\n\nThe smudge filter needs to be prepared to work in such scenarios.\n\nPerhaps \"Depending on ...\" talks too much without giving readers\nenough benefit.  A shorter description like this one ...\n\n    Note that the purpose of %f is to tell the filter for what output\n    path it is asked to smudge the clean blob data, and should not be\n    used for anything else.\n\n... may be less confusing, perhaps?\n\n"},{"id":"531255","messageId":"DEHLKBB96BBI.3V74A5NGVTVZA@gmail.com","threadId":"64525","inReplyTo":"xmqqms4bw7f7.fsf@gitster.g","subject":"Re: Filter smudge for secret restoration: no disk access?","fromName":"Kache Hit","fromEmail":"kache.hit@gmail.com","sentAt":"2025-11-25T07:28:42Z","receivedAt":"2025-11-25T07:28:44Z","isPatch":false,"sender":{"key":"kache.hit@gmail.com","avatar":null},"body":"On Mon Nov 24, 2025 at 11:35 AM PST, Junio C Hamano wrote:\n> I do not think it is a total dogmatic prohibition, but is a\n> practical piece of advice to be prepared in a situation where the\n> file %f does not exist on the disk in the working tree.  Also even\n> when the file %f does exist, its contents would not match (because\n> it was smudged when it was checked out, and the user may have\n> further modified it) what in the tree of the commit you are\n> switching out of.\n\nYou're right, it can be tricky as there are several cases to handle. I\ntry covering this and other cases in the script's tests.\n\nHowever, isn't properly handling different scenarios a separate issue?\nSimplying my concept to \"ignoring\" instead of \"redacting\":\n\n * Clean: ignore certain lines, preventing them from being committed\n * Smudge: don't overwrite working copy of ignored lines on checkout\n\nThen the functionality becomes line-wise analogous to gitignore working\non whole files. My local copy of gitignored `.env` isn't overwritten\nwhen I checkout. I'm looking for the same, just line-wise.\n\nOn Mon Nov 24, 2025 at 11:35 AM PST, Junio C Hamano wrote:\n> ... one of the things the checkout needs to do may be to remove the\n> existing file from the working tree, and then create a file anew\n> (probably in a temporary file) and move it to the final place, in\n> which case, your smudge filter may be called during \"create a file\n> anew\" phase, where the old file F or G may be missing from the working\n> tree.\n\nThe old file being missing, being wholly removed right away, is exactly\nwhat I'm running into. If the working copy was kept around for `smudge`,\nI could achive a basic implementation of line-wise ignore/redact.\n\nAs-is, git's clean -> smudge filters can:\n * idempotent op -> no-op, e.g. identing or formatting\n * perfect mapping -> map back, e.g. git-lfs\n * add info -> remove info, e.g. expand RCS keyword -> unexpand\n\nBut not:\n * remove info -> restore info, e.g. ignoring lines, redacting\n\n\n- Kache\n\nPS\n\nI've just found a case I'm not yet handling: at the end of `smudge()`,\nany unused secrets from the \"previous working copy\" that haven't been\nrestored into the template would be lost. It is analogous to having\nlocal changes to a file at commit `X` and checking out `Y` where that\nfile has been deleted. Git avoids overwriting local changes by aborting\nthe checkout.\n"},{"id":"531257","messageId":"CAPx1GveYzEs_iAo2oV2OgoGbJJfw4Q0VVzRApEwCOMUAAY_v1Q@mail.gmail.com","threadId":"64525","inReplyTo":"DEH58DEF5MGO.2CFIKCM2CAQY2@gmail.com","subject":"Re: Filter smudge for secret restoration: no disk access?","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2025-11-25T08:55:12Z","receivedAt":"2025-11-25T08:55:26Z","isPatch":false,"sender":{"key":"chris.torek@gmail.com","avatar":"https://avatars.githubusercontent.com/u/16826774?v=4"},"body":"On Mon, Nov 24, 2025 at 10:40 AM Kache Hit <kache.hit@gmail.com> wrote:\n> I'm familiar with this practice, e.g. committing an `.env.template`\n> which is used to create an `.env` file with secrets within.\n>\n> However, this is my dotfiles repo that includes `~/.config`. There are\n> config files that store credentials right next to configuration, managed\n> by software that I don't control.\n\nMy technique for this is that my dotfiles are in a repository where\nthey are named \"profile\", \"bashrc\", \"gitconfig\", and so on. These\nget installed by my dotfiles-installer as $HOME/.profile, etc. The\ninstaller (my own creation, tuned to my personal needs and not really\nsuitable for anyone else) builds the target files as needed.\n\n(The thing probably needs a redesign and rewrite since newer\nsoftware messes with these files more dynamically at this point,\nbut I have not had to do that yet. So far I haven't needed to\ndo the \"update repository from active files\" part, which would\nbe harder.)\n\nThe reason for naming them without the leading dot is to\nmake it abundantly obvious during editing whether I'm on the\ntemplate or the actual config file.\n\nAs you've seen, there are more issues with going back in\nhistory (to points where various files didn't exist yet). This\nsidesteps most of these.\n\nChris\n"}]}