{"thread":{"id":"59446","subject":"Supporting automated removal of the UTF-8 BOM","startedAt":"2023-03-22T17:24:13Z","lastAt":"2023-03-22T17:24:13Z","messageCount":1,"participants":["mqudsi@neosmart.net"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"473990","messageId":"010001870a53c24c-52985ae7-dcb7-4422-96cd-ef88a01f138d-000000@email.amazonses.com","threadId":"59446","inReplyTo":null,"subject":"Supporting automated removal of the UTF-8 BOM","fromName":"","fromEmail":"mqudsi@neosmart.net","sentAt":"2023-03-22T17:17:54Z","receivedAt":"2023-03-22T17:24:13Z","isPatch":false,"sender":{"key":"mqudsi@neosmart.net","avatar":"https://gravatar.com/avatar/c2643dd7c6df61aed49d9f3d917ac6d61cafbbda5f9b1619f50d3b749dca415a?d=mp&s=160"},"body":"Hello team,\n\nI am curious if you would be open to an effort to extend git with the\nability to actively manage the presence of a UTF-8 BOM in the index (and\nworking tree), probably via the .gitattributes interface. I'm aware that\nthe existing encoding mechanism is already BOM-aware can deduce from its\npresence the charset of a file, but unfortunately it doesn't seem to be\npossible to have git strip the BOM from *UTF-8* content (or particular\ncontent) before storing it in the index without the use of precommit\nhooks.\n\nAfter giving it some thought and assuming that you would be open to the\nidea in principle, I can see several different approaches or\npossibilities.\n\nOne option would be to add a new charset named UTF-8-BOM for the express\npurpose of allowing particular filetypes to always be stored BOM-free in\nthe index (as iconv does not itself recognize this as a separate\ncharset). This would be along the same vein in which support was added\nfor UTF-16XX-BOM such that content can be converted to (BOM-free) UTF-8\nfor storage in the index and then converted back to UTF-16XX-BOM when\nchecked out.\n\nWhile it's true that most editors that emit UTF-8 files w/ a BOM will\nstill work with them if the BOM is removed (obviating the *need* to\nconvert back to UTF-8-BOM on checkout, unlike how editors expecting\nUTF-16LE-BOM will fail if the file is checked out as UTF-8), this would\nprevent users on other platforms from having to deal with the BOM when a\nWindows user checks the file in, and would also prevent the needless\nchurn from carelessly committed diffs adding or removing the BOM.\n\nAn alternative option that wouldn't involve adding a new UTF-8-BOM\ncharset would be to make core git aware of the BOM and able to treat its\npresence/absence as consequential or to be ignored in the same fashion\nas how git currently has an option to transparently convert line endings\nbetween lf and crlf at commit/checkout, but I was under the impression\nthat this is largely considered a dirty hack to handle a very common\nproblem and not something that the team would be eager to expand on by\nadding a similar option for BOM markers.\n\nOne other option would be to add support for BOM removal only through\n.gitattributes but not via an explicit charset conversion, i.e. by\nadding a \"bom\" option like the \"eol\" option that could be specified for\nparticular file types, for example,\n\n    *.csproj text=auto eol=crlf bom=[add|strip]\n\nWith bom=add, the file would be stored in the index without a BOM but it\nwould be added on checkout (as a parallel to how eol=crlf acts). With\nbom=strip, the file would again be stored in the index without a BOM and\ngit would prevent addition of the BOM on checkout.\n\nCurious about your thoughts on the matter. I've gone to great lengths to\npurge the BOM from my machine and have managed to hack most of\nMicrosoft's tools to refrain from adding it but alas Visual Studio still\ninsists on inserting the BOM to machine-generated content (for the\nreset, I use a VS extension to strip the BOM from regular code files\nopened/saved through Visual Studio).\n\nThanks for the time,\n\nMahmoud\n"}]}