{"thread":{"id":"6772","subject":"mingw, windows, crlf/lf, and git","startedAt":"2007-02-11T23:13:16Z","lastAt":"2007-02-14T20:24:09Z","messageCount":83,"participants":["Mark Levedahl","Johannes Schindelin","Robin Rosenberg","Jakub Narebski","Theodore Tso","David Lang","Linus Torvalds","Junio C Hamano","Shawn O. Pearce","Alexander Litvinov","Jeff King","Nicolas Pitre","Sam Ravnborg","Johannes Sixt"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"34206","messageId":"45CFA30C.6030202@verizon.net","threadId":"6772","inReplyTo":null,"subject":"mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mlevedahl@verizon.net","sentAt":"2007-02-11T23:13:16Z","receivedAt":"2007-02-11T23:13:16Z","isPatch":false,"sender":{"key":"mlevedahl@verizon.net","avatar":null},"body":"I am NOT intending to start a flamewar O:-) , so please don't turn this \ninto one.\n\nThe recent threads on a mingw git port are explicit in the intent to \nprovide a Windows native git. I believe there is a fundamental conflict \nhere with the position, clearly stated by Linus, that git does not alter \ncontent in any way. Windows suffers the curse of DOS line endings (\\r\\n \nvs \\n), and a true port to Windows *must* allow for \\r\\n and \\n to be \nsemantically the same thing as most large projects end up with a mixture \nof such files and/or are targeting cross-platform capabilities. The \nmajor competing solutions git seeks to supplant (cvs, cvsnt, svn, hg) \nhave capability to recognize \"text\" files and transparently replace \\r\\n \nwith \\n on input, the reverse on output, and ignore all such differences \non diff operations. To be relevant on native Windows, git must do the \nsame. Otherwise, git will be deemed \"too wierd\" and dismissed in favor \nof a tool \"that works.\"\n\nThere is no use to debating the technical merits of \\r\\n vs \\n vs \\r vs \nwhatever, nor of not converting. Really. Just accept that there is a \nfundamental requirement that any version control tool on Windows be able \nto silently convert between \\r\\n and \\n. To believe otherwise is to \nexpect that the conversion be pushed elsewhere into the tool chain in \nuse, and that won't happen as the competition already provide this \nconversion capability.\n\nSo, I think the git project needs to come to an explicit position on \nthis, basically being:\n\n1) git is a POSIX only tool (i.e., there will be no \\r\\n munging), or\n2) a Windows port of git will handle and mung \\r\\n and \\n line endings.\n\nIf the answer is 1, the mingw port is a waste of time as it simply won't \nbe usable by its target audience. If the answer is 2, then I think a \nvery careful design of this capability is in order.\n\nComments?\n\nBTW, I have addressed this in my own world using a pre-commit script \nthat converts textfile line endings into \\n, recognizing that our \nWindows tool chain handles such files perfectly well, while our Linux \ntoolchain requires it.\n\nMark Levedahl\n"},{"id":"34212","messageId":"Pine.LNX.4.63.0702120028340.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"45CFA30C.6030202@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-11T23:34:27Z","receivedAt":"2007-02-11T23:34:27Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 11 Feb 2007, Mark Levedahl wrote:\n\n> The major competing solutions git seeks to supplant (cvs, cvsnt, svn, \n> hg) have capability to recognize \"text\" files and transparently replace \n> \\r\\n with \\n on input, the reverse on output, and ignore all such \n> differences on diff operations.\n\nAgree with transformations on input and output; disagree on diff.\n\nThe problem is that it really is a transformtion. Since most Windows tools \n(at least those used in portable software) handle \\n without \\r quite \nwell, thank you, I'd tend towards the view point: do not mess with line \nendings pre-commit/post-checkout.\n\nEven MacOSX uses \\n now, instead of \\r.\n\nOf course, for those projects which _use_ CRLF: they can continue with it. \nGit has no problem with those line endings.\n\nThe only problem CVS tried to solve (badly) was to be able to checkout \ntext files on DOS, Unix _and_ MacOS. In practice, though, this use case \ndoes not matter anymore IMHO.\n\nCiao,\nDscho\n"},{"id":"34222","messageId":"200702120114.21477.robin.rosenberg.lists@dewire.com","threadId":"6772","inReplyTo":"45CFA30C.6030202@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-02-12T00:14:21Z","receivedAt":"2007-02-12T00:14:21Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"måndag 12 februari 2007 00:13 skrev Mark Levedahl:\n> The recent threads on a mingw git port are explicit in the intent to \n> provide a Windows native git. I believe there is a fundamental conflict \n> here with the position, clearly stated by Linus, that git does not alter \n> content in any way. Windows suffers the curse of DOS line endings (\\r\\n \n> vs \\n), and a true port to Windows *must* allow for \\r\\n and \\n to be \n> semantically the same thing as most large projects end up with a mixture \n> of such files and/or are targeting cross-platform capabilities. The \n> major competing solutions git seeks to supplant (cvs, cvsnt, svn, hg) \n> have capability to recognize \"text\" files and transparently replace \\r\\n \n> with \\n on input, the reverse on output, and ignore all such differences \n> on diff operations. To be relevant on native Windows, git must do the \n> same. Otherwise, git will be deemed \"too wierd\" and dismissed in favor \n> of a tool \"that works.\"\n> \nAs of today git is a posix tool simply because it's not fully ported to\nother enviromnents. I brought this up quite a time ago, and didn't face heavy artillery\nthen, and wouldn't today either. The code is still missing though. I didn't \nwrite it then, because it's my #1 priority and nobody else did. Linus even did a \nrough scetch, but that's it. \n\nI guess git will get this feature when someone does the code for it.\n\n-- robin\n"},{"id":"34227","messageId":"eqodam$r17$1@sea.gmane.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702120028340.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-02-12T00:46:56Z","receivedAt":"2007-02-12T00:46:56Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin wrote:\n> On Sun, 11 Feb 2007, Mark Levedahl wrote:\n> \n>> The major competing solutions git seeks to supplant (cvs, cvsnt, svn, \n>> hg) have capability to recognize \"text\" files and transparently replace \n>> \\r\\n with \\n on input, the reverse on output, and ignore all such \n>> differences on diff operations.\n> \n> Agree with transformations on input and output; disagree on diff.\n\nI wonder if this could/should be solved with adding some option to git-diff,\nsimilar to --ignore-space-change and --ignore-all-space...\n\nJust a [idle] thought.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"34232","messageId":"45CFD2AA.5090509@verizon.net","threadId":"6772","inReplyTo":"eqodam$r17$1@sea.gmane.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-12T02:36:26Z","receivedAt":"2007-02-12T02:36:26Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Jakub Narebski wrote:\n> Johannes Schindelin wrote:\n>   \n>> On Sun, 11 Feb 2007, Mark Levedahl wrote:\n>>\n>>     \n>>> The major competing solutions git seeks to supplant (cvs, cvsnt, svn, \n>>> hg) have capability to recognize \"text\" files and transparently replace \n>>> \\r\\n with \\n on input, the reverse on output, and ignore all such \n>>> differences on diff operations.\n>>>       \n>> Agree with transformations on input and output; disagree on diff.\n>>     \n>\n> I wonder if this could/should be solved with adding some option to git-diff,\n> similar to --ignore-space-change and --ignore-all-space...\n>\n> Just a [idle] thought.\n>   \nThat would work. Assuming blobs are stored in with \\n, diff just has to \nopen files in 'rt' mode rather than just 'r' and the \\r\\n are \ntransformed  on read so are never seen by git code. That is basically \nwhat Windows native tools do, but they also write files opened in 'wt' \nmode so \\n become \\r\\n on output. Of course, if this were an option, \nusers could look for line ending differences if they cared.\n\nMark\n"},{"id":"34233","messageId":"45CFD2E4.5030806@verizon.net","threadId":"6772","inReplyTo":"200702120114.21477.robin.rosenberg.lists@dewire.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-12T02:37:24Z","receivedAt":"2007-02-12T02:37:24Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Robin Rosenberg wrote:\n> As of today git is a posix tool simply because it's not fully ported to\n> other enviromnents. I brought this up quite a time ago, and didn't face heavy artillery\n> then, and wouldn't today either. The code is still missing though. I didn't \n> write it then, because it's my #1 priority and nobody else did. Linus even did a \n> rough scetch, but that's it.\nSo, the basic design for this feature exists where? I would assume this \nwould include a file mode indicator set in the blob or tree designating \nthe blob is \"text\", along with mechanism to specify for a project what \nfiles are \"text\", along with some safety valve to check and not do \ntransformation when the file does not look text-ish.\n\nMark\n"},{"id":"34260","messageId":"20070212042425.GB18010@thunk.org","threadId":"6772","inReplyTo":"45CFA30C.6030202@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-02-12T04:24:25Z","receivedAt":"2007-02-12T04:24:25Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sun, Feb 11, 2007 at 06:13:16PM -0500, Mark Levedahl wrote:\n> I am NOT intending to start a flamewar O:-) , so please don't turn this \n> into one.\n> \n> The recent threads on a mingw git port are explicit in the intent to \n> provide a Windows native git. I believe there is a fundamental conflict \n> here with the position, clearly stated by Linus, that git does not alter \n> content in any way. Windows suffers the curse of DOS line endings (\\r\\n \n> vs \\n), and a true port to Windows *must* allow for \\r\\n and \\n to be \n> semantically the same thing as most large projects end up with a mixture \n> of such files and/or are targeting cross-platform capabilities. The \n> major competing solutions git seeks to supplant (cvs, cvsnt, svn, hg) \n> have capability to recognize \"text\" files and transparently replace \\r\\n \n> with \\n on input, the reverse on output, and ignore all such differences \n> on diff operations. To be relevant on native Windows, git must do the \n> same. Otherwise, git will be deemed \"too wierd\" and dismissed in favor \n> of a tool \"that works.\"\n\nSo this is something that I've tried proposing to the Mercurial\ndevelopers, but it's never been implemented in hg.  It'll be\ninteresting to see what the git community thinks.  :-)\n\nMy proposal does require adding a file type to each file, as tracked\nmetadata, which may doom it from the start.  If you add a file type,\nthen you have to support mutating the file type, and some way of\nhandling merge conflicts (generally, picking one type or another).\n\nThen for each file type, we implement a set of interfaces (perhaps as\nsimple as a series of executables named git-<type>-<operation>) which\nif present, transforms the file from its live format to the canonical\nformat which is actually checked in and back again.  Besides using\nthis for the DOS CR/LF problem, it also allows for an efficient\nstorage of things like OpenOffice files which are a zipped set of .xml\nfiles.  By decompressing them before pushing them into the SCM, it\nmeans that if the user makes a tiny spelling correction in their\nOpenOffice file, the delta stored in the git repository can be much\nmore efficiently stored (since the diff of the .xml tree will be\nsmall, where as the diff of the entire compressed file is likely going\nto be close to the entire size of the .odt file).\n\nAnother nice thing to provide for each file type would be a\npretty-printer for the diffs, so it becomes easier to see the delta\nbetween two versions of an OpenOffice file in a textual window.\n\nSo, is this idea sane or completely insane?  Hopefully it passes\nLinus's it-solves-multiple-problems-at-once test, at least.  :-)\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"34264","messageId":"Pine.LNX.4.63.0702112325560.6287@qynat.qvtvafvgr.pbz","threadId":"6772","inReplyTo":"20070212042425.GB18010@thunk.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-02-12T07:28:02Z","receivedAt":"2007-02-12T07:28:02Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Sun, 11 Feb 2007, Theodore Tso wrote:\n\n> Then for each file type, we implement a set of interfaces (perhaps as\n> simple as a series of executables named git-<type>-<operation>) which\n> if present, transforms the file from its live format to the canonical\n> format which is actually checked in and back again.  Besides using\n> this for the DOS CR/LF problem, it also allows for an efficient\n> storage of things like OpenOffice files which are a zipped set of .xml\n> files.  By decompressing them before pushing them into the SCM, it\n> means that if the user makes a tiny spelling correction in their\n> OpenOffice file, the delta stored in the git repository can be much\n> more efficiently stored (since the diff of the .xml tree will be\n> small, where as the diff of the entire compressed file is likely going\n> to be close to the entire size of the .odt file).\n>\n> Another nice thing to provide for each file type would be a\n> pretty-printer for the diffs, so it becomes easier to see the delta\n> between two versions of an OpenOffice file in a textual window.\n>\n> So, is this idea sane or completely insane?  Hopefully it passes\n> Linus's it-solves-multiple-problems-at-once test, at least.  :-)\n\nthere have been other things discussed that could use the 'do this on checkout' \nhooks, specificly on the issue of useing git to manage /etc the need to \nsave/restore permissions requires a hook on checkout that doesn't exist yet. \nthis sounds like it would solve that problem as well.\n\nDavid Lang\n"},{"id":"34273","messageId":"Pine.LNX.4.63.0702121219550.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"eqodam$r17$1@sea.gmane.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-12T11:21:29Z","receivedAt":"2007-02-12T11:21:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[Cc'ing git list, which I sometimes have to do when Jakub replies]\n\nOn Mon, 12 Feb 2007, Jakub Narebski wrote:\n\n> Johannes Schindelin wrote:\n> > On Sun, 11 Feb 2007, Mark Levedahl wrote:\n> > \n> >> The major competing solutions git seeks to supplant (cvs, cvsnt, svn, \n> >> hg) have capability to recognize \"text\" files and transparently replace \n> >> \\r\\n with \\n on input, the reverse on output, and ignore all such \n> >> differences on diff operations.\n> > \n> > Agree with transformations on input and output; disagree on diff.\n> \n> I wonder if this could/should be solved with adding some option to git-diff,\n> similar to --ignore-space-change and --ignore-all-space...\n\nIt could be done, but those options were introduced for CRLF breakage in \nthe first place.\n\nYou need --ignore-crlf-breakage? Just holler.\n\nCiao,\nDscho\n"},{"id":"34275","messageId":"Pine.LNX.4.63.0702121232120.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"20070212042425.GB18010@thunk.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-12T11:36:26Z","receivedAt":"2007-02-12T11:36:26Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 11 Feb 2007, Theodore Tso wrote:\n\n> My proposal does require adding a file type to each file, as tracked\n> metadata, which may doom it from the start.\n\nI'd rather do that a la .gitignore, i.e. make this handling dependent on \nfile name patterns. It is not only backwards compatible (from the \nviewpoint of the repository format), it also avoids having to specify over \nand over again that yes, this new .odt file _is_ an OpenOffice document.\n\n> Then for each file type, we implement a set of interfaces (perhaps as\n> simple as a series of executables named git-<type>-<operation>) which\n> if present, transforms the file from its live format to the canonical\n> format which is actually checked in and back again.\n\nAgain, I propose a slight change: Let's add a transformation driver like \nthe merge driver: this allows inlining common operations like unzipping, \nCRLF->LF conversion, etc.\n\nCiao,\nDscho\n"},{"id":"34299","messageId":"Pine.LNX.4.64.0702120839490.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"20070212042425.GB18010@thunk.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-12T17:20:43Z","receivedAt":"2007-02-12T17:20:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 11 Feb 2007, Theodore Tso wrote:\n> \n> So this is something that I've tried proposing to the Mercurial\n> developers, but it's never been implemented in hg.  It'll be\n> interesting to see what the git community thinks.  :-)\n> \n> My proposal does require adding a file type to each file, as tracked\n> metadata, which may doom it from the start.  If you add a file type,\n> then you have to support mutating the file type, and some way of\n> handling merge conflicts (generally, picking one type or another).\n\nI agree that a file-type approch would work, but I personally think it's \ntoo inflexible (just cr/lf vs lf? There are tons of other interesting \nissues that are valid). I also think it falls down on another (and in some \nways much more fundamental problem): these things exist EVEN WHEN THE FILE \nITSELF DOES NOT EXIST!\n\nIn other words, a policy about cr/lf is *not* a policy about actual \ncontent. It's something much more: it's a policy about representation in \ngeneral, which includes *potential* content. It should obviously take \neffect on \"git add\" even with content that didn't exist before, and to \nwork well, it should do so without the user having to think about it.\n\nEqually importantly, this happens with content that was added by people \nwho simply DO NOT CARE. In other words, I think a \"file type\" thing \nfundamentally cannot work, because under UNIX, it would be stupid and \npointless, so any project that is maintained under UNIX might _add_ the \nfile types, but since they won't matter, they'll inevitably be wrong (ie \npeople forgot to mark a binary thing binary, or a text thing as text).\n\nSo: file types or attributes are broken. They cannot work well.\n\nBut enough on the negative rambling, I do have a positive and constructive \nsuggestion, because I actually think I have a great model for it. But I've \nnever cared enough (and since the main target would be some windows issue, \nI suspect I never really _will_ care enough) to really worry about it.\n\nAnyway, if somebody really wants to look at this, and wants to create \nsomething that is actually _usable_, my suggestion is to simply extend on \nthe \".gitignore\" file approach. The great thing about .gitignore is that\n\n (a) you can track it like you track any other file\n\n     This makes merges a *lot* easier. You see it as conflicts, you can \n     fix it up, and in general, you can use all the same tools with it as \n     you use with anything else. In contrast, explicit per-file filetypes \n     are _horrible_ for maintenance.\n\n (b) you can add to it with *patterns*, which is exactly what you want for \n     file types.\n\n     You can do things like\n\n\t*.bin: binary\n\t*: text\n\n     to say \"everythgn that matches *.bin is binary, the rest is text\", \n     and solves the maintenance issue trivially. Everybody will like it. \n     For the kernel, for example, we'd have a really easy\n\n\tDocumentation/logo.gif: binary\n\t*: text\n\n     and that would probably take care of it.\n\n     You can also have a few default file patterns built in, which would \n     take care of it for 99% of all projects without anybody ever having \n     to even think about it - even under DOS.\n\n (c) it doesn't actually affect database representation, it only changes \n     behaviour for programs, which is also  exactly what you want (if you \n     have per-file \"file types\", you end up having serious problems at \n     merge time: when I say \"affect database representation\", I don't mean \n     that I think git cannot change its database, I literally mean at a \n     \"higher\" level: represening per-file attributes is a DISASTER from a \n     merge situation)\n\n     So not only is it backwards-compatible with traditional git usage, \n     it's much more fundamentally simple: it doesn't add any new core data \n     structures or rules. All the core stays exactly as it is, and it just \n     affects higher-level behaviour. And that's important: one reason git \n     has been so stable is that the really core data structures are really \n     really stable and simple.\n\n     Even when we did *really* core changes like the whole packfile thing, \n     the fundamental data structures didn't change at all *conceptually*.\n\n (d) it's actually a lot more flexible than file types.\n\n     Merge stategies, anybody? We can easily have the default merge \n     strategy be the normal three-way merge (which is obviously the right \n     thing for almost anything), but how about something like\n\n\t*.doc: binary,merge=doc-merge\n\n     which tells git that it should use a separate \"doc-merge\" program to \n     merge those kinds of files when it needs to do a nontrivial merge..\n\n (e) exactly like \".gitignore\", you should also be able to have a \n     \".git/info/exclude\" file that is your _private_ rules, and \n     per-directory \".gitignore\" files that are the _hierarchical_ rules.\n\n     This just makes maintenance much simpler. Not one big file that has \n     everything, and that clashes. Make the top-level one contain all the \n     generic default rules, and then lower down we can have more specific \n     rules for very specific things, exactly like the kernel .gitignore \n     files do. The top-level file should *not* have to know all the \n     details of some architecture- or sub-project specific file behaviour.\n\n     Similarly, having an untracked file (.git/info/exclude) allows people \n     to have rules that make sense for *them*, but that might not make \n     sense for the upstream developers (say, somebody crazy enough to \n     develop Linux under Windows). So people can have their purely local \n     rules without forcing them on others.\n\nAnyway, that would be my suggestion. Call it \".gitattributes\" or \nsomething. Make it a nice ASCII format, exactly like .gitignore, and make \nall the rules exactly the same, except it has a \": <attributelist>\" at the \nend for each line.\n\nStart off supporting just \"binary\" and \"text\", but keep in mind that \npeople may want other things. Individualized merge strategies etc.\n\nAlso, keep in mind that a *lot* of git operations will work purely on a \nSHA1 level, and those operations fundamentally *will*not*care* about file \ntypes. So when you merge a file, for example, the initial merge will be \ndone purely on SHA1's, and git would do all the normal \"if it didn't \nchange in branch 1, take the branch 2 version directly\" without ever even \n*looking* at any file rules.\n\nThis is important, because this is what makes git efficient for large \nprojects, and which would allow git to _remain_ efficient even in the face \nof having to read all those comples .gitattributes files. When we merge \ntwo repositories with 20,000+ files, we usually really only \"merge\" a \ncouple of the files. \n\nSame goes for \"text\" mode. The \"text\" thing would only affect things like \n\"git add\" etc that use \"git-update-index\" to calculate the new SHA1. We'd \nnever use it \"normally\". \"git diff\" would still be instantaneous, because \nthe git index shows the file still matches, and that is all done on a SHA1 \nonly level. So only when you do a \"git add\" or when it needs to refresh \nthe index because the file changed, and it reads in the file, will it \nactually care about whether it's a text or a binary file.\n\nThis is actually *exactly* what you want. Not just for performance, but \nsimply because this is also how you can take something like the Linux \narchive, and \"just use it\" under Windows, even if your editor adds (or \nwants) CR/LF.\n\nBtw, how would I implement this? If I really were energetic enough to \nimplement it, I would do:\n\n (a) Add a flag to \"git-ls-files\" logic to add \"type information\" in \n     front.\n\n     Not only do you want this *anyway* for other reasons, but for\n     binary/text, the thing you actually care most about is \"git add\", and \n     it already basically just does \"take this file pattern, feed it \n     through git-ls-files, and add those files\". So you'd get it basically \n     for free.\n\n     It is also fairly easy to add at this stage, because you can simply \n     look for all the places that work with \"info/exclude\" and \n     \".gitignore\", and you know that \"Ahh, I need to teach these exact \n     places to understand about attributes\". So you'd add an \n     \"add_attributes_from_file()\" function etc etc.\n\n     Quite straightforward. In fact, you might be able to use the \n     gitignore parsing *as*is*, and just teach it about more flags that \n     just \"ignore\": both in \"struct dir_entry\" and in \"struct exclude\".\n\n (b) Teach the git-update-index logic about hashing text blobs.\n\n (c) Profit!\n\nIt really should be fairly straightforward. I'm sure it wouldn't be \n*entirely* trivial, but I'm also fairly sure that somebody reasonably \ncompetent could do it in a couple of days (with testing) if they were just \nsufficiently motivated to get started.\n\nAnybody?\n\n\t\tLinus\n"},{"id":"34327","messageId":"Pine.LNX.4.63.0702122332180.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702120839490.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-12T22:37:58Z","receivedAt":"2007-02-12T22:37:58Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[I agree on the .gitignore approach; see my other mail in this thread]\n\nOn Mon, 12 Feb 2007, Linus Torvalds wrote:\n\n> Btw, how would I implement this? If I really were energetic enough to \n> implement it, I would do:\n> \n>  (a) Add a flag to \"git-ls-files\" logic to add \"type information\" in \n>      front.\n> \n>      Not only do you want this *anyway* for other reasons, but for\n>      binary/text, the thing you actually care most about is \"git add\", and \n>      it already basically just does \"take this file pattern, feed it \n>      through git-ls-files, and add those files\". So you'd get it basically \n>      for free.\n> \n>      It is also fairly easy to add at this stage, because you can simply \n>      look for all the places that work with \"info/exclude\" and \n>      \".gitignore\", and you know that \"Ahh, I need to teach these exact \n>      places to understand about attributes\". So you'd add an \n>      \"add_attributes_from_file()\" function etc etc.\n> \n>      Quite straightforward. In fact, you might be able to use the \n>      gitignore parsing *as*is*, and just teach it about more flags that \n>      just \"ignore\": both in \"struct dir_entry\" and in \"struct exclude\".\n> \n>  (b) Teach the git-update-index logic about hashing text blobs.\n> \n>  (c) Profit!\n\nNot so fast.\n\nIn order for this to be _useful_, you also have to have a way to _extract_ \nthe text blobs. Not only for read-tree, but _also_ for diff. It makes no \nsense at all to have this transformation one-way. For diff, you _might_ \nwant to have a diff beautifier (for example the .odt thing), but read-tree \nis _really_ important.\n\nCiao,\nDscho\n"},{"id":"34328","messageId":"7vps8f6l81.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702120839490.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-12T22:54:06Z","receivedAt":"2007-02-12T22:54:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Btw, how would I implement this? If I really were energetic enough to \n> implement it, I would do:\n>\n>  (a) Add a flag to \"git-ls-files\" logic to add \"type information\" in \n>      front.\n>\n>      Not only do you want this *anyway* for other reasons, but for\n>      binary/text, the thing you actually care most about is \"git add\", and \n>      it already basically just does \"take this file pattern, feed it \n>      through git-ls-files, and add those files\". So you'd get it basically \n>      for free.\n>\n>      It is also fairly easy to add at this stage, because you can simply \n>      look for all the places that work with \"info/exclude\" and \n>      \".gitignore\", and you know that \"Ahh, I need to teach these exact \n>      places to understand about attributes\". So you'd add an \n>      \"add_attributes_from_file()\" function etc etc.\n>\n>      Quite straightforward. In fact, you might be able to use the \n>      gitignore parsing *as*is*, and just teach it about more flags that \n>      just \"ignore\": both in \"struct dir_entry\" and in \"struct exclude\".\n>\n>  (b) Teach the git-update-index logic about hashing text blobs.\n\nI agree that we can assume editors can grok files with LF\nend-of-line just fine and we would not need to do the reverse\nconversion on checkout paths (e.g. \"read-tree -u\", \"checkout-index\").\n\nTextual diff generation needs to learn the CRLF-to-LF conversion\nin diff_populate_filespec(); this needs to be done even when the\ncaller wants size_only.\n\nOops.\n\nNot so fast.  What's your plan for st_size?\n\n>  (c) Profit!\n>\n> It really should be fairly straightforward. I'm sure it wouldn't be \n> *entirely* trivial, but I'm also fairly sure that somebody reasonably \n> competent could do it in a couple of days (with testing) if they were just \n> sufficiently motivated to get started.\n>\n> Anybody?\n\nNot me.\n"},{"id":"34330","messageId":"Pine.LNX.4.64.0702121455560.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702122332180.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-12T23:02:42Z","receivedAt":"2007-02-12T23:02:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 12 Feb 2007, Johannes Schindelin wrote:\n>\n> >  (c) Profit!\n> \n> Not so fast.\n\nAww! And just when I _finally_ had a \"step 2\".\n\n> In order for this to be _useful_, you also have to have a way to _extract_ \n> the text blobs. Not only for read-tree, but _also_ for diff.\n\nActually, my argument is that we don't need it all that much.\n\nFor example, your \"read-tree\" argument is actually wrong. Anything that is \nin a tree is _already_ fixed to be '\\n'. So as long as we keep to things \nlike\n\n\tgit diff version1..version2\n\nwe'll actually always get the right version.\n\nAlso, the index will make sure that we don't even *try* to diff normal \nchecked out files.\n\nSo the only time you actually really need to test the .gitattributes file \nis when you do an \"open blob in working tree\". And once you do that \nfunction right, and just make sure both git-update-index and yes, the \n\"diff against working tree\" cases use it, you really should be mostly \ndone.\n\nBoth git-update-index and git-diff-files want basically the same \ninterface:\n\n\tstruct file_buf {\n\t\tconst char *buf;\n\t\tunsigned long size;\n\t\tint flags;\n\t}\n\n\tint read_file(const char *path, struct file_buf *);\n\tclose_file(struct file_buf *);\n\nand we should use that instead of the current \"open + stat + mmap/read + \nclose\" sequences.\n\nIt really shouldn't be too nasty.\n\n\t\tLinus\n"},{"id":"34329","messageId":"7vk5yn6ktl.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"7vps8f6l81.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-12T23:02:46Z","receivedAt":"2007-02-12T23:02:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> Btw, how would I implement this? If I really were energetic enough to \n>> implement it, I would do:\n>> ...\n>>  (b) Teach the git-update-index logic about hashing text blobs.\n>\n> I agree that we can assume editors can grok files with LF\n> end-of-line just fine and we would not need to do the reverse\n> conversion on checkout paths (e.g. \"read-tree -u\", \"checkout-index\").\n>\n> Textual diff generation needs to learn the CRLF-to-LF conversion\n> in diff_populate_filespec(); this needs to be done even when the\n> caller wants size_only.\n>\n> Oops.\n>\n> Not so fast.  What's your plan for st_size?\n\nIf I were to do this, I would say the cache should store the\nsize on the filesystem in stat fields.  Which means that the\nobject name recorded is text blob _after_ line endings are\nnormalized to LF, and its exploded size does not necessarily\nmatch the cached size.\n\nSo this means that whoever does the diff_populate_filespec()\nchange needs to be careful, but it is not such a big deal.\n"},{"id":"34332","messageId":"Pine.LNX.4.64.0702121505560.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"7vps8f6l81.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-12T23:09:03Z","receivedAt":"2007-02-12T23:09:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 12 Feb 2007, Junio C Hamano wrote:\n> \n> Not so fast.  What's your plan for st_size?\n\nUmm. There's two (very distinct) uses for st_size.\n\nThe one that we actually use to validate the current index obviously must \nmatch the \"OS returned value\". It contains all the CR/LF stuff.\n\nThe one where we actually read the file and run SHA1 on the result must \nobviously be the post-conversion one.\n\nBut it shouldn't be a problem. We'll always know which one matters: the \nindex case is always about pure stat information (and has no meaning \noutside of that, really - after all, it's no different from st_mode etc, \nand we actually keep it in a special binary format that is endian-safe!) \nand the \"real object\" case is always about the *data* we use to compare \nwith.\n\nI don't think we ever mix the two anyway.\n\n\t\tLinus\n"},{"id":"34338","messageId":"Pine.LNX.4.63.0702121513350.6630@qynat.qvtvafvgr.pbz","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702121514500.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-02-12T23:23:00Z","receivedAt":"2007-02-12T23:23:00Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Mon, 12 Feb 2007, Linus Torvalds wrote:\n\n> So we'd just need to pass in the information about whether it's binary or\n> not, and then do something like\n>\n> \t@@ -2091,6 +2091,10 @@ int index_fd(unsigned char *sha1, int fd, struct stat *st, int write_object, con\n>\n> \t \tif (!type)\n> \t \t\ttype = blob_type;\n> \t+#ifndef __UNIX__\n> \t+\tif (text && !strcmp(type, blob_type))\n> \t+\t\tconvert_crlf_to_lf(&buf, &size);\n> \t+#endif\n> \t \tif (write_object)\n> \t \t\tret = write_sha1_file(buf, size, type, sha1);\n> \t \telse\n>\n> and that would take care of a lot of things (yeah, I'd not do it that way\n> in practice, but really doesn't look that nasty - it's actually much\n> nastier to have to look up the text/binary type in the first place).\n\nyou could do something like this and it would deal with the srlf/lf problem, but \nif you instead put in the conversion hooks like Ted suggested then you can \nactually gain a LOT more.\n\nhis example of openoffice documents that are gziped xml files is a very good \none. if the 'conversion' is to gunzip on checkin and gzip on checkout then the \ncore git logic will work on the nice diffable xml instead of the compressed \nbinary blob.\n\nif this is extensable to arbatrary helper functions to do the conversions I'll \nbet that there are many other cases that can use this.\n\nI think the big questions needs to be, is this helper app a filter, or can it be \npassed a filename as the destination (which would let it do things like set \npermissions on the files it creates), or should it be both?\n\nDavid Lang\n"},{"id":"34335","messageId":"Pine.LNX.4.63.0702130020450.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"7vps8f6l81.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-12T23:24:55Z","receivedAt":"2007-02-12T23:24:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 12 Feb 2007, Junio C Hamano wrote:\n\n> I agree that we can assume editors can grok files with LF end-of-line \n> just fine and we would not need to do the reverse conversion on checkout \n> paths (e.g. \"read-tree -u\", \"checkout-index\").\n\nIn that case, a simple pre-commit hook would suffice.\n\nNo, the problem mentioned by Mark was a very real one: you _cannot_ rely \non Windows' editors not to fsck up with line endings. The worst case is if \nthe file contains _some_ CRLF and _some _LF_. Almost always I had the \nproblem that it now converted _all_ LFs to CRLFs. Even those which already \nwere converted.\n\nSo, if we are to support text mode, it is not one-way. If we do one-way, \nwe really do _not_ support text mode, but pre-commit conversion to LF \nstyle text. And in this case, core git does not need _any_ change.\n\nCiao,\nDscho\n"},{"id":"34336","messageId":"Pine.LNX.4.64.0702121514500.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702121505560.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-12T23:25:00Z","receivedAt":"2007-02-12T23:25:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 12 Feb 2007, Linus Torvalds wrote:\n> \n> But it shouldn't be a problem. We'll always know which one matters: the \n> index case is always about pure stat information (and has no meaning \n> outside of that, really - after all, it's no different from st_mode etc, \n> and we actually keep it in a special binary format that is endian-safe!) \n> and the \"real object\" case is always about the *data* we use to compare \n> with.\n\nIn fact, for git-update-index, I think it's *literally* as easy as just \nchanging \"index_fd()\" to convert the buffer on-the-fly as needed, before \nwe actually call \"write_sha1_file()\" or \"hash_sha1_file()\".\n\nSo we'd just need to pass in the information about whether it's binary or \nnot, and then do something like\n\n\t@@ -2091,6 +2091,10 @@ int index_fd(unsigned char *sha1, int fd, struct stat *st, int write_object, con\n\t \n\t \tif (!type)\n\t \t\ttype = blob_type;\n\t+#ifndef __UNIX__\n\t+\tif (text && !strcmp(type, blob_type))\n\t+\t\tconvert_crlf_to_lf(&buf, &size);\n\t+#endif\n\t \tif (write_object)\n\t \t\tret = write_sha1_file(buf, size, type, sha1);\n\t \telse\n\nand that would take care of a lot of things (yeah, I'd not do it that way \nin practice, but really doesn't look that nasty - it's actually much \nnastier to have to look up the text/binary type in the first place).\n\nSomething similar looks to be true in diff generation. The core \"compare \ntwo SHA1's at a time\" doesn't need any changes, but the code that actually \nreads in the temporary file from disk obviously does. But even that is \njust _one_ point, afaik - diff_populate_filespec()\":\n\n\t@@ -1362,6 +1362,10 @@ int diff_populate_filespec(struct diff_filespec *s, int size_only)\n\t \t\tif (fd < 0)\n\t \t\t\tgoto err_empty;\n\t \t\ts->data = xmmap(NULL, s->size, PROT_READ, MAP_PRIVATE, fd, 0);\n\t+#ifndef __UNIX__\n\t+\t\tif (text)\n\t+\t\t\tconvert_crlf_to_lf(&s->data, &s->size);\n\t+#endif\n\t \t\tclose(fd);\n\t \t\ts->should_munmap = 1;\n\t \t}\n\n(and again, that's not real code, it would also need to change the \n\"should_munmap\" flag to indicate the state of the _new_ \"data\" thing.\n\n\t\tLinus\n"},{"id":"34341","messageId":"7vfy9b6iyt.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702130020450.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-12T23:42:50Z","receivedAt":"2007-02-12T23:42:50Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> On Mon, 12 Feb 2007, Junio C Hamano wrote:\n>\n>> I agree that we can assume editors can grok files with LF end-of-line \n>> just fine and we would not need to do the reverse conversion on checkout \n>> paths (e.g. \"read-tree -u\", \"checkout-index\").\n>\n> In that case, a simple pre-commit hook would suffice.\n>\n> No, the problem mentioned by Mark was a very real one: you _cannot_ rely \n> on Windows' editors not to fsck up with line endings. The worst case is if \n> the file contains _some_ CRLF and _some _LF_. Almost always I had the \n> problem that it now converted _all_ LFs to CRLFs. Even those which already \n> were converted.\n>\n> So, if we are to support text mode, it is not one-way. If we do one-way, \n> we really do _not_ support text mode, but pre-commit conversion to LF \n> style text. And in this case, core git does not need _any_ change.\n\nWell I disagree in two counts.\n\n - I do not see how you propose to solve some CRLF and some LF\n   case with both-ways conversion.\n\n - Pre-commit hook would not be sufficient.  In a edit, diff,\n   test and then commit cycle, diff and test step needs to look\n   at whatever the editor left on the filesystem, so the changes\n   to populate-filespec is needed to make diff part work.\n\n\n   \n"},{"id":"34344","messageId":"Pine.LNX.4.63.0702121544550.6630@qynat.qvtvafvgr.pbz","threadId":"6772","inReplyTo":"7vfy9b6iyt.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-02-12T23:46:26Z","receivedAt":"2007-02-12T23:46:26Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Mon, 12 Feb 2007, Junio C Hamano wrote:\n\n>> Hi,\n>>\n>> On Mon, 12 Feb 2007, Junio C Hamano wrote:\n>>\n>>> I agree that we can assume editors can grok files with LF end-of-line\n>>> just fine and we would not need to do the reverse conversion on checkout\n>>> paths (e.g. \"read-tree -u\", \"checkout-index\").\n>>\n>> In that case, a simple pre-commit hook would suffice.\n>>\n>> No, the problem mentioned by Mark was a very real one: you _cannot_ rely\n>> on Windows' editors not to fsck up with line endings. The worst case is if\n>> the file contains _some_ CRLF and _some _LF_. Almost always I had the\n>> problem that it now converted _all_ LFs to CRLFs. Even those which already\n>> were converted.\n>>\n>> So, if we are to support text mode, it is not one-way. If we do one-way,\n>> we really do _not_ support text mode, but pre-commit conversion to LF\n>> style text. And in this case, core git does not need _any_ change.\n>\n> Well I disagree in two counts.\n>\n> - I do not see how you propose to solve some CRLF and some LF\n>   case with both-ways conversion.\n\nthe expectation is that the some-of-each situation is unlikly to happen if you \nconvert all the time.\n\nand if you do end up with a mixed ending file, the next time you check it in \nfrom a windows box it should clean it up.\n\nDavid Lang\n"},{"id":"34342","messageId":"Pine.LNX.4.63.0702130046450.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"7vfy9b6iyt.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-12T23:50:57Z","receivedAt":"2007-02-12T23:50:57Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 12 Feb 2007, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > Hi,\n> >\n> > On Mon, 12 Feb 2007, Junio C Hamano wrote:\n> >\n> >> I agree that we can assume editors can grok files with LF end-of-line \n> >> just fine and we would not need to do the reverse conversion on checkout \n> >> paths (e.g. \"read-tree -u\", \"checkout-index\").\n> >\n> > In that case, a simple pre-commit hook would suffice.\n> >\n> > No, the problem mentioned by Mark was a very real one: you _cannot_ rely \n> > on Windows' editors not to fsck up with line endings. The worst case is if \n> > the file contains _some_ CRLF and _some _LF_. Almost always I had the \n> > problem that it now converted _all_ LFs to CRLFs. Even those which already \n> > were converted.\n> >\n> > So, if we are to support text mode, it is not one-way. If we do one-way, \n> > we really do _not_ support text mode, but pre-commit conversion to LF \n> > style text. And in this case, core git does not need _any_ change.\n> \n> Well I disagree in two counts.\n> \n>  - I do not see how you propose to solve some CRLF and some LF\n>    case with both-ways conversion.\n\nVery easy. Forward: s/\\r\\n/\\n/. Backward: s/\\(^\\|[^\\r]\\)\\n/\\r\\n/.\n\n>  - Pre-commit hook would not be sufficient.  In a edit, diff,\n>    test and then commit cycle, diff and test step needs to look\n>    at whatever the editor left on the filesystem, so the changes\n>    to populate-filespec is needed to make diff part work.\n\nYes, you are right.\n\nHowever, since this is all post-1.5.0 (right? Right?) why not go with more \nof Ted's proposal, and make this whole mess also usable for other things \nthan just crlf issues?\n\nAnd I _really_ think that you do not help Windows people by doing this \none-way thing.\n\nCiao,\nDscho\n"},{"id":"34347","messageId":"45D10726.2070501@verizon.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702130020450.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mlevedahl@verizon.net","sentAt":"2007-02-13T00:32:38Z","receivedAt":"2007-02-13T00:32:38Z","isPatch":false,"sender":{"key":"mlevedahl@verizon.net","avatar":null},"body":"Johannes Schindelin wrote:\n> Hi,\n>\n> On Mon, 12 Feb 2007, Junio C Hamano wrote:\n>\n>   \n>> I agree that we can assume editors can grok files with LF end-of-line \n>> just fine and we would not need to do the reverse conversion on checkout \n>> paths (e.g. \"read-tree -u\", \"checkout-index\").\n>>     \n>\n> In that case, a simple pre-commit hook would suffice.\n>\n> No, the problem mentioned by Mark was a very real one: you _cannot_ rely \n> on Windows' editors not to fsck up with line endings. The worst case is if \n> the file contains _some_ CRLF and _some _LF_. Almost always I had the \n> problem that it now converted _all_ LFs to CRLFs. Even those which already \n> were converted.\n>\n> So, if we are to support text mode, it is not one-way. If we do one-way, \n> we really do _not_ support text mode, but pre-commit conversion to LF \n> style text. And in this case, core git does not need _any_ change.\n>\n> Ciao,\n> Dscho\nIn my work flow, I am using a pre-commit script that (among other \nthings) rewrites all text files to have \\n endings. This is a one-way \nconversion, and does work well for the set of tools I am using. The \nconverters I use I wrote years ago, and are smart enough to deal with \nmixtures of \\n, \\r\\n, and \\r line endings in one file, transforming all \ninto one unified form. d2u / u2d were not that robust when I last tried \nthem (years ago), but this is an absolute necessity.\n\nHowever, I don't think the one-way conversion is acceptable across the \nboard. While the only Windows editor I am aware of that doesn't grok \\n \nis Notepad (the moral equivalent of edlin), I suspect that undo reliance \nupon this will still lead to grief. If nothing else, someone, somewhere \nwill find that their beloved crlf's are missing and will complain. \nLoudly. And in the lore, git will become known for being \"wierd.\"  So, I \nsuspect a checkout script is necessary.\n\nMark\n"},{"id":"34351","messageId":"45D10D86.3030508@verizon.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702130046450.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-13T00:59:50Z","receivedAt":"2007-02-13T00:59:50Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Johannes Schindelin wrote:\n> However, since this is all post-1.5.0 (right? Right?) why not go with more \n> of Ted's proposal, and make this whole mess also usable for other things \n> than just crlf issues\nWhatever is done, it needs to be robust to the notion that people will \nfail to set the correct file type somewhere. Current cvsnt is fairly \ngood at autodetecting and setting text vs binary file type, and enforces \nthis across all platforms, so things don't go awry too often. It is in \nmy experience more reliable than subversion, which basically relies upon \nfile extensions mapping to mime types to identify content. All of which \nis a very much too low standard of accuracy for a version control \nsystem: I lost many files per year due to the above nonsense, so I worry \nabout trying to create a very general transform solution and not making \nit really, really failsafe. Having projects define individual globbing \npatterns is good, double checking the content for sanity is an absolute \nmust, but I don't think that is enough. I suspect the solution should \ninclude round-trip conversion when creating blobs to assure that the \ninput can be exactly reconstructed by the inverse transformation (and \ntherefore possibly rejecting input with mixed line endings). A similar \ncheck could be applied on checkout.\n\nPerhaps I'm too paranoid, but I've been burnt way too many times by \ntext/binary mode stuff to let this part be trivialized. Maybe it only \ngets enabled by core.ImReallyParanoid, but I want that option.\n\nMark\n"},{"id":"34352","messageId":"Pine.LNX.4.63.0702130204120.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"45D10D86.3030508@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-13T01:06:09Z","receivedAt":"2007-02-13T01:06:09Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 12 Feb 2007, Mark Levedahl wrote:\n\n> Perhaps I'm too paranoid, but I've been burnt way too many times by \n> text/binary mode stuff to let this part be trivialized. Maybe it only \n> gets enabled by core.ImReallyParanoid, but I want that option.\n\nBe aware that what you proposed costs many CPU cycles. I am totally \nopposed to enabling that option by default on all platforms. I am okay \nwith .gitattributes (but I would call it .gitfiletypes), but I am _not_ \nokay with git being _too much_ fscked up by Windows. Microsoft has done \nenough harm already.\n\nCiao,\nDscho\n"},{"id":"34353","messageId":"20070213011329.GB31377@spearce.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702130204120.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-02-13T01:13:29Z","receivedAt":"2007-02-13T01:13:29Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Mon, 12 Feb 2007, Mark Levedahl wrote:\n> \n> > Perhaps I'm too paranoid, but I've been burnt way too many times by \n> > text/binary mode stuff to let this part be trivialized. Maybe it only \n> > gets enabled by core.ImReallyParanoid, but I want that option.\n> \n> Be aware that what you proposed costs many CPU cycles. I am totally \n> opposed to enabling that option by default on all platforms. I am okay \n> with .gitattributes (but I would call it .gitfiletypes), but I am _not_ \n> okay with git being _too much_ fscked up by Windows. Microsoft has done \n> enough harm already.\n\nIndeed; this type of checking should only occur if there is a filter\napplied to a file.  Most files in most projects would hopefully\njust be considered to be byte streams to Git, like they are today,\nand thus not incur any additional overhead, beyond matching their\ntype to determine they are in fact just a byte stream.\n\nThe type could be cached in the index; or at least a single bit\nwhich says \"I'm just a byte stream, thanks\" so that the matching\nonly needs to occur during an initial read-tree.\n\n-- \nShawn.\n"},{"id":"34359","messageId":"Pine.LNX.4.63.0702121715220.6630@qynat.qvtvafvgr.pbz","threadId":"6772","inReplyTo":"20070213011329.GB31377@spearce.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-02-13T01:20:15Z","receivedAt":"2007-02-13T01:20:15Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Mon, 12 Feb 2007, Shawn O. Pearce wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>> On Mon, 12 Feb 2007, Mark Levedahl wrote:\n>>\n>>> Perhaps I'm too paranoid, but I've been burnt way too many times by\n>>> text/binary mode stuff to let this part be trivialized. Maybe it only\n>>> gets enabled by core.ImReallyParanoid, but I want that option.\n>>\n>> Be aware that what you proposed costs many CPU cycles. I am totally\n>> opposed to enabling that option by default on all platforms. I am okay\n>> with .gitattributes (but I would call it .gitfiletypes), but I am _not_\n>> okay with git being _too much_ fscked up by Windows. Microsoft has done\n>> enough harm already.\n>\n> Indeed; this type of checking should only occur if there is a filter\n> applied to a file.  Most files in most projects would hopefully\n> just be considered to be byte streams to Git, like they are today,\n> and thus not incur any additional overhead, beyond matching their\n> type to determine they are in fact just a byte stream.\n>\n> The type could be cached in the index; or at least a single bit\n> which says \"I'm just a byte stream, thanks\" so that the matching\n> only needs to occur during an initial read-tree.\n\nfor the limited case of line endings it may be reasonable to define the internal \ngit format to be lf, and if you are running on a platform that uses this nativly \nno transition is needed\n\none possible way to make this be a general feture is to have the helper script \nhave a --needed flag that tells git if it would do anything on the current \nplatform or not. this way you don't need to run it (and sanity check it) if it's \nnot needed.\n\nDavid Lang\n"},{"id":"34360","messageId":"45D1161C.6040805@verizon.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702130204120.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-13T01:36:28Z","receivedAt":"2007-02-13T01:36:28Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Mon, 12 Feb 2007, Mark Levedahl wrote:\n> \n>> Perhaps I'm too paranoid, but I've been burnt way too many times by \n>> text/binary mode stuff to let this part be trivialized. Maybe it only \n>> gets enabled by core.ImReallyParanoid, but I want that option.\n> \n> Be aware that what you proposed costs many CPU cycles. I am totally \n> opposed to enabling that option by default on all platforms. I am okay \n> with .gitattributes (but I would call it .gitfiletypes), but I am _not_ \n> okay with git being _too much_ fscked up by Windows. Microsoft has done \n> enough harm already.\n\nI would assume that none of this crlf stuff exists at all on Linux / \nUnix / Posix, so if done right has zero impact outside of the Windows \nnuthouse. Inside that, folks are already so used to incredible slowness \nin file I/O that I'm not sure the round tripping I suggest as a check \nwould be very noticeable, but in any case I fully agree it should be \noptional even there. However, if git could support something that never \nscrews up, absolutely guaranteeing data integrity in the presence of \nthese transforms, that would be a first in this arena and I believe a \nsignificant selling point.\n\nMark\n"},{"id":"34361","messageId":"7vodny4xxo.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"45CFA30C.6030202@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-13T02:02:27Z","receivedAt":"2007-02-13T02:02:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mark Levedahl <mlevedahl@verizon.net> writes:\n\n> I am NOT intending to start a flamewar O:-) , so please don't turn\n> this into one.\n\nHeh, a lofty goal.  And I am glad to see that a thread full of\nconstructive suggestion is already going on.\n\nSo now I do not have to fear starting a flamewar; I can safely\nvent.\n\n> The recent threads on a mingw git port are explicit in the intent to\n> provide a Windows native git. I believe there is a fundamental\n> conflict here with the position, clearly stated by Linus, that git\n> does not alter content in any way. Windows suffers the curse of DOS\n> line endings (\\r\\n vs \\n), and a true port to Windows *must* allow for\n> \\r\\n and \\n to be semantically the same thing as most large projects\n> end up with a mixture of such files and/or are targeting\n> cross-platform capabilities. The major competing solutions git seeks\n> to supplant (cvs, cvsnt, svn, hg) have capability to recognize \"text\"\n> files and transparently replace \\r\\n with \\n on input, the reverse on\n> output, and ignore all such differences on diff operations. To be\n> relevant on native Windows, git must do the same. Otherwise, git will\n> be deemed \"too wierd\" and dismissed in favor of a tool \"that works.\"\n> \n> There is no use to debating the technical merits of \\r\\n vs \\n vs \\r\n> vs whatever, nor of not converting. Really. Just accept that there is\n> a fundamental requirement that any version control tool on Windows be\n> able to silently convert between \\r\\n and \\n. To believe otherwise is\n> to expect that the conversion be pushed elsewhere into the tool chain\n> in use, and that won't happen as the competition already provide this\n> conversion capability.\n\nI think there is a fundamental misconception in the above.  I do\nnot know about others, but to me personally, I do not see any\n\"seeking to supplant\", nor \"competition\".  It's not like I or\npeople who raised git into the current shape are begging to\nwindows users to consider using git and bending backwards to\nplease them.  You should hone your diplomacy.\n\nCurrent git may or may not match what they need, and if it does\nnot match what they need, making it match what they need is\nprimarily the responsibility of them.  If Windows users find\nsomething in git that is interesting and useful, but if they\nfind something else lacking in it to be truly useful for them,\nthey can submit patches, or if they cannot implement the changes\nthemselves but only have wishlist items, then _they_ can do the\nbegging.\n\nPeople in git community are certainly friendly and helpful\nbunch, and some (including me) are unfortunate enough that\nsometimes they have to touch Windows, so some degree of need is\nfelt to support Windows better even within the community, but it\nhas never been high priority.  Making it higher priority by\nbringing in better ideas and starting the fire must come from\npeople who care more about Windows than me and Linus.\n\n> So, I think the git project needs to come to an explicit position on\n> this, basically being:\n>\n> 1) git is a POSIX only tool (i.e., there will be no \\r\\n munging), or\n> 2) a Windows port of git will handle and mung \\r\\n and \\n line endings.\n\nI do not think git project needs to do any such thing.  The\nproject evolves reflecting the needs of its users, and the\ndesign is not decided upfront without doing any feasibility\nstudy.  I would certainly not say our position is (1), IOW, I\nwould not say we will rule out Windows support.  If it can be\nreasonably done without harming the code, why not?\n\nDepending on how cleanly a change Windows users want is done\nwithout negatively affecting the existing users, it may or may\nnot be judged acceptable.  We will know only when we see at\nleast the design and preferably the code.  I feel no need to\ndecide between (1) and (2) upfront before that happens.\n\n> If the answer is 1, the mingw port is a waste of time as it simply\n> won't be usable by its target audience. If the answer is 2, then I\n> think a very careful design of this capability is in order.\n>\n> Comments?\n\nThis is not just you, and fortunately it does not happen very\noften in git community, but I find it _very_ irritating when\nsomebody says: \"here is a patch, I'll do the doc, test, and\ntidying up if this patch is accepted\".  I usually pretend to be\na nice person and accept the patch when it is obviously good,\nor pretend that I was too busy and did not notice such a\nmessage, but I feel _very_ tempted to say: \"if you care deeply\nenough that what you did is useful, I expect you'd perfect it\nwhether or not I apply your patch to my tree right now.  If even\nthe original author, you, do not find it worth perfecting, then\nI am not interested at all.\"\n\nEven if all existing git community members felt (1) above and\nwere unwilling to accept line-end conversions (which by now you\nalready know is not the case -- and that is why I waited until\nnow to address this as a separate \"attitude\" issue), if somebody\nwho works on Windows is motivated enough to make git work better\nfor him, he can fork (and forking is very easy with git).  If\nthe forked git works well both on Windows and on non Windows,\npeople who initially felt (1) will realize that they were wrong\nand then the codebase can be merged back together (and merging\nthe forked projects is very easy with git).\n\nIt's open source.  People shouldn't worry too much about what\nthey have done \"wasted\".  You are not even talking about what\nyou've already done -- you are talking about what you _might_\ndo.\n\nAnd your saying \"If 2, then we need to think carefully\" was VERY\ngood.  My point is that you did not have to say \"Is it 1, or is\nit 2, and if 2 then\" part.\n"},{"id":"34366","messageId":"45D12EA4.8090306@verizon.net","threadId":"6772","inReplyTo":"7vodny4xxo.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-13T03:21:08Z","receivedAt":"2007-02-13T03:21:08Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Junio C Hamano wrote:\n > Mark Levedahl <mlevedahl@verizon.net> writes:\n >\n >> I am NOT intending to start a flamewar O:-) , so please don't turn\n >> this into one.\n >\n > Heh, a lofty goal.  And I am glad to see that a thread full of\n > constructive suggestion is already going on.\n >\n > So now I do not have to fear starting a flamewar; I can safely\n > vent.\n\nJunio,\n\nI meant absolutely no offense in anything I wrote, and sincerely \napologize if any was taken. My past experiences caused me to be \nskeptical that a significant change to accommodate a very bad design of \nWindows would be accepted here. Happily, that skepticism was misplaced. \nI am much heartened by the responses, and am optimistic a good solution \nwill be found that is acceptable to all. It is very clear that the group \nis open and supportive of working through the issues to help this, and I \nintend to contribute to that solution. (If nothing else, I would like to \nbe known for something besides some modest ability to hack around Tk bugs).\n\nSo, I trust the flamethrowers can remain buried.\n\nMark\n"},{"id":"34368","messageId":"200702130932.51601.litvinov2004@gmail.com","threadId":"6772","inReplyTo":"45CFA30C.6030202@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-02-13T03:32:51Z","receivedAt":"2007-02-13T03:32:51Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Monday 12 February 2007 05:13 Mark Levedahl написал(a):\n> 1) git is a POSIX only tool (i.e., there will be no \\r\\n munging), or\n> 2) a Windows port of git will handle and mung \\r\\n and \\n line endings.\n>\n> If the answer is 1, the mingw port is a waste of time as it simply won't\n> be usable by its target audience. If the answer is 2, then I think a\n> very careful design of this capability is in order.\n\nI am strongly object this statement. I develop one project under Windows and \nuse Cygwin git for this. Yes, I have a problem with git's thinking line \nending is a \\n but most of troubles are diff and rebase. In general git works \nwell with \\r\\n line endings.\n\nWhen I have file that was converted from dos to unix format (or from unix to \ndos) git genereta big diff. But anyway, c++ compiler works well with both \nformats and in this case I simply convert file to dos format and git shows \nagain nice diff. If unix format was commited to git I simply change the \nformat and commit that file again.\n\nThe only trouble is the rebase, it does not like \\r\\n ending and othen produce \nunexpected merge conflict. But I don't use rebse to othen to realy \ninvestigate and try to solve the problem.\n"},{"id":"34373","messageId":"20070213051816.GB328@coredump.intra.peff.net","threadId":"6772","inReplyTo":"45D10D86.3030508@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-02-13T05:18:16Z","receivedAt":"2007-02-13T05:18:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 12, 2007 at 07:59:50PM -0500, Mark Levedahl wrote:\n\n> fail to set the correct file type somewhere. Current cvsnt is fairly \n> good at autodetecting and setting text vs binary file type, and enforces \n> this across all platforms, so things don't go awry too often. It is in \n\nThere is obviously much sentiment that this should _not_ be the default\n(and I agree). But if arbitrary filters are possible, then you can\ntheoretically write an 'autocrlf' filter which will try to do the right\nthing, and you could set it for some or all files:\n\n  echo '*: autocrlf' >.gitattributes\n\nbut it would be off by default. If we implement this, everyone has to\n\"pay\" for .gitattributes (even if you don't use it, we have to look it\nup to make sure you're not using it!), but nobody has to pay for any\nfilters they don't use.\n\n-Peff\n"},{"id":"34379","messageId":"7v3b5a384v.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"45D12EA4.8090306@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-13T06:05:04Z","receivedAt":"2007-02-13T06:05:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mark Levedahl <mdl123@verizon.net> writes:\n\n> I meant absolutely no offense in anything I wrote, and sincerely\n> apologize if any was taken.\n\nNone taken, although I admit that I was somewhat annoyed, having\nto write the first part of my response.\n"},{"id":"34389","messageId":"Pine.LNX.4.63.0702131105240.1300@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"200702130932.51601.litvinov2004@gmail.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-13T10:06:17Z","receivedAt":"2007-02-13T10:06:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 13 Feb 2007, Alexander Litvinov wrote:\n\n> When I have file that was converted from dos to unix format (or from \n> unix to dos) git genereta big diff. But anyway, c++ compiler works well \n> with both formats and in this case I simply convert file to dos format \n> and git shows again nice diff. If unix format was commited to git I \n> simply change the format and commit that file again.\n\nThat's awful!\n\n> The only trouble is the rebase, it does not like \\r\\n ending and othen \n> produce unexpected merge conflict. But I don't use rebse to othen to \n> realy investigate and try to solve the problem.\n\nWell, if everybody thinks like you, maybe we do not have to change \nanything for Windows after all?\n\nCiao,\nDscho\n"},{"id":"34400","messageId":"200702131816.27705.litvinov2004@gmail.com","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702131105240.1300@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-02-13T12:16:27Z","receivedAt":"2007-02-13T12:16:27Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Tuesday 13 February 2007 16:06 Johannes Schindelin написал(a):\n> Hi,\n>\n> On Tue, 13 Feb 2007, Alexander Litvinov wrote:\n> > When I have file that was converted from dos to unix format (or from\n> > unix to dos) git genereta big diff. But anyway, c++ compiler works well\n> > with both formats and in this case I simply convert file to dos format\n> > and git shows again nice diff. If unix format was commited to git I\n> > simply change the format and commit that file again.\n>\n> That's awful!\nIf you are tring to build history that looks good - you are right this is a \nterrible workflow.\n\n> > The only trouble is the rebase, it does not like \\r\\n ending and othen\n> > produce unexpected merge conflict. But I don't use rebse to othen to\n> > realy investigate and try to solve the problem.\n>\n> Well, if everybody thinks like you, maybe we do not have to change\n> anything for Windows after all?\nI still wish to have working rebase so if git will hanle somehow \\r\\n it would \nbe nice. But please do not produce the same behavior as cvs does: under \ncygwin it still use \\n !\n\nBy the way, most windows programmers I work with says 'git is cool but is \nthere gui like tortoise or wincvs ?' :-)\n"},{"id":"34401","messageId":"Pine.LNX.4.63.0702131328440.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"200702131816.27705.litvinov2004@gmail.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-13T12:37:22Z","receivedAt":"2007-02-13T12:37:22Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 13 Feb 2007, Alexander Litvinov wrote:\n\n> Tuesday 13 February 2007 16:06 Johannes Schindelin:\n> > At some stage, Alexander wrote this:\n> \n> > > The only trouble is the rebase, it does not like \\r\\n ending and othen\n> > > produce unexpected merge conflict. But I don't use rebse to othen to\n> > > realy investigate and try to solve the problem.\n> >\n> > Well, if everybody thinks like you, maybe we do not have to change\n> > anything for Windows after all?\n>\n> I still wish to have working rebase so if git will hanle somehow \\r\\n it \n> would be nice. But please do not produce the same behavior as cvs does: \n> under cygwin it still use \\n !\n\nYou really should teach format-patch to output \\n patches, and keep all \nyour blobs CR free.\n\n> By the way, most windows programmers I work with says 'git is cool but \n> is there gui like tortoise or wincvs ?' :-)\n\nSome time ago, I started playing with a shell extension. Now that MinGW \ngit is almost there, I might clean it up... Would you be interested in \nworking on it, or is this just wishtalk?\n\nCiao,\nDscho\n"},{"id":"34423","messageId":"Pine.LNX.4.64.0702130845330.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702131105240.1300@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T16:52:04Z","receivedAt":"2007-02-13T16:52:04Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Johannes Schindelin wrote:\n> \n> On Tue, 13 Feb 2007, Alexander Litvinov wrote:\n> \n> > When I have file that was converted from dos to unix format (or from \n> > unix to dos) git genereta big diff. But anyway, c++ compiler works well \n> > with both formats and in this case I simply convert file to dos format \n> > and git shows again nice diff. If unix format was commited to git I \n> > simply change the format and commit that file again.\n> \n> That's awful!\n> \n> > The only trouble is the rebase, it does not like \\r\\n ending and othen \n> > produce unexpected merge conflict. But I don't use rebse to othen to \n> > realy investigate and try to solve the problem.\n> \n> Well, if everybody thinks like you, maybe we do not have to change \n> anything for Windows after all?\n\nNo no no.\n\nIt's going to be _horrible_ if people start interesting projects in \nWindows, and there are files in a git repository that are encoded with \nCRLF. \n\nI'd much rather just get this right, and that means \"no hooks\". If people \nstart using commit hooks etc, that will just mean that they won't use them \nfor all-windows environments (why use it? Everybody hass CRLF, and \neverybody _wants_ CRLF), or it will just be relatively expensive to have a \ncomplex hook anyway.\n\nSo I think we should plan on something like .gitattributes or similar, so \nthat we _can_ handle mixed environments well, without any real setup or \nany real costs.\n\nThe costs really shouldn't be too high - we tend to avoid doing any \nexpensive working tree changes *anyway*. For example, even \"git checkout\" \nhas a huge optimization to avoid rewriting files that are already ok, so \ndoing things like switching whole branches usually wouldn't even need any \nconversion for most files - even on platforms like Windows that need the \nconversion in the first place.\n\nSo considering that it looks _trivial_ for git-update-index, fairly easy \nfor diff generation, and I doubt \"git checkout\" is really likely to be any \nworse either, this should just be somethign we do.\n\nThe *ONLY* case where we may not be able to do things automatically is \nactually a much more subtle one: \"git cat-file\". If we just get a SHA1, we \ndon't know what the path to look it up was like, and thus we can never \nknow whether it's a binary or a text object. With \"-p\" we can trivially \nguess, of course, but \"git cat-file blob\" simply must not do that!\n\nBut that really doesn't sound like a big problem to me ;)\n\n\t\tLinus\n"},{"id":"34428","messageId":"Pine.LNX.4.64.0702130919100.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702130845330.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T17:23:24Z","receivedAt":"2007-02-13T17:23:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Linus Torvalds wrote:\n> \n> I'd much rather just get this right, and that means \"no hooks\". If people \n> start using commit hooks etc, that will just mean that they won't use them \n> for all-windows environments (why use it? Everybody hass CRLF, and \n> everybody _wants_ CRLF), or it will just be relatively expensive to have a \n> complex hook anyway.\n> \n> So I think we should plan on something like .gitattributes or similar, so \n> that we _can_ handle mixed environments well, without any real setup or \n> any real costs.\n\nHere's a patch that I think we can merge right now. There may be other \nplaces that need this, but this at least points out the three places that \nread/write working tree files for git update-index, checkout and diff \nrespectively. That should cover a lot of it.\n\nSome day we can actually implement it. In the meantime, this points out a \nplace for people to start. We *can* even start with a really simple \"we do \nCRLF conversion automatically, regardless of filename\" kind of approach, \nthat just look at the data (all three cases have the _full_ file data \nalready in memory) and says \"ok, this is text, so let's convert to/from \nDOS format directly\".\n\nTHAT somebody can write in ten minutes, and it would already make git much \nnicer on a DOS/Windows platform, I suspect.\n\nAnd it would be totally zero-cost if you just make it a config option \n(but please make it dynamic with the _default_ just being 0/1 depending \non whether it's UNIX/Windows, just so that UNIX people can _test_ it \neasily).\n\n\t\tLinus\n"},{"id":"34429","messageId":"Pine.LNX.4.64.0702130923310.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702130919100.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T17:23:59Z","receivedAt":"2007-02-13T17:23:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Linus Torvalds wrote:\n> \n> Here's a patch [...]\n\nNo. HERE's the trivial stupid patch that just marks the core places.\n\n\t\tLinus\n---\ndiff --git a/diff.c b/diff.c\nindex aaab309..13b9b6c 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1364,6 +1364,7 @@ int diff_populate_filespec(struct diff_filespec *s, int size_only)\n \t\ts->data = xmmap(NULL, s->size, PROT_READ, MAP_PRIVATE, fd, 0);\n \t\tclose(fd);\n \t\ts->should_munmap = 1;\n+\t\t/* FIXME! CRLF -> LF conversion goes here, based on \"s->path\" */\n \t}\n \telse {\n \t\tchar type[20];\ndiff --git a/entry.c b/entry.c\nindex 0ebf0f0..c2641dd 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -89,6 +89,7 @@ static int write_entry(struct cache_entry *ce, char *path, struct checkout *stat\n \t\t\treturn error(\"git-checkout-index: unable to create file %s (%s)\",\n \t\t\t\tpath, strerror(errno));\n \t\t}\n+\t\t/* FIXME: LF -> CRLF conversion goes here, based on \"ce->name\" */\n \t\twrote = write_in_full(fd, new, size);\n \t\tclose(fd);\n \t\tfree(new);\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 0d4bf80..8ad7fad 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2091,6 +2091,7 @@ int index_fd(unsigned char *sha1, int fd, struct stat *st, int write_object, con\n \n \tif (!type)\n \t\ttype = blob_type;\n+\t/* FIXME: CRLF -> LF conversion here for blobs! We'll need the path! */\n \tif (write_object)\n \t\tret = write_sha1_file(buf, size, type, sha1);\n \telse\n"},{"id":"34430","messageId":"Pine.LNX.4.64.0702131219220.1757@xanadu.home","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702130845330.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-13T17:25:23Z","receivedAt":"2007-02-13T17:25:23Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 13 Feb 2007, Linus Torvalds wrote:\n\n> The *ONLY* case where we may not be able to do things automatically is \n> actually a much more subtle one: \"git cat-file\". If we just get a SHA1, we \n> don't know what the path to look it up was like, and thus we can never \n> know whether it's a binary or a text object. With \"-p\" we can trivially \n> guess, of course, but \"git cat-file blob\" simply must not do that!\n\ngit-cat-file, and its counter part git-hash-object, are fairly low level \nplumbing.  Anyone using them should be aware of the issue and apply the \nneeded conversion.  And actually, since we're going to have the \nconversion routines in the core, we'd only need to add a --crlf argument \nto both of them to optionally perform the conversion since the user of \nthose commands is more likely to know if the conversion is needed.\n\n\nNicolas\n"},{"id":"34441","messageId":"7v7iumx7hu.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702130919100.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-13T18:00:45Z","receivedAt":"2007-02-13T18:00:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Here's a patch that I think we can merge right now. There may be other \n> places that need this, but this at least points out the three places that \n> read/write working tree files for git update-index, checkout and diff \n> respectively. That should cover a lot of it.\n\nThanks, applied.  I think git-apply has separate codepaths for\nboth reading and writing; I won't look into them before 1.5.0\nbut people are welcome to help advancing the cause before I get\nto it ;-).\n"},{"id":"34443","messageId":"Pine.LNX.4.63.0702131901040.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702130845330.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-13T18:04:26Z","receivedAt":"2007-02-13T18:04:26Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 13 Feb 2007, Linus Torvalds wrote:\n\n> On Tue, 13 Feb 2007, Johannes Schindelin wrote:\n> > \n> > On Tue, 13 Feb 2007, Alexander Litvinov wrote:\n> > \n> > > The only trouble is the rebase, it does not like \\r\\n ending and \n> > > othen produce unexpected merge conflict. But I don't use rebse to \n> > > othen to realy investigate and try to solve the problem.\n> > \n> > Well, if everybody thinks like you, maybe we do not have to change \n> > anything for Windows after all?\n> \n> No no no.\n> \n> It's going to be _horrible_ if people start interesting projects in \n> Windows, and there are files in a git repository that are encoded with \n> CRLF.\n> \n> I'd much rather just get this right, and that means \"no hooks\".\n\nNo hooks means something like cvsnt does, and that means no .gitattributes \neither. (BTW I really hate .gitattributes, as it does not at all say what \nthis is about; it's about file _conversions_, not attributes).\n\nCVSNT analyzes the files, and guesses if they are text, and only then \nactivates the text mode.\n\nI am strongly opposed to including something like that. (It was already \nproposed, and your \"no hooks\" suggests the same.)\n\nHowever, I am slightly positive about the .gitfiletypes approach, _iff_ we \nthink about more than just text/binary from the start. If we do it right, \nit will buy us more.\n\nCiao,\nDscho\n"},{"id":"34444","messageId":"Pine.LNX.4.63.0702131904330.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702130919100.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-13T18:05:55Z","receivedAt":"2007-02-13T18:05:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 13 Feb 2007, Linus Torvalds wrote:\n\n> Here's a patch that I think we can merge right now.\n\nWhy the haste all of a sudden? Your patch is easily applyable for anyone \nwho wants to work on text/binary or arbitrary file types. No need to rush \na developers-only patch into git.git.\n\nCiao,\nDscho\n"},{"id":"34451","messageId":"7vlkj2vset.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702131901040.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-13T18:11:54Z","receivedAt":"2007-02-13T18:11:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> No hooks means something like cvsnt does, and that means no .gitattributes \n> either. (BTW I really hate .gitattributes, as it does not at all say what \n> this is about; it's about file _conversions_, not attributes).\n\n> However, I am slightly positive about the .gitfiletypes approach, _iff_ we \n> think about more than just text/binary from the start. If we do it right, \n> it will buy us more.\n\nWe might start with only binary/text attributes, but we may add\nmore later, e.g. chmod=o-rwx.  I do not see much differnece\nbetween attributes vs filetypes.\n"},{"id":"34459","messageId":"Pine.LNX.4.64.0702131037330.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702131901040.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T18:39:19Z","receivedAt":"2007-02-13T18:39:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Johannes Schindelin wrote:\n> \n> No hooks means something like cvsnt does, and that means no .gitattributes \n> either. (BTW I really hate .gitattributes, as it does not at all say what \n> this is about; it's about file _conversions_, not attributes).\n\nNo, it *is* about attributes.\n\nIn order to know how to convert, you need to know the attributes of the \nfile.\n\nSo it's not about conversion: we would ALWAYS do conversion. It's about \nthe fact that in order to do the conversion, we need to know what the \nattributes of the file is - is it text, or what.\n\nAnd the equal point is that there are _other_ attributes that git might \ncare about. The \"merge strategy\" attribute, for example. Or \"owner\" \nattributes for files etc.\n\n\t\tLinus\n"},{"id":"34461","messageId":"Pine.LNX.4.63.0702131942290.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131037330.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-13T18:42:50Z","receivedAt":"2007-02-13T18:42:50Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 13 Feb 2007, Linus Torvalds wrote:\n\n> On Tue, 13 Feb 2007, Johannes Schindelin wrote:\n> > \n> > No hooks means something like cvsnt does, and that means no .gitattributes \n> > either. (BTW I really hate .gitattributes, as it does not at all say what \n> > this is about; it's about file _conversions_, not attributes).\n> \n> No, it *is* about attributes.\n> \n> In order to know how to convert, you need to know the attributes of the \n> file.\n> \n> So it's not about conversion: we would ALWAYS do conversion. It's about \n> the fact that in order to do the conversion, we need to know what the \n> attributes of the file is - is it text, or what.\n> \n> And the equal point is that there are _other_ attributes that git might \n> care about. The \"merge strategy\" attribute, for example. Or \"owner\" \n> attributes for files etc.\n\nYes, you're right. Colour me converted (pun intended).\n\nCiao,\nDscho\n"},{"id":"34462","messageId":"Pine.LNX.4.64.0702131053110.8424@woody.linux-foundation.org","threadId":"6772","inReplyTo":"7v7iumx7hu.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T19:07:23Z","receivedAt":"2007-02-13T19:07:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Junio C Hamano wrote:\n> \n> Thanks, applied.  I think git-apply has separate codepaths for\n> both reading and writing; I won't look into them before 1.5.0\n> but people are welcome to help advancing the cause before I get\n> to it ;-).\n\nActually, I did it myself.\n\nThis is a \"lazy man's auto-CRLF\", and it really is pretty simple.\n\nIt currently does NOT know about file attributes, so it does its \nconversion purely based on content. Maybe that is more in the \"git \nphilosophy\" anyway, since content is king, but I think we should try to do \nthe file attributes to turn it off on demand.\n\nAnyway, BY DEFAULT it is off regardless, because it requires a\n\n\t[core]\n\t\tAutoCRLF = true\n\nin your config file to be enabled. We could make that the default for \nWindows, of course, the same way we do some other things (filemode etc).\n\nBut you can actually enable it on UNIX, and it will cause:\n\n - \"git update-index\" will write blobs without CRLF\n - \"git diff\" will diff working tree files without CRLF\n - \"git checkout\" will write files to the working tree _with_ CRLF\n\nand things work fine.\n\nFunnily, it actually shows an odd file in git itself:\n\n\tgit clone -n git test-crlf\n\tcd test-crlf\n\tgit config core.autocrlf true\n\tgit checkout\n\tgit diff\n\nshows a diff for \"Documentation/docbook-xsl.css\". Why? Because we have \nactually checked in that file *with* CRLF! So when \"core.autocrlf\" is \ntrue, we'll always generate a *different* hash for it in the index, \nbecause the index hash will be for the content _without_ CRLF.\n\nIs this complete? I dunno. It seems to work for me. It doesn't use the \nfilename at all right now, and that's probably a deficiency (we could \ncertainly make the \"is_binary()\" heuristics also take standard filename \nheuristics into account).\n\nI don't pass in the filename at all for the \"index_fd()\" case \n(git-update-index), so that would need to be passed around, but this \nactually works fine.\n\nNOTE NOTE NOTE! The \"is_binary()\" heuristics are totally made-up by yours \ntruly. I will not guarantee that they work at all reasonable. Caveat \nemptor. But it _is_ simple, and it _is_ safe, since it's all off by \ndefault.\n\nThe patch is pretty simple - the biggest part is the new \"convert.c\" file, \nbut even that is really just basic stuff that anybody can write in \n\"Teaching C 101\" as a final project for their first class in programming. \nNot to say that it's bug-free, of course - but at least we're not talking \nabout rocket surgery here.\n\n\t\tLinus\n\n---\ncommit f0731319497ac8121bd901a91fc33d715745d3af\nAuthor: Linus Torvalds <torvalds@osdl.org>\nDate:   Tue Feb 13 10:56:50 2007 -0800\n\n    Add \"auto-CRLF\" conversion logic\n    \n    It's simple and it's stupid.  But it actually seems to work.  What more\n    can you want?\n    \n    It's not enabled by default: you need to add a\n    \n    \t[core]\n    \t\tAutoCRLF = true\n    \n    to your .git/config file to enable it universally.\n    \n    Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n Makefile      |    3 +-\n cache.h       |    5 ++\n config.c      |    5 ++\n convert.c     |  179 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n diff.c        |   16 +++++\n entry.c       |   15 +++++\n environment.c |    1 +\n sha1_file.c   |   22 +++++++-\n 8 files changed, 244 insertions(+), 2 deletions(-)\n create mode 100644 convert.c\n\ndiff --git a/Makefile b/Makefile\nindex 40bdcff..60496ff 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -262,7 +262,8 @@ LIB_OBJS = \\\n \trevision.o pager.o tree-walk.o xdiff-interface.o \\\n \twrite_or_die.o trace.o list-objects.o grep.o \\\n \talloc.o merge-file.o path-list.o help.o unpack-trees.o $(DIFF_OBJS) \\\n-\tcolor.o wt-status.o archive-zip.o archive-tar.o shallow.o utf8.o\n+\tcolor.o wt-status.o archive-zip.o archive-tar.o shallow.o utf8.o \\\n+\tconvert.o\n \n BUILTIN_OBJS = \\\n \tbuiltin-add.o \\\ndiff --git a/cache.h b/cache.h\nindex c62b0b0..9c019e8 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -201,6 +201,7 @@ extern const char *apply_default_whitespace;\n extern int zlib_compression_level;\n extern size_t packed_git_window_size;\n extern size_t packed_git_limit;\n+extern int auto_crlf;\n \n #define GIT_REPO_VERSION 0\n extern int repository_format_version;\n@@ -468,4 +469,8 @@ extern int nfvasprintf(char **str, const char *fmt, va_list va);\n extern void trace_printf(const char *format, ...);\n extern void trace_argv_printf(const char **argv, int count, const char *format, ...);\n \n+/* convert.c */\n+extern int convert_to_git(const char *path, char **bufp, unsigned long *sizep);\n+extern int convert_to_working_tree(const char *path, char **bufp, unsigned long *sizep);\n+\n #endif /* CACHE_H */\ndiff --git a/config.c b/config.c\nindex d821071..ffe0212 100644\n--- a/config.c\n+++ b/config.c\n@@ -324,6 +324,11 @@ int git_default_config(const char *var, const char *value)\n \t\treturn 0;\n \t}\n \n+\tif (!strcmp(var, \"core.autocrlf\")) {\n+\t\tauto_crlf = git_config_bool(var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"user.name\")) {\n \t\tstrlcpy(git_default_name, value, sizeof(git_default_name));\n \t\treturn 0;\ndiff --git a/convert.c b/convert.c\nnew file mode 100644\nindex 0000000..c04b6c2\n--- /dev/null\n+++ b/convert.c\n@@ -0,0 +1,179 @@\n+#include \"cache.h\"\n+/*\n+ * convert.c - convert a file when checking it out and checking it in.\n+ *\n+ * This should use the pathname to decide on whether it wants to do some\n+ * more interesting conversions (automatic gzip/unzip, general format\n+ * conversions etc etc), but by default it just does automatic CRLF<->LF\n+ * translation when the \"auto_crlf\" option is set.\n+ */\n+\n+struct text_stat {\n+\t/* CR, LF and CRLF counts */\n+\tunsigned cr, lf, crlf;\n+\n+\t/* These are just approximations! */\n+\tunsigned printable, nonprintable;\n+};\n+\n+static void gather_stats(const char *buf, unsigned long size, struct text_stat *stats)\n+{\n+\tunsigned long i;\n+\n+\tmemset(stats, 0, sizeof(*stats));\n+\n+\tfor (i = 0; i < size; i++) {\n+\t\tunsigned char c = buf[i];\n+\t\tif (c == '\\r') {\n+\t\t\tstats->cr++;\n+\t\t\tif (i+1 < size && buf[i+1] == '\\n')\n+\t\t\t\tstats->crlf++;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (c == '\\n') {\n+\t\t\tstats->lf++;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (c == '\\t' || (c >= 32 && c < 127)) {\n+\t\t\tstats->printable++;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tstats->nonprintable++;\n+\t}\n+}\n+\n+/*\n+ * This is just a heuristic!\n+ *\n+ * We do allow nonprintable characters (utf-8 and latin1 etc), but we\n+ * require that they are just a fairly small percentage of the total\n+ * file. \n+ */\n+static int is_binary(unsigned long size, struct text_stat *stats)\n+{\n+\tif (stats->nonprintable > (size >> 3))\n+\t\treturn 1;\n+\t/*\n+\t * Other heuristics? Average line length might be relevant,\n+\t * as might LF vs CR vs CRLF counts..\n+\t *\n+\t * NOTE! It might be normal to have a low ratio of CRLF to LF\n+\t * (somebody starts with a LF-only file and edits it with an editor\n+\t * that adds CRLF only to lines that are added..). But do  we\n+\t * want to support CR-only? Probably not.\n+\t */\n+\treturn 0;\n+}\n+\n+int convert_to_git(const char *path, char **bufp, unsigned long *sizep)\n+{\n+\tchar *buffer, *nbuf;\n+\tunsigned long size, nsize;\n+\tstruct text_stat stats;\n+\n+\t/*\n+\t * FIXME! Other pluggable conversions should go here,\n+\t * based on filename patterns. Right now we just do the\n+\t * stupid auto-CRLF one.\n+\t */\n+\tif (!auto_crlf)\n+\t\treturn 0;\n+\n+\tsize = *sizep;\n+\tif (!size)\n+\t\treturn 0;\n+\tbuffer = *bufp;\n+\n+\tgather_stats(buffer, size, &stats);\n+\n+\t/* No CR? Nothing to convert, regardless. */\n+\tif (!stats.cr)\n+\t\treturn 0;\n+\n+\t/*\n+\t * We're currently not going to even try to convert stuff\n+\t * that has bare CR characters. Does anybody do that crazy\n+\t * stuff?\n+\t */\n+\tif (stats.cr != stats.crlf)\n+\t\treturn 0;\n+\n+\t/*\n+\t * And add some heuristics for binary vs text, of course.. \n+\t */\n+\tif (is_binary(size, &stats))\n+\t\treturn 0;\n+\n+\t/*\n+\t * Ok, allocate a new buffer, fill it in, and return true\n+\t * to let the caller know that we switched buffers on it.\n+\t */\n+\tnsize = size - stats.crlf;\n+\tnbuf = xmalloc(nsize);\n+\t*bufp = nbuf;\n+\t*sizep = nsize;\n+\tdo {\n+\t\tunsigned char c = *buffer++;\n+\t\tif (c != '\\r')\n+\t\t\t*nbuf++ = c;\n+\t} while (--size);\n+\n+\treturn 1;\n+}\n+\n+int convert_to_working_tree(const char *path, char **bufp, unsigned long *sizep)\n+{\n+\tchar *buffer, *nbuf;\n+\tunsigned long size, nsize;\n+\tstruct text_stat stats;\n+\tunsigned char last;\n+\n+\t/*\n+\t * FIXME! Other pluggable conversions should go here,\n+\t * based on filename patterns. Right now we just do the\n+\t * stupid auto-CRLF one.\n+\t */\n+\tif (!auto_crlf)\n+\t\treturn 0;\n+\n+\tsize = *sizep;\n+\tif (!size)\n+\t\treturn 0;\n+\tbuffer = *bufp;\n+\n+\tgather_stats(buffer, size, &stats);\n+\n+\t/* No LF? Nothing to convert, regardless. */\n+\tif (!stats.lf)\n+\t\treturn 0;\n+\n+\t/* Was it already in CRLF format? */\n+\tif (stats.lf == stats.crlf)\n+\t\treturn 0;\n+\n+\t/* If we have any bare CR characters, we're not going to touch it */\n+\tif (stats.cr != stats.crlf)\n+\t\treturn 0;\n+\n+\tif (is_binary(size, &stats))\n+\t\treturn 0;\n+\n+\t/*\n+\t * Ok, allocate a new buffer, fill it in, and return true\n+\t * to let the caller know that we switched buffers on it.\n+\t */\n+\tnsize = size + stats.lf - stats.crlf;\n+\tnbuf = xmalloc(nsize);\n+\t*bufp = nbuf;\n+\t*sizep = nsize;\n+\tlast = 0;\n+\tdo {\n+\t\tunsigned char c = *buffer++;\n+\t\tif (c == '\\n' && last != '\\r')\n+\t\t\t*nbuf++ = '\\r';\n+\t\t*nbuf++ = c;\n+\t\tlast = c;\n+\t} while (--size);\n+\n+\treturn 1;\n+}\ndiff --git a/diff.c b/diff.c\nindex aaab309..561587c 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1332,6 +1332,9 @@ int diff_populate_filespec(struct diff_filespec *s, int size_only)\n \t    reuse_worktree_file(s->path, s->sha1, 0)) {\n \t\tstruct stat st;\n \t\tint fd;\n+\t\tchar *buf;\n+\t\tunsigned long size;\n+\n \t\tif (lstat(s->path, &st) < 0) {\n \t\t\tif (errno == ENOENT) {\n \t\t\terr_empty:\n@@ -1364,6 +1367,19 @@ int diff_populate_filespec(struct diff_filespec *s, int size_only)\n \t\ts->data = xmmap(NULL, s->size, PROT_READ, MAP_PRIVATE, fd, 0);\n \t\tclose(fd);\n \t\ts->should_munmap = 1;\n+\n+\t\t/*\n+\t\t * Convert from working tree format to canonical git format\n+\t\t */\n+\t\tbuf = s->data;\n+\t\tsize = s->size;\n+\t\tif (convert_to_git(s->path, &buf, &size)) {\n+\t\t\tmunmap(s->data, s->size);\n+\t\t\ts->should_munmap = 0;\n+\t\t\ts->data = buf;\n+\t\t\ts->size = size;\n+\t\t\ts->should_free = 1;\n+\t\t}\n \t}\n \telse {\n \t\tchar type[20];\ndiff --git a/entry.c b/entry.c\nindex 0ebf0f0..472a9ef 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -78,6 +78,9 @@ static int write_entry(struct cache_entry *ce, char *path, struct checkout *stat\n \t\t\tpath, sha1_to_hex(ce->sha1));\n \t}\n \tswitch (ntohl(ce->ce_mode) & S_IFMT) {\n+\t\tchar *buf;\n+\t\tunsigned long nsize;\n+\n \tcase S_IFREG:\n \t\tif (to_tempfile) {\n \t\t\tstrcpy(path, \".merge_file_XXXXXX\");\n@@ -89,6 +92,18 @@ static int write_entry(struct cache_entry *ce, char *path, struct checkout *stat\n \t\t\treturn error(\"git-checkout-index: unable to create file %s (%s)\",\n \t\t\t\tpath, strerror(errno));\n \t\t}\n+\n+\t\t/*\n+\t\t * Convert from git internal format to working tree format\n+\t\t */\n+\t\tbuf = new;\n+\t\tnsize = size;\n+\t\tif (convert_to_working_tree(ce->name, &buf, &nsize)) {\n+\t\t\tfree(new);\n+\t\t\tnew = buf;\n+\t\t\tsize = nsize;\n+\t\t}\n+\n \t\twrote = write_in_full(fd, new, size);\n \t\tclose(fd);\n \t\tfree(new);\ndiff --git a/environment.c b/environment.c\nindex 54c22f8..2fa0960 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -28,6 +28,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n int pager_in_use;\n int pager_use_color = 1;\n+int auto_crlf = 0;\n \n static const char *git_dir;\n static char *git_object_dir, *git_index_file, *git_refs_dir, *git_graft_file;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 0d4bf80..6ec67b2 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2082,7 +2082,7 @@ int index_fd(unsigned char *sha1, int fd, struct stat *st, int write_object, con\n {\n \tunsigned long size = st->st_size;\n \tvoid *buf;\n-\tint ret;\n+\tint ret, re_allocated = 0;\n \n \tbuf = \"\";\n \tif (size)\n@@ -2091,10 +2091,30 @@ int index_fd(unsigned char *sha1, int fd, struct stat *st, int write_object, con\n \n \tif (!type)\n \t\ttype = blob_type;\n+\n+\t/*\n+\t * Convert blobs to git internal format\n+\t */\n+\tif (!strcmp(type, blob_type)) {\n+\t\tunsigned long nsize = size;\n+\t\tchar *nbuf = buf;\n+\t\tif (convert_to_git(NULL, &nbuf, &nsize)) {\n+\t\t\tif (size)\n+\t\t\t\tmunmap(buf, size);\n+\t\t\tsize = nsize;\n+\t\t\tbuf = nbuf;\n+\t\t\tre_allocated = 1;\n+\t\t}\n+\t}\n+\n \tif (write_object)\n \t\tret = write_sha1_file(buf, size, type, sha1);\n \telse\n \t\tret = hash_sha1_file(buf, size, type, sha1);\n+\tif (re_allocated) {\n+\t\tfree(buf);\n+\t\treturn ret;\n+\t}\n \tif (size)\n \t\tmunmap(buf, size);\n \treturn ret;\n"},{"id":"34463","messageId":"eqt40c$5ov$1@sea.gmane.org","threadId":"6772","inReplyTo":"200702131816.27705.litvinov2004@gmail.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-13T19:36:44Z","receivedAt":"2007-02-13T19:36:44Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Alexander Litvinov wrote:\n\n> ? ????????? ?? Tuesday 13 February 2007 16:06 Johannes Schindelin\n> ???????(a):\n>> Hi,\n>>\n>> On Tue, 13 Feb 2007, Alexander Litvinov wrote:\n>> > When I have file that was converted from dos to unix format (or from\n>> > unix to dos) git genereta big diff. But anyway, c++ compiler works well\n>> > with both formats and in this case I simply convert file to dos format\n>> > and git shows again nice diff. If unix format was commited to git I\n>> > simply change the format and commit that file again.\n>>\n>> That's awful!\n> If you are tring to build history that looks good - you are right this is\n> a terrible workflow.\n> \n>> > The only trouble is the rebase, it does not like \\r\\n ending and othen\n>> > produce unexpected merge conflict. But I don't use rebse to othen to\n>> > realy investigate and try to solve the problem.\n>>\n>> Well, if everybody thinks like you, maybe we do not have to change\n>> anything for Windows after all?\n> I still wish to have working rebase so if git will hanle somehow \\r\\n it\n> would be nice. But please do not produce the same behavior as cvs does:\n> under cygwin it still use \\n !\n\nCygwin != Windows, Cygwin is a POSIX emulation layer with the explicit goal\nof providing user tools behaving exactly as they do under Linux, and this\nincludes line ending style.\n\nSo, the Cygwin ports of various Linux tools are not expected to satisfy\nusers who want native Win32 behavior. This is where the mingw port of git\nfits in. Yes, under Cygwin git can track files with \\r\\n endings, but: \n1) Those projects are not portable to non-windows platforms, and \n2) As you noted, git will have trouble with rebase, merge, etc. as there is\nan assumption of \\n endings throughout.\n\nA proper win32 port will accept any of \\n, \\r\\n as valid line endings (add\n\\r to support Mac pre-OSX if anyone cares, I still occasionally see such\nfiles), treat any of them as semantically equal, and enforce the user's\nchosen style (\\n or \\r\\n) on output. cvsnt and svn under Windows do this\ntoday, serving up \"text\" files from the same repository with \\n endings or\n\\r\\n endings depending upon the client, and is what we need a win32 git to\ndo as well.\n\nMark\n"},{"id":"34464","messageId":"Pine.LNX.4.64.0702131225070.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"eqt40c$5ov$1@sea.gmane.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T20:32:14Z","receivedAt":"2007-02-13T20:32:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Mark Levedahl wrote:\n> \n> A proper win32 port will accept any of \\n, \\r\\n as valid line endings (add\n> \\r to support Mac pre-OSX if anyone cares, I still occasionally see such\n> files), treat any of them as semantically equal, and enforce the user's\n> chosen style (\\n or \\r\\n) on output.\n\nThe patch I sent out does that, except right now the \"autocrlf\" flag is \njust a pure boolean.\n\nI could easily make it take a ternary value:\n - off (normal UNIX semantics - never change anything)\n - on (turn CRLF->LF on input, turn LF->CRLF on output)\n - input-only (turn CRLF->LF on input, leave LF alone on output)\n\nthat would be just a couple of extra lines (almost all of them in the \nconfig file parsing logic).\n\n[ The \"output-only\" case is obviously possible, but insane. It would turn \n  a LF-only file into CRLF on output, and then not turn it back on input, \n  so doing any \"git commit -a\" would basically turn every single lines \n  into CRLF, which you do NOT want. \n\n  So hopefully that explains the three - not four - cases ]\n\nAnd the patch already leaves files that the user doesn't touch alone (ie \nif you check something out with CRLF turned off, and then turn it on in \nthe config, nobody will care - the checked-out copy will have LF-only even \nif explicitly re-checking it out would turn it into CRLF, but that's fine.\n\nIt would be interesting to hear if the patch works for the MinGW people in \nparticular. People using git with a Cygnus environment are probably used \nto try to keep files with just LF, since they are really trying to do a \nUNIX environment on top of Windows. But I suspect that WinGW people are \nmore likely to use native Windows tools for things, and then perhaps just \na smattering of UNIXy tools..\n\n\t\t\tLinus\n"},{"id":"34465","messageId":"20070213204248.GA21046@uranus.ravnborg.org","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131053110.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Sam Ravnborg","fromEmail":"sam@ravnborg.org","sentAt":"2007-02-13T20:42:48Z","receivedAt":"2007-02-13T20:42:48Z","isPatch":false,"sender":{"key":"sam@ravnborg.org","avatar":"https://gravatar.com/avatar/168a912606ed0742d840bb365e3cc21db390c36531a58341dc7a069cc1f15f62?d=mp&s=160"},"body":"> Anyway, BY DEFAULT it is off regardless, because it requires a\n> \n> \t[core]\n> \t\tAutoCRLF = true\n> \n> in your config file to be enabled. We could make that the default for \n> Windows, of course, the same way we do some other things (filemode etc).\n\nThis whole auto CRLF things seems to deal with DOS issues that I personally\nhave not encountered since looong time ago.\nGranted notepad in Windows does not understand UNIX files but that a bug\nin notepad and everyone knows that wordpad can be used.\n\nI wonder what we are really trying to address here. Or in other words\ncould the original poster maybe tell what Windows IDE's that does\nnot handle UNIX files properly?\n\ncore git today should not care about CRLF as opposed to LF end-of-line\nas long as the end-of-line is consistent - correct?\n\nSo defaulting to autoCRLF in Windows/DOS environments was maybe\nsane 10 years ago but today that seems to be the wrong thing to do.\nFor certain project the option could be useful if the tool-set in\nthe project *requires* CRLF, but if the toolset like all modern toolset\nsupports both CRLF and LF then git better avoid changing end-of-line marker.\n\n\tSam\n"},{"id":"34466","messageId":"Pine.LNX.4.64.0702131557430.1757@xanadu.home","threadId":"6772","inReplyTo":"20070213204248.GA21046@uranus.ravnborg.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-02-13T21:08:15Z","receivedAt":"2007-02-13T21:08:15Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 13 Feb 2007, Sam Ravnborg wrote:\n\n> This whole auto CRLF things seems to deal with DOS issues that I personally\n> have not encountered since looong time ago.\n\nMaybe you didn't share a work environment with Windows users since \nlooong time ago.\n\n> Granted notepad in Windows does not understand UNIX files but that a bug\n> in notepad and everyone knows that wordpad can be used.\n> \n> I wonder what we are really trying to address here. Or in other words\n> could the original poster maybe tell what Windows IDE's that does\n> not handle UNIX files properly?\n\nWindows IDE's can _create_files.  Those files will be CRLF infected.\n\nAlso some of them read UNIX files just fine but they will use CRLF to \nend new added lines despite the rest of the file using only LF.\n\n> core git today should not care about CRLF as opposed to LF end-of-line\n> as long as the end-of-line is consistent - correct?\n\nConsistency won't come alone if not enforced in some way.\n\n> So defaulting to autoCRLF in Windows/DOS environments was maybe\n> sane 10 years ago but today that seems to be the wrong thing to do.\n> For certain project the option could be useful if the tool-set in\n> the project *requires* CRLF, but if the toolset like all modern toolset\n> supports both CRLF and LF then git better avoid changing end-of-line marker.\n\nRather git better enforce consistency otherwise it'll be only a mix of \npossible combination as soon as Windows and UNIX users work on the same \nproject.\n\n\nNicolas\n"},{"id":"34468","messageId":"200702132258.20278.robin.rosenberg.lists@dewire.com","threadId":"6772","inReplyTo":"eqt40c$5ov$1@sea.gmane.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-02-13T21:58:19Z","receivedAt":"2007-02-13T21:58:19Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"tisdag 13 februari 2007 20:36 skrev Mark Levedahl:\n> Alexander Litvinov wrote:\n> \n> > ? ????????? ?? Tuesday 13 February 2007 16:06 Johannes Schindelin\n> > ???????(a):\n> >> Hi,\n> >>\n> >> On Tue, 13 Feb 2007, Alexander Litvinov wrote:\n> >> > When I have file that was converted from dos to unix format (or from\n> >> > unix to dos) git genereta big diff. But anyway, c++ compiler works well\n> >> > with both formats and in this case I simply convert file to dos format\n> >> > and git shows again nice diff. If unix format was commited to git I\n> >> > simply change the format and commit that file again.\n> >>\n> >> That's awful!\n> > If you are tring to build history that looks good - you are right this is\n> > a terrible workflow.\n> > \n> >> > The only trouble is the rebase, it does not like \\r\\n ending and othen\n> >> > produce unexpected merge conflict. But I don't use rebse to othen to\n> >> > realy investigate and try to solve the problem.\n> >>\n> >> Well, if everybody thinks like you, maybe we do not have to change\n> >> anything for Windows after all?\n> > I still wish to have working rebase so if git will hanle somehow \\r\\n it\n> > would be nice. But please do not produce the same behavior as cvs does:\n> > under cygwin it still use \\n !\n> \n> Cygwin != Windows, Cygwin is a POSIX emulation layer with the explicit goal\n> of providing user tools behaving exactly as they do under Linux, and this\n> includes line ending style.\n\nLine ending style is selectable in cygwin, both on a global level and path level (cygwin \nmounts). If you use CVS for windows development using CRLF works well and\nis the only option if you want to use the same working are with both native CVS clients\nlike TortoiseCVS and the cygwin client. I use the CRLF style by default and LF only\nfor selected directories. The only annoying thing I see is that files transformed by patch end \nup with LF-only line endings.\n\n> So, the Cygwin ports of various Linux tools are not expected to satisfy\n> users who want native Win32 behavior. This is where the mingw port of git\n> fits in. Yes, under Cygwin git can track files with \\r\\n endings, but: \n> 1) Those projects are not portable to non-windows platforms, and \n> 2) As you noted, git will have trouble with rebase, merge, etc. as there is\n> an assumption of \\n endings throughout.\n\nEven if there is a native port, I'm inclined to want to use the cygwin version \nanyway because of the nice shell and scripting capabilities and large selection of packages\nthat match what I'm used to in Linux. Git under cygwin should do CRLF transformations \naccording to the same rules that apply to text files in cygwin.\n\n-- robin\n"},{"id":"34482","messageId":"Pine.LNX.4.63.0702131518290.8035@qynat.qvtvafvgr.pbz","threadId":"6772","inReplyTo":"20070213204248.GA21046@uranus.ravnborg.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2007-02-13T23:19:39Z","receivedAt":"2007-02-13T23:19:39Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Tue, 13 Feb 2007, Sam Ravnborg wrote:\n\n>\n> I wonder what we are really trying to address here. Or in other words\n> could the original poster maybe tell what Windows IDE's that does\n> not handle UNIX files properly?\n>\n> core git today should not care about CRLF as opposed to LF end-of-line\n> as long as the end-of-line is consistent - correct?\n>\n> So defaulting to autoCRLF in Windows/DOS environments was maybe\n> sane 10 years ago but today that seems to be the wrong thing to do.\n> For certain project the option could be useful if the tool-set in\n> the project *requires* CRLF, but if the toolset like all modern toolset\n> supports both CRLF and LF then git better avoid changing end-of-line marker.\n\nI've actually run into grief on this subject with perl scripts within the last \nyear (files from windows systems with crlf not working cleanly on a linux system \nwith just lf)\n\nthis is real, not just historic\n\nDavid Lang\n"},{"id":"34480","messageId":"Pine.LNX.4.64.0702131524130.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"20070213204248.GA21046@uranus.ravnborg.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-13T23:28:15Z","receivedAt":"2007-02-13T23:28:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Sam Ravnborg wrote:\n> \n> This whole auto CRLF things seems to deal with DOS issues that I personally\n> have not encountered since looong time ago.\n\nMaybe you stopped using DOS a loong time ago ;)\n\nIt's definitely an issue. Yes, all windows programs basically *understand* \nfiles that have just LF. But almost all of them will *write* files with \nCRLF.\n\n(Which means that I suspect I made the default for \"auto_crlf\" be wrong in \nmy patch: I probably should not default to checking out with CRLF, but \nchecking out with just LF, and only do the CRLF->LF conversion on input).\n\nAnybody who has ever worked with _any_ Windows people have long since \nlearnt that they always end up having to convert CRLF to just LF when they\nget files. Even _I_ know it, and I seldom have to work with people who use \nWindows ;)\n\nSo it's a good idea to try to make sure that Windows users don't corrupt \nfiles by adding CRLF where there is no need for them into a git archive. \nWe hope to convert those people to a real OS some day (\"here's a nickel, \nboy\"), and to make it easier for them to do it, making sure that their \nprojects in -git are already in a sane format is probably a good idea.\n\n\t\t\tLinus\n"},{"id":"34487","messageId":"45D2637C.8030905@verizon.net","threadId":"6772","inReplyTo":"200702132258.20278.robin.rosenberg.lists@dewire.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-14T01:18:52Z","receivedAt":"2007-02-14T01:18:52Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Robin Rosenberg wrote:\n> \n> Even if there is a native port, I'm inclined to want to use the cygwin version \n> anyway because of the nice shell and scripting capabilities and large selection of packages\n> that match what I'm used to in Linux. Git under cygwin should do CRLF transformations \n> according to the same rules that apply to text files in cygwin.\n> \n> -- robin\n\nThe cygwin project is explicitly trying to bury the \"text\" mount option \nand drive towards binary (= \\n line endings) only. They once had a rule \nthat all cygwin programs fully grok \\r\\n, but that ethic disappeared a \ncouple of years ago, it was just too hard. The cygwin git port itself \nwill not operate on a text mount, it requires a binary mount, so crlf \ntranslations are simply not available with git under cygwin.\n\nMark\n"},{"id":"34490","messageId":"45D2691C.4090005@verizon.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131225070.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-14T01:42:52Z","receivedAt":"2007-02-14T01:42:52Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Linus Torvalds wrote:\n> \n> On Tue, 13 Feb 2007, Mark Levedahl wrote:\n>> A proper win32 port will accept any of \\n, \\r\\n as valid line endings (add\n>> \\r to support Mac pre-OSX if anyone cares, I still occasionally see such\n>> files), treat any of them as semantically equal, and enforce the user's\n>> chosen style (\\n or \\r\\n) on output.\n> \n> The patch I sent out does that, except right now the \"autocrlf\" flag is \n> just a pure boolean.\n> \n> I could easily make it take a ternary value:\n>  - off (normal UNIX semantics - never change anything)\n>  - on (turn CRLF->LF on input, turn LF->CRLF on output)\n>  - input-only (turn CRLF->LF on input, leave LF alone on output)\n> \n> \n> \t\t\tLinus\n\nWow, this is an incredible response: I expected I was going to be \nstudying git internals for a while to get to this point. Thank you!\n\nThe ternary value is definitely useful. As noted elsewhere, most tools \non windows are very happy with \\n ending, few honor those line endings \nwhen files are modified, and fewer still allow the user to specify use \nof \\n for new files. However, cygwin tools in particular are not \ntolerant of crlf, so for that environment it makes sense to banish crlf \nand the input-only option is most likely the best default setting there.\n\nMark\n"},{"id":"34491","messageId":"Pine.LNX.4.64.0702131813560.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"45D2691C.4090005@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T02:16:12Z","receivedAt":"2007-02-14T02:16:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Mark Levedahl wrote:\n> \n> The ternary value is definitely useful. As noted elsewhere, most tools on\n> windows are very happy with \\n ending, few honor those line endings when files\n> are modified, and fewer still allow the user to specify use of \\n for new\n> files. However, cygwin tools in particular are not tolerant of crlf, so for\n> that environment it makes sense to banish crlf and the input-only option is\n> most likely the best default setting there.\n\nHere's a UNTESTED patch on top of the patch I already sent, which allows \nyou to do\n\n\t[core]\n\t\tAutoCRLF = input\n\nand it should do only the CRLF->LF translation (ie it simplifies CRLF only \nwhen reading working tree files, but when checking out files, it leaves \nthe LF alone, and doesn't turn it into a CRLF).\n\nAnd by \"untested\" I mean that it looks ok and seems to compile, but I \nreally didn't do anything else.\n\n\t\tLinus\n---\ndiff --git a/config.c b/config.c\nindex ffe0212..e8ae919 100644\n--- a/config.c\n+++ b/config.c\n@@ -325,6 +325,10 @@ int git_default_config(const char *var, const char *value)\n \t}\n \n \tif (!strcmp(var, \"core.autocrlf\")) {\n+\t\tif (value && !strcasecmp(value, \"input\")) {\n+\t\t\tauto_crlf = -1;\n+\t\t\treturn 0;\n+\t\t}\n \t\tauto_crlf = git_config_bool(var, value);\n \t\treturn 0;\n \t}\ndiff --git a/convert.c b/convert.c\nindex c04b6c2..b5a47c2 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -133,7 +133,7 @@ int convert_to_working_tree(const char *path, char **bufp, unsigned long *sizep)\n \t * based on filename patterns. Right now we just do the\n \t * stupid auto-CRLF one.\n \t */\n-\tif (!auto_crlf)\n+\tif (auto_crlf <= 0)\n \t\treturn 0;\n \n \tsize = *sizep;\ndiff --git a/environment.c b/environment.c\nindex 2fa0960..570e32a 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -28,7 +28,7 @@ size_t packed_git_window_size = DEFAULT_PACKED_GIT_WINDOW_SIZE;\n size_t packed_git_limit = DEFAULT_PACKED_GIT_LIMIT;\n int pager_in_use;\n int pager_use_color = 1;\n-int auto_crlf = 0;\n+int auto_crlf = 0;\t/* 1: both ways, -1: only when adding git objects */\n \n static const char *git_dir;\n static char *git_object_dir, *git_index_file, *git_refs_dir, *git_graft_file;\n"},{"id":"34495","messageId":"200702140947.43527.litvinov2004@gmail.com","threadId":"6772","inReplyTo":"20070213204248.GA21046@uranus.ravnborg.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-02-14T03:47:42Z","receivedAt":"2007-02-14T03:47:42Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Wednesday 14 February 2007 02:42 Sam Ravnborg написал(a):\n> I wonder what we are really trying to address here. Or in other words\n> could the original poster maybe tell what Windows IDE's that does\n> not handle UNIX files properly?\nMS VC has text file for project file but don't like \\n line endings, only \n\\r\\n.\n"},{"id":"34496","messageId":"7v8xf1uxme.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131053110.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-14T05:16:57Z","receivedAt":"2007-02-14T05:16:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> NOTE NOTE NOTE! The \"is_binary()\" heuristics are totally made-up by yours \n> truly. I will not guarantee that they work at all reasonable. Caveat \n> emptor. But it _is_ simple, and it _is_ safe, since it's all off by \n> default.\n\nIt might be safe for some definition of safe, but it is very\nAsian unfriendly.\n\nI'd probably suggest replacing it with what GNU diff uses, which\nwe stolen and implemented in diff.c::mmfile_is_binary().\n"},{"id":"34499","messageId":"Pine.LNX.4.64.0702132127330.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"7v8xf1uxme.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T05:36:21Z","receivedAt":"2007-02-14T05:36:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 13 Feb 2007, Junio C Hamano wrote:\n> \n> It might be safe for some definition of safe, but it is very\n> Asian unfriendly.\n> \n> I'd probably suggest replacing it with what GNU diff uses, which\n> we stolen and implemented in diff.c::mmfile_is_binary().\n\nWell, the thing is, mmfile_is_binary() doesn't really have a big downside \nif it's wrong one way or the other.\n\nIn contrast CR->CRLF conversion, if wrong, actually corrupts binary files. \nSo I felt it was better to be really safe than sorry. It's *much* better \nto miss some CRLF translation than to do too much of it.\n\nThat said, I'm sure it could be improved a lot. In particular, characters \nin the range 0x00 - 0x1f are clearly \"more binary\" than the 0x7f+ range, \nwith the obvious exceptions (tab, cr, lf).\n\n0x00 - which is the only one mmfile_is_binart() uses - is arguably the \n\"most binary\" one, of course, but it might be interesting to give \ndifferent weights to the whole range.. In particular, especially for small \nfiles, the fact that there is no 0x00 byte in no way indicates that it's \nnot \"binary\".\n\nThis whole issue is obviously one reason I'd like to involve the filename \nitself, and make it use a \".gitattributes\" file - exactly because that \nallows you to be much more aggressive and more precise.\n\n(0x00 may be one of the more _common_ characters in many binary files, \nwhich makes it a good character to search for too, so I don't really have \nany hugely strong opinions here. After all, the whole heuristic is off by \ndefault anyway, so it's \"really safe\" ;^)\n\n\t\t\tLinus\n"},{"id":"34504","messageId":"20070214084121.GB25617@uranus.ravnborg.org","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131524130.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Sam Ravnborg","fromEmail":"sam@ravnborg.org","sentAt":"2007-02-14T08:41:21Z","receivedAt":"2007-02-14T08:41:21Z","isPatch":false,"sender":{"key":"sam@ravnborg.org","avatar":"https://gravatar.com/avatar/168a912606ed0742d840bb365e3cc21db390c36531a58341dc7a069cc1f15f62?d=mp&s=160"},"body":"> > This whole auto CRLF things seems to deal with DOS issues that I personally\n> > have not encountered since looong time ago.\n> \n> Maybe you stopped using DOS a loong time ago ;)\nUnfortunately not. (Sitting with a Windows 2000 laptop atm but saved by ssh).\n\n> \n> It's definitely an issue. Yes, all windows programs basically *understand* \n> files that have just LF. But almost all of them will *write* files with \n> CRLF.\n\nSo the issue with git supporting CRLF -> LF is to make interoperability between\nUNIX* programs and Windows programs which is anohter domain.\n\nMy main objective is the proposal to make a conversion default when many users\ndo not need it. For the UNIX* compatibility thing having conversion at lowest\nlayer make sense.\n\n> (Which means that I suspect I made the default for \"auto_crlf\" be wrong in \n> my patch: I probably should not default to checking out with CRLF, but \n> checking out with just LF, and only do the CRLF->LF conversion on input).\nExpect that it seems a few br0ken programs yet does not support LF as\nend-of-line marker - so .gitattriutes make take special care here.\n\n\tSam\n"},{"id":"34514","messageId":"Pine.LNX.4.63.0702141208020.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702132127330.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-14T11:10:08Z","receivedAt":"2007-02-14T11:10:08Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 13 Feb 2007, Linus Torvalds wrote:\n\n> 0x00 - which is the only one mmfile_is_binart() uses - is arguably the \n> \"most binary\" one, of course, but it might be interesting to give \n> different weights to the whole range.. In particular, especially for \n> small files, the fact that there is no 0x00 byte in no way indicates \n> that it's not \"binary\".\n\nLast time I checked, the text files never had lines longer than 200 \ncharacters (I chose this intentionally large). So, it might be a good \nheuristic to check the maximal line length, and refuse to believe that \nit's text once a certain (configurable) threshold is reached.\n\nCiao,\nDscho\n"},{"id":"34518","messageId":"200702141736.57521.litvinov2004@gmail.com","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131053110.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Alexander Litvinov","fromEmail":"litvinov2004@gmail.com","sentAt":"2007-02-14T11:36:57Z","receivedAt":"2007-02-14T11:36:57Z","isPatch":false,"sender":{"key":"litvinov2004@gmail.com","avatar":null},"body":"В сообщении от Wednesday 14 February 2007 01:07 Linus Torvalds написал:\n> Actually, I did it myself.\n>\n> This is a \"lazy man's auto-CRLF\", and it really is pretty simple.\n\nWow ! Thanks. \n\nI just tried this patch and it works! From now I can use git-cvsimport under \nLinux and then clone it to cygwin and work there with full history. Nice, \nvery nice. In my case text file detection work well as far most of our files \nare .cpp and .h\n"},{"id":"34532","messageId":"45D31C0E.2040206@verizon.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702141208020.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mdl123@verizon.net","sentAt":"2007-02-14T14:26:22Z","receivedAt":"2007-02-14T14:26:22Z","isPatch":false,"sender":{"key":"mdl123@verizon.net","avatar":"https://avatars.githubusercontent.com/u/5302462?v=4"},"body":"Johannes Schindelin wrote:\n> Last time I checked, the text files never had lines longer than 200 \n> characters (I chose this intentionally large). So, it might be a good \n> heuristic to check the maximal line length, and refuse to believe that \n> it's text once a certain (configurable) threshold is reached.\n>\n> Ciao,\n> Dsch\nUnfortunately, on my program we have folks using text files with single \nlines over 60,000 characters long, these are data files. Think for \nexample of a comma or tab separated data file saved from a spreadsheet. \nIn this case, the files are pure ascii. So, the line length could be \nsomething else to take into account, but is not decisive by itself.\n\nTo recap, we have the following various suggestions to determine textness:\n\n1) ratio of ascii to non-ascii characters, possibly weighting some chars \nmore than others\n2) line length\n3) existence of a null (\\0)\n4) file name globbing\n5) roundtrip ( lf(crlf(file) ) == file\n\nI don't think any one suggestion is completely adequate for all uses, \nall need to be available, somehow configurable. This suggests to me a \ncore.AutoCRLFstrategy variable that is a comma separated list of methods \nto use (set to a reasonable default of course that does not cause \nruntime headaches on Unix): a file would be deemed binary unless all \nlisted methods declare the file as text (with an empty list disabling \nAutoCRLF detection).\n\nMark\n"},{"id":"34537","messageId":"Pine.LNX.4.64.0702140742280.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702141208020.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T15:44:00Z","receivedAt":"2007-02-14T15:44:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Johannes Schindelin wrote:\n> \n> Last time I checked, the text files never had lines longer than 200 \n> characters (I chose this intentionally large). So, it might be a good \n> heuristic to check the maximal line length,\n\nNo, some broken editor programs and people use \"flowing text\" files, where \na newline is actually a _paragraph_ end. You have lines in the hundreds \n(and thousands) of characters, and the program will just flow the text for \nyou.\n\nUgh. Horrible, I know. \n\n\t\tLinus\n"},{"id":"34539","messageId":"Pine.LNX.4.64.0702140745110.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"45D31C0E.2040206@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T15:51:00Z","receivedAt":"2007-02-14T15:51:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Mark Levedahl wrote:\n> \n> To recap, we have the following various suggestions to determine textness:\n> \n> 1) ratio of ascii to non-ascii characters, possibly weighting some chars more\n> than others\n> 2) line length\n> 3) existence of a null (\\0)\n> 4) file name globbing\n> 5) roundtrip ( lf(crlf(file) ) == file\n\nActually, my patch already had one that you didn't mention: \n 6) CR never shows up alone.\n\nSo the patch I sent out basicallyhad the following rules:\n - no more than ~10% of all characters being other than regular printable \n   ASCII (where any control character except for newline/cr/tab was deemed \n   nonprintable)\n - any \"lonely\" CR automatically means it's binary, and I would refuse \n   to convert that to a LF (the test in the code is that CRLF count must \n   match CR count)\n\nbut the \"roundtrip\" rule is much too strict (it's actually perfectly \npossible for an editor to add CRLF characters only to new _lines_, leaving \nold lines with just LF - or the other way around. In fact, the editor I \nuse under Linux does exactly that in reverse - if I add new lines, it will \nadd those without CR, but will leave old lines with CRLF alone).\n\nI think that to help asian languages (or strange text-files in utf8 or \nLatin1 too, for that matter: test-files with _just_ special characters), I \nshould probably make the rule be that only the 0-31 range is special.\n\n\t\t\tLinus\n"},{"id":"34541","messageId":"Pine.LNX.4.63.0702141653160.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702140742280.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-14T15:53:34Z","receivedAt":"2007-02-14T15:53:34Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Feb 2007, Linus Torvalds wrote:\n\n> On Wed, 14 Feb 2007, Johannes Schindelin wrote:\n> > \n> > Last time I checked, the text files never had lines longer than 200 \n> > characters (I chose this intentionally large). So, it might be a good \n> > heuristic to check the maximal line length,\n> \n> No, some broken editor programs and people use \"flowing text\" files, where \n> a newline is actually a _paragraph_ end.\n\nGood point.\n\nCiao,\nDscho\n"},{"id":"34542","messageId":"Pine.LNX.4.63.0702141653440.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6772","inReplyTo":"45D31C0E.2040206@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-14T15:56:01Z","receivedAt":"2007-02-14T15:56:01Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Feb 2007, Mark Levedahl wrote:\n\n> This suggests to me a core.AutoCRLFstrategy variable that is a comma \n> separated list of methods to use (set to a reasonable default of course \n> that does not cause runtime headaches on Unix): a file would be deemed \n> binary unless all listed methods declare the file as text (with an empty \n> list disabling AutoCRLF detection).\n\nThis sounds regretfully complex. Somebody (you?) mentioned that cvsnt does \na kick-ass job here. Does cvsnt need strategies? I don't think so. Neither \ndo we. Someone who cares enough should just rip^H^H^Hlook at cvsnt's text \ndetection.\n\nCiao,\nDscho\n"},{"id":"34545","messageId":"45D335C3.E28D28E0@eudaptics.com","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702131053110.8424@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2007-02-14T16:16:03Z","receivedAt":"2007-02-14T16:16:03Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Linus Torvalds wrote:\n> \n> On Tue, 13 Feb 2007, Junio C Hamano wrote:\n> >\n> > Thanks, applied.  I think git-apply has separate codepaths for\n> > both reading and writing; I won't look into them before 1.5.0\n> > but people are welcome to help advancing the cause before I get\n> > to it ;-).\n> \n> Actually, I did it myself.\n> \n> This is a \"lazy man's auto-CRLF\", and it really is pretty simple.\n\nThanks a lot, busy beaver! I gave this a quick spin with a few\ninteresting operations: merges and rebase. Merges leave the merge\nresults with only LFs behind. Rebasing seems to work as expected\n(working files have CRLFs), except when merges are needed.\n\nDoesn't git-unpack-file also need to call into the converter?\n\n-- Hannes\n"},{"id":"34547","messageId":"Pine.LNX.4.64.0702140813190.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702141653440.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T16:23:04Z","receivedAt":"2007-02-14T16:23:04Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Johannes Schindelin wrote:\n> \n> This sounds regretfully complex. Somebody (you?) mentioned that cvsnt does \n> a kick-ass job here. Does cvsnt need strategies? I don't think so. Neither \n> do we. Someone who cares enough should just rip^H^H^Hlook at cvsnt's text \n> detection.\n\nWell, one thing to keep in mind is that for source code in particular, \nthis really very seldom is an issue.\n\nSo you can do a really *bad* job in theory, and in practice it really \nworks very very well.\n\nVery few people keep binary blobs in any SCM archive _anyway_, partly \nbecause they've always been told that it's unsafe (and with a lot of SCM's \nit is), but even more because binary blobs are almost always generated by \nsome build method, so normally you'd never version them in the first \nplace, or versioning isn't all that helpful.\n\nAnd most binary blobs are so *obviously* binary that even the stupidest \nalgorithm on earth will get it right. The only hard cases actually tend to \nbe really tiny files, or literally test-sequences.\n\nTiny files are hard because:\n\n - they (by being tiny) have so few characters that they can easily lack \n   a \"fingerprint\" character (eg a NUL character or similar). \n\n - tiny files are a lot more likely than bigger files to have strange \n   statistics that throw some more \"sophisticated\" rule off the scent. \n   Something like a \"10% rule\" tends to work fine if you have a big text, \n   and ten percent is still a reasonable number to average things out \n   over, but what if you only had ten characters to begin with?\n\nThe good news is that tiny files can usually be considered text, since \nyou'd seldom use a binary format for something really small anyway.\n\nSo I suspect that IN PRACTICE, especially if you come as a CVS replacement \n(where binary files are just damn hard to get right even under the best of \ncircumstances!), you can do just about anything, including just saying \n\"everything is text\", and you'd be fine.\n\nIt's entirely possible that that is exactly what CVSNT does ;)\n\n\t\tLinus\n"},{"id":"34550","messageId":"Pine.LNX.4.64.0702140823220.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"20070214084121.GB25617@uranus.ravnborg.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T16:28:24Z","receivedAt":"2007-02-14T16:28:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Sam Ravnborg wrote:\n>\n> > (Which means that I suspect I made the default for \"auto_crlf\" be wrong in \n> > my patch: I probably should not default to checking out with CRLF, but \n> > checking out with just LF, and only do the CRLF->LF conversion on input).\n>\n> Expect that it seems a few br0ken programs yet does not support LF as\n> end-of-line marker - so .gitattriutes make take special care here.\n\nYes, but I also think that even without .gitattributes, you just want to \nhave a default for what \"text\" actually means, and it's entirely possible \nthat the default should be: \"check out with just LF, and on check-in turn \nCRLF into LF\".\n\nBut exactly because _some_ programs might want to always see CRLF on input \ntoo, it should be overridable. \n\nOr maybe the default should be \"turn into CRLF\", and there should just be \nan option to make it check out as LF-only.\n\nRegardless, I think that is independent of \".gitattributes\". The \n_attribute_ should be \"text\", but what it then means in practice is a \nseparate flag.\n\nAnd yes, we *could* have a per-file attribute (\"text,crlf-checkout\") which \ncould be used to say \"I want to always check out as crlf regardless of any \nother policy\") and the same for lf-only, but I seriously doubt that \nanybody really needs that kind of knob-tweaking. At some point it's just \nfine to say \"you're crazy\".\n\n\t\t\tLinus\n"},{"id":"34555","messageId":"Pine.LNX.4.64.0702140835440.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"200702141736.57521.litvinov2004@gmail.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T16:37:33Z","receivedAt":"2007-02-14T16:37:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Alexander Litvinov wrote:\n> \n> I just tried this patch and it works! From now I can use git-cvsimport under \n> Linux and then clone it to cygwin and work there with full history. Nice, \n> very nice.\n\nBtw, it didn't do any commit message conversion etc, so you'll still \nalways see commit messages with LF-only, and if you _create_ commits, you \nneed to make sure that whatever program you use will do the right thing.\n\n> In my case text file detection work well as far most of our files \n> are .cpp and .h\n\nYeah, considering that it worked in my testing for \"git\" itself, I'm not \nsurprised. Source code tends to look the same..\n\n\t\tLinus\n"},{"id":"34556","messageId":"7v64a4snfo.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702140745110.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-14T16:39:55Z","receivedAt":"2007-02-14T16:39:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Actually, my patch already had one that you didn't mention: \n>  6) CR never shows up alone.\n\nOlder Macs ;-)?\n\n> So the patch I sent out basicallyhad the following rules:\n>  - no more than ~10% of all characters being other than regular printable \n>    ASCII (where any control character except for newline/cr/tab was deemed \n>    nonprintable)\n>  - any \"lonely\" CR automatically means it's binary, and I would refuse \n>    to convert that to a LF (the test in the code is that CRLF count must \n>    match CR count)\n> ...\n> I think that to help asian languages (or strange text-files in utf8 or \n> Latin1 too, for that matter: test-files with _just_ special characters), I \n> should probably make the rule be that only the 0-31 range is special.\n\nI would agree.  0-31 except HT, CR, LF and ESC would be a good\nidea; that would not harm text in UTF-8, EUC based various\nlocales nor ISO 2022.\n\nPatch is relative to 'pu'.\n-- >8 --\n\ndiff --git a/convert.c b/convert.c\nindex ebcf717..b6b7c66 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -13,7 +13,7 @@ struct text_stat {\n \tunsigned cr, lf, crlf;\n \n \t/* These are just approximations! */\n-\tunsigned printable, nonprintable, nul;\n+\tunsigned printable, nonprintable;\n };\n \n static void gather_stats(const char *buf, unsigned long size, struct text_stat *stats)\n@@ -34,13 +34,11 @@ static void gather_stats(const char *buf, unsigned long size, struct text_stat *\n \t\t\tstats->lf++;\n \t\t\tcontinue;\n \t\t}\n-\t\tif (c == '\\t' || (c >= 32 && c < 127)) {\n-\t\t\tstats->printable++;\n+\t\tif ((c < 32) && (c != '\\t' && c != '\\033')) {\n+\t\t\tstats->nonprintable++;\n \t\t\tcontinue;\n \t\t}\n-\t\tif (!c)\n-\t\t\tstats->nul++;\n-\t\tstats->nonprintable++;\n+\t\tstats->printable++;\n \t}\n }\n \n@@ -50,7 +48,7 @@ static void gather_stats(const char *buf, unsigned long size, struct text_stat *\n static int is_binary(unsigned long size, struct text_stat *stats)\n {\n \n-\tif (stats->nul)\n+\tif (stats->nonprintable)\n \t\treturn 1;\n \t/*\n \t * Other heuristics? Average line length might be relevant,\n"},{"id":"34559","messageId":"20070214164735.GA28359@uranus.ravnborg.org","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702140823220.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Sam Ravnborg","fromEmail":"sam@ravnborg.org","sentAt":"2007-02-14T16:47:35Z","receivedAt":"2007-02-14T16:47:35Z","isPatch":false,"sender":{"key":"sam@ravnborg.org","avatar":"https://gravatar.com/avatar/168a912606ed0742d840bb365e3cc21db390c36531a58341dc7a069cc1f15f62?d=mp&s=160"},"body":"On Wed, Feb 14, 2007 at 08:28:24AM -0800, Linus Torvalds wrote:\n> \n> \n> On Wed, 14 Feb 2007, Sam Ravnborg wrote:\n> >\n> > > (Which means that I suspect I made the default for \"auto_crlf\" be wrong in \n> > > my patch: I probably should not default to checking out with CRLF, but \n> > > checking out with just LF, and only do the CRLF->LF conversion on input).\n> >\n> > Expect that it seems a few br0ken programs yet does not support LF as\n> > end-of-line marker - so .gitattriutes make take special care here.\n> \n> Yes, but I also think that even without .gitattributes, you just want to \n> have a default for what \"text\" actually means, and it's entirely possible \n> that the default should be: \"check out with just LF, and on check-in turn \n> CRLF into LF\".\nThe definition of what is \"text\" and what action to take upon check-in /\ncheck-out of text is two sepearate things.\n\nI could see it as beneficial as a per-project or even as an overall\ngit-policy to say \"checkin-as-LF\" - \"checkout-as-LF\" to overcome\ninteroperability issues when more tools gets UNIX* based.\n\n> \n> But exactly because _some_ programs might want to always see CRLF on input \n> too, it should be overridable. \nWhich is where I see .gitattributes come into play.\n-> A rule that says files with extension .prj and of type \"text\" shall not see\nany conversion.\n\nIn this way almost all \"text\" over time get a proper format and the remaining\nbrain-dead tools that continue to save in CRLF format will not destroy the sane\nLF format.\n\nIf anything gets defualt I would vote for LF. But overrideable.\n\nMy editor-of-choice does eol auto-sense. If I recall correct it scans the\nfirst 200 lines and counts number of CR,LF,CRLF and based on this judge the\nactual eol character used. But not all editors are that sensible :-(\n\n\tSam\n"},{"id":"34561","messageId":"Pine.LNX.4.64.0702140847020.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"45D335C3.E28D28E0@eudaptics.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T16:53:14Z","receivedAt":"2007-02-14T16:53:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Johannes Sixt wrote:\n>\n> Thanks a lot, busy beaver! I gave this a quick spin with a few\n> interesting operations: merges and rebase. Merges leave the merge\n> results with only LFs behind.\n\nYes. Merge uses \"git-cat-file\" (well, it historically did, now that it's \nbuilt-in it still does the equivalent operation).\n\nI already talked about how git-cat-file was special ;)\n\n> Rebasing seems to work as expected (working files have CRLFs), except \n> when merges are needed.\n\nWell, it always \"merges\", but yes, you mean three-way data merges. The \nnormal SHA1-direct merges will just use the normal git-read-tree thing \nwhich is the same as checkout.\n\n> Doesn't git-unpack-file also need to call into the converter?\n\nSee earlier discussions. git-cat-file (and git-unpack-file, which is just \na version of it, really) don't have the original filename, so we'll need \nto extend on it some way in order to support file attributes even in \ntheory. So before we do that, I'd hate to do any format conversion there.\n\nYes, yes, right now it ignores the filename *anyway*, but the point is, \nright now that's a \"small implementation detail\". I would NOT want to do \nthis if I couldn't know the filename at all!\n\nThe merge algorithms actually obviously *do* know the filename fo the \nthings that they are going to merge, so the filename information does \nexists. It's just not passed on far enough.\n\nFinally, one comment: if you use \"autocrlf = input\" (my second patch), all \nof this works even now, since the default is to just leave things as \nLF-only anyway. In fact, even with \"autocrlf = on\", nothing should really \n*break* except for silly editors that actuall *require* CRLF.\n\nIOW, it's more important to do the CRLF->LF conversion than it is to do \nthe LF->CRLF one ;)\n\n\t\tLinus\n"},{"id":"34566","messageId":"Pine.LNX.4.64.0702140855080.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"7v64a4snfo.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T17:01:41Z","receivedAt":"2007-02-14T17:01:41Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > Actually, my patch already had one that you didn't mention: \n> >  6) CR never shows up alone.\n> \n> Older Macs ;-)?\n\nYeah, I think we can ignore them..\n\nLet's see if anybody ever complains ;)\n\n> I would agree.  0-31 except HT, CR, LF and ESC would be a good\n> idea; that would not harm text in UTF-8, EUC based various\n> locales nor ISO 2022.\n\nYou could possibly add 127 to the list too (it's ascii DEL, I don't know \nif you should ever see it in anything that has anything to do with text).\n\n> -\tif (stats->nul)\n> +\tif (stats->nonprintable)\n\nBut this is too harsh.\n\nIt's quite common to have the occasional FF character. Some things really \ndo use it for page breaks. So saying that *any* nonprintable character is \nbad is not a good idea.\n\nSame goes for BS (some programs use it to show bold and underlined text: \nman-pages, for example).\n\n\t\tLinus\n"},{"id":"34574","messageId":"7vbqjwr727.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702140835440.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-14T17:18:56Z","receivedAt":"2007-02-14T17:18:56Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 14 Feb 2007, Alexander Litvinov wrote:\n>> \n>> I just tried this patch and it works! From now I can use git-cvsimport under \n>> Linux and then clone it to cygwin and work there with full history. Nice, \n>> very nice.\n>\n> Btw, it didn't do any commit message conversion etc, so you'll still \n> always see commit messages with LF-only, and if you _create_ commits, you \n> need to make sure that whatever program you use will do the right thing.\n\nI think stripspace removes CR so we should be Ok.\n"},{"id":"34578","messageId":"45D346B6.5020802@verizon.net","threadId":"6772","inReplyTo":"Pine.LNX.4.63.0702141653440.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Mark Levedahl","fromEmail":"mlevedahl@verizon.net","sentAt":"2007-02-14T17:28:22Z","receivedAt":"2007-02-14T17:28:22Z","isPatch":false,"sender":{"key":"mlevedahl@verizon.net","avatar":null},"body":"Johannes Schindelin wrote:\n> Hi,\n>\n> On Wed, 14 Feb 2007, Mark Levedahl wrote:\n>   \n> This sounds regretfully complex. Somebody (you?) mentioned that cvsnt does \n> a kick-ass job here. Does cvsnt need strategies? I don't think so. Neither \n> do we. Someone who cares enough should just rip^H^H^Hlook at cvsnt's text \n> detection.\n>\n> Ciao,\n> Dscho\n>   \nI agree that is complex, I started thinking of PAM when I wrote that, \nleading to, \"this aint gonna work.\" But in the modern day let's all feel \ngood spirit of \"there are no stupid ideas, just some are better\" I threw \nit out anyway.\n\nAs to cvsnt, my actual feeling is I'd like to kick it in the ass, it has \ndestroyed too many files for me over the years, binary and text, so I \ndon't think its strategies are very good. That is why I'm kicking these \nideas around, if I thought I knew the \"right\" way I would have written \nit already.\n\nMark\n"},{"id":"34579","messageId":"7v3b58r6kc.fsf@assigned-by-dhcp.cox.net","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702140855080.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-14T17:29:39Z","receivedAt":"2007-02-14T17:29:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n>> -\tif (stats->nul)\n>> +\tif (stats->nonprintable)\n>\n> But this is too harsh.\n>\n> It's quite common to have the occasional FF character. Some things really \n> do use it for page breaks. So saying that *any* nonprintable character is \n> bad is not a good idea.\n>\n> Same goes for BS (some programs use it to show bold and underlined text: \n> man-pages, for example).\n\nOk.  How about adding BS and FF to the Ok set, and checking if\nbad ones are less than 1% of the good ones?\n\ndiff --git a/convert.c b/convert.c\nindex b6b7c66..b0c7641 100644\n--- a/convert.c\n+++ b/convert.c\n@@ -34,11 +34,22 @@ static void gather_stats(const char *buf, unsigned long size, struct text_stat *\n \t\t\tstats->lf++;\n \t\t\tcontinue;\n \t\t}\n-\t\tif ((c < 32) && (c != '\\t' && c != '\\033')) {\n+\t\tif (c == 127)\n+\t\t\t/* DEL */\n \t\t\tstats->nonprintable++;\n-\t\t\tcontinue;\n+\t\telse if (c < 32) {\n+\t\t\tswitch (c) {\n+\t\t\t\t/* BS, HT, ESC and FF */\n+\t\t\tcase '\\b': case '\\t': case '\\033': case '\\014':\n+\t\t\t\tstats->printable++;\n+\t\t\t\tbreak;\n+\t\t\tdefault:\n+\t\t\t\tstats->nonprintable++;\n+\t\t\t}\n+\t\t\t\n \t\t}\n-\t\tstats->printable++;\n+\t\telse\n+\t\t\tstats->printable++;\n \t}\n }\n \n@@ -48,7 +59,7 @@ static void gather_stats(const char *buf, unsigned long size, struct text_stat *\n static int is_binary(unsigned long size, struct text_stat *stats)\n {\n \n-\tif (stats->nonprintable)\n+\tif ((stats->printable >> 7) < stats->nonprintable)\n \t\treturn 1;\n \t/*\n \t * Other heuristics? Average line length might be relevant,\n"},{"id":"34581","messageId":"Pine.LNX.4.64.0702140941590.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"7v3b58r6kc.fsf@assigned-by-dhcp.cox.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T17:43:10Z","receivedAt":"2007-02-14T17:43:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Junio C Hamano wrote:\n> \n> Ok.  How about adding BS and FF to the Ok set, and checking if\n> bad ones are less than 1% of the good ones?\n\nI think that looks fine.\n\n\t\tLinus\n"},{"id":"34598","messageId":"200702141917.51341.robin.rosenberg.lists@dewire.com","threadId":"6772","inReplyTo":"45D346B6.5020802@verizon.net","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-02-14T18:17:50Z","receivedAt":"2007-02-14T18:17:50Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"onsdag 14 februari 2007 18:28 skrev Mark Levedahl:\n> As to cvsnt, my actual feeling is I'd like to kick it in the ass, it has \n> destroyed too many files for me over the years, binary and text, so I \n> don't think its strategies are very good. That is why I'm kicking these \n> ideas around, if I thought I knew the \"right\" way I would have written \n> it already.\n\nThat may be why an excellent piece of software, TortoiseCVS,  doesn't trust \ncvs or cvsnt to do the job. Here is how they do the binary detection (and \nsome more):\n\nhttp://tortoisecvs.cvs.sourceforge.net/tortoisecvs/TortoiseCVS/src/CVSGlue/CVSStatus.cpp?revision=1.172&view=markup\n\n-- robin\n"},{"id":"34603","messageId":"Pine.LNX.4.64.0702141027380.3604@woody.linux-foundation.org","threadId":"6772","inReplyTo":"200702141917.51341.robin.rosenberg.lists@dewire.com","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-14T18:31:39Z","receivedAt":"2007-02-14T18:31:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 14 Feb 2007, Robin Rosenberg wrote:\n> \n> That may be why an excellent piece of software, TortoiseCVS,  doesn't trust \n> cvs or cvsnt to do the job. Here is how they do the binary detection (and \n> some more):\n> \n> http://tortoisecvs.cvs.sourceforge.net/tortoisecvs/TortoiseCVS/src/CVSGlue/CVSStatus.cpp?revision=1.172&view=markup\n\nWell, it does seem to boil down to what Junio already got to:\n\n - 0-31 and 127 are never in text, except for BEL, BS, HT, LF, FF, CR and \n   ESC.\n - 128-255 can all be in either iso-8859 or extended ascii (or they \n   explicitly add NEL but not 128+27 to \"normal ASCII\", which is strange)\n\nSo they've effectively added BEL and ESC to the listof characters that \nJunio has now. But they also make it an absolute error to have anything \nelse (no \"1% rule\").\n\nBut they also do the filename tests, and I think that's more important in \nmany ways.\n\n\t\tLinus\n"},{"id":"34619","messageId":"200702142124.10996.robin.rosenberg.lists@dewire.com","threadId":"6772","inReplyTo":"Pine.LNX.4.64.0702141027380.3604@woody.linux-foundation.org","subject":"Re: mingw, windows, crlf/lf, and git","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-02-14T20:24:09Z","receivedAt":"2007-02-14T20:24:09Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"onsdag 14 februari 2007 19:31 skrev Linus Torvalds:\n> \n> On Wed, 14 Feb 2007, Robin Rosenberg wrote:\n> > \n> > That may be why an excellent piece of software, TortoiseCVS,  doesn't trust \n> > cvs or cvsnt to do the job. Here is how they do the binary detection (and \n> > some more):\n> > \n> > http://tortoisecvs.cvs.sourceforge.net/tortoisecvs/TortoiseCVS/src/CVSGlue/CVSStatus.cpp?revision=1.172&view=markup\n> \n> Well, it does seem to boil down to what Junio already got to:\n> \n>  - 0-31 and 127 are never in text, except for BEL, BS, HT, LF, FF, CR and \n>    ESC.\n>  - 128-255 can all be in either iso-8859 or extended ascii (or they \n>    explicitly add NEL but not 128+27 to \"normal ASCII\", which is strange)\n>\n> So they've effectively added BEL and ESC to the listof characters that \nEspecially ESC used to be common in DOS/Windows and quite a few hang around in\nolder code.\n\n> Junio has now. But they also make it an absolute error to have anything \n> else (no \"1% rule\").\nCan this 1%-rule be motivated from real cases, rather that hypotetical ones? It makes \nit harder to understand  why the tools makes a particular decision.\n\n> But they also do the filename tests, and I think that's more important in \n> many ways.\n\nA unixy tool like git should maybe use magic too :).\n\nBtw the filename (like .gitignore or similar) test in practice would give us \n the binary flag. Just list a filename instead of a pattern.\n\n-- robin\n"}]}