Hi,
Thanks for your reply.
I think we can come to the conclusion there are no real text files, there is no real text standard.
All text files are ultimately binary files, some are ascii, some are ansi, some are unicode, etc.
Some have codepage encodings, some have LF, some have CR, some have both, some have vice versa.
Some have BOM some don't.
Basically it is a GIGANTIC MESS.
This is probably the biggest flaw of git, assuming that there is such a thing is a text standard.
Perhaps it's better to start treating everything as binary, and also creating, yet again a new true text standard ! LOL :)
PDF, DOCS ? I once heard "top demo coders" use Microsoft Words to do their coding in. I am beginning to understand why that might be ! ;)
Perhaps a new text format where there is no such thing as nil terminator and carriage returns and line feeds, but everything pre-fixed-lengths or so....
This would also solve the "nil" character frustration you shared, thanks for that !
Downside for this new idea would be text length limited to what the number of length bits can hold. Which would be plenty for 32 or 64 bits.
Alternatively, Skybuck's Universal Code or another flexible coding technique could be used as well.
However, the alphabet itself is an encoding as well... Unicode feels a bit over done, with emotion smileys etc and other strange things, but it is a big world wide standard.
Perhaps it could function as the encoding for the characters.
This would leave some binary format for text to be developed which would be suited for coding and editors.
Editor could would become a bit more complex I suppose, to handle the prefix length fields and can no longer inject/delete characters, I am not sure how code editors work internally, maybe a doubled linked list of characters.
Perhaps line numbers could be hard coded as well... or inferred/counted a bit more quicker... right now AI would have to count CR/LF characters which might make AI processing more expensive to find actual line numbers...
I wonder...
Plus some code could also be stored in "line number segments/ranges"... like lines 51 to 56... and perhaps line segments could be stored on disk directly, even randomly... and could be stitched together later sequentially for rendering purposes.
Historical changes might also be kept a bit more easy that way... like some kind of diff form... so I do see some potential for this...
Bye for now,
Skybuck.