When Compilers Disagree About UTF‑8(nemanjatrifunovic.substack.com) |
When Compilers Disagree About UTF‑8(nemanjatrifunovic.substack.com) |
Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.
Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.