If you skip this initial validation, you open yourself to security problems caused by differences in the undefined behavior of these various routines that are assuming valid input. As an example, imagine if your code for escaping HTML saw an invalid byte sequence immediately before a "<" and failed to escape it because of that, but then the HTML parser you pass the result to treats the invalid byte sequence differently and successfully sees and parses the "<" that should have been escaped.
e.g. the developer might either deny or allow /admin/, or /admin/.*
But if a given Unicode string [1] has two representations, then that check may not work as intended for one of them.
---
UTF-8 represents code points with 1-4 bytes, which are roughly similar to base 64 digits (because a continuation byte is 10xx_xxxx, it has 6 bits of freedom)
This opens up the possibility of "overlong encodings", which a UTF-8 validator must reject.
That is, you can do the equivalent of representing a number as "09" instead of "9" -- that's an overlong encoding. A decoder may understand "09" and "9" as both being the digit 9, but it's a different sequence of bytes.
So a naive decoding algorithm can accept a spelling of "admin" as 6, 7, 8, ... 20 bytes, not 5. There are many overlong encodings!
The only valid UTF-8 spelling has 5 bytes, but decoders that don't validate will "naturally" accept more (in fact I think I even wrote one of these :-/ )
In summary, if you don't do UTF-8 validation, then one layer doesn't reject the invalid representation, and another layer may decode it into a valid Unicode string.
Also, some regex engines work on UTF-8 encoded bytes, and some work on arrays of code points.
---
Again I'd be interested in more real examples.
[1] A sequence of "Unicode scalars" -- code points not in the surrogate range
https://www.brainonfire.net/blog/2022/04/11/what-is-parser-m...
> This C++ library is part of the JavaScript package utf-8-validate. The utf-8-validate package is routinely downloaded more than a million times per week.
> If you are using Node JS (19.4.0 or better), you already have access to this function as buffer.isUtf8(input).
And at https://github.com/websockets/utf-8-validate/tree/master/dep...
> # Reference
> John Keiser, Daniel Lemire, [Validating UTF-8 In Less Than One Instruction Per Byte](https://arxiv.org/abs/2010.03090), Software: Practice & Experience 51 (5), 2021
In fact, yes there is! Your application may vary. :)