Code points, bytes and entities of any text.
Breaks a string into its code points — U+ notation, decimal, UTF-8 and UTF-16 bytes, HTML entity — the way the parser sees it, not the way it looks.