Unicode Inspector
See the real code points, UTF-8 bytes and normalization forms behind any pasted text.
Type or paste text on the left.
How it works
Reads a string by Unicode code point (correctly pairing a UTF-16 surrogate pair into one character, unlike
.length or .split(''), both of which operate on UTF-16 code units), shows the UTF-8 byte
encoding of each, and reports the four Unicode normalization forms.
Text that reads differently from what it is
Bidirectional controls such as RIGHT-TO-LEFT OVERRIDE (U+202E) change the order characters are displayed in without changing the order they are stored in. In source code that lets a reviewer see one program while the compiler reads another — the Trojan Source technique, CVE-2021-42574. An override or isolate still open at the end of a line is flagged as an alert; balanced pairs, which legitimate right-to-left text uses, are flagged more quietly.
Zero-width characters, tag characters (which can carry a whole hidden sentence), stray control characters and
look-alike spaces are flagged too, and every one of them is shown in the output as a visible marker such as
[U+200B ZWSP], so the report itself cannot be reordered or hide anything.
Mixed scripts, not a confusables table
A word that mixes Latin, Cyrillic and Greek letters is flagged, because that is the usual shape of a homoglyph
spoof — a Cyrillic а (U+0430) standing in for a Latin a (U+0061). This tool does
not bundle Unicode's confusables data, so a look-alike drawn entirely from one other script is not detected.
No character-name lookup
Ordinary code points are shown by number (e.g. U+0041) only, not a name like "LATIN CAPITAL LETTER A"
— a name database is several megabytes and this tool does not bundle one. The flagged characters above are the
exception: each carries its Unicode name.
Example
é can be one composed code point (NFC) or an "e" followed by a combining acute accent (NFD) — both render identically but compare unequal in code, or in a database that does not normalize before matching. The normalization panel shows both forms explicitly so the difference is visible rather than invisible.
Frequently asked questions
Why does .length in JavaScript disagree with the code point count here?
string.length counts UTF-16 code units, and a code point outside the Basic Multilingual Plane (most emoji, for instance) is encoded as a surrogate pair — two units for one character. This tool iterates by code point instead, so an emoji correctly counts as one, matching what a person would call "one character".
Why do NFC and NFD show the same character but report "changed: true"?
They can look identical while being different underlying code point sequences — é as one composed code point (NFC) versus e followed by a combining acute accent (NFD) renders the same but compares unequal in code or in a database that does not normalize before matching. This is exactly the case that "changed: true" is flagging.
Related tools
Base64 Encoder
Encode text or bytes to Base64 or Base64URL.
LocalHex ↔ ASCII Converter
Convert between hex byte sequences and readable text.
LocalHTML Entity Encoder & Decoder
Convert between raw characters and HTML entity references.
LocalBase64 Decoder
Decode Base64 and Base64URL back to text or raw bytes.
LocalURL Encoder
Percent-encode text for a URL path, query value or component.
LocalURL Decoder
Decode percent-encoded text, including repeatedly-encoded strings.
Local