Skip to main content
QUIETLYTIC
Encoding & Data

HTML Entity Encoder & Decoder

Convert between raw characters and HTML entity references.

Local · nothing leaves this browser Waiting for input
Esc Clear
Output

Paste text on the left.

How it works

An entity reference lets a character appear in a document without being read as part of it. Five characters need that treatment because they change how a parser reads the surrounding markup: &, <, >, the double quote and the apostrophe. Everything else is a question of transport, not safety.

Why this does not use the browser to decode

The usual one-line solution assigns the string to innerHTML and reads textContent back. That hands a stranger's input to the HTML parser, which is an XSS sink in any page that does it — the string being decoded is exactly the string you do not trust. Decoding here runs off a table, so nothing is ever parsed as markup.

What that costs, stated plainly

The table holds 198 named references: the five markup characters, the Latin-1 block from HTML 4.01, the common symbol names, and HTML5's upper-case spellings such as &AMP;. HTML5 defines 2,231. Every numeric reference decodes regardless — decimal &#8212; and hexadecimal &#x2014; alike — and any name outside the table is listed rather than silently dropped.

Decoded the way a browser decodes text

HTML5 still honours some pre-HTML4 habits, and a decoder that ignores them shows you something other than what a browser renders. The Latin-1 names and &amp;, &lt;, &gt; and &quot; decode without their semicolon, taking the longest name that fits — so &copy2026 is ©2026 and &notit; is ¬it;. Numeric references in the C1 range are read as Windows-1252 (&#150; is an en dash), and a reference to NUL, a surrogate or anything past U+10FFFF becomes U+FFFD, the replacement character. This follows the rules for text content; inside an attribute value a browser is stricter about semicolon-less names.

Unknown names come back unchanged

&bogusentity; is literal text in a document. A decoder that swallowed it would change the content it was asked to reveal, so it is returned exactly as written and reported separately.

Example

Encoding <a href="x"> gives &lt;a href=&quot;x&quot;&gt;. Note what happens to &amp;lt; on the way back: it decodes to &lt;, not to <. Decoding runs one pass, so a double-encoded string stays visibly double-encoded instead of quietly resolving to markup.

Frequently asked questions

Why not just use the browser to decode entities?

The usual one-liner assigns the string to innerHTML and reads textContent back, which hands a stranger’s input to the HTML parser — an XSS sink in any page that does it. Decoding here runs off a table, so nothing is ever parsed as markup.

Which characters does the default encoding escape?

The five that change how a parser reads a document: ampersand, less-than, greater-than, double quote and apostrophe. Escaping ordinary accented text buys no safety and makes it unreadable, so it is left alone unless you switch to the non-ASCII scope.

Why does &copy2026 decode when it has no semicolon?

Because a browser decodes it. HTML5 keeps 106 legacy names semicolon-optional — the Latin-1 block plus amp, lt, gt and quot — and takes the longest one that fits, so &copy2026 renders as ©2026 and &notit; as ¬it;. Showing anything else would misrepresent what the page actually displays. Numeric references follow the browser too: &#150; is an en dash, not an invisible control character.

What happens to an entity it does not recognise?

It comes back exactly as written. A name outside the table might be literal text in the document, and a decoder that swallowed it would change the content it was asked to reveal. Unrecognised names are listed separately so you can see what was skipped.

Related tools

From the intelligence desk