DevTools Hub

Search tools

Search for a developer tool

Why the Same Emoji Produces Three Different Escape Sequences

Part of the Encoding Toolkit

Hex-escape the same emoji for a SQL string, a JavaScript string, and an HTML attribute, and you get three sequences that don't just look different — they're built on three completely different ideas about what "the character" even is. None of them is wrong. Each one is exactly correct for its own format's spec, and that spec quietly picks a different unit to count.

Three real models for the same character

A character like 😀 has a single Unicode code point, U+1F600 — but that code point is outside the Basic Multilingual Plane, so representing it depends on which encoding or unit system a format actually committed to:

encoding 😀 (U+1F600) — one character, five modes, three different models

plain / sql

UTF-8 bytes

f09f9880

4 bytes — one hex pair per UTF-8 byte

javascript / json

UTF-16 code units

😀

2 units — split into a surrogate pair

html / xml

Unicode code point

😀

1 reference — the raw code point, no splitting

Plain hex and SQL's 0x... literal encode the character's raw UTF-8 bytes — this emoji takes four bytes in UTF-8, so the hex output is four byte-pairs, eight hex digits. JavaScript and JSON string escapes are defined in terms of UTF-16 code units, and this code point is too large to fit in one unit — so it's split into a surrogate pair, two separate \u escapes that only mean something read together. HTML and XML numeric character references are defined directly on Unicode code points, with no intermediate encoding step at all — so the same character becomes exactly one reference, no splitting, no byte count.

Encoding and decoding have to agree on which model they're using

Each mode's decoder has to reverse the exact operation its encoder performed, not a generic one:

plain/sql:        parse hex byte pairs → decode as UTF-8
javascript/json:  parse \x/\u escapes  → String.fromCharCode (UTF-16 units)
html/xml:         parse &#x...; refs   → String.fromCodePoint (code points)

Feed a surrogate-pair \u escape through code expecting single code points, or feed a single HTML numeric reference through code expecting UTF-16 units one at a time, and the reconstruction silently uses the wrong model — which is exactly why a decoder has to know which format produced its input rather than guessing from the hex digits alone.

Why the JavaScript/JSON split specifically looks surprising

Most escaped characters — accented letters, symbols, anything inside the Basic Multilingual Plane — fit in one UTF-16 unit, so \u escaping looks like a clean one-escape- per-character rule right up until a character doesn't fit. Emoji are the common case that breaks it: almost every emoji people actually type lives outside the BMP, so almost every emoji in a JS or JSON string escape comes out as two \u escapes, not one — not a special case, just the normal rule applied to a code point too large for a single 16-bit unit.

Try it yourself

Hex Encode runs all five modes on real text you type in — plain hex, SQL's 0x literal, JavaScript's \x/\u escapes, JSON's \u escapes, and HTML/XML numeric character references — so you can see exactly how many units each one produces for the same input. Hex Decode reverses any of them back to plain text. Both run entirely in your browser.

Related tools