Two strings render identically on screen. A user copies one, pastes it into a search box or a login form, and the exact-match check against the other one silently fails — no typo, no visible difference, nothing a human could spot by looking. This isn't a bug in the comparison code. It's because "the same text" isn't one thing in Unicode — the same character can be encoded more than one way, and nothing about how it looks tells you which encoding you're holding.
A character, a code point, and a byte sequence are three different counts
Every character is a code point — a single number, written U+XXXX. How that number gets stored as bytes depends on the encoding, and for anything outside the first 65,536 code points (the Basic Multilingual Plane), UTF-16 can't fit it in one unit — it splits the code point into a pair of units called a surrogate pair. UTF-8 handles the same code point differently again, using between one and four bytes depending on its size:
| char | code point | UTF-16 units | UTF-8 bytes |
|---|---|---|---|
| A | U+0041 | 0041 | 41 |
| é | U+00E9 | 00e9 | c3a9 |
| 😀 | U+1F600 | d83dde00 | f09f9880 |
one visual character, one code point — but 😀 needs two UTF-16 units (a surrogate pair) and four UTF-8 bytes
A is simple all the way down: one code point, one UTF-16 unit, one UTF-8 byte. é is still one code point, but needs two UTF-8 bytes — ASCII doesn't cover it. The emoji needs a full surrogate pair in UTF-16 and four bytes in UTF-8, despite being exactly one character on screen. This is exactly why "😀".length is 2 in JavaScript, not 1 — .length counts UTF-16 units, not visual characters or code points. Iterating with something that respects code points (like a for...of loop, or Array.from) gets the count right; indexing by .length doesn't.
The same character, two different code point sequences
é can be stored as a single precomposed code point, U+00E9. Or it can be stored as two code points — a plain e (U+0065) followed by a combining acute accent mark (U+0301) — which render as the exact same glyph. Nothing about looking at the rendered text tells you which one you have; you have to actually inspect the code points. Unicode defines canonical normalization forms precisely to resolve this ambiguity. Run both versions through all four:
input: "é" (precomposed, U+00E9) — 1 code point
NFC → "é" (1 code point) — same as input
NFD → "é" (2 code points) — different from input
input: "é" (decomposed, U+0065 U+0301) — 2 code points
NFC → "é" (1 code point) — different from input
NFD → "é" (2 code points) — same as inputNFC (Normalization Form Canonical Composition) always composes down to the shortest form — one code point when a precomposed one exists. NFD does the opposite, always decomposing into base character plus combining marks. Two strings that look pixel-identical can disagree on which of these forms they're already in, which means === between them can return false even though a human reading both would never notice a difference. A username field, a search index, or a database uniqueness constraint that compares raw strings instead of normalizing first inherits that exact failure mode.
Not every 16-bit value is a valid character on its own
Surrogate pairs work because the two halves — a high surrogate in the range U+D800–U+DBFF and a low surrogate in U+DC00–U+DFFF — only mean something as a pair. Neither half is a valid character by itself, which is exactly why rebuilding text from a list of U+XXXX code points has to reject a lone surrogate outright rather than silently emitting broken output:
buildFromCodePoints("U+0048 U+0065 U+006C U+006C U+006F")
→ "Hello" (five valid code points)
buildFromCodePoints("U+D800")
→ error: "U+D800 is not a valid Unicode code point."That range is reserved specifically so surrogate pairs and genuine standalone code points never collide — a lone surrogate isn't a rare character, it's not a character at all.
Try it yourself
Unicode Converter runs every example above on real text you type in — per-character code point, UTF-8 bytes, UTF-16 units, category and script, all four normalization forms with a flag for whenever one disagrees with your input, and a code-point builder that rejects invalid input exactly like the lone-surrogate example. Runs entirely in your browser.