DevTools Hub

Search tools

Search for a developer tool

Why Two Identical-Looking Strings Aren't Equal

Part of the Encoding Toolkit

Two strings render identically on screen. A user copies one, pastes it into a search box or a login form, and the exact-match check against the other one silently fails — no typo, no visible difference, nothing a human could spot by looking. This isn't a bug in the comparison code. It's because "the same text" isn't one thing in Unicode — the same character can be encoded more than one way, and nothing about how it looks tells you which encoding you're holding.

A character, a code point, and a byte sequence are three different counts

Every character is a code point — a single number, written U+XXXX. How that number gets stored as bytes depends on the encoding, and for anything outside the first 65,536 code points (the Basic Multilingual Plane), UTF-16 can't fit it in one unit — it splits the code point into a pair of units called a surrogate pair. UTF-8 handles the same code point differently again, using between one and four bytes depending on its size:

charcode pointUTF-16 unitsUTF-8 bytes
AU+0041
0041
41
éU+00E9
00e9
c3a9
😀U+1F600
d83dde00
f09f9880

one visual character, one code point — but 😀 needs two UTF-16 units (a surrogate pair) and four UTF-8 bytes

A is simple all the way down: one code point, one UTF-16 unit, one UTF-8 byte. é is still one code point, but needs two UTF-8 bytes — ASCII doesn't cover it. The emoji needs a full surrogate pair in UTF-16 and four bytes in UTF-8, despite being exactly one character on screen. This is exactly why "😀".length is 2 in JavaScript, not 1 — .length counts UTF-16 units, not visual characters or code points. Iterating with something that respects code points (like a for...of loop, or Array.from) gets the count right; indexing by .length doesn't.

The same character, two different code point sequences

é can be stored as a single precomposed code point, U+00E9. Or it can be stored as two code points — a plain e (U+0065) followed by a combining acute accent mark (U+0301) — which render as the exact same glyph. Nothing about looking at the rendered text tells you which one you have; you have to actually inspect the code points. Unicode defines canonical normalization forms precisely to resolve this ambiguity. Run both versions through all four:

input: "é" (precomposed, U+00E9) — 1 code point
  NFC  → "é" (1 code point) — same as input
  NFD  → "é" (2 code points) — different from input

input: "é" (decomposed, U+0065 U+0301) — 2 code points
  NFC  → "é" (1 code point) — different from input
  NFD  → "é" (2 code points) — same as input

NFC (Normalization Form Canonical Composition) always composes down to the shortest form — one code point when a precomposed one exists. NFD does the opposite, always decomposing into base character plus combining marks. Two strings that look pixel-identical can disagree on which of these forms they're already in, which means === between them can return false even though a human reading both would never notice a difference. A username field, a search index, or a database uniqueness constraint that compares raw strings instead of normalizing first inherits that exact failure mode.

Not every 16-bit value is a valid character on its own

Surrogate pairs work because the two halves — a high surrogate in the range U+D800–U+DBFF and a low surrogate in U+DC00–U+DFFF — only mean something as a pair. Neither half is a valid character by itself, which is exactly why rebuilding text from a list of U+XXXX code points has to reject a lone surrogate outright rather than silently emitting broken output:

buildFromCodePoints("U+0048 U+0065 U+006C U+006C U+006F")
  → "Hello"  (five valid code points)

buildFromCodePoints("U+D800")
  → error: "U+D800 is not a valid Unicode code point."

That range is reserved specifically so surrogate pairs and genuine standalone code points never collide — a lone surrogate isn't a rare character, it's not a character at all.

Try it yourself

Unicode Converter runs every example above on real text you type in — per-character code point, UTF-8 bytes, UTF-16 units, category and script, all four normalization forms with a flag for whenever one disagrees with your input, and a code-point builder that rejects invalid input exactly like the lone-surrogate example. Runs entirely in your browser.

Related tools