DevTools Hub

Search tools

Search for a developer tool

Why You Can't Minify SQL with Find-and-Replace

Part of the SQL Toolkit

Stripping SQL comments and collapsing whitespace looks like a job for two regular expressions and a few minutes. It is, right up until a comment marker shows up inside a string literal that was never a comment at all — at which point a regex-based minifier doesn't just fail to strip it, it corrupts the query by truncating a value mid-string.

A comment marker inside a string isn't a comment

WHERE note = 'contains -- not a comment' contains the exact two characters a line-comment regex looks for, sitting inside a perfectly ordinary string value. A regex that strips everything from -- to the end of the line has no way to know it's inside quotes at that point in the text — regular expressions don't track that kind of state:

input: WHERE note = 'contains -- not a comment'

naive: strip /--.*$/gm with a regex

...WHERE note = 'contains

the string is truncated mid-value — everything after -- inside it is gone

real: tokenize, then strip only comment tokens

...WHERE note = 'contains -- not a comment'

the tokenizer already consumed the whole string as one unit — nothing inside it looks like a comment

The fix isn't a smarter regex. It's not using regular expressions for this at all. A real minifier has to tokenize the query first — walk it character by character, recognize when it has entered a string literal, and consume the entire string as one unit before it ever looks for the next token. By the time the tokenizer would otherwise notice a --, it has already moved past the whole string in a single step — there's nothing left inside it to misinterpret.

The same problem, and the same fix, applies to block comments

'contains /* not a comment */ either' survives a real tokenizer completely intact for the identical reason — the string-literal rule runs before any comment-detection rule gets a chance to fire, because the whole string was already consumed as one token. A /\*.*?\*\//-style regex has no equivalent protection; it matches the pattern wherever it appears, string or not.

SQL's own quoting rules have to be followed exactly, not approximated

Properly tokenizing a string means handling SQL's actual escape rule for a literal quote inside one — doubling it:

input:  'O''Brien'
output: 'O''Brien'   — one string, not two broken ones

A naive "find the next quote character" approach would see the second quote in O''Brien and conclude the string ended there, leaving Brien' dangling as unparseable leftover text. The tokenizer has to check for the doubled-quote case specifically and treat it as an escaped character, not a terminator.

Minification has to make the same judgment call as a human reading the query

Beyond comments and strings, collapsing whitespace safely means keeping exactly the separators that still mean something: a number written in exponent notation (1.5E-3) has to stay one token, not get split on the embedded -, and identifiers quoted three different ways across database engines — double quotes, backticks, square brackets — all have to pass through untouched rather than being mistaken for string literals or stripped as punctuation. Get any of these wrong and the output isn't just ugly, it's a different query.

Try it yourself

SQL Minifier tokenizes a real query the way described above — comments stripped, whitespace collapsed to single spaces, string literals and every quoting style left exactly as written — and reports the actual byte savings alongside the result. Runs entirely in your browser.

Related tools