Stripping SQL comments and collapsing whitespace looks like a job for two regular expressions and a few minutes. It is, right up until a comment marker shows up inside a string literal that was never a comment at all — at which point a regex-based minifier doesn't just fail to strip it, it corrupts the query by truncating a value mid-string.
A comment marker inside a string isn't a comment
WHERE note = 'contains -- not a comment' contains the exact two characters a line-comment regex looks for, sitting inside a perfectly ordinary string value. A regex that strips everything from -- to the end of the line has no way to know it's inside quotes at that point in the text — regular expressions don't track that kind of state:
input: WHERE note = 'contains -- not a comment'
naive: strip /--.*$/gm with a regex
...WHERE note = 'contains
the string is truncated mid-value — everything after -- inside it is gone
real: tokenize, then strip only comment tokens
...WHERE note = 'contains -- not a comment'
the tokenizer already consumed the whole string as one unit — nothing inside it looks like a comment
The fix isn't a smarter regex. It's not using regular expressions for this at all. A real minifier has to tokenize the query first — walk it character by character, recognize when it has entered a string literal, and consume the entire string as one unit before it ever looks for the next token. By the time the tokenizer would otherwise notice a --, it has already moved past the whole string in a single step — there's nothing left inside it to misinterpret.
The same problem, and the same fix, applies to block comments
'contains /* not a comment */ either' survives a real tokenizer completely intact for the identical reason — the string-literal rule runs before any comment-detection rule gets a chance to fire, because the whole string was already consumed as one token. A /\*.*?\*\//-style regex has no equivalent protection; it matches the pattern wherever it appears, string or not.
SQL's own quoting rules have to be followed exactly, not approximated
Properly tokenizing a string means handling SQL's actual escape rule for a literal quote inside one — doubling it:
input: 'O''Brien'
output: 'O''Brien' — one string, not two broken onesA naive "find the next quote character" approach would see the second quote in O''Brien and conclude the string ended there, leaving Brien' dangling as unparseable leftover text. The tokenizer has to check for the doubled-quote case specifically and treat it as an escaped character, not a terminator.
Minification has to make the same judgment call as a human reading the query
Beyond comments and strings, collapsing whitespace safely means keeping exactly the separators that still mean something: a number written in exponent notation (1.5E-3) has to stay one token, not get split on the embedded -, and identifiers quoted three different ways across database engines — double quotes, backticks, square brackets — all have to pass through untouched rather than being mistaken for string literals or stripped as punctuation. Get any of these wrong and the output isn't just ugly, it's a different query.
Try it yourself
SQL Minifier tokenizes a real query the way described above — comments stripped, whitespace collapsed to single spaces, string literals and every quoting style left exactly as written — and reports the actual byte savings alongside the result. Runs entirely in your browser.