Why your regex works in the tester but fails in your code
A pattern that matches everywhere except in your code is usually not a skill problem. Regex syntax has
no single standard: JavaScript follows ECMAScript, Python has its own re module, and PHP, Ruby, Java
and most CLI tools sit near PCRE. Copying a pattern from a tester or another language copies its
dialect with it.
$ does not mean the same thing in two engines
Without m, JavaScript’s $ matches only at the very end of the input, so /^a$/.test("a\n") is
false. In PCRE and Python the same pattern succeeds, because their $ also matches just before a
trailing newline: a plain JavaScript $ is PCRE’s \z, absolute end with no tolerance. Put the
optional newline in the pattern instead — /^a(?:\r?\n)?$/ matches "a\n". With m, $ matches
before every line break and at the end, so /^a$/m.test("a\n") is true, but m redefines every
anchor in the pattern.
JavaScript has no \A and no \Z. Without u, /\A/ is just /A/, an identity escape matching the
letter A. With u, /\A/u is a SyntaxError. Start-of-string is what bare ^ already means.
\d does not mean “a digit”
In JavaScript \d is exactly [0-9], ten ASCII digits, with or without the u flag. Python 3’s \d
matches Unicode decimal digits by default. 1 matches in both, while ٥ (Arabic-Indic five, U+0665)
and 5 (full-width five, U+FF15) match in Python and not in JavaScript — why a validator passes in a
tester and refuses a number typed with a CJK input method. JavaScript’s Unicode-aware class is the
property escape \p{Nd}, which needs u: /^\d+$/u.test("٥") is false and /^\p{Nd}+$/u.test("٥")
is true.
\w and \b stay ASCII under u too. A Chinese character is not a word character, so /\b中文\b/
never matches inside 这是中文测试: no boundary exists there. \s is the exception, matching U+3000
ideographic space and non-breaking space.
Syntax that JavaScript does not have
| Syntax | Where it comes from | In JavaScript |
|---|---|---|
(?<name>…) | ECMAScript, PCRE, .NET | Supported; Python’s (?P<name>…) and Perl’s (?'name'…) do not compile |
(?<=…), (?<!…) | PCRE, Python, .NET | V8 and Firefox parse them; Safari was last, and older browsers may not |
\K | PCRE only | Matches a literal K, or fails to compile under u |
(?>…), *+ | PCRE, Java | Not supported: a syntax error even under v in current V8 |
(?i), (?s) | PCRE, Python | Bare inline flags never exist; scoped (?i:…) only arrived in very recent engines |
These fail at build time: an unsupported construct throws when the regex is built rather than matching nothing. A pattern that “worked” in a PCRE tester and broke your page was never portable.
The replacement string is a second language
Passing the match test says nothing about the substitution: in the replacement given to
String.replace, $ is an escape character.
$&is the whole match,$`the text before it,$'the text after it.$1to$9are numbered groups,$<name>a named one, and$$collapses to a single$.$0is not special and stays literal.- With one group,
$10is group 1 plus the character0, so"a1b".replace(/a(\d)/, "$10")gives10b.
Python’s re.sub uses \1 and \g<name> and treats $ as text, so pasting a replacement in either
direction yields wrong output with no error. When the replacement is data, pass a function:
"abc".replace(/b/, () => "$$1") returns $$1, because a callback’s result is never reinterpreted.
A regex object keeps state
With g or y, .test() and .exec() read and write lastIndex, so a shared object does not answer
the same way twice:
const re = /a/g;
re.test("ab"); // true
re.test("ab"); // false — the first call moved lastIndex
re.test("ab"); // true
The answer depends on how often the object has already been used, hence the intermittent bug. A regex
literal inside a function creates a fresh object on each evaluation, so the trouble comes from patterns
hoisted to module scope, cached in a variable or stored on a class field. Drop g when you only ask
whether it matches, or set re.lastIndex = 0 first. .match() and .replace() with g, and
.matchAll(), ignore stale state and leave lastIndex at 0.
A manual exec loop has one more trap: a zero-width match leaves lastIndex where it was, so unless
you push it — if (re.lastIndex === m.index) re.lastIndex++ — the loop never terminates.
Code points, code units and the u flag
Strings are UTF-16 code units, and an emoji outside the basic plane is two of them:
"\u{1F600}".length is 2, and without u one . consumes one unit, so /^.{2}$/ matches it as two
characters. With u the dot and quantifiers count code points, so /^.$/u matches it as one. The
family sequence "\u{1F468}\u200D\u{1F469}\u200D\u{1F467}" is 8 code units, 5 code points, one thing
on screen — neither .length nor . under u agrees with the reader. Use Array.from(s).length for
code points, Intl.Segmenter for grapheme clusters.
The u flag also turns leniency into errors: \A and \K fail to compile instead of matching
letters. For a dot that crosses a line break use s, since (?s) is not JavaScript.
Collecting every match
matchAll is the sane way to take every match: it requires g, clones the regex so your lastIndex is
untouched, exposes named groups as match.groups, and advances past empty matches —
"ab".matchAll(/(?:)/g) yields three positions instead of hanging.
A workflow that survives the paste
Test the pattern in the engine you will run, with the flags you intend to ship already on it: the same
pattern with and without u is two matchers. Test the substitution separately; it has its own escape
language. If the pattern must also work server-side, port it by rewriting anchors, digit classes and
group syntax instead of assuming one tester speaks for both. The regex tester linked below runs your
browser’s own engine, the only dialect your page can use.
Engine behaviour moves: lookbehind, scoped modifiers and property escapes landed on different schedules across browsers and are still uneven on older devices. Where a detail matters, verify it in the runtime you deploy to rather than trusting any table, including this one.