⚙️ Coding

Regex for People Who Keep Googling Regex

You do not need to memorise regex. You need about nine building blocks and an understanding of why your pattern matches too much.

Regular expressions have a reputation for being write-only — easy to produce, impossible to read later. That reputation is mostly earned by people trying to do too much in one pattern.

In practice, the regex you will write day to day uses about nine constructs. Learn those and you can build and read most patterns without looking anything up.

The nine pieces

Character classes — what kind of character.

  • \d a digit · \w a word character (letter, digit, underscore) · \s whitespace
  • Capitalised versions negate: \D is any non-digit.
  • [abc] any one of these · [^abc] anything but these · [a-z] a range
  • . any character at all (except newline, by default)

Quantifiers — how many.

  • + one or more · * zero or more · ? zero or one
  • {3} exactly three · {2,5} between two and five · {2,} two or more

Anchors and groups — where, and what to capture.

  • ^ start of the string · $ end of the string
  • (…) a capture group · (?:…) a group that does not capture
  • | alternation — this or that

That is genuinely most of it.

Greedy matching: the trap everyone falls into

If your pattern matches far more than you expected, this is almost certainly why.

Quantifiers are greedy by default — they consume as much as possible, then give back only as much as needed for the rest of the pattern to succeed.

Run <.+> against <b>bold</b> and it matches the entire string, not just the opening tag. The .+ eats everything to the end, then backtracks just far enough to find a final > — which is the one closing </b>.

Adding ? makes a quantifier lazy: it takes as little as possible.

<.+>   → matches <b>bold</b>
<.+?>  → matches <b>

Whenever a pattern over-matches, try adding ? to the quantifier before anything else.

Flags change everything

  • g — global. Find every match rather than stopping at the first. Without it, a replace operation changes only the first occurrence, which is a very common source of "my replace did not work".
  • i — case-insensitive.
  • m — multiline. Makes ^ and $ match at line boundaries instead of only at the very start and end of the whole string. Essential when working line by line.
  • s — dotall. Makes . match newlines too. Needed for patterns that span lines.

Escaping

These characters have special meaning and must be escaped with a backslash to match literally:

. * + ? ( ) [ ] { } | ^ $ \

So a literal dot is \., and matching a price like $4.99 needs \$4\.99.

Forgetting to escape a dot is the second most common regex bug after greediness. example.com as a pattern matches exampleXcom, because the dot means "any character".

Patterns worth keeping

A few that come up constantly:

  • \s+$ — trailing whitespace at the end of lines
  • ^\s*$ — blank lines (with the m flag)
  • \n{3,} — runs of three or more newlines, for collapsing to two
  • \d{4}-\d{2}-\d{2} — ISO dates
  • #[0-9a-fA-F]{6}\b — hex colours
  • https?://[^\s]+ — URLs, roughly

Capture groups in the search pattern are referenced in the replacement as $1, $2 and so on. That is how you reformat rather than merely substitute:

find:    (\d{4})-(\d{2})-(\d{2})
replace: $3/$2/$1

2026-10-01  →  01/10/2026

Do not validate email addresses with regex

This deserves its own section because it is attempted constantly.

The specification for a valid email address permits forms almost nobody implements — quoted local parts, comments in parentheses, IP-literal domains. A regex that accepts every valid address is enormous and famously unreadable. A regex that is readable rejects valid addresses.

More importantly, syntactic validity tells you nothing useful. definitely-not-real@gmail.com is perfectly well-formed and will bounce.

The practical approach: check loosely for something resembling x@y.z, then send a verification email. The email is the validation.

Build patterns by testing, not by reasoning

Regex is difficult to get right by reading alone — the difference between a pattern that works and one that quietly matches too much is often a single character, and you cannot see it by staring.

Build incrementally against real sample text. Start with the simplest thing that matches one case, check it, then extend. Our regex tester highlights matches live as you type, which turns the whole exercise from guesswork into observation.

Include your edge cases in the sample text from the start: the empty value, the one with unusual characters, the one that is nearly but not quite a match. A pattern that works on three tidy examples and fails on real data is worse than no pattern.

Flavours differ

Regex is not one language. JavaScript, Python, PCRE (PHP, Perl), Go's RE2 and POSIX all differ in meaningful ways.

The common constructs above work everywhere. Where they diverge: lookbehind support, named groups syntax, Unicode property escapes, and whether backtracking is permitted at all — Go's RE2 deliberately excludes features that allow catastrophic backtracking, so some patterns simply will not compile there.

If a pattern you found online does not work, flavour mismatch is a likely cause. Our tester uses JavaScript's engine, since it runs in your browser.

Frequently asked questions

Why does my pattern match more than I expected?

Greedy quantifiers. By default + and * take as much as they can. Add ? to make them lazy — .+? instead of .+ — and they take as little as possible.

Why does my replace only change the first match?

You are missing the g (global) flag. Without it, the operation stops after the first occurrence.

How do I match a literal dot?

Escape it: \. An unescaped dot means 'any character', so example.com as a pattern also matches exampleXcom.

Can I validate an email address with regex?

Not meaningfully. The specification permits forms almost nobody implements, and syntactic validity does not mean the address exists. Match loosely and send a verification email.

Why does a pattern from Stack Overflow not work in my language?

Regex flavours differ. Lookbehind, named groups and Unicode property escapes vary between JavaScript, Python, PCRE and Go's RE2. Check which flavour the example was written for.

Everything on ToolYard runs in your browser. No uploads, no accounts, no limits.

Browse all tools →