Regular expressions intimidate people because the syntax looks like line noise, but the building blocks are few. Learn a dozen primitives and you can read almost any pattern. This is a practical tour, not a grammar specification, with the parts that bite in real code called out explicitly.
The metacharacter cheat sheet
These are the symbols you will meet most. Keep this table open the first ten times you read a pattern:
| Symbol | Meaning | Example |
|---|---|---|
. | any char except newline (with s flag, includes newline) | a.c matches "abc" |
^ $ | start / end anchor | ^a starts with a |
\b | word boundary (zero-width) | \bcat\b whole word cat |
\d \D | digit / non-digit | \d{3} three digits |
\w \W | word char / non-word char | \w+ identifier |
\s \S | whitespace / non-whitespace | \s+ runs of space |
* + ? | 0+ / 1+ / 0-or-1 quantifier | a+ one or more a |
{n,m} | n to m times | a{2,4} |
[ ] | character set | [a-z0-9] |
[^ ] | negated set | [^0-9] non-digit |
( ) | group and capture | (ab)+ |
(?: ) | group only, no capture | (?:ab)+ |
(?= ) (?! ) | lookahead yes / no | \d(?=px) |
(<= ) (<! ) | lookbehind yes / no | (<=\$)\d+ |
| | alternation (or) | cat|dog |
\ | escape a metachar | \. literal dot |
Character classes
A class matches one character from a set. [aeiou] matches any vowel; [^0-9] matches any character that is not a digit. Predefined shortcuts: \d digit, \w word character (letter, digit, underscore), \s whitespace. Uppercase inverts: \D, \W, \S.
/[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}/
Quantifiers and greed
* zero or more, + one or more, ? zero or one, {n,m} between n and m. By default they are greedy: they grab as much as possible. Append ? to make them lazy.
const s = 'a x b y b';
s.match(/a.*b/); // "a x b y b" (greedy: eats to the last b)
s.match(/a.*?b/); // "a x b" (lazy: stops at the first b)
Anchors
^ matches the start of the string (or line, with the m flag), $ the end. \b is a word boundary. Anchors do not consume characters; they assert position.
/^\d{4}-\d{2}-\d{2}$/ // a strictly-formatted date, whole string
Groups and capturing
Parentheses both group and capture. (\d{4})-(\d{2}) captures the year and month separately, available in the match result or in a replacement as $1, $2. Use (?:...) when you only need grouping, not capture — it is cheaper.
'2026-08-26'.replace(/(\d{4})-(\d{2})-(\d{2})/, '$3/$2/$1');
// "26/08/2026"
Backreferences
Inside the pattern, \1 refers to the first captured group. Useful for "the same word twice":
/(\w+)\s+\1/ // matches a repeated word like "the the"
Lookahead and lookbehind
These are zero-width assertions about what is ahead or behind, without consuming it. (?=...) is positive lookahead, (?!...) negative; (<=...) and (<!...) are the lookbehinds. They are how you say "followed by" or "not preceded by" without including that text in the match.
/\d(?=\.)/ // a digit that is followed by a dot
/(<=@)\w+/ // word characters right after an @
/foo(?!bar)/ // "foo" not followed by "bar"
/(<!\$)\d+/ // digits not preceded by a dollar sign
A practical use: validate a password that must contain a digit without putting the digit in the captured match. Combine anchors and lookahead:
/^(?=.*\d).{8,}$/ // at least 8 chars and at least one digit
Flags
g— global, find all matches instead of stopping at the first.i— case-insensitive.m— multiline, so^and$match line breaks, not just the whole string.s— dotall, lets.match newline characters too.u— unicode, treats the pattern as UTF-16 code points; use it whenever you handle non-ASCII.y— sticky, matches only at the last index (lastIndex).
Catastrophic backtracking, reproduced
Some patterns explode in cost on certain inputs. The classic culprit is nested quantifiers over overlapping content. Take ^(a+)+b$ matched against a long run of a's that does not end in b:
const re = /^(a+)+b$/;
console.time('re');
re.test('aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa'); // no match, but...
console.timeEnd('re');
// with a few more a's this goes from instant to seconds
The engine tries every possible way to split the a's between the inner and outer plus, then backtracks when the trailing b is missing. Each extra a multiplies the work. On a server this is a denial-of-service: an attacker sends a long string and your regex hangs the event loop. The fix is to make the parts non-overlapping — here a single a+ and an anchor already imply the structure, so drop the outer plus:
// Safer: no nested quantifier on the same atom
const re = /^a+b$/;
re.test('aaaa...a'); // fails fast
If you genuinely need grouping, anchor it and avoid letting the inner and outer repeat the same character class. Where the engine supports atomic groups (PCRE, Java, .NET, Python regex), (?>a+)+ prevents the backtrack entirely. JavaScript lacks them, so the practical rule is: never nest a quantifier inside a quantifier over the same characters.
Common real-world patterns and why the textbook regex fails
The "standard" regex you find on Stack Overflow is usually a format check, not a validator. Here is what goes wrong in production:
- Email. The pragmatic pattern above catches the vast majority of valid addresses, but no regex fully implements the RFC 5322 grammar — it allows comments in parentheses, quoted local parts, and more. Worse, many "email regexes" reject perfectly valid addresses like
[email protected]or[email protected]. In a real signup form, use a shallow pattern to block obvious typos, then send a confirmation link to prove the address works. A regex cannot tell you whethera@bis a real mailbox. - URL. A regex can recognise a URL-shaped string, but it will happily accept
https://not a real hostshaped text and reject legitimate but unusual schemes. Crucially, a pattern that "validates" a URL does nothing about the open-redirect and SSRF risks from the previous article — those need parsing and host checks. Treat URL regex as a shallow filter for "does this look like a link in free text", never as authorization. - Date.
^\d{4}-\d{2}-\d{2}$checks format, not validity —2026-13-40passes. Even adding(0[1-9]|1[0-2])for the month still permits2026-02-30. Real date validation means parsing into a date object and checking it round-trips. A regex is the wrong tool for "is this a real calendar date". - Phone. International numbers have wildly different formats; a single regex either rejects valid numbers or accepts impossible ones. Prefer a dedicated library (
libphonenumber) and just check "contains mostly digits and a few punctuation marks".
When to reach for a real parser instead
Do not parse HTML or JSON with regex. Both are nested, recursive structures; a flat pattern cannot handle arbitrary nesting and will break on the first quoted attribute or escaped character. Use a real parser. The cost of pulling in a library once is far lower than the bug that ships when your regex misses an edge case. Likewise, for CSV, use a CSV parser; for dates, use the language's date library; for paths, use the URL/Path APIs.
Takeaway
Regex is a sharp tool: fast for matching and extracting, dangerous for parsing structured formats. Start with the metacharacter table, then character classes, quantifiers, anchors and groups; add lookaround and flags when you need them; and respect the backtracking cliff — never nest quantifiers over the same atoms. When you want to experiment, the regex tester shows matches live, and the URL encoding tool is handy for the examples above.