The everyday email regex
Most applications do not need a perfect email validator. They need a pattern that catches typos at signup and rejects the empty string, a stray space, or a missing @. The pattern below is the one I reach for when a teammate asks for "an email regex" in a code review:
^[^\s@]+@[^\s@]+\.[^\s@]{2,}$
It is short, it survives a paste into almost any engine (PCRE, Python, JavaScript, Go), and it does not pretend to be a standard. Treat it as a gate, not a verdict.
What this pattern actually matches
Read it left to right and every piece earns its place:
^anchors the match to the start of the string so a valid address cannot hide inside a longer sentence.[^\s@]+is the local part. It allows one or more characters that are not whitespace and not the @ sign. That coversuser,name.surname+tag, anddev.team.@is the literal separator. Exactly one is required.[^\s@]+is the domain. Again, anything but whitespace or another @.\.matches a single literal dot. Note the backslash: an unescaped dot would match any character and silently accepta@b_c.[^\s@]{2,}is the top-level domain, forced to two or more non-space, non-@ characters. This is what rejects[email protected]with a single-letter TLD.$anchors the end, so trailing junk fails.
The result accepts [email protected] and [email protected], and rejects @nope.com, no-at-sign.com, and a@b c.
How RFC 5322 blows up the idea of "valid"
If you go looking for the "correct" email regex, you will eventually meet RFC 5322. It permits address forms that almost no signup form should accept, and a fully compliant pattern is hundreds of characters long and still wrong in practice. A few highlights:
- A quoted local part may contain spaces, the @ sign, and most punctuation:
"john doe"@example.comand even"a\b@c"@example.comare legal. - Comments in parentheses can be inserted almost anywhere:
user(comment)@example.com. - The domain may be a literal IP address in brackets:
user@[192.168.1.1]. - Internationalized mailboxes use UTF-8 in the local part, which a plain ASCII pattern cannot describe at all.
Building a regex that accepts every one of these is possible and pointless. You will reject addresses your users actually have, while accepting addresses your mail server will bounce.
Why no regex is ever the right answer
The core argument is simple: a regex can only check shape, never deliverability. An address that matches every rule on earth is still useless if nobody reads the inbox. The only validation that has ever mattered is sending a message with a confirmation link and waiting for the click. That single step proves three things a pattern never can: the address is real, the person controls it, and they want your email.
So the correct architecture is boring. Use a loose pattern to block obvious typos at the form level, store whatever passes, then send a verification mail. If the click never comes, the account stays unverified. No regex, however clever, removes that step.
Common mistakes that break real addresses
These are the errors I see most often in production code:
| Bad habit | Why it hurts | Fix |
|---|---|---|
Forbidding + in the local part | Gmail users lose name+shop tagging, a feature they rely on | Allow + and most ASCII punctuation |
| Rejecting subdomains | [email protected] is valid and common | Use @ then one or more dot-separated labels |
Writing . instead of \. | An unescaped dot matches any character | Escape it: \. |
| Hard-coding a TLD whitelist | New TLDs appear constantly; you will age out | Require length {2,} instead |
Valid and invalid, at a glance
Use this table as a quick test of whatever pattern you are considering. A good practical regex should accept the first group and reject the second.
| Input | Expected | Note |
|---|---|---|
[email protected] | accept | the happy path |
[email protected] | accept | plus tag and subdomain |
@nope.com | reject | empty local part |
spaces [email protected] | reject | space is not allowed |
[email protected] | reject | single-character TLD |
Better than a regex: built-in validators
Every serious platform already ships an email checker, and it is almost always better than what you would write:
- HTML5 gives you
<input type="email">, which applies the browser's own rules and native UI for free. - Python has
email.utils.parseaddrand theemailheader parser. - JavaScript has no built-in validator, but libraries such as those wrapping the RFC are a one-line install.
- Most backend frameworks (Django, Rails, Laravel) validate email on the model layer.
Lean on these, keep your custom regex loose, and reserve the strict patterns for log parsing where you need a fast heuristic.
A stricter variant, and when to use it
If you must reject anything beyond plain ASCII, the variant below forces the local part into a safe character set and demands an alphabetic TLD:
^[A-Za-z0-9._%+\-]+@[A-Za-z0-9.\-]+\.[A-Za-z]{2,}$
Use it for legacy systems that choke on quotes or Unicode. Do not use it as proof that an address works, and do not use it to block international users who have every right to a non-ASCII mailbox.
Testing your patterns
Paste your candidate into a live tester with both the happy-path and the trap cases above. Watch for two failures in particular: a pattern that accepts [email protected] (TLD too short) and a pattern that rejects [email protected] (plus sign forbidden). If it survives those, it is good enough for form-level checking. Remember the rule that ends this page: the regex decides shape, the confirmation mail decides truth.