The everyday domain regex
A domain name is a chain of labels joined by dots. The rules for each label are stricter than people expect, and the pattern below encodes them. Use it to sanity-check a hostname a user typed before you do anything risky with it.
^(?![0-9]+$)(?!-)[A-Za-z0-9-]{1,63}(?<!-)(\.(?!-)[A-Za-z0-9-]{1,63}(?<!-))*$
It is longer than the "simple" versions for a reason: the cheap versions let through the exact strings that break resolvers.
The label rules, one by one
DNS constrains every label and the whole name. Here is what the pattern enforces:
- Each label is
[A-Za-z0-9-]{1,63}: letters, digits, and hyphens, one to 63 characters. (?!-)at the start of a label forbids a leading hyphen; a label may not begin with-.(?<!-)at the end forbids a trailing hyphen; a label may not end with-.(?![0-9]+$)on the first label rejects a name made only of digits, which would collide with an IP address.- The whole name must be 253 characters or fewer, a check you still do in code, not in the regex alone.
Together these reject -example.com, bad..name.com, and a 64-character label, while accepting sub.example.co.uk.
Internationalized domain names need punycode
A user may type 例子.测试 or münchen.de. Those are valid domains, but only after they are converted to punycode: xn--fsqu00a.xn--g6w251d and xn--mnchen-3ya.de. A plain ASCII pattern rejects them as written. The correct pipeline is: detect non-ASCII, convert with your platform's punycode encoder (for example idna.encode in Python), then validate the ASCII form. Do not try to stuff UTF-8 into the regex; you will miss edge cases the encoder already solved.
What a regex can never prove
The core argument: a regex confirms the string looks like a domain. It does not confirm the domain is registered, resolvable, or yours to use. this-domain-does-not-exist-anywhere.example matches the pattern perfectly and resolves to nothing. So a passing regex is necessary but useless as proof of existence. Pair it with a DNS lookup (getaddrinfo, an SOA query, or a registrar check) when existence actually matters.
Valid and invalid, at a glance
| Input | Expected | Why |
|---|---|---|
example.com | accept | classic two-label name |
sub.example.co.uk | accept | multiple labels are fine |
-example.com | reject | label starts with hyphen |
bad..name.com | reject | empty label between dots |
a-label-63-chars-long-............................x | reject if over 63 | label length cap |
Common mistakes that break real domains
| Mistake | Consequence | Fix |
|---|---|---|
| Allowing underscores | SRV records use them, but hostnames should not | Keep _ out unless you mean SRV |
| Allowing consecutive dots | a..b is not a valid name | Require a label between every dot |
| Ignoring the trailing dot | example.com. means the root, and is valid | Optionally allow one final dot |
| Hard-coding a TLD whitelist | New TLDs break your list constantly | Validate length and shape, not membership |
A stricter variant with a forced TLD
If you only care about classic public domains and want to reject IP-literal or IDN forms up front, force an alphabetic final label:
^(?![0-9]+$)(?!-)[A-Za-z0-9-]{1,63}(?<!-)(\.(?!-)[A-Za-z0-9-]{1,63}(?<!-))*\.[A-Za-z]{2,}$
This is stricter and rejects more legal input, so only use it when your downstream system truly cannot handle anything else.
Beyond the regex: verify existence
When the domain must actually work, step past the pattern:
- Convert IDN to punycode before checking.
- Run a DNS resolution or SOA lookup to confirm registration.
- For signup, send a verification link to a mailbox at that domain.
- Enforce the 253-character total in code, since regex length math gets unwieldy.
The regex is the doorman; DNS is the bouncer. Both have a job, and only one of them knows who is really inside.
Testing your domain pattern
Throw the table cases at your pattern and add a 253-character name and a punycode name. If your pattern accepts a..b or a label that starts with a hyphen, tighten the anchors. If it rejects example.com. and you need root-domain support, allow the final dot. And remember: matching is not owning.
Subdomains, depth, and the 253 ceiling
Real names can be long and deeply nested. A pattern that only checks each label misses the whole-name limit of 253 characters and the practical depth most resolvers handle. a.b.c.d.e.f.g.h.example.com is legal in shape but may exceed limits or hit a resolver cap in practice. Enforce the 253-character total in code after the regex passes, and decide consciously whether you accept arbitrary depth. For signup forms, capping subdomain levels is reasonable; for a DNS tool, accept what the grammar allows and let the lookup be the final word on whether the name lives.