The everyday URL regex
When someone asks for a "URL regex", first decide what they actually want. Do they want to pull links out of a blob of text, or do they want to prove a user-typed string is a valid URL? Those are different jobs, and the patterns differ. For extraction, this is the line I use:
https?://[^\s]+
It grabs an http:// or https:// scheme and then everything up to the next whitespace. That is deliberately dumb: it is a net, not a judge.
Why a universal URL regex does not exist
The URI grammar in RFC 3986 is deliberately permissive. A URL can carry userinfo, IPv6 literals in brackets, opaque schemes, relative references, and percent-encoded bytes that mean anything. A single regex that accepted exactly the set of strings a browser would accept, and rejected exactly the set it would reject, would be longer than the spec and still wrong at the edges. So stop hunting for "the one URL regex". Pick a regex for a narrow job, or use a parser.
The real answer in JavaScript: new URL()
If you are validating a URL a user typed, do not use a regex in JavaScript. Use the built-in parser:
function isValidUrl(s) {
try { new URL(s); return true; }
catch { return false; }
}
new URL() follows the platform rules, rejects malformed input by throwing, handles percent-encoding, and gives you parsed components for free. A regex cannot do any of that reliably. The only place the regex still wins is scanning free text where a parser would choke on the surrounding sentence.
Breaking the pattern into parts
When you do write a URL pattern, understand each piece. A structured example:
https?://(?:[\w\-]+\.)+[a-z]{2,}(?::\d+)?(?:/[^\s]*)?
https?://is the scheme:httporhttpsplus the double slash.(?:[\w\-]+\.)+[a-z]{2,}is the host: one or more labels separated by dots, ending in an alphabetic TLD of two or more characters.(?::\d+)?is an optional port, the colon plus digits.(?:/[^\s]*)?is an optional path that starts with a slash and runs to the next whitespace.
This split makes it obvious what you are and are not checking: scheme, host shape, port, and path. It says nothing about whether the host resolves.
What this regex matches and misses
| Input | Matched? | Comment |
|---|---|---|
https://example.com | yes | the common case |
http://sub.example.com:8080/p?q=1 | yes | port and path captured |
https://[2001:db8::1] | no | IPv6 literal not handled |
ftp://example.com | no | scheme restricted to http(s) |
example.com | no | no scheme, correctly rejected |
Common mistakes that bite you
| Mistake | Consequence | Fix |
|---|---|---|
Using .* greedily | Swallows trailing punctuation and the next sentence | Stop at whitespace or a safe set |
| Ignoring userinfo | https://user@host slips or breaks inconsistently | Decide explicitly whether to allow it |
| Forgetting IPv6 literals | https://[::1] is rejected though valid | Add a bracketed host branch or use a parser |
Requiring www | Blocks every bare domain URL | Never assume a subdomain |
A stricter variant for logs
For parsing access logs where you trust the source, force a dotted host with an alphabetic TLD and keep the path optional:
https?://(?:[\w\-]+\.)+[a-z]{2,}(?::\d+)?(?:/[\w\-./?%&=]*)?
This still cannot see whether the host resolves or the port is open. It only describes a shape, which is exactly what a regex is for and exactly what it should stop claiming to do.
When to reach for a library instead
If you need more than shape, a parser or library is the honest choice:
- JavaScript and Node: the built-in
new URL()andURLSearchParams. - Python:
urllib.parse.urlparseandurlsplit. - Go: the
net/urlpackage. - Any language: a well-maintained validation library beats a regex you copied from a forum.
Reserve the regex for the one job it owns: pulling candidate URLs out of prose so the parser can judge them afterward.
Testing and trusting your pattern
Run your pattern against the table above and add the nasty cases: a URL at the end of a sentence with a period, a URL with a comma after it, an IPv6 literal, and a URL with userinfo. If a greedy .* is eating the trailing period, switch to a whitespace-bounded class. And for any value that must be valid before you act on it, parse it rather than match it.
Scheme-relative and relative URLs
You will occasionally meet //example.com/path, a scheme-relative URL, or /just/a/path, a relative one. Neither matches an http or https pattern, and that is the right outcome: on their own they are not absolute URLs. A regex that tries to "helpfully" accept them usually ends up accepting noise instead. If your input might be relative, resolve it against a known base using the platform URL parser first, confirm the scheme afterward, and keep the two checks separate. Bolting relative and absolute handling into a single pattern is exactly how brittle validators get shipped into production.