Skip to content
d.devtul.fun
中文
Regex · URL

URL Regex

A practical http and https URL pattern, plus the explanation of why a universal URL regex is a myth.

The pattern

https?://[^\s]+

Matches an http or https scheme followed by any non-whitespace characters. It is meant for extracting URLs from text, not for proving a URL is well formed.

Test this pattern

Edit the text below — matching happens locally, nothing is uploaded.

Variants

Scheme, host, optional port
https?://[\w\-]+(\.[\w\-]+)+(:\d+)?(/[^\s]*)?

Tighter: requires at least one dot in the host and allows an optional port and path. Still allows some invalid hosts.

Full structure
https?://(?:[\w\-]+\.)+[a-z]{2,}(?::\d+)?(?:/[^\s]*)?

Forces an alphabetic TLD and a path. Good for logs, weak for intranets and IP-literal hosts.

The everyday URL regex

When someone asks for a "URL regex", first decide what they actually want. Do they want to pull links out of a blob of text, or do they want to prove a user-typed string is a valid URL? Those are different jobs, and the patterns differ. For extraction, this is the line I use:

https?://[^\s]+

It grabs an http:// or https:// scheme and then everything up to the next whitespace. That is deliberately dumb: it is a net, not a judge.

Why a universal URL regex does not exist

The URI grammar in RFC 3986 is deliberately permissive. A URL can carry userinfo, IPv6 literals in brackets, opaque schemes, relative references, and percent-encoded bytes that mean anything. A single regex that accepted exactly the set of strings a browser would accept, and rejected exactly the set it would reject, would be longer than the spec and still wrong at the edges. So stop hunting for "the one URL regex". Pick a regex for a narrow job, or use a parser.

The real answer in JavaScript: new URL()

If you are validating a URL a user typed, do not use a regex in JavaScript. Use the built-in parser:

function isValidUrl(s) {
  try { new URL(s); return true; }
  catch { return false; }
}

new URL() follows the platform rules, rejects malformed input by throwing, handles percent-encoding, and gives you parsed components for free. A regex cannot do any of that reliably. The only place the regex still wins is scanning free text where a parser would choke on the surrounding sentence.

Breaking the pattern into parts

When you do write a URL pattern, understand each piece. A structured example:

https?://(?:[\w\-]+\.)+[a-z]{2,}(?::\d+)?(?:/[^\s]*)?
  • https?:// is the scheme: http or https plus the double slash.
  • (?:[\w\-]+\.)+[a-z]{2,} is the host: one or more labels separated by dots, ending in an alphabetic TLD of two or more characters.
  • (?::\d+)? is an optional port, the colon plus digits.
  • (?:/[^\s]*)? is an optional path that starts with a slash and runs to the next whitespace.

This split makes it obvious what you are and are not checking: scheme, host shape, port, and path. It says nothing about whether the host resolves.

What this regex matches and misses

InputMatched?Comment
https://example.comyesthe common case
http://sub.example.com:8080/p?q=1yesport and path captured
https://[2001:db8::1]noIPv6 literal not handled
ftp://example.comnoscheme restricted to http(s)
example.comnono scheme, correctly rejected

Common mistakes that bite you

MistakeConsequenceFix
Using .* greedilySwallows trailing punctuation and the next sentenceStop at whitespace or a safe set
Ignoring userinfohttps://user@host slips or breaks inconsistentlyDecide explicitly whether to allow it
Forgetting IPv6 literalshttps://[::1] is rejected though validAdd a bracketed host branch or use a parser
Requiring wwwBlocks every bare domain URLNever assume a subdomain

A stricter variant for logs

For parsing access logs where you trust the source, force a dotted host with an alphabetic TLD and keep the path optional:

https?://(?:[\w\-]+\.)+[a-z]{2,}(?::\d+)?(?:/[\w\-./?%&=]*)?

This still cannot see whether the host resolves or the port is open. It only describes a shape, which is exactly what a regex is for and exactly what it should stop claiming to do.

When to reach for a library instead

If you need more than shape, a parser or library is the honest choice:

  • JavaScript and Node: the built-in new URL() and URLSearchParams.
  • Python: urllib.parse.urlparse and urlsplit.
  • Go: the net/url package.
  • Any language: a well-maintained validation library beats a regex you copied from a forum.

Reserve the regex for the one job it owns: pulling candidate URLs out of prose so the parser can judge them afterward.

Testing and trusting your pattern

Run your pattern against the table above and add the nasty cases: a URL at the end of a sentence with a period, a URL with a comma after it, an IPv6 literal, and a URL with userinfo. If a greedy .* is eating the trailing period, switch to a whitespace-bounded class. And for any value that must be valid before you act on it, parse it rather than match it.

Scheme-relative and relative URLs

You will occasionally meet //example.com/path, a scheme-relative URL, or /just/a/path, a relative one. Neither matches an http or https pattern, and that is the right outcome: on their own they are not absolute URLs. A regex that tries to "helpfully" accept them usually ends up accepting noise instead. If your input might be relative, resolve it against a known base using the platform URL parser first, confirm the scheme afterward, and keep the two checks separate. Bolting relative and absolute handling into a single pattern is exactly how brittle validators get shipped into production.

Frequently asked

What is the best URL regex?

There is none. For validation use your language parser, such as new URL() in JavaScript; use a regex only to extract links from text.

Why should I not validate URLs with regex?

The URI grammar is too loose to capture in one pattern, and a parser handles encoding and components a regex cannot.

How do I match a URL in text?

Use something like https?://[^\s]+ to grab candidates, then parse each one to confirm it is well formed.

Does my URL regex need to handle IPv6?

Only if you actually serve IPv6-literal hosts. Otherwise use a parser, which handles them natively.

Related regex guides

Open the full Regex Tester