You will hear "URL" in everyday conversation and "URI" in specifications, and it is natural to assume they are two words for the same thing. They are not exactly, but the difference matters far less often than pedants imply. Here is the honest version, with the details that actually change how you write code.
URI is the superset
A URI (Uniform Resource Identifier) is the broad category: a string that identifies a resource. It has two familiar sub-types:
- URL (Uniform Resource Locator) — a URI that also tells you how to reach the resource (it has a scheme like
httpsorftp). This is the "web address". - URN (Uniform Resource Name) — a URI that names a resource by a persistent identifier, with no location implied, e.g.
urn:isbn:0451450523.
So every URL is a URI, but not every URI is a URL. A URN is a URI that is neither a URL nor a locator. The Venn diagram is: URI is the outer circle, URL and URN are two non-overlapping circles inside it.
URN in practice
A URN looks like urn:<namespace>:<value>. You meet them as ISBNs for books, as urn:uuid:... identifiers, and inside some XML and SOAP tooling. The key point is that a URN is a name, not an address: knowing the URN tells you what the thing is, not where to fetch it. If your system needs "a stable handle that survives the resource moving", a URN-style identifier is the right tool; if it needs "a clickable link", you want a URL.
The RFC 3986 generic syntax
A URI is parsed into five parts:
scheme://authority/path?query#fragment
- scheme —
http,https,mailto… ends at the first:. - authority — the
//part:user@host:port. - path — the hierarchical part after the authority.
- query — everything after
?, the familiar key=value pairs. - fragment — after
#, handled only by the client, never sent to the server.
Knowing these parts matters because encoding rules differ per part: a / in the path is a separator, but a / inside a query value must be %2F. It also matters because some parts are sent to the server and some are not.
What the server never sees: the fragment
The fragment (everything after #) is processed only by the browser. It is used for in-page anchors, and for single-page apps that route client-side. If you put sensitive data after a # thinking "it is in the URL so it is secure", remember that the fragment is never transmitted in the HTTP request — but it does show up in document.location and in any analytics or logging that copies the full URL, and it can be sent to other origins via the Referer header in some configurations. Treat it as public, not secret.
Parsing a URI in code
Do not split a URL with string.split('/') and friends. Every language ships a real parser:
// JavaScript: the URL object
const u = new URL('https://[email protected]:8080/a/b?x=1#frag');
u.protocol; // "https:"
u.host; // "ex.com:8080"
u.hostname; // "ex.com"
u.port; // "8080"
u.pathname; // "/a/b"
u.search; // "?x=1"
u.hash; // "#frag"
# Python
from urllib.parse import urlsplit
p = urlsplit('https://[email protected]:8080/a/b?x=1#frag')
p.scheme, p.netloc, p.path, p.query, p.fragment
The parser correctly handles edge cases you will get wrong by hand: userinfo, IPv6 literals in brackets, and percent-decoding only where it belongs.
Relative references and base URI resolution
A link like ../img/logo.png is a relative reference. To turn it into an absolute URI you resolve it against a base URI — usually the URL of the document containing it. The rules are mechanical: strip the current path's last segment, apply .., and append. Mis-resolved bases are a common cause of broken assets when a site moves domains.
new URL('../img/logo.png', 'https://ex.com/blog/post/').href;
// "https://ex.com/blog/../img/logo.png" resolved -> "https://ex.com/img/logo.png"
data: and other non-locator URIs
Not every URI is a network address. A data: URI embeds the resource inline, e.g. data:text/plain;base64,SGVsbG8=. It is a valid URI and a valid URL (it has a scheme and tells you how to obtain the resource — by decoding the payload). It is also a favourite XSS vector when built from untrusted input, so never put user text into a data: URI without strict allow-listing.
Why we say URL but the standard says URI
The WHATWG and RFC documents say "URL" today, but older RFCs (and many textbooks) say "URI" because the term was meant to cover locators, names, and everything in between. In practice, when a developer says "URL" they mean the web address in the bar. The terminology drift is harmless except in specs, where precision is the point.
When the distinction actually matters
- Signing algorithms. If you sign a canonical request string, you must decide whether to sign the URI or the URL form. Signing the wrong normalisation lets an attacker tweak the path and keep a valid signature.
- Redirect validation. After a login you often redirect to a
Locationsupplied by the user. If you only check the scheme and forget that//evil.comis a protocol-relative URL, you open an open-redirect. - SSRF protection. Validating a user-supplied URI means parsing it fully — scheme, host, port — not string-matching. A raw substring check misses
https://[email protected]/style tricks. - Logging and storage. If you store a "URI" but later treat it as a "URL" you can fetch, you may try to fetch a URN and fail, or fetch a
data:URI you did not expect.
A redirect bug, reproduced
The protocol-relative trap is worth seeing concretely. A naive validator that only blocks http:// and https:// prefixes still lets //evil.com through, because that is a valid relative URL resolved against the current scheme:
// naive check
function isSafe(target) {
return !/^https?:\/\//i.test(target); // blocks http(s):// but...
}
isSafe('//evil.com/x'); // returns true -- WRONG, it is still an absolute-ish URL
// better: parse and require an http/https scheme explicitly
function isSafe2(target) {
try { return new URL(target).protocol === 'http:' || new URL(target).protocol === 'https:'; }
catch { return false; }
}
When it is just pedantry
If you are writing a blog post link or explaining to a colleague where to click, "URL" is correct and "URI" would sound affected. Reach for "URI" only when the distinction — locator versus name versus identifier — is load-bearing.
Scheme normalisation and case
The scheme is case-insensitive by spec, so HTTPS:// and https:// are the same URL. Most parsers lowercase it for you, and any comparison or signature you do should normalise the scheme first. The same goes for the host: EX.COM and ex.com are the same authority. If your security check compares hosts as raw strings without lowercasing, an attacker can slip Ex.com past it. Normalise scheme and host before any comparison.
Internationalised domain names and punycode
When a host contains non-ASCII characters, browsers convert it to punycode (the xn-- form) before sending the request. So https://例子.com actually travels as https://xn--fsqu00a.com. The Unicode form you see in the bar is presentation only. If you parse a URL and inspect hostname, you get the punycode form, not the pretty one — keep that in mind when logging or matching hosts, or you will be confused why your "example.com" check never matches.
Why this ties back to encoding
Every component of the five-part syntax has its own encoding rules, which is exactly why the URL/URI split matters: the path, query and fragment are encoded differently, and a parser that understands the parts will percent-encode each one correctly. Hand-built URLs that ignore the parts are the source of every bug in the encoding articles — so treat the syntax as a contract, not decoration.
Takeaway
URI is the family, URL is the member you use daily, URN is the naming cousin. Learn the five-part syntax because encoding and parsing depend on it, care about the URL/URI wording only where a security or signing decision hangs on it, and always parse with a real URL/URI parser instead of splitting strings by hand.