Skip to content
d.devtul.fun
中文
HTTP · 2026-08-26

What Is URL Encoding? A Developer's Guide to Percent-Encoding

Every URL you type, click, or fetch is not really "a string". It is a structured address with rules about which characters are allowed to appear and which ones have to be disguised first. Percent-encoding — the thing written as %20, %E4 and so on — is how the web keeps those rules from breaking. This article explains where the rules come from, what actually happens at the byte level, and the mistakes that quietly corrupt URLs in production.

Where the reserved and unsafe characters come from

A URL has a grammar. In https://user@host:443/path?query#fragment, the colon separates the scheme, the @ separates user from host, the slash divides the path, the question mark introduces the query, and the hash marks the fragment. Each of those punctuation marks is a delimiter with a job. The problem is that the data you want to put in a URL — a search term, a username, a file name — might itself contain those very characters. If a path segment contains a literal /, how does the parser know it is data and not a separator?

That tension is the entire reason encoding exists. The spec (RFC 3986) had to draw a line: some characters are reserved because they might act as delimiters, and everything else falls into unreserved or is simply unsafe to send raw. Reserved characters are the ones the grammar cares about. Unsafe characters are the ones that might get mangled by older transports, proxies, or just careless string handling.

The byte-level mechanism of percent-encoding

Percent-encoding does not operate on characters directly. It operates on bytes. The rule is: take the byte you want to send, write it in hexadecimal, and prefix it with a percent sign. One byte becomes exactly three characters: % plus two hex digits.

This is where non-ASCII text gets interesting. The character does not get encoded as "the character"; it is first turned into bytes by a character encoding — almost always UTF-8 — and then each byte is percent-encoded independently.

const s = '你好';
const bytes = new TextEncoder().encode(s);
// bytes = [0xE4, 0xBD, 0xA0, 0xE5, 0xA5, 0xBD]
console.log(encodeURIComponent(s));
// "%E4%BD%A0%E5%A5%BD"

The word 你好 is two Chinese characters. In UTF-8 each character is three bytes, so you end up with six bytes, and each byte becomes a %XX triplet. That is why a single Chinese character — which is just one "letter" to you — turns into nine characters in a URL. It is not wasteful; it is the only way to carry arbitrary text through a channel that was designed around ASCII.

Reserved versus unreserved characters

RFC 3986 defines the unreserved set precisely: uppercase and lowercase letters, digits, and four symbols — hyphen -, period ., underscore _, and tilde ~. These never need encoding. They cannot be mistaken for delimiters and they survive transport intact.

Everything else is either reserved or simply not allowed raw:

  • Reserved: : / ? # [ ] @ ! $ & ' ( ) * + , ; =. These are allowed to appear in their own component, but if you want to use one of them as literal data instead of as its delimiter meaning, you must encode it.
  • Unsafe / never raw: spaces, control characters, and anything outside the ASCII printable range. These have to be percent-encoded everywhere.

The subtle point: a reserved character is perfectly legal in a URL — as long as it is playing its grammatical role. The / between path segments is fine. But the moment you want a / to be part of a value, it has to become %2F.

Why a space becomes %20 and not +

This is the single most confusing corner of URL encoding, and it comes from history. There are two different conventions:

  • In the general URI grammar, a space is an unsafe character and is encoded as %20.
  • In the older application/x-www-form-urlencoded media type — the format browsers use when submitting an HTML form — a space is encoded as a literal plus sign +, and only a handful of characters use percent-encoding.

So the plus sign is a form-encoding artifact, not a URL-encoding rule. When you call encodeURIComponent('a b') in JavaScript you get a%20b, not a+b. If you then paste that into a form decoder that treats + as a space, fine — but if you build a query string by hand and write q=a+b meaning "a plus b", a form parser will read it as "a b". That mismatch is responsible for a surprising number of bug reports.

// The URL-spec way (what fetch and most APIs expect):
const q = 'hello world';
'https://example.com/search?q=' + encodeURIComponent(q);
// -> https://example.com/search?q=hello%20world

// The form-decoder way (what some servers still assume):
'hello world'.replace(/ /g, '+');
// -> "hello+world"  -- only correct if the server uses form decoding

What must be encoded, and what you can leave alone

Practical guidance, in order of strictness:

  • Always encode anything that is not in the unreserved set before you drop it into a component you control (query values, path segments you build yourself).
  • Never encode the structural delimiters of a URL you did not build yourself. If you already have a valid URL, do not run it through an encoder again.
  • Be careful with % itself. A lone % that is not followed by two hex digits is invalid and will break decoding.

Two mistakes that quietly corrupt URLs

The first is double encoding. You encode a value once to put it in a URL, then a layer further down encodes it again, turning %20 into %2520. The server decodes once, sees %20, and treats the literal percent-twenty as data — so a space becomes the string "%20" instead of a space. This happens constantly with frameworks that encode parameters automatically and application code that also encodes them.

The second is encoding the entire URL as one blob. Do not take a full URL string and pass it through encodeURIComponent. You will encode the ://, the slashes, and the query delimiters, producing a string that is no longer a URL at all. Encode the pieces — the values — and assemble the URL from those.

// Wrong: this destroys the URL structure
fetch(encodeURIComponent('https://example.com/a?b=c'));
// -> fetches "https%3A%2F%2Fexample.com%2Fa%3Fb%3Dc"

// Right: encode only the value, keep the structure
const url = 'https://example.com/search?q=' + encodeURIComponent(userInput);

Takeaway

Percent-encoding is a byte-level escape hatch, not a text transformation. Know which characters are reserved where, remember that + for spaces is a form quirk, and encode values rather than whole URLs. Get those three things right and most "my URL is broken" tickets disappear.

Keep reading