A reference you can keep open while debugging: here is how the same character behaves under different encoders. Memorising the table saves you from a dozen "why is my parameter wrong" sessions, and knowing the few rows that differ between languages saves you from the subtler class of bug where the same code produces a different string on the server than it did in the browser.
The three contexts
Before the table, fix the three situations in your head. They are not the same tool and they do not produce the same string:
encodeURIComponent— encodes a single value in JavaScript. It assumes the input is data, so it escapes every reserved character.encodeURI— encodes an already-complete URL. It spares the structural characters so the URL stays navigable.- Form submission (
application/x-www-form-urlencoded, whatURLSearchParamsand HTML forms emit) — the only one of the three that turns a space into a literal+.
The first two are JavaScript functions; the third is a wire format. Conflating them is the root of most surprises below.
The full reference table
This is the table to bookmark. Read it column by column: the same character can change form depending only on which encoder touched it.
| Character | encodeURIComponent | encodeURI | Form submit | Note |
|---|---|---|---|---|
| space | %20 | %20 | + | form uses +, others use %20 |
| ! | %21 | ! | %21 | encodeURI keeps the bang |
| " | %22 | %22 | %22 | double quote |
| # | %23 | # | %23 | fragment delimiter |
| $ | %24 | $ | %24 | encodeURI keeps it |
| % | %25 | %25 | %25 | the percent sign itself must be encoded |
| & | %26 | %26 | %26 | query separator |
| ' | %27 | ' | %27 | single quote |
| ( ) | %28 %29 | ( ) | %28 %29 | encodeURI keeps parens |
| * | %2A | * | %2A | encodeURI keeps the star |
| + | %2B | + | %2B | a literal plus must be encoded |
| , | %2C | , | %2C | encodeURI keeps the comma |
| / | %2F | / | %2F | path separator |
| : | %3A | : | %3A | scheme separator |
| ; | %3B | ; | %3B | encodeURI keeps the semicolon |
| = | %3D | %3D | %3D | key=value separator |
| ? | %3F | ? | %3F | query start |
| @ | %40 | @ | %40 | user separator |
| [ ] | %5B %5D | %5B %5D | %5B %5D | used for IPv6 literals |
| ~ | ~ | ~ | %7E | unreserved, but forms still encode it |
| 你好 | %E4%BD%A0%E5%A5%BD | %E4%BD%A0%E5%A5%BD | %E4%BD%A0%E5%A5%BD | six UTF-8 bytes |
| 😀 | %F0%9F%98%80 | %F0%9F%98%80 | %F0%9F%98%80 | four-byte emoji |
| newline | %0A | %0A | %0A | control char, always encoded |
| tab | %09 | %09 | %09 | control char |
The rows that surprise people are ?, #, /, @, :, and the punctuation ! $ ( ) * , ;. encodeURI leaves them alone because they are structural; encodeURIComponent does not, because it assumes the input is a value, not a URL. The ~ row is the quiet trap: it is in the unreserved set so the two JS functions pass it through, but form encoding still turns it into %7E.
Why the space row differs
A space is an unsafe character in the URI grammar, so both JS functions emit %20. But the form media type predates the modern URI spec and chose + for spaces to keep query strings readable. The server's job is to decode the form format, where + means space. If you build a query by hand with encodeURIComponent you get %20, and a standards-compliant form decoder still reads that as a space — so those two are usually compatible. The mismatch appears only when one side expects the form + convention and the other does not.
The plus-sign trap: a+b vs a%2Bb
The single most reported encoding bug is the plus sign. In a form-decoded query string, a+b means "a space b". Only a%2Bb means "a plus b". If your data can contain a literal plus, you must encode it, or the server will silently turn it into a space.
// server receives q="a b" (plus read as space)
?q=a+b
// server receives q="a+b" (plus preserved)
?q=a%2Bb
This bites hardest in base64 and URL-safe tokens: a base64 value like ab+cD/ef== run through a form decoder on the server becomes ab cD/ef==, and the token no longer verifies. Always encode the token with encodeURIComponent (or use a URL-safe base64 alphabet with - and _).
Cross-language encoding differences
The same logical operation gives different output depending on the language, and that is exactly where cross-system bugs live. Here are the ones to memorise:
# Python: quote keeps slashes, quote_plus turns them into %2F AND uses +
from urllib.parse import quote, quote_plus
quote('a/b c') # 'a/b%20c'
quote_plus('a/b c') # 'a%2Fb+c'
// Go: QueryEscape is form-style, space becomes +
import "net/url"
url.QueryEscape("a b") // "a+b"
// Java: URLEncoder is form-style, space becomes +, slashes become %2F
URLEncoder.encode("a/b c", "UTF-8") // "a%2Fb%2Bc"
# Ruby: URI.encode_www_form_component is form-style (space -> +)
require 'uri'
URI.encode_www_form_component('a b') # "a+b"
The takeaway: in Python, Go, Java and Ruby the "default" encoder is the form encoder, which produces + and mangles slashes. JavaScript's encodeURIComponent is the outlier that produces %20 and leaves slashes as %2F only when you call it on a whole path. When your frontend encodes with encodeURIComponent and your backend decodes with a form decoder, you have a mismatch waiting to happen.
The double-encoding accident
The classic "it works but the value is wrong" bug. One layer encodes, then another layer encodes the already-encoded string. A space that was correctly %20 becomes %2520. The server decodes once, sees the literal string %20, and treats it as data — so the space you wanted is gone and the user sees "%20" on screen.
// Frontend already encodes the value
const sent = 'q=' + encodeURIComponent('a b'); // "q=a%20b"
// Bug: a proxy or framework encodes AGAIN before sending
const double = encodeURIComponent(sent); // "q%3Da%2520b"
// Server decodes once: q=a%20b -> the value is literally "a%20b", not "a b"
The fix is discipline: encode at exactly one layer. If your framework or HTTP client already encodes parameters, do not pre-encode them in application code. If you must pre-encode, configure the client not to.
The "should have encoded but didn't" accident
The mirror image: a value with a reserved character is dropped raw into a URL because someone assumed "it is just a string". The parser then splits the value at the wrong delimiter.
// Bug: an ampersand in the value is read as a NEW query parameter
const name = 'Tom & Jerry';
const url = 'https://example.com/search?q=' + name;
// Result: https://example.com/search?q=Tom & Jerry
// Server parses TWO params: q="Tom " and Jerry="" (Jerry is now a key!)
// Right: encode the value
const url2 = 'https://example.com/search?q=' + encodeURIComponent(name);
// https://example.com/search?q=Tom%20%26%20Jerry
The same hazard applies to # (everything after it is dropped into the fragment), ? (starts a second query), and / (starts a path segment). Any of these in unencoded user input reshapes the URL.
Encoding non-ASCII and emoji
Percent-encoding operates on bytes, not characters. Non-ASCII text is first converted to UTF-8, then each byte becomes a %XX triple. That is why 你好 is %E4%BD%A0%E5%A5%BD (six bytes, nine visible characters) and the grinning emoji 😀 is %F0%9F%98%80 (four bytes). The same character can encode differently only if the source used a different charset — which is why you should pin UTF-8 everywhere and never let a framework fall back to Latin-1.
The percent sign and malformed sequences
A lone % not followed by two hex digits is invalid. If you concatenate strings that already contain percent signs (logs, hashes, formatted numbers like 50%) without encoding, you produce a sequence that a strict decoder will reject. Always treat % as data and encode it to %25 before it enters a URL.
A real scenario: building a safe search URL
Put it together. A user types a free-text query that may contain spaces, ampersands, Chinese, and emoji. The correct construction encodes only the value pieces:
function buildSearch(term, page) {
const p = new URLSearchParams();
p.set('q', term); // URLSearchParams encodes each value
p.set('page', String(page));
return 'https://example.com/search?' + p.toString();
}
// term = "C++ & 你好" -> "...?q=C%2B%2B+%26+%E4%BD%A0%E5%A5%BD&page=2"
// Note: spaces became + (form style), + and & were encoded, Chinese is UTF-8 bytes.
If the server expects strict %20 rather than +, replace URLSearchParams with manual encodeURIComponent joins.
When to reach for a tool
If you are staring at a long encoded string trying to work out what it really says, paste it into the URL encoding tool and decode it. Encoding by hand for anything beyond a short value is how mistakes slip in, and a decoder is the fastest way to confirm what a string actually contains.
Takeaway
The encoding of a character is not a fact about the character; it is a fact about which encoder touched it. Keep the three contexts separate, bookmark the table, and remember that + means space only in form encoding. Encode every value exactly once, and when a server and a client disagree, the table above is where you find out why.