Skip to content
d.devtul.fun
中文
Encoding · 2026-09-12

What Is Base64 Encoding? A Developer's Guide

Base64 shows up in more places than almost any other encoding in daily development — email attachments, JSON APIs, JWTs, data URLs, Kubernetes secrets. Most developers use it for years without sitting down to understand what it actually does. This article fixes that. By the end you should know what the 64 characters are, why the output is always a third larger, where the three common variants differ, and exactly why btoa('中文') blows up in your browser.

What Base64 actually is

Base64 is a way of representing arbitrary binary data as text. That is the whole story. It takes a sequence of bytes and turns it into a string built only from 64 specific characters, plus one padding character. There is no key, no secret, no compression — given the string, anyone can reverse it instantly.

This matters because a lot of systems were designed to carry text, not bytes. Early email (SMTP) was specified to move 7-bit ASCII. JSON and XML are text formats. URLs are text. If you try to stuff raw image bytes or a ZIP file into any of those, you hit characters that mean something else, get dropped, or corrupt the channel. Base64 gives those bytes a safe, text-only disguise so they survive the trip.

The 64-character alphabet

The standard Base64 alphabet is:

A B C D ... Z   (26 uppercase letters)
a b c d ... z   (26 lowercase letters)
0 1 2 ... 9     (10 digits)
+ /             (2 symbols)
=               (padding, not part of the 64)

That is 26 + 26 + 10 + 2 = 64 symbols, which is exactly 2 to the 6th power. Each symbol therefore encodes 6 bits of information. The padding = is added only at the end when the input length is not a multiple of 3, and it carries no data of its own.

Why the output is 4/3 larger

Because each output character holds 6 bits, but a byte holds 8, the overhead is forced by arithmetic:

  • 3 bytes = 24 bits.
  • 24 bits ÷ 6 bits per character = 4 characters.
  • So 3 bytes become 4 characters — a 4/3 expansion, or roughly 33% overhead.

In practice a 3 MB file becomes about 4 MB once encoded, and a small JSON field grows the same way. This is pure information theory, not a quirk of one library. If you need to shrink data, reach for compression (gzip, zstd) — Base64 will only make it bigger.

Three common variants

VariantChangeWhere you see it
StandardUses + and /, keeps = paddingGeneric use, PEM files, crypto
URL-safe+ → -, / → _, padding often droppedJWTs, query strings, filenames
MIMEInserts a line break every 76 charactersEmail bodies (RFC 2045)

The URL-safe variant exists for a concrete reason: + and / are reserved in URLs. A + in a query string is frequently read as a space, and / is the path separator. Swapping them for - and _ — symbols on no URL special-meaning list — lets the text survive a round trip through a URL untouched. Note the padding also changes: many URL-safe systems drop the = entirely and infer it on decode.

The trap: multi-byte characters

This is the single most common Base64 bug in browser JavaScript. New developers write:

btoa('中文')   // throws InvalidCharacterError
btoa('😀')     // also throws

The reason is that btoa treats its input as a Latin-1 string — every character must fit in a single byte (0–255). JavaScript strings are UTF-16, so anything outside Latin-1, such as a Chinese character or an emoji, has nowhere to go and the call throws.

The correct path is to convert the text to raw bytes first with TextEncoder, then feed those bytes to btoa one at a time:

function toBase64(str) {
  const bytes = new TextEncoder().encode(str);
  let binary = '';
  for (const b of bytes) binary += String.fromCharCode(b);
  return btoa(binary);
}

function fromBase64(b64) {
  const binary = atob(b64);
  const bytes = Uint8Array.from(binary, c => c.charCodeAt(0));
  return new TextDecoder().decode(bytes);
}

Decoding reverses the same road: run atob, build a Uint8Array from the result, and hand it to TextDecoder. Skip the TextDecoder step and any non-ASCII content comes back as mojibake. In Node.js the same problem is already solved for you: Buffer.from(str, 'utf8').toString('base64') handles the encoding properly, which is why server-side code rarely hits this wall.

Where Base64 ends: its boundaries

Base64 is simple, but it has edges that bite if you forget them.

  • Truncation. Chop a Base64 string in the middle and decoding fails or silently loses the tail. You cannot "preview the first half" of a Base64 blob.
  • Line breaks. The MIME variant adds newlines; if you decode a MIME string with a strict decoder that rejects whitespace, it errors. Strip whitespace before decoding when the source is email or PEM.
  • Dropped padding. Some encoders omit =. Most decoders tolerate this, but strict ones do not. When interoperating with an unknown system, normalize padding first.
  • Case sensitivity. Base64 is case-sensitive; AB and ab decode to different bytes.

It is encoding, not encryption

Repeat after me: Base64 hides nothing. It is a reversible transform with no secret component. Anyone who can read the string can decode it. Using Base64 to "protect" a password, an API key, or a session token is equivalent to writing it in slightly larger, more readable letters.

For real confidentiality use TLS in transit, AES-GCM for symmetric encryption, and bcrypt or argon2 (which are hashes, strictly speaking) for stored passwords. Base64's job is purely mechanical: move bytes through a text-only world.

Want to experiment? Paste some accented text or an emoji into the Base64 tool and flip the URL-safe option to see how the alphabet changes.

Keep reading