Base64 exists to solve one problem: moving arbitrary bytes through a channel that only reliably carries printable text. Email headers, JSON string values, URL query parameters, XML attributes — none of these can hold a raw 0x00 byte safely. Base64 re-expresses the bytes using 64 characters that every system agrees on.
That is the whole purpose. It is worth being precise about it, because three mistaken beliefs about Base64 cause real problems: that it hides anything, that it is a reasonable way to store files, and that any Base64 string can be fed to any Base64 decoder.
How it works, in one pass
Base64 reads the input three bytes at a time. Three bytes is 24 bits, which divides evenly into four 6-bit groups. Each 6-bit group indexes into a 64-character alphabet: A–Z, a–z, 0–9, +, /.
Take Man:
- Bytes:
77 97 110 - Bits:
01001101 01100001 01101110 - Regrouped into six:
010011 010110 000101 101110 - As numbers:
19 22 5 46 - Alphabet lookup:
T W F u
So Man becomes TWFu. Three bytes in, four characters out — every time. That ratio is where the size increase comes from.
The 33% overhead, and why it is exactly that
Four output characters per three input bytes is a ratio of 4/3, so the encoded form is 33.3% larger than the original, plus up to two padding characters.
This is the single most important practical fact about Base64, because it is the reason not to use it for bulk data. A 3 MB image becomes a 4 MB data URI. Inline it in your CSS and that is 4 MB of your stylesheet that cannot be cached separately, cannot be loaded in parallel, and must be parsed as text before the image can be decoded.
Where inlining does pay off is small assets where the HTTP round trip costs more than the bytes: a 200-byte SVG icon, a 1 KB font subset. Below roughly 1–2 KB, inlining usually wins; above about 10 KB it reliably loses. In between, measure.
Note also that Base64 does not compress. If you gzip a Base64 string you will recover most of the overhead — but you would have done better gzipping the original bytes. Encoding then compressing is strictly worse than compressing then encoding.
Padding, and the length rule
If the input length is not a multiple of three, the last group is short. Base64 pads the output with = so the total length is always a multiple of four:
- 3 bytes → 4 characters, no padding.
Man→TWFu - 2 bytes → 3 characters + one
=.Ma→TWE= - 1 byte → 2 characters + two
=.M→TQ==
Which gives a useful invariant: a valid unpadded Base64 string has a length of 0, 2 or 3 modulo 4 — never 1. A length of exactly 1 more than a multiple of four is proof the string is corrupt, and it is the cheapest validation you can run.
The padding is technically redundant, since the length already tells a decoder how many bytes to produce. Many decoders accept unpadded input for that reason, but not all do, and the specification requires it in contexts where strings get concatenated.
URL-safe Base64, and why JWTs use it
The standard alphabet ends in + and /, both of which mean something in a URL. A + in a query string decodes to a space; a / is a path separator. Put standard Base64 in a URL and it will be mangled somewhere between the browser and your handler.
RFC 4648 defines a URL-safe variant that substitutes:
+becomes-/becomes_- padding
=is usually dropped, since it also needs escaping
This is what JWTs use for all three segments, and what you will see in OAuth state parameters, password reset tokens and S3 pre-signed URL components. Converting between the two is a pair of character substitutions — but you have to actually do it. Feeding a URL-safe string to a standard decoder produces either an error or, worse, silently wrong bytes.
There are other variants in the wild: Base64url without padding, the MIME variant that inserts a line break every 76 characters, and the crypt alphabet with a different ordering entirely. If a decoder rejects a string that looks like valid Base64, the alphabet is the first thing to check.
The UTF-8 trap in the browser
JavaScript's btoa() and atob() are a long-standing source of confusion because they operate on Latin-1, not UTF-8. One character in, one byte out.
So btoa('héllo') throws, because é is U+00E9 and the function wants a byte — and btoa('日本') throws outright. Meanwhile atob() returns a string of char codes 0–255, which you cannot simply treat as text if the original bytes were UTF-8.
The correct round trip encodes to UTF-8 bytes first:
- Encode:
btoa(String.fromCharCode(...new TextEncoder().encode(s))) - Decode:
new TextDecoder().decode(Uint8Array.from(atob(b), c => c.charCodeAt(0)))
The older idiom btoa(unescape(encodeURIComponent(s))) does the same thing and still works, though unescape is deprecated. Either way, the symptom of getting this wrong is distinctive: ASCII works perfectly and accented characters come back as mojibake, which means the bug ships because the test data was all ASCII.
It is not encryption. It is not even obfuscation.
Base64 is a public, reversible, keyless transformation. Decoding it requires no secret and takes one function call. Yet Base64-encoded credentials in config files and API payloads remain genuinely common, presumably because the output looks opaque.
It is not. cGFzc3dvcmQxMjM= is password123 to anyone who has seen a Base64 string before, and the trailing = is a giveaway that invites exactly that check.
Two places this matters in practice:
HTTP Basic auth sends base64(username:password) in a header. That is not protection — it is purely an encoding so the colon-separated pair survives the header. Basic auth is only safe over TLS, and the Base64 contributes nothing to that.
JWT payloads are Base64url, not encrypted. Anyone holding a token can read every claim in it. The signature stops them changing the claims; it does not stop them reading them. Never put anything in a JWT payload that the bearer should not see.
Encoding and decoding without a server
The Base64 encoder and Base64 decoder handle UTF-8 correctly and offer the URL-safe variant, so you can round-trip a JWT segment or an accented string without hitting either trap above. For files, the image to Base64 tool produces a ready-made data URI and reports the size overhead so you can judge whether inlining is worth it.
All three run in your browser. That is the point: pasting a token or a credential into a web form that posts it somewhere defeats the purpose of checking it at all.
Frequently asked questions
Why is Base64 33% larger than the original?
It encodes every three bytes as four characters, a ratio of 4/3. That is a fixed 33.3% increase, plus up to two padding characters. Base64 never compresses — if size matters, compress the bytes before encoding, never after.
Is Base64 encryption?
No. It is a public, keyless, reversible encoding that anyone can decode in one function call. Base64-encoded passwords are plaintext passwords with extra steps, and the trailing = is an invitation to check.
What are the = signs at the end for?
Padding, so the output length is always a multiple of four. One = means the input had two bytes left over, two = means one byte. A valid Base64 length is never 1 more than a multiple of four — that is the cheapest corruption check available.
What is URL-safe Base64?
A variant that replaces + with - and / with _, and usually drops the padding, because those characters have meaning in a URL. JWTs and OAuth tokens use it. A standard decoder will fail or return wrong bytes if given a URL-safe string unconverted.
Why does btoa() fail on accented characters?
btoa expects one byte per character — Latin-1, not UTF-8. Encode to UTF-8 bytes first: btoa(String.fromCharCode(...new TextEncoder().encode(s))). The giveaway for this bug is that ASCII works and everything else comes back as mojibake.
Everything on ToolYard runs in your browser. No uploads, no accounts, no limits.
Browse all tools →