Base64 and UTF-8: why btoa fails on Chinese and what mojibake means
btoa('中文') throws InvalidCharacterError. The workaround that makes the error go away often returns
garbage on the way back. Both come from one fact: Base64 encodes bytes, and a string in JavaScript is not bytes.
What Base64 does
It reads 3 bytes (24 bits), splits them into four 6-bit groups, and writes each as one character from a
64-character alphabet. If the input does not divide by 3, the last group is padded with =.
| Input | Output |
|---|---|
Man | TWFu |
Ma | TWE= |
M | TQ== |
Hello | SGVsbG8= |
Output length is ceil(n / 3) * 4, so Base64 is always about 4/3 the size of its input. It is an
encoding for carrying bytes through text-only channels, not a compression and not encryption.
Why btoa throws on Chinese
btoa takes a string and treats each character as one byte, so it accepts only characters up to U+00FF. é
(U+00E9) passes and becomes the single byte E9: btoa('é') is 6Q==. 中 is U+4E2D, far above
255, so it throws. Emoji are worse still, because JavaScript stores 🚀 as two UTF-16 code units.
The fix is to turn the string into bytes first, with a defined encoding, and encode the bytes. In UTF-8:
| Text | UTF-8 bytes | Base64 |
|---|---|---|
é | C3 A9 | w6k= |
中 | E4 B8 AD | 5Lit |
中文 | E4 B8 AD E6 96 87 | 5Lit5paH |
🚀 | F0 9F 9A 80 | 8J+agA== |
function toBase64(text) {
const bytes = new TextEncoder().encode(text);
let bin = '';
for (const b of bytes) bin += String.fromCharCode(b);
return btoa(bin);
}
function fromBase64(b64) {
const bin = atob(b64);
return new TextDecoder().decode(Uint8Array.from(bin, (c) => c.charCodeAt(0)));
}
The String.fromCharCode loop is not a mistake: btoa wants a “binary string”, where each character
stands for one byte, and that is the form the loop builds. In Node, Buffer.from(text, 'utf8').toString('base64')
and Buffer.from(b64, 'base64').toString('utf8') do the same in one call.
Why atob returns mojibake
Decode 5Lit5paH with atob alone and you get "䏿\u0096\u0087". atob returns one character per byte, and
those six bytes are being read as Latin-1: E4 is ä, B8 is ¸, AD is an invisible soft hyphen. The
data was never damaged. It was decoded one step short, missing the UTF-8 pass in the fromBase64 function above.
If you see 䏿 or å¼ ä¸ in a decoded string, that is the signature: valid UTF-8 read as single-byte.
When the text is not UTF-8
Mojibake in the other direction is a different problem. If the data was produced by a system that used
GBK, 中 is D6 D0 and encodes to 1tA=. Decoded as UTF-8, that is invalid: D6 starts a two-byte sequence
that D0 cannot complete. A strict decoder reports an error, a lenient one substitutes U+FFFD (�). Python’s
.decode('gbk') or Java’s new String(bytes, "GBK") reads it correctly, and the general rule in Java is to name the
charset every time instead of relying on the platform default. The base64 converter linked below decodes as
UTF-8 only. When the bytes are not valid UTF-8 it says so, shows where the first invalid byte is and prints a hex dump, so you can tell GBK data from a binary file.
Another UTF-8 trap: a file saved by some Windows editors starts with the byte-order mark EF BB BF,
which Base64-encodes to 77u/. If every decoded string starts with an invisible U+FEFF, look for that prefix.
URL-safe, padding and line breaks
Standard Base64 uses + and /. Both are special in URLs, so base64url (RFC 4648) replaces them with - and
_, and JWTs and many web APIs also drop the = padding. Three things then go wrong:
atobrejects-and_. Map them back to+and/before decoding.- Padding. The HTML spec’s
atobtolerates missing padding, but many other decoders do not. Python’sbase64.b64decoderaisesIncorrect padding. Pad the string to a multiple of 4 with=, or usebase64.urlsafe_b64decodeafter padding. +in a query string. A form-style query decoder turns+into a space, so a standard Base64 value passed unescaped in a URL arrives corrupted. Percent-encode the+,/and=, or use base64url.
MIME wraps lines at 76 characters (RFC 2045) and PEM at 64. Decoders typically skip the line breaks, but code
that does a naive length check will not. A pasted data:image/png;base64,... is also not Base64 on its own: the prefix
before the comma must go.
Telling what is wrong
A decoder that only says “invalid input” wastes time when the string is 900 characters long. A length
that leaves exactly one character over a multiple of 4 can never be valid, because one character holds
6 bits and cannot make a whole byte; that is usually a truncated paste. A = in the middle means two
strings were joined. The tool names the rule broken and the position of the first bad character.
When encoding it also decodes its own result and tells you if the round trip returned your text exactly. The one
case where it cannot is a lone surrogate half, which UTF-8 has no way to represent, so
the encoder substitutes U+FFFD. Encoding, decoding, and the file option all run in the tab with
TextEncoder and TextDecoder, with nothing uploaded.