Base64 and UTF-8: why btoa fails on Chinese and what mojibake means

Updated 4 min read

btoa('中文') throws InvalidCharacterError. The workaround that makes the error go away often returns garbage on the way back. Both come from one fact: Base64 encodes bytes, and a string in JavaScript is not bytes.

What Base64 does

It reads 3 bytes (24 bits), splits them into four 6-bit groups, and writes each as one character from a 64-character alphabet. If the input does not divide by 3, the last group is padded with =.

InputOutput
ManTWFu
MaTWE=
MTQ==
HelloSGVsbG8=

Output length is ceil(n / 3) * 4, so Base64 is always about 4/3 the size of its input. It is an encoding for carrying bytes through text-only channels, not a compression and not encryption.

Why btoa throws on Chinese

btoa takes a string and treats each character as one byte, so it accepts only characters up to U+00FF. é (U+00E9) passes and becomes the single byte E9: btoa('é') is 6Q==. 中 is U+4E2D, far above 255, so it throws. Emoji are worse still, because JavaScript stores 🚀 as two UTF-16 code units.

The fix is to turn the string into bytes first, with a defined encoding, and encode the bytes. In UTF-8:

TextUTF-8 bytesBase64
éC3 A9w6k=
中E4 B8 AD5Lit
中文E4 B8 AD E6 96 875Lit5paH
🚀F0 9F 9A 808J+agA==
function toBase64(text) {
  const bytes = new TextEncoder().encode(text);
  let bin = '';
  for (const b of bytes) bin += String.fromCharCode(b);
  return btoa(bin);
}

function fromBase64(b64) {
  const bin = atob(b64);
  return new TextDecoder().decode(Uint8Array.from(bin, (c) => c.charCodeAt(0)));
}

The String.fromCharCode loop is not a mistake: btoa wants a “binary string”, where each character stands for one byte, and that is the form the loop builds. In Node, Buffer.from(text, 'utf8').toString('base64') and Buffer.from(b64, 'base64').toString('utf8') do the same in one call.

Why atob returns mojibake

Decode 5Lit5paH with atob alone and you get "中æ\u0096\u0087". atob returns one character per byte, and those six bytes are being read as Latin-1: E4 is ä, B8 is ¸, AD is an invisible soft hyphen. The data was never damaged. It was decoded one step short, missing the UTF-8 pass in the fromBase64 function above. If you see 中æ or å¼ ä¸ in a decoded string, that is the signature: valid UTF-8 read as single-byte.

When the text is not UTF-8

Mojibake in the other direction is a different problem. If the data was produced by a system that used GBK, 中 is D6 D0 and encodes to 1tA=. Decoded as UTF-8, that is invalid: D6 starts a two-byte sequence that D0 cannot complete. A strict decoder reports an error, a lenient one substitutes U+FFFD (�). Python’s .decode('gbk') or Java’s new String(bytes, "GBK") reads it correctly, and the general rule in Java is to name the charset every time instead of relying on the platform default. The base64 converter linked below decodes as UTF-8 only. When the bytes are not valid UTF-8 it says so, shows where the first invalid byte is and prints a hex dump, so you can tell GBK data from a binary file.

Another UTF-8 trap: a file saved by some Windows editors starts with the byte-order mark EF BB BF, which Base64-encodes to 77u/. If every decoded string starts with an invisible U+FEFF, look for that prefix.

URL-safe, padding and line breaks

Standard Base64 uses + and /. Both are special in URLs, so base64url (RFC 4648) replaces them with - and _, and JWTs and many web APIs also drop the = padding. Three things then go wrong:

  • atob rejects - and _. Map them back to + and / before decoding.
  • Padding. The HTML spec’s atob tolerates missing padding, but many other decoders do not. Python’s base64.b64decode raises Incorrect padding. Pad the string to a multiple of 4 with =, or use base64.urlsafe_b64decode after padding.
  • + in a query string. A form-style query decoder turns + into a space, so a standard Base64 value passed unescaped in a URL arrives corrupted. Percent-encode the +, / and =, or use base64url.

MIME wraps lines at 76 characters (RFC 2045) and PEM at 64. Decoders typically skip the line breaks, but code that does a naive length check will not. A pasted data:image/png;base64,... is also not Base64 on its own: the prefix before the comma must go.

Telling what is wrong

A decoder that only says “invalid input” wastes time when the string is 900 characters long. A length that leaves exactly one character over a multiple of 4 can never be valid, because one character holds 6 bits and cannot make a whole byte; that is usually a truncated paste. A = in the middle means two strings were joined. The tool names the rule broken and the position of the first bad character.

When encoding it also decodes its own result and tells you if the round trip returned your text exactly. The one case where it cannot is a lone surrogate half, which UTF-8 has no way to represent, so the encoder substitutes U+FFFD. Encoding, decoding, and the file option all run in the tab with TextEncoder and TextDecoder, with nothing uploaded.

Open the tool: Base64 Encoder & Decoder

Back to guides

More guides