Base64 Pitfalls: btoa Errors, Missing Padding, URL-Safe Variants and Size Surprises
Solutions for the Base64 problems developers hit most: InvalidCharacterError with non-ASCII text, decoding failures, URL-safe alphabets, line breaks and the 33 percent size increase.
Published September 16, 2026 · By Sudip Bhowmick
Base64 looks trivial: call an encode function, get a string. Then a user with an accented name breaks it, a token fails to decode because it lost its equals signs, or a file doubles in size inside a JSON payload. All of these have simple explanations. This guide walks through the real behavior of Base64 and the fixes for its most common failures.
What Base64 Is and What It Is Not
Base64 turns arbitrary bytes into 64 safe, printable characters so that binary data can pass through systems designed for text, such as JSON, email and URLs. Every three bytes become four characters. It is an encoding, not encryption. There is no key, and anyone can decode it, so never use it to hide a password or a secret.
Problem: btoa Throws an Error on Non-ASCII Text
In browsers, btoa accepts only characters whose code is below 256. Passing a string with é might work by accident, and passing a euro sign, an emoji or Chinese text throws an InvalidCharacterError. The function treats a string as a sequence of bytes, so you must turn the text into UTF-8 bytes yourself first.
- ▸Encode: convert the string to bytes with TextEncoder, turn each byte into a character, and give the result to btoa.
- ▸Decode: call atob, turn each character into its byte value, and decode the bytes with TextDecoder using UTF-8.
- ▸Without that step, decoded text shows as garbled characters like é, which is the same family of bug covered in our guide to broken encodings.
- ▸Newer runtimes also offer direct conversion helpers for byte arrays, but the TextEncoder approach works everywhere.
Problem: Decoding Fails or Returns Nonsense
- ▸Padding removed. Base64 output is padded with equals signs to a multiple of four characters. Some systems strip the padding, notably JSON Web Tokens. Add equals signs until the length is a multiple of four before decoding.
- ▸URL-safe alphabet. The standard alphabet uses plus and slash, which are special in URLs. The URL-safe variant uses hyphen and underscore instead. Feeding one alphabet to a strict decoder for the other fails. Convert the characters, or use a decoder that accepts both.
- ▸Whitespace and line breaks. Email formats wrap Base64 at 76 characters. Strip all whitespace before decoding.
- ▸Data URL prefix. A string like data:image/png;base64, followed by the data must have the prefix removed before decoding.
- ▸The result is binary. If the original was an image or a compressed file, decoding to text gives garbage by design. Treat the bytes as bytes and write them to a file or a blob.
Problem: The Data Became One Third Larger
Four output characters for every three input bytes means a 33 percent increase, plus padding and any line breaks. A 3 MB image becomes about 4 MB of text, and if that text is placed in JSON that is then compressed with gzip, the gain from compression is smaller than for the raw file. For large files, send them as binary with multipart upload or a direct storage upload, and keep Base64 for small payloads such as icons and tokens.
Where Base64 Appears and What to Watch
- ▸HTTP Basic authentication: the header carries username and password joined by a colon and encoded. It is not protection, so Basic authentication must always travel over HTTPS.
- ▸JSON Web Tokens: Base64URL without padding for the header, payload and signature.
- ▸Data URLs: small images inlined in HTML or CSS. Fine for tiny assets, but they cannot be cached separately and inflate the page.
- ▸Binary fields in JSON APIs: acceptable for small blobs, with the cost described above.
- ▸Email attachments: the standard way to carry binary in a text protocol.
Quick Checks When Something Fails
Look at the length: is it a multiple of four? Look at the characters: are there plus or slash characters, or hyphen and underscore? Is there a prefix or whitespace? Decode a small sample in the Base64 Encoder and Decoder, which accepts both alphabets, adds missing padding and decodes as UTF-8. If the result is unreadable, the data is probably binary and the next step is a file, not a string.
Conclusion
Base64 problems come down to text encoding, padding, alphabet variants, stray whitespace and the fact that it enlarges data. Always convert text to UTF-8 bytes before encoding, normalize the alphabet and padding before decoding, never rely on it for secrecy and move large files out of Base64 text altogether.
Free Tool
Open the Base64 Encoder and Decoder