Decodex

DECODEX · BASE64 · UTF-8

JavaScript Base64 with Unicode: fix btoa() errors using UTF-8

Encode and decode Korean, Hindi and emoji in JavaScript with TextEncoder, btoa, atob and TextDecoder. Includes tested Unicode-safe code and browser examples.

btoa() expects a binary string whose code units each fit in one byte. Passing Korean, Hindi or many emoji directly can throw InvalidCharacterError. For Unicode text, first encode it as UTF-8 bytes, then convert those bytes to Base64. To recover the text, decode Base64 to bytes and read those bytes as UTF-8.

Why btoa("안녕하세요") fails

JavaScript strings are not raw UTF-8 bytes. btoa() does not automatically encode a string as UTF-8: characters outside the byte range fail, and some characters inside that range may produce bytes that are not the UTF-8 representation you intended. Use TextEncoder for multilingual text, including accents, rather than relying on whether a particular call happens to avoid an exception.

A Unicode-safe encoder and decoder

The code below converts text to bytes explicitly. The encoder processes bytes in chunks so a large spread operation does not exceed the JavaScript argument limit. The decoder accepts Base64URL characters and whitespace, validates the format, restores missing padding, and uses a fatal UTF-8 decoder so invalid bytes produce an error rather than replacement characters. This example is for text, not a file preview.

Check the round trip with Korean and emoji

For the exact text 안녕하세요 😀, the UTF-8 Base64 result is 7JWI64WV7ZWY7IS47JqUIPCfmIA=. Decoding it must restore the same text, including the space and emoji. The sample below is checked against the browser conversion code. Use the example link to load it into Decodex without putting user-entered text into the URL.

Two different errors need different fixes

An atob() error usually points to malformed Base64, incorrect padding or an unexpected prefix. A TextDecoder error with fatal mode means the bytes are not valid UTF-8, even if the Base64 alphabet is valid. Images, PDFs and legacy encodings can produce that second error. Knowing which step failed is more useful than repeatedly changing padding.

Browser code and Node.js are different environments

In Node.js, Buffer.from(text, "utf8").toString("base64") is a standard way to encode UTF-8 text. Buffer decoding can be permissive, so do not treat successful conversion alone as validation of untrusted input. In a browser, the TextEncoder/TextDecoder approach below avoids depending on Node’s Buffer API. Decodex supports UTF-8 text up to 5 MB and performs conversion locally.

Try this example

// Browser JavaScript: text → UTF-8 bytes → Base64.
function utf8ToBase64(text) {
  const bytes = new TextEncoder().encode(text);
  let binary = '';
  for (let i = 0; i < bytes.length; i += 8192) {
    binary += String.fromCharCode(...bytes.subarray(i, i + 8192));
  }
  return btoa(binary);
}

// Accept standard Base64 or Base64URL, with complete or omitted padding.
function base64ToUtf8(value) {
  let normalized = value.replace(/\s/g, '').replace(/-/g, '+').replace(/_/g, '/');
  if (!/^[A-Za-z0-9+/]*={0,2}$/.test(normalized) ||
      normalized.length % 4 === 1 ||
      (normalized.includes('=') && normalized.length % 4 !== 0)) {
    throw new Error('Invalid Base64 characters, length or padding');
  }
  normalized += '='.repeat((4 - normalized.length % 4) % 4);
  const binary = atob(normalized);
  const bytes = Uint8Array.from(binary, character => character.charCodeAt(0));
  return new TextDecoder('utf-8', { fatal: true }).decode(bytes);
}

const text = '안녕하세요 😀';
const encoded = utf8ToBase64(text);
console.log(encoded); // 7JWI64WV7ZWY7IS47JqUIPCfmIA=
console.log(base64ToUtf8(encoded)); // 안녕하세요 😀

Try this example →

Base64 → UTF-8 · UTF-8 → Base64

Common questions

Will atob() return readable Korean text?

Not directly. It returns a binary string. Convert its byte values to Uint8Array and decode with TextDecoder("utf-8").

Should I use btoa(unescape(encodeURIComponent(text)))?

Prefer explicit TextEncoder and TextDecoder code. It makes the byte encoding clear and avoids depending on the legacy unescape() function.

References

More Base64 guides