convertCASEpro

MODE

UTF-8 Encoder and Decoder

See exactly which bytes a string occupies.

INPUT
CHARS: 0WORDS: 0LINES: 0
OUTPUT
CHARS: 0WORDS: 0LINES: 0

How the UTF-8 Encoder and Decoder works

UTF-8 stores ASCII characters in one byte and uses two to four bytes for everything else. The tool shows those bytes in hexadecimal, decimal, binary or percent encoded form so that you can inspect them, count them or paste them into code.

To decode, paste bytes in the chosen format. Every value must be between 0 and 255, and the sequence must form valid UTF-8, otherwise a specific error explains the problem.

When it is useful

  • •Finding out why a string is longer in bytes than in characters.
  • •Debugging garbled text caused by an encoding mismatch.
  • •Preparing byte arrays for code or protocol work.
  • •Learning how Unicode characters map to bytes.

Frequently Asked Questions

How many bytes does a character use?

One for ASCII, two for most European accented letters, three for most Asian scripts and the euro sign, and four for emoji and rare characters.

How is this different from the Unicode Text Converter?

This shows the encoded bytes. The Unicode Text Converter shows the code points, which are the abstract numbers assigned to characters.

Why do I see strange characters like é instead of an accent?

That is mojibake, caused by reading UTF-8 bytes as another encoding. Decoding the bytes as UTF-8 here shows what the text was meant to be.

Related tools

[ Browse all tools ]