Skip to content
Case Converter Online
en

UTF-8 Encoding and Decoding

See exactly which bytes UTF-8 uses for any text, decode bytes back into characters, and repair garbled text such as café. Each character is listed with its code point.

Works both ways. Pick the byte format that matches what you paste on the right.

Character by character

CharacterCode pointUTF-8 bytesBytes

What UTF-8 is

Every piece of text on a computer is stored as numbers. Unicode is the master list that gives each character a number, called a code point. The letter A is U+0041, é is U+00E9 and the euro sign € is U+20AC. UTF-8 is the rule that turns those code points into bytes.

UTF-8 is the standard encoding of the web, of JSON, of most programming languages and of modern operating systems. Its best feature is that plain ASCII text is unchanged: A is still the single byte 41. Other characters use two, three or four bytes, and each of those bytes is 128 or higher, so they can never be confused with ASCII.

How to use the UTF-8 tool

Encode. Type in the Text box. The UTF-8 bytes appear in the UTF-8 bytes box. Press Copy to take them.

Decode. Paste bytes into the UTF-8 bytes box. The characters appear on the left. Make sure the Byte format matches what you paste.

Byte format offers six ways to write the bytes:

  • Hex bytes: C3 A9 for reading and debugging
  • Escapes: \xC3\xA9 for string literals in code
  • Percent: %C3%A9 as used in web addresses
  • Decimal: 195 169 for byte arrays
  • Binary: 11000011 10101001 for learning how the bits are laid out
  • Garbled text: é (fix mojibake) for repairing text that was read with the wrong encoding

Character by character. The table under the tool lists every character with its code point, its UTF-8 bytes and the byte count. The summary shows the total number of characters and bytes. Spaces, tabs and new lines are named so you can see them.

How UTF-8 works

The number of bytes depends on the size of the code point:

Code points Bytes Bit pattern
U+0000 to U+007F 1 0xxxxxxx
U+0080 to U+07FF 2 110xxxxx 10xxxxxx
U+0800 to U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000 to U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The first byte says how many bytes follow. Every following byte starts with 10. This makes UTF-8 self-synchronising: a program can always find the start of the next character, even in the middle of a stream.

Examples

Character Code point UTF-8 hex Decimal
A U+0041 41 65
é U+00E9 C3 A9 195 169
ñ U+00F1 C3 B1 195 177
ß U+00DF C3 9F 195 159
€ U+20AC E2 82 AC 226 130 172
日 U+65E5 E6 97 A5 230 151 165
thumbs-up emoji U+1F44D F0 9F 91 8D 240 159 145 141

Fixing mojibake

Mojibake is the jumble you see when UTF-8 bytes are read as Windows-1252 or Latin-1. Each byte turns into its own character, so two-byte letters become two odd symbols:

Correct text Seen as
café café
naïve naïve
It’s It’s
€ €

The Garbled text format undoes this. It turns each garbled character back into its byte and decodes the bytes as UTF-8. If the text contains characters that could not come from this mistake, the tool tells you instead of guessing.

Tips and limits

  • Save files as UTF-8. Most mojibake starts with a file saved in one encoding and opened in another. In editors, choose UTF-8 when saving.
  • Declare the encoding. In HTML, put <meta charset="utf-8"> at the top of the page.
  • Excel and CSV files. Older Excel versions guess the encoding of a CSV. Saving as “CSV UTF-8” avoids broken accents.
  • Fixing works once. Text that was garbled twice needs the fix applied twice.
  • The table shows up to 500 characters to keep the page fast. The byte count always covers the whole text.

Other ways to do it

Python. 'café'.encode('utf-8') returns b'caf\xc3\xa9'. To fix mojibake, 'café'.encode('cp1252').decode('utf-8') returns café.

JavaScript. new TextEncoder().encode('é') returns the bytes 195 and 169. new TextDecoder().decode(bytes) turns them back.

Command line. echo -n é | xxd shows the bytes c3a9 on a UTF-8 terminal. iconv -f WINDOWS-1252 -t UTF-8 converts whole files.

To decode plain hex bytes, use the hex to text converter. Percent-encoded bytes in links are handled by the URL decoder. To see the bits in full, try the binary code translator.

UTF-8 Encoder: questions and answers

How many bytes does a character take in UTF-8?

Between one and four. English letters, digits and basic punctuation take one byte. Most accented Latin letters, Greek, Cyrillic, Hebrew and Arabic take two. Most Chinese, Japanese and Korean characters and symbols such as the euro sign take three. Emoji usually take four.

How do I fix text like café or It’s?

Set Byte format to Garbled text: é (fix mojibake) and paste the broken text into the right-hand box. The repaired text appears on the left: café becomes café and It’s becomes It’s.

What is the difference between Unicode and UTF-8?

Unicode gives every character a number, called a code point, such as U+20AC for the euro sign. UTF-8 is one way to store those numbers as bytes. The euro sign is stored as E2 82 AC in UTF-8.

Why does my text have more bytes than characters?

Because characters outside basic ASCII need several bytes each. Café €5 plus a thumbs-up emoji is 9 characters but 15 bytes. The summary above the character table shows both counts.

More code and data tools

See all tools