What UTF-8 is
Every piece of text on a computer is stored as numbers. Unicode is the master list that gives each character a number, called a code point. The letter A is U+0041, é is U+00E9 and the euro sign € is U+20AC. UTF-8 is the rule that turns those code points into bytes.
UTF-8 is the standard encoding of the web, of JSON, of most programming languages and of modern operating systems. Its best feature is that plain ASCII text is unchanged: A is still the single byte 41. Other characters use two, three or four bytes, and each of those bytes is 128 or higher, so they can never be confused with ASCII.
How to use the UTF-8 tool
Encode. Type in the Text box. The UTF-8 bytes appear in the UTF-8 bytes box. Press Copy to take them.
Decode. Paste bytes into the UTF-8 bytes box. The characters appear on the left. Make sure the Byte format matches what you paste.
Byte format offers six ways to write the bytes:
- Hex bytes: C3 A9 for reading and debugging
- Escapes: \xC3\xA9 for string literals in code
- Percent: %C3%A9 as used in web addresses
- Decimal: 195 169 for byte arrays
- Binary: 11000011 10101001 for learning how the bits are laid out
- Garbled text: é (fix mojibake) for repairing text that was read with the wrong encoding
Character by character. The table under the tool lists every character with its code point, its UTF-8 bytes and the byte count. The summary shows the total number of characters and bytes. Spaces, tabs and new lines are named so you can see them.
How UTF-8 works
The number of bytes depends on the size of the code point:
| Code points | Bytes | Bit pattern |
|---|---|---|
| U+0000 to U+007F | 1 | 0xxxxxxx |
| U+0080 to U+07FF | 2 | 110xxxxx 10xxxxxx |
| U+0800 to U+FFFF | 3 | 1110xxxx 10xxxxxx 10xxxxxx |
| U+10000 to U+10FFFF | 4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
The first byte says how many bytes follow. Every following byte starts with 10. This makes UTF-8 self-synchronising: a program can always find the start of the next character, even in the middle of a stream.
Examples
| Character | Code point | UTF-8 hex | Decimal |
|---|---|---|---|
A |
U+0041 | 41 |
65 |
é |
U+00E9 | C3 A9 |
195 169 |
ñ |
U+00F1 | C3 B1 |
195 177 |
ß |
U+00DF | C3 9F |
195 159 |
€ |
U+20AC | E2 82 AC |
226 130 172 |
日 |
U+65E5 | E6 97 A5 |
230 151 165 |
| thumbs-up emoji | U+1F44D | F0 9F 91 8D |
240 159 145 141 |
Fixing mojibake
Mojibake is the jumble you see when UTF-8 bytes are read as Windows-1252 or Latin-1. Each byte turns into its own character, so two-byte letters become two odd symbols:
| Correct text | Seen as |
|---|---|
café |
café |
naïve |
naïve |
It’s |
It’s |
€ |
€ |
The Garbled text format undoes this. It turns each garbled character back into its byte and decodes the bytes as UTF-8. If the text contains characters that could not come from this mistake, the tool tells you instead of guessing.
Tips and limits
- Save files as UTF-8. Most mojibake starts with a file saved in one encoding and opened in another. In editors, choose UTF-8 when saving.
- Declare the encoding. In HTML, put
<meta charset="utf-8">at the top of the page. - Excel and CSV files. Older Excel versions guess the encoding of a CSV. Saving as “CSV UTF-8” avoids broken accents.
- Fixing works once. Text that was garbled twice needs the fix applied twice.
- The table shows up to 500 characters to keep the page fast. The byte count always covers the whole text.
Other ways to do it
Python. 'café'.encode('utf-8') returns b'caf\xc3\xa9'. To fix mojibake, 'café'.encode('cp1252').decode('utf-8') returns café.
JavaScript. new TextEncoder().encode('é') returns the bytes 195 and 169. new TextDecoder().decode(bytes) turns them back.
Command line. echo -n é | xxd shows the bytes c3a9 on a UTF-8 terminal. iconv -f WINDOWS-1252 -t UTF-8 converts whole files.
Related tools
To decode plain hex bytes, use the hex to text converter. Percent-encoded bytes in links are handled by the URL decoder. To see the bits in full, try the binary code translator.