UTF-8 Decoder and Encoder
Decode UTF-8 bytes to text, encode text to UTF-8, and repair garbled characters like é or ’ back to what they should have been. Includes Unicode NFC/NFD normalization and precise reporting of invalid byte sequences.
Whether you call it utf-8 decoder or fix mojibake, this free tool handles it instantly in your browser.
Options
Example usage
Example input
43 61 66 C3 A9
Example output
Decoded text: Café Read as hex. Bytes in: 5 Characters out: 4
Batch Processing
Enter multiple inputs, one per line. Each line will be processed separately.
More like this
Discover related tools and step-by-step guides for this kind of task.
How it works
- Decode mode accepts hex ("43 61 66 C3 A9" or packed "436166C3A9") or decimal bytes, and walks the sequence the way a real UTF-8 decoder does — reporting the exact byte index and reason when a sequence is invalid, rather than silently producing question marks.
- Mojibake repair reverses a specific accident: UTF-8 bytes that were read as Latin-1 or Windows-1252. Each character is mapped back to the byte it was mis-read from, and that byte string is decoded as UTF-8 again. Text that went through the mistake twice is repaired in two passes.
- Normalization converts between composed and decomposed forms. NFC stores é as one code point, NFD as e plus a combining accent — identical on screen, different in a database comparison.
- Because the base is ambiguous, the decoder chooses hex or decimal once for the whole input rather than per value, and tells you which it used.
Use cases
Why é appears instead of é
This is the most common encoding failure there is, and it has one cause. In UTF-8 the character é is stored as two bytes, C3 and A9. If a program reads those bytes one at a time as Latin-1 or Windows-1252 — one byte per character — it displays them as the two separate characters à and ©. Nothing has been lost. The bytes are still exactly what they were; only the interpretation is wrong, which is why the repair is deterministic rather than guesswork.
The same mechanism explains the other familiar artefacts. A right single quotation mark becomes ’ because its three UTF-8 bytes map to those three Windows-1252 characters. An em dash becomes —. A non-breaking space becomes  followed by a space, which is why mangled text often has stray  characters sprinkled through it. Recognising the pattern is enough to know it is recoverable.
Repair works by reversing the steps: map each character back to the single byte it was mis-read from, then decode that byte string as UTF-8 properly. Text that passed through the mistake twice needs two passes, which happens more often than you would expect when data moves between several systems. The tool reports how many passes it applied.
When mojibake cannot be repaired
The repair depends on the original bytes surviving. They usually do, because the mis-decoding is a display and storage problem rather than a lossy conversion. But there are cases where the information is genuinely gone, and it is worth recognising them rather than hunting for a better tool.
If the text passed through a step that replaced unmappable characters with question marks or with the replacement character U+FFFD, the original bytes were discarded at that point. A database column with a restrictive collation, an export that forced ASCII, or a transcoding step configured to substitute rather than fail will all do this. Once a character has become "?" there is no way to know what it was.
The practical implication is to fix the pipeline before re-exporting. If you can regenerate the file from its source with the encoding declared correctly, do that instead of repairing the output — and if you cannot, repair it once and store the result, because running a repair pass over already-correct text can corrupt legitimate characters that happen to look like artefacts.
NFC and NFD: two spellings of the same word
Unicode allows some characters to be written more than one way. The letter é can be the single code point U+00E9, or it can be a plain e followed by a combining acute accent at U+0301. Both render identically. Neither is wrong. But they are different strings, so they do not compare as equal, do not match the same search, and do not produce the same hash.
Normalization resolves this by picking a canonical spelling. NFC composes: it uses the single code point wherever one exists, and this is what you want for storing text, comparing strings, and sending data between systems, because it is what the web and most software expect. NFD decomposes into base characters plus marks, which is useful when stripping accents — decompose, drop the combining marks, and you have the unaccented text — and is what macOS has historically used for filenames.
The K forms, NFKC and NFKD, go further and apply compatibility mappings: a fullwidth A becomes A, a ligature fi becomes fi, and the superscript ² becomes 2. That is destructive in the sense that it loses formatting distinctions, which makes it right for search indexing and wrong for storing what the user typed.
Unicode
Explore all unicode tools
Browse every tool for this kind of task, along with helpful guides that show you how to get the most out of them.
Related tools
Supporting guides
Conversion History
Loading...
FAQ
Why does my text show é instead of é?
Because the text was written as UTF-8 and then read as Latin-1. In UTF-8, é is the two bytes C3 A9; read one byte at a time as Latin-1, those same bytes display as à and ©. Nothing is lost — the original bytes are still there, just interpreted wrongly — which is why the repair is reliable. Paste it into "fix mojibake" mode.
What is mojibake and can it always be fixed?
Mojibake is text made unreadable by being decoded with the wrong character encoding. It is fixable when the original bytes survived the round trip, which is the usual case. It is not fixable when the mis-decoding was itself saved through a lossy step — if characters were replaced with question marks or U+FFFD along the way, the original bytes are gone and no tool can recover them. This tool says so rather than guessing.
What is the difference between Unicode and UTF-8?
Unicode is the catalogue: it assigns every character a number. UTF-8 is one encoding of those numbers into bytes, using one byte for ASCII characters and two to four for everything else. A character has exactly one Unicode code point but a different byte representation in UTF-8, UTF-16 and UTF-32.
When should I use NFC rather than NFD?
Use NFC for almost everything: storing text, comparing strings, and sending data between systems. Most software and most of the web expect composed form. NFD is useful when you want to strip accents — decompose, drop the combining marks, recompose — and you will meet it on macOS, which stores filenames in a decomposed form. Two strings that look identical but use different forms will not compare equal, which is the bug this usually shows up as.
Why does my file open correctly in one program and not another?
Because the file does not say which encoding it uses, so each program guesses. One guesses UTF-8 and is right; another assumes the system default and produces mojibake. Adding a byte order mark makes some programs guess correctly, though it introduces an invisible character at the start of the file, which causes its own problems.
How many bytes is a UTF-8 character?
Between one and four. ASCII characters take one byte, most Latin, Greek, Cyrillic, Hebrew and Arabic letters take two, most CJK characters and the rest of the Basic Multilingual Plane take three, and emoji and rarer scripts take four. Encode mode shows the exact byte count for any text you paste.

