Garbled Text Like é and ’: How to Diagnose and Fix Broken Encoding
Why café turns into café, how to tell whether the damage can be undone, and exact fixes for Excel, databases, web pages and code. Includes a symptom lookup table and a prevention checklist.
Published October 3, 2026 · By Sudip Bhowmick
You open a CSV and the name José has become José. A customer email shows an apostrophe as ’. A database export is full of question marks where Chinese characters used to be. These all look like the same bug, but they are three different problems, and only some of them can be repaired. The fastest way to fix broken text is to work out which kind of damage you have before you touch anything, because the wrong repair can make the data worse.
What Is Actually Going Wrong
Computers store text as bytes, and an encoding is the table that says which bytes mean which characters. The letter é is stored in UTF-8 as two bytes, C3 and A9. If the program reading the file believes the file is Windows-1252 instead, it looks up C3 and finds Ã, looks up A9 and finds ©, and shows you é. Nothing in the file changed. The reader simply used the wrong table.
That is why this kind of damage is called mojibake, a Japanese word for transformed characters, and why it is often reversible: the original bytes are still there, just displayed through the wrong lens. Damage becomes permanent only when something converts the garbled text and saves it, so that the wrong characters themselves become the data.
Symptom Lookup: What You See and What It Means
Match your problem to one of these patterns before trying a fix.
- ▸à followed by a symbol (é, è, ü, ñ): UTF-8 text read as Windows-1252 or Latin-1. Reversible. The à is the first byte C3 of a two byte UTF-8 character.
- ▸’, “, †followed by a quote-like mark: curly quotes, apostrophes and dashes from UTF-8 read as Windows-1252. Reversible. Typographic punctuation takes three bytes, which is why you see three junk characters per mark.
- ▸Ð or Ñ in front of other symbols (Привет): Cyrillic text read through the same wrong table. Reversible.
- ▸é or similar with à and  repeating: the text was misread and then saved, and then misread again. Double encoded. Reversible with two passes.
- ▸A black diamond with a question mark (the replacement character): a program met bytes that were invalid in the encoding it expected and substituted this symbol. If the file was saved after that, the original bytes are gone.
- ▸Plain question marks where letters used to be: the text was converted to a character set that cannot hold those letters, such as ASCII or old Latin-1, and each unsupported character was replaced by a question mark. Permanent.
- ▸Empty boxes: the data is fine, but the font has no glyph for the character. Switch fonts. Nothing needs repairing.
Confirm the Diagnosis in Two Minutes
Do not guess. Take one broken word and prove what happened. Paste it into the Unicode Text Converter on this site and look at the code points. For José, you will see U+004A U+006F U+0073 U+00C3 U+00A9. Now take the correct word José into the UTF-8 Encoder and Decoder and look at its bytes: 4a 6f 73 c3 a9. The last two bytes, c3 and a9, match the code points of the junk characters exactly. That match proves UTF-8 bytes were displayed as Latin-1, and that the fix is to read the same bytes as UTF-8.
If the broken text contains symbols such as € or ™, the code points will not match the byte values directly, because Windows-1252 puts those characters in positions 80 to 9F where Latin-1 has control codes. The idea is the same: each junk character stands for one original byte, and you can map it back.
If the converter shows U+FFFD, or the letters have already become question marks, stop here. Repair is impossible from this copy, and you should go back to the original source and export again with the correct settings.
Fix 1: A CSV That Looks Wrong in Excel
The file is usually fine and Excel is guessing the encoding wrong. Many versions of Excel assume the system's legacy code page when a CSV has no byte order mark, so UTF-8 files show mojibake while the same file looks perfect in a text editor.
- ▸Do not double click the file. Open a blank workbook, choose Data, then From Text/CSV, and set File Origin to 65001: Unicode (UTF-8). Excel then reads the same bytes correctly.
- ▸If you generate the CSV yourself and others will open it in Excel, write a UTF-8 byte order mark (the bytes EF BB BF) at the start of the file. Excel recognizes it and chooses UTF-8 automatically.
- ▸Be careful with the BOM in other contexts. Some command line tools and parsers treat it as a stray character at the start of the first column name, so use it for files meant for Excel and leave it out of files meant for scripts.
- ▸Never fix it by saving over the original from Excel after it displayed the junk. That saves the junk as data.
Fix 2: Reversing Mojibake That Is Already in Your Data
When the damaged text has already been stored, you can often reverse it by running the wrong decoding backwards: turn the characters back into the bytes they represented, then decode those bytes as UTF-8.
- ▸Python: encode the string with the same wrong encoding, then decode as UTF-8, for example text.encode('cp1252').decode('utf-8'). Use 'latin-1' if cp1252 raises an error. For double encoding, run it twice.
- ▸JavaScript: map each character back to its byte value, build a Uint8Array and pass it to new TextDecoder('utf-8'). This works directly when all characters are below 256. If the text contains € or ™, convert those through the Windows-1252 table first.
- ▸Mixed data is the hard case. If only some rows are damaged, apply the reversal only to values that contain the telltale à or †patterns. Running it on healthy text such as a real é will either fail or damage it.
- ▸For large or messy datasets, a library built for the job such as ftfy for Python detects and repairs many kinds of mojibake automatically. Review a sample of the output before trusting it on everything.
- ▸Always work on a copy and compare a handful of rows before and after.
Fix 3: Databases and Question Marks
Question marks in a database almost always come from a character set that cannot represent the data. The classic trap is MySQL's utf8, which is really a three byte subset of UTF-8. It cannot store emoji or some rarer characters, which need four bytes. Depending on the SQL mode, such a value is rejected, cut off at the first unsupported character, or stored with replacement characters.
- ▸Use utf8mb4 for the database, tables and columns in MySQL and MariaDB, and a matching collation such as utf8mb4_0900_ai_ci or utf8mb4_unicode_ci.
- ▸Set the connection encoding as well. A correct table with a Latin-1 connection still converts text on the way in and out. In MySQL this is the connection character set, in PostgreSQL it is client_encoding, and most drivers have an option for it.
- ▸In PostgreSQL, create the database with UTF8 encoding. It cannot be changed later without recreating the database.
- ▸After changing a column's character set, existing garbled values are not repaired. They are converted as they are, so repair mojibake first, then convert.
Fix 4: Web Pages, Emails and APIs
When a browser decides on the encoding of a page, it checks a byte order mark first, then the charset in the Content-Type HTTP header, then a meta tag in the HTML. A single disagreement between these is enough to break the page.
- ▸Serve pages as UTF-8 in the Content-Type header, and add a meta charset tag with utf-8 inside the first 1024 bytes of the document, before any text that contains special characters.
- ▸Make sure the file is really saved as UTF-8. An editor that saves in Windows-1252 while the page claims UTF-8 produces exactly the à patterns above.
- ▸JSON should always be UTF-8. If an API response is garbled, check the Content-Type charset and whether an intermediate layer, such as a logging proxy or a legacy gateway, is re-encoding the body.
- ▸For email, the message headers must declare the charset of the body, and the sending library must encode the text in the charset it declares. Broken apostrophes in newsletters usually mean a template saved as Windows-1252 and sent as UTF-8.
- ▸Inside forms, add accept-charset utf-8 if you support very old systems that submit forms in the page's legacy encoding.
Prevent It: One Encoding From End to End
Encoding bugs appear at the boundary where text crosses between systems, so the cure is to make every boundary say UTF-8 explicitly instead of relying on defaults.
- ▸Editors and source files: UTF-8 without BOM for code and data files. Very old versions of Windows Notepad saved ANSI by default, so check the status bar before saving.
- ▸Files you exchange: specify the encoding in the filename, documentation or the file's own header whenever the format allows it.
- ▸Code: pass the encoding to every open, read, write and request call. Avoid functions that fall back to the platform default, which differs between Windows, Linux and containers.
- ▸Database: utf8mb4 or UTF8 at the database level, the column level and the connection level.
- ▸HTTP and HTML: charset in the header and in the first bytes of the document.
- ▸Test with a string that stresses every layer, such as Zażółć ñ ü € ✓ 😀 日本語, and send it through your whole pipeline once. If it survives, ordinary names will too.
When You Only Need Plain ASCII
Sometimes the receiving system truly cannot handle anything outside ASCII, such as an old mainframe feed or a strict file naming rule. In that case, do not let the system silently turn letters into question marks. Convert on purpose: remove accents so that José becomes Jose, replace curly quotes with straight ones, and expand characters such as ß into ss. The Remove Accents and Plain Text Converter tools on this site do exactly that and show you the result before anything is sent. Keep the original text somewhere, because stripping accents loses information that cannot be restored.
Conclusion
Broken characters are not random. A pattern such as é, ’ or é tells you exactly which wrong table was used, and that means the original is usually recoverable by reading the same bytes the right way. Replacement characters and question marks tell you the opposite: the information is gone and you must return to the source. Diagnose first by looking at the code points and bytes of one bad example, repair a copy, and then fix the boundary that caused it by declaring UTF-8 at every step from the editor to the database to the HTTP header.
Free Tool
Open the UTF-8 Encoder and Decoder