Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

Unicode and UTF-8: Why Your Text Turns Into Question Marks

About Post

A user signs up with their name in Arabic. The form accepts it, the success message appears, everyone's happy. Then you open the admin panel and their name is ????. Not garbled. Not weird symbols. Just four polite question marks where a name used to be.

If you build for users in the Gulf, or anywhere with more than one language, you will meet this bug. Usually in production, usually with someone's name, which is the one thing people really notice.

The good news: text encoding bugs look mysterious, but there are only a handful of them, and each one leaves a recognisable fingerprint.

Characters, code points and bytes

Three ideas, and most confusion comes from mixing them up.

  • Unicode is a giant catalogue. It gives every character a number called a code point. A is U+0041, the Arabic letter meem م is U+0645, and the "face with tears of joy" emoji is U+1F602.
  • An encoding decides how those numbers are stored as bytes. UTF-8 is the one that won the web.
  • UTF-8 is variable-length. Plain English letters take 1 byte, Arabic letters take 2, most Chinese and Japanese characters take 3, and emoji take 4.

That last point is the root of almost everything below. Text is not "one byte per character", and any code, column or connection that assumes it is will eventually damage someone's data.

Reading the fingerprints

The shape of the damage tells you where it happened.

You seeWhat happenedRecoverable?
????Text was converted to a character set that has no way to represent it, like a latin1 column or connection. Each unknown character became ?.No. The original bytes are gone.
محمدCorrect UTF-8 bytes were read as Latin-1 or Windows-1252. Classic "mojibake", often from a connection charset mismatch.Usually, by reinterpreting the bytes.
�Invalid UTF-8 bytes. Often a string cut in the middle of a character, or a file saved in another encoding.Depends on where it was cut.
Incorrect string value: '\xF0\x9F...'A 4-byte character (an emoji) hit a MySQL utf8 column in strict mode.Nothing saved yet; fix the column.

The question marks are the scary one, because nothing will bring that data back. Fix the cause before more names disappear.

MySQL's utf8 is not UTF-8

This is the most famous trap in the stack. In MySQL, the character set called utf8 is really utf8mb3: it stores at most 3 bytes per character. Arabic fits, so everything looks fine during testing. Then the first user types an emoji, which needs 4 bytes, and the insert fails or the character is mangled.

The real UTF-8 is called utf8mb4. Laravel uses it by default in config/database.php, with utf8mb4_unicode_ci as the collation. Problems come from older tables, databases created by hand, or imports from legacy systems. Check with SHOW CREATE TABLE users; and convert what's wrong:

-- Back up first. This rewrites the whole table.
ALTER TABLE users
  CONVERT TO CHARACTER SET utf8mb4
  COLLATE utf8mb4_unicode_ci;

On old MySQL versions you might hit an index length error after converting, because indexed VARCHAR(255) in 4-byte characters exceeds the old 767-byte key limit. That's why some older Laravel apps have Schema::defaultStringLength(191) in a service provider.

The connection matters as much as the column

Your columns can be perfect utf8mb4 and you can still produce mojibake if the connection speaks a different charset. The client sends UTF-8 bytes, MySQL thinks they're Latin-1, "converts" them, and stores double-encoded nonsense.

Laravel sets the connection charset from the same config, so you're covered there. In plain PDO code, put it in the DSN: mysql:host=localhost;dbname=app;charset=utf8mb4. Legacy scripts that skip this are a common source of mojibake that only appears for data written by one part of the system.

PHP strings are bytes

PHP's classic string functions count and cut bytes, not characters:

strlen('محمد');     // 8  (bytes)
mb_strlen('محمد');  // 4  (characters)

substr('محمد', 0, 3);     // "م" plus half a letter: shows �
mb_substr('محمد', 0, 3);  // "محم"

Truncating a name for a notification, an SMS or a PDF with substr is a reliable way to produce �. Use the mb_* functions, or Laravel's Str::limit(), which is multibyte-safe.

Two more PHP gotchas:

  • Regexes need the u flag. preg_match('/^[\p{L} ]+$/u', $name) accepts letters in any script. Without u, the pattern works on bytes and an Arabic name fails a "letters only" check.
  • Know your validation rule. Laravel's alpha rule accepts Unicode letters by default. If you only want A to Z, use alpha:ascii, and think twice about whether you really do.

Emoji add one more twist. A single visible emoji can be several code points glued together (skin tones, family emoji, flags). mb_strlen counts those code points separately. If you need "what a human sees as one character", the intl extension's grapheme_strlen() is the right tool. JavaScript has its own version of the problem: .length counts UTF-16 units, so most emoji have a length of 2.

Collations: when different text counts as equal

The _ci in utf8mb4_unicode_ci means case-insensitive, and Unicode collations also treat many accented letters as equal to their plain versions. So a unique index on a name or username may reject "José" because "Jose" already exists. That's often what you want for searching and rarely what you expect for uniqueness. If a column needs exact matching, give it a binary collation like utf8mb4_bin.

The rule: UTF-8 everywhere, with no exceptions. Source files, HTML <meta charset="utf-8">, HTTP headers, database columns as utf8mb4, the database connection, and multibyte-safe string functions. Encoding bugs live in the one layer you forgot.

Bonus: the Excel CSV problem

You export a perfect UTF-8 CSV with Arabic names. A colleague opens it in Excel and sees mojibake. Excel often guesses the encoding wrong unless the file starts with a byte order mark (BOM). Writing "\xEF\xBB\xBF" at the start of the file usually fixes it. Ugly, but it saves a lot of confused emails.

The quick checklist

  • Tables and columns are utf8mb4, never MySQL's utf8.
  • The connection charset is utf8mb4 too.
  • mb_* functions for length and truncation; u flag on regexes.
  • Test forms with Arabic text and emoji, not just "John Smith".
  • Collations chosen on purpose, especially on unique columns.

What's the worst thing you've seen an encoding bug do to real data? I suspect everyone has a story involving a customer's name.

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close