Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

Me vs the Bug That Only Happens in Production

About Post

Tests: green. Staging: green. My laptop: flawless, a picture of health.

Production: "Hi, a few users are saying the invoice page is blank?"

A few users. Not all of them. Not me. Not anyone I can phone. Welcome to the most humbling genre of software bug, the one that only exists where real people are using real data on a real server. Every developer meets it eventually, and we all go through the same stages.

The five stages of a production-only bug

Stage 1: Denial

❌ "Works for me." You open the page. It works. You open it in incognito. It works. You ask the user to clear their cache, with the quiet confidence of someone who has solved nothing.

Stage 2: Archaeology

❌ You open the logs. There are thousands of lines. Most of them say Something went wrong, with no user, no request and no clue. One says Undefined array key "address", which is both the answer and completely useless, because it doesn't say whose address.

Stage 3: Bargaining

❌ "What if I just add a dd() on production for one minute?" You know this is wrong. You think about it anyway. You compromise with a Log::info('HERE 1'), and then 'HERE 2', and deploy twice.

Stage 4: The discovery

❌ It turns out the bug only happens for customers created before a migration last year, whose address field is null instead of an empty string. Your local database has fifty tidy factory users, all with perfect addresses. Production has years of history, imports and edge cases. Of course it worked on your machine. Your machine has never met a real customer.

Stage 5: Acceptance

❌ The fix is one line. Finding it took the whole afternoon. You whisper "never again", knowing full well there will be an again.

Why production is different

Production-only bugs are rarely mysterious once you find them. They almost always come from one of three gaps:

  • Data: old records, nulls, huge values, odd characters, users with permissions nobody planned for.
  • Config: different environment variables, cache drivers, queue drivers, timezones, PHP extensions.
  • Scale and timing: concurrency, real traffic, slow third-party APIs, jobs overlapping.

You can't remove these gaps completely. But you can make them visible, and that's what turns an afternoon into ten minutes.

What actually helps

✅ Logs with context, not just messages. A log line is only useful if it tells you who, what and where. In Laravel, add context once and every log line in that request carries it:

// In a middleware, early in the request
Log::withContext([
    'request_id' => (string) Str::uuid(),
    'user_id' => $request->user()?->id,
    'route' => $request->route()?->getName(),
]);

Now "Undefined array key" comes with a user ID you can look up, and a request ID you can follow across every line it wrote. An error tracker that groups exceptions and shows the request details turns the archaeology stage into a quick search.

✅ Config parity. Run the same PHP version, extensions, database engine, cache and queue drivers locally as in production. Docker makes this much easier. If production uses Redis and your laptop uses the array cache driver, you're testing a different app.

✅ Feature flags for risky changes. Release new code paths to a small group first, staff accounts or a handful of users, and widen it once it behaves. If something breaks, you turn the flag off instead of rolling back a deploy. Laravel Pennant gives you this out of the box:

if (Feature::active('new-invoice-page')) {
    return view('invoices.v2', compact('invoice'));
}

return view('invoices.show', compact('invoice'));

✅ Reproduce with real-like data. Factories should create messy users too: nulls, very long names, Arabic and accented characters, old statuses, missing relationships. For tricky bugs, an anonymised copy of production data in a safe environment beats any amount of guessing. Just make sure personal information is scrubbed before it leaves production.

The rule: when you can't reproduce a bug, don't guess harder. Add the visibility that would have told you the answer, then wait for it to happen again. The second time, it can't hide.

The real lesson

"It only happens in production" isn't bad luck. It's production telling you exactly where your test data and environment are lying to you. Each one of these bugs is a free lesson about the gap, if you fix the gap and not just the line.

So yes, add the null check. Then add the messy user to your factories, the context to your logs, and the flag to your next risky feature.

What's the strangest production-only bug you've chased? Bonus points if the cause was a single character.

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close