Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

Observability Explained: Logs, Metrics and Traces (and Where to Start)

About Post

The message arrives: "The app is slow." No screenshot, no time, no idea which screen. Just vibes.

What happens next depends entirely on what you set up before that message. Either you open a dashboard and see the problem in two minutes, or you SSH into a server, tail a log file, and start guessing.

The difference between those two afternoons has a name: observability. It sounds like a buzzword, but the idea is simple, and you don't need a big budget to get most of the value.

Monitoring vs observability: what's the difference?

Monitoring answers questions you knew to ask in advance: is the server up, is the disk full, is the error rate above the threshold?

Observability is being able to answer questions you didn't know you'd need to ask, by looking at what your system tells you from the outside. "Why are payments slow only for users of the mobile app, only since this morning?" No one built an alert for that. Good observability lets you find it anyway.

The raw material comes in three forms, often called the three pillars: logs, metrics and traces.

Logs: what happened?

A log is a timestamped record of an event. "User 42 signed contract 918." "Payment webhook rejected: bad signature." They're the most detailed signal and the one every developer already uses.

The upgrade that matters most is going from text to structured logs:

// Hard to search
Log::info("Contract $contract->id signed by user $user->id");

// Easy to search, filter and count
Log::info('contract.signed', [
    'contract_id' => $contract->id,
    'user_id' => $user->id,
    'channel' => 'mobile',
]);

With structured fields (and a JSON log format in production), "show me every failed signature from the mobile app today" becomes a query instead of a grep marathon.

Logs are great for: the details of a specific event. Logs are bad at: trends. Counting a million log lines to draw a graph is slow and expensive.

Metrics: how much, how often, how fast?

A metric is a number measured over time: requests per second, error rate, queue length, 95th-percentile response time, free disk space. Metrics are cheap to store, quick to graph, and perfect for alerts.

If you're not sure which metrics to start with, the RED method is a good default for every service or endpoint:

  • Rate: how many requests per second.
  • Errors: how many of them fail.
  • Duration: how long they take, as percentiles, not averages. An average hides the slow requests that users actually complain about.

Add a few business-flavoured ones (queue backlog, failed jobs, webhooks received) and you'll notice most problems before users do.

The classic mistake: adding high-cardinality labels like user_id or a full URL with IDs to a metric. Every unique value creates a new time series, and your metrics system slowly melts. Put that kind of detail in logs and traces instead.

Traces: where did the time go?

A trace follows one request through your whole system. Each step is a span with a start and an end: the HTTP request, the three database queries, the call to the payment gateway, the job it put on the queue. Spans nest, so you get a timeline that looks like a waterfall chart.

This is where "the app is slow" finally becomes specific: the request took 2.4 seconds, and 2.1 of them were spent waiting on one external API. Or one endpoint fires 80 nearly identical queries, which is an N+1 problem with a timestamp on it.

Traces become essential once requests cross service boundaries: a mobile app calling an API that calls a microservice that publishes to a queue. Without a trace ID passed along each hop, every service only sees its own piece.

So what is OpenTelemetry?

OpenTelemetry (often shortened to OTel) is an open-source, vendor-neutral standard for producing telemetry: APIs and SDKs for many languages to create logs, metrics and traces, a common protocol (OTLP) to send them, and a Collector that can receive, process and forward them.

The key benefit is that you instrument your code once, and choose or change the backend later: Grafana's stack, Jaeger, Datadog, Honeycomb, a cloud provider's tools, or something self-hosted. Your code doesn't care where the data ends up. It also standardises how trace context travels between services in HTTP headers, so a Node service and a PHP service can share one trace.

How do the three fit together?

LogsMetricsTraces
AnswersWhat happened?Is something wrong?Where is it slow or broken?
DetailHighLow (aggregated)High, per request
Cost at scaleGrows with trafficCheapUsually sampled
Best forInvestigating one eventDashboards and alertsLatency and cross-service flows

The real power comes from linking them. A metric alert fires, you click through to slow traces from that minute, and each trace shows the log lines for that exact request because they share a trace ID. That's the "two-minute afternoon" from the beginning.

I'm a small team. What do I set up first?

Not everything at once. In order of value for effort:

  1. Structured logs with a request ID, shipped somewhere searchable (not just a file on one server). Most of your debugging improves on day one.
  2. Error tracking that groups exceptions and shows stack traces and context, so you hear about errors before the email arrives.
  3. A few RED metrics and alerts on what users feel: error rate, slow responses, queue backlog, failed jobs. Alert on symptoms, not on every CPU spike.
  4. Tracing when you have several services, external API calls you need to understand, or performance problems you can't explain from logs.

In the Laravel world, Telescope is great for local debugging and Pulse gives a lightweight production dashboard of slow requests, slow queries and job activity, which is a sensible first step before a full observability stack.

The rule of thumb: logs for the story, metrics for the alarm, traces for the map. Start with good logs, add metrics for alerts, and add traces when "where did the time go?" becomes a weekly question.

One last thing: don't log secrets

Observability data gets copied, shipped and kept. Never log passwords, tokens, full card details or personal documents. Mask or drop sensitive fields at the source, because cleaning them out of a log platform later is miserable.

What's in your observability setup today, and which signal do you reach for first when someone says "it's slow"?

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close