- Knowledge
- technology
- OOP
- Tips
- Programming
- Tips
- Tutorial
- SEO
- Ranking
- Knowledge
- Special Day
- Seo
- Bug
- Data science
- Seo
- artificial intelligence
- Machine Learning
- Robotics
- happyNewYear2021
- newYearEve
- 2021
- Automation
- Smart Home
- Career
- Best Practices
- Git
- Logging
- Web Fundamentals
- DNS
- HTTPS
- Performance
- AI Tools
- ChatGPT
- Claude
- Gemini
- Laravel
- Eloquent
- MySQL
- HTTPS
- TLS
- Web Security
- Certificates
- Developer Life
- Debugging
- Docker
- DevOps
- Transactions
- Queues
- LLMs
- AI
- AI Coding
- Developer Tools
- React Native
- Expo
- Kate PMS
- Mobile Apps
- Laravel
- Authentication
- Sanctum
- Cookies
- API Design
- Payments
- Idempotency
- DeepSeek
- Open Source AI
- LLMs
- AI News
- Git
- Version Control
- AI Coding
- Prompting
- PHP
- Checklist
- MCP
- AI Agents
- OpenAI
- Architecture
- Microservices
- Modular Monolith
- Estimation
- Developer Life
- Project Planning
- Humour
- OAuth
- OpenID Connect
- Authentication
- Embeddings
- Vector Search
- RAG
- pgvector
- OpenAI
- GPT-4.1
- Codex CLI
- Events
- Testing
- Clean Code
- Maintainability
- Code Review
- Webhooks
- API
- Security
- Claude Code
- Workflow
- AI
- LLM
- Prompt Injection
- Mobile
- React
- Networking
- TCP
- UDP
- HTTP/3
- CLAUDE.md
- AWS
- Cloud Security
- Backups
- PHPUnit
- Software Engineering
- Leadership
- Communication
- RAG
- Embeddings
- AI Engineering
- IT Infrastructure
- Networking
- Access Control
- CI/CD
- GitHub Actions
- Gemini CLI
- Claude Code
- JavaScript
- Async/Await
- Node.js
- Promises
- Security
- Cryptography
- Passwords
- MySQL
- Database
- Vibe Coding
- Software Quality
- DNS
- Code Reading
- Onboarding
- Productivity
- Background Jobs
- Developer Humour
- Estimates
- Dev Life
- JWT
- o3-mini
- DeepSeek R1
- Rate Limiting
- Kate PMS
- E-Signing
- Audit Trail
- REST
- GraphQL
- API Design
- Laravel 12
- Upgrade Guide
- Open Source
- Self-Hosting
- Task Scheduling
- Cron
- Secrets
- CORS
- PHP
- PHP-FPM
- OPcache
- GitHub Copilot
- Software Architecture
- Engineering
- TypeScript
- JavaScript
- Type Safety
- AI Security
- React Native
- Product Design
- AI Agents
- Kiro
- Queues
- Redis
- RabbitMQ
- AWS SQS
- Nginx
- Apache
- GPT-5
- gpt-oss
- Clean Code
- Architecture
- Naming
- Documentation
- Career
- ADR
- Teamwork
- Supply Chain
- Kate HRM
- HR Software
- Permissions
- System Design
- Pagination
- SSH
- Linux
- Big O
- Databases
- Laravel Boost
- MCP
- Developer Skills
- Validation
- Databases
- Indexes
- Code Quality
- Deployment
- Developer Humour
- Feature Flags
- Code Review
- Pull Requests
- Docker
- Cursor
- Authorization
- RBAC
- Gemini
- Long Context
- PHP 8.4
- Caching
- Dependency Injection
- Web Performance
- Browser
- CSS
- Database
- Migrations
- ChatGPT
- AI for Developers
- Monitoring
- On-Call
- REST
- Backend
- SQL
- NoSQL
- Database Design
- Coding Agents
- Claude 4
- API Resources
- REST API
- Load Balancing
- Scaling
- AWS
- AI Tools
- Claude
- Sora 2
- CTE
- 2FA
- TOTP
- Programming Languages
- Prompts
- Developer Workflow
- API Gateway
- APIs
- Passport
- API Auth
- Learning
- Burnout
- Developer Growth
- Web Development
- SEO
- Kate Mall
- ChatGPT Atlas
- Agent Skills
- Middleware
- Laravel 12
- Collections
- Context Window
- Monitoring
- Commit Messages
- Self Review
- Growth
- Regex
- Programming Basics
- Text Processing
- Database Design
- Normalization
- Linux
- Server Security
- Linux Foundation
- Open Standards
- Legacy Code
- Documentation
- AI Workflow
- File Uploads
- Test Data
- Hashing
- Performance
- Caching
- Enums
- Scope Creep
- Estimation
- Codex
- Gemini CLI
- Timezones
- Carbon
- Bugs
- PHP 8.5
- Gemini 3
- GPT-5.1
- Data Integrity
- Event Loop
- Async
- Opus 4.5
- AI Models
- React
- Forms
- Frontend
- Backups
- AI Images
- DALL-E
- Midjourney
- Race Conditions
- Concurrency
- Legacy Code
- Refactoring
- Senior Engineer
- Scope
- LLM
- CDN
- Web
- Sub-Agents
- Soft Deletes
- Audit Log
- Concurrency
- AI Learning
- NestJS
- AI Evals
- Policies
- SPF DKIM DMARC
- Unicode
- UTF-8
- Knowledge Graph
- Value Objects
- Technical Debt
- Feature Flags
- Laravel Pennant
- Deployment
- Copilot
- Composer
- Dependencies
- Artisan
- Automation
- AWS S3
- Object Storage
- Cloud
- Small Language Models
- Ollama
- Production
- Sessions
- HTTP
- Mentoring
- SQL
- Virtual Machines
- Web Development
- HTTP/2
- QUIC
- Web Performance
- AI Integration
- LLM API
- SOLID
- OOP
- Hosting
- Serverless
- Merge Conflicts
- Temperature
- AI Development
- Reverse Proxy
- Nginx
- Infrastructure
- Verification
- Passkeys
- WebAuthn
- Teams
- Communication
- Stakeholders
- Monorepo
- CI/CD
- Versioning
- JSON Schema
- Livewire
- Inertia
- Meetings
- Distributed Systems
- Privacy
- Full-Stack
- T-Shaped Skills
- Money
- Notifications
- Web Security
- HTTP Headers
- CSP
- Function Calling
- Load Testing
- k6
- Data Extraction
- Debugging
- WebSockets
- SSE
- Real-Time
- Laravel Reverb
- Infrastructure as Code
- Terraform
- Side Projects
- Laravel Pint
- OpenAPI
- Swagger
- UX
- Multimodal
- Jest
- Pair Programming
- APIs
- Rate Limiting
- Resilience
- Dev Humour
- Design Tokens
- JWT
- API Keys
- Sessions
- PHPStan
- Rector
- Incidents
- Reporting
- Dashboards
- Zero Trust
- IAM
- Search
- Laravel Scout
- Junior Developers
- Mentoring
- Images
- WebP
- AVIF
- Bug Reports
- Let's Encrypt
- Design Docs
- Software Design
- Observers
- Replication
- Accountability
- Data Structures
- Reliability
- LLM Memory
- Error Handling
- Payments
- Payment Gateway
- Webhooks
- PCI DSS
- Observability
- OpenTelemetry
- Personal Brand
- Writing
- Conventions
- Dates
- Scheduling
- Disaster Recovery
- Compression
- Brotli
- Deadlines
- Developer Habits
- State Machines
- Tech Roles
- UUID
- ULID
- Horizon
- Planning
- Engineering Culture
- Ownership
- Soft Skills
- Socialite
- Cost Control
- Collations
- Unicode
- Octane
- PostgreSQL
Observability Explained: Logs, Metrics and Traces (and Where to Start)
About Post
The message arrives: "The app is slow." No screenshot, no time, no idea which screen. Just vibes.
What happens next depends entirely on what you set up before that message. Either you open a dashboard and see the problem in two minutes, or you SSH into a server, tail a log file, and start guessing.
The difference between those two afternoons has a name: observability. It sounds like a buzzword, but the idea is simple, and you don't need a big budget to get most of the value.
Monitoring vs observability: what's the difference?
Monitoring answers questions you knew to ask in advance: is the server up, is the disk full, is the error rate above the threshold?
Observability is being able to answer questions you didn't know you'd need to ask, by looking at what your system tells you from the outside. "Why are payments slow only for users of the mobile app, only since this morning?" No one built an alert for that. Good observability lets you find it anyway.
The raw material comes in three forms, often called the three pillars: logs, metrics and traces.
Logs: what happened?
A log is a timestamped record of an event. "User 42 signed contract 918." "Payment webhook rejected: bad signature." They're the most detailed signal and the one every developer already uses.
The upgrade that matters most is going from text to structured logs:
// Hard to search
Log::info("Contract $contract->id signed by user $user->id");
// Easy to search, filter and count
Log::info('contract.signed', [
'contract_id' => $contract->id,
'user_id' => $user->id,
'channel' => 'mobile',
]);
With structured fields (and a JSON log format in production), "show me every failed signature from the mobile app today" becomes a query instead of a grep marathon.
Logs are great for: the details of a specific event. Logs are bad at: trends. Counting a million log lines to draw a graph is slow and expensive.
Metrics: how much, how often, how fast?
A metric is a number measured over time: requests per second, error rate, queue length, 95th-percentile response time, free disk space. Metrics are cheap to store, quick to graph, and perfect for alerts.
If you're not sure which metrics to start with, the RED method is a good default for every service or endpoint:
- Rate: how many requests per second.
- Errors: how many of them fail.
- Duration: how long they take, as percentiles, not averages. An average hides the slow requests that users actually complain about.
Add a few business-flavoured ones (queue backlog, failed jobs, webhooks received) and you'll notice most problems before users do.
The classic mistake: adding high-cardinality labels like user_id or a full URL with IDs to a metric. Every unique value creates a new time series, and your metrics system slowly melts. Put that kind of detail in logs and traces instead.
Traces: where did the time go?
A trace follows one request through your whole system. Each step is a span with a start and an end: the HTTP request, the three database queries, the call to the payment gateway, the job it put on the queue. Spans nest, so you get a timeline that looks like a waterfall chart.
This is where "the app is slow" finally becomes specific: the request took 2.4 seconds, and 2.1 of them were spent waiting on one external API. Or one endpoint fires 80 nearly identical queries, which is an N+1 problem with a timestamp on it.
Traces become essential once requests cross service boundaries: a mobile app calling an API that calls a microservice that publishes to a queue. Without a trace ID passed along each hop, every service only sees its own piece.
So what is OpenTelemetry?
OpenTelemetry (often shortened to OTel) is an open-source, vendor-neutral standard for producing telemetry: APIs and SDKs for many languages to create logs, metrics and traces, a common protocol (OTLP) to send them, and a Collector that can receive, process and forward them.
The key benefit is that you instrument your code once, and choose or change the backend later: Grafana's stack, Jaeger, Datadog, Honeycomb, a cloud provider's tools, or something self-hosted. Your code doesn't care where the data ends up. It also standardises how trace context travels between services in HTTP headers, so a Node service and a PHP service can share one trace.
How do the three fit together?
| Logs | Metrics | Traces | |
|---|---|---|---|
| Answers | What happened? | Is something wrong? | Where is it slow or broken? |
| Detail | High | Low (aggregated) | High, per request |
| Cost at scale | Grows with traffic | Cheap | Usually sampled |
| Best for | Investigating one event | Dashboards and alerts | Latency and cross-service flows |
The real power comes from linking them. A metric alert fires, you click through to slow traces from that minute, and each trace shows the log lines for that exact request because they share a trace ID. That's the "two-minute afternoon" from the beginning.
I'm a small team. What do I set up first?
Not everything at once. In order of value for effort:
- Structured logs with a request ID, shipped somewhere searchable (not just a file on one server). Most of your debugging improves on day one.
- Error tracking that groups exceptions and shows stack traces and context, so you hear about errors before the email arrives.
- A few RED metrics and alerts on what users feel: error rate, slow responses, queue backlog, failed jobs. Alert on symptoms, not on every CPU spike.
- Tracing when you have several services, external API calls you need to understand, or performance problems you can't explain from logs.
In the Laravel world, Telescope is great for local debugging and Pulse gives a lightweight production dashboard of slow requests, slow queries and job activity, which is a sensible first step before a full observability stack.
The rule of thumb: logs for the story, metrics for the alarm, traces for the map. Start with good logs, add metrics for alerts, and add traces when "where did the time go?" becomes a weekly question.
One last thing: don't log secrets
Observability data gets copied, shipped and kept. Never log passwords, tokens, full card details or personal documents. Mask or drop sensitive fields at the source, because cleaning them out of a log platform later is miserable.
What's in your observability setup today, and which signal do you reach for first when someone says "it's slow"?

Be first to comment it...