- Knowledge
- technology
- OOP
- Tips
- Programming
- Tips
- Tutorial
- SEO
- Ranking
- Knowledge
- Special Day
- Seo
- Bug
- Data science
- Seo
- artificial intelligence
- Machine Learning
- Robotics
- happyNewYear2021
- newYearEve
- 2021
- Automation
- Smart Home
- Career
- Best Practices
- Git
- Logging
- Web Fundamentals
- DNS
- HTTPS
- Performance
- AI Tools
- ChatGPT
- Claude
- Gemini
- Laravel
- Eloquent
- MySQL
- HTTPS
- TLS
- Web Security
- Certificates
- Developer Life
- Debugging
- Docker
- DevOps
- Transactions
- Queues
- LLMs
- AI
- AI Coding
- Developer Tools
- React Native
- Expo
- Kate PMS
- Mobile Apps
- Laravel
- Authentication
- Sanctum
- Cookies
- API Design
- Payments
- Idempotency
- DeepSeek
- Open Source AI
- LLMs
- AI News
- Git
- Version Control
- AI Coding
- Prompting
- PHP
- Checklist
- MCP
- AI Agents
- OpenAI
- Architecture
- Microservices
- Modular Monolith
- Estimation
- Developer Life
- Project Planning
- Humour
- OAuth
- OpenID Connect
- Authentication
- Embeddings
- Vector Search
- RAG
- pgvector
- OpenAI
- GPT-4.1
- Codex CLI
- Events
- Testing
- Clean Code
- Maintainability
- Code Review
- Webhooks
- API
- Security
- Claude Code
- Workflow
- AI
- LLM
- Prompt Injection
- Mobile
- React
- Networking
- TCP
- UDP
- HTTP/3
- CLAUDE.md
- AWS
- Cloud Security
- Backups
- PHPUnit
- Software Engineering
- Leadership
- Communication
- RAG
- Embeddings
- AI Engineering
- IT Infrastructure
- Networking
- Access Control
- CI/CD
- GitHub Actions
- Gemini CLI
- Claude Code
- JavaScript
- Async/Await
- Node.js
- Promises
- Security
- Cryptography
- Passwords
- MySQL
- Database
- Vibe Coding
- Software Quality
- DNS
- Code Reading
- Onboarding
- Productivity
- Background Jobs
- Developer Humour
- Estimates
- Dev Life
- JWT
- o3-mini
- DeepSeek R1
- Rate Limiting
- Kate PMS
- E-Signing
- Audit Trail
- REST
- GraphQL
- API Design
- Laravel 12
- Upgrade Guide
- Open Source
- Self-Hosting
- Task Scheduling
- Cron
- Secrets
- CORS
- PHP
- PHP-FPM
- OPcache
- GitHub Copilot
- Software Architecture
- Engineering
- TypeScript
- JavaScript
- Type Safety
- AI Security
- React Native
- Product Design
- AI Agents
- Kiro
- Queues
- Redis
- RabbitMQ
- AWS SQS
- Nginx
- Apache
- GPT-5
- gpt-oss
- Clean Code
- Architecture
- Naming
- Documentation
- Career
- ADR
- Teamwork
- Supply Chain
- Kate HRM
- HR Software
- Permissions
- System Design
- Pagination
- SSH
- Linux
- Big O
- Databases
- Laravel Boost
- MCP
- Developer Skills
- Validation
- Databases
- Indexes
- Code Quality
- Deployment
- Developer Humour
- Feature Flags
- Code Review
- Pull Requests
- Docker
- Cursor
- Authorization
- RBAC
- Gemini
- Long Context
- PHP 8.4
- Caching
- Dependency Injection
- Web Performance
- Browser
- CSS
- Database
- Migrations
- ChatGPT
- AI for Developers
- Monitoring
- On-Call
- REST
- Backend
- SQL
- NoSQL
- Database Design
- Coding Agents
- Claude 4
- API Resources
- REST API
- Load Balancing
- Scaling
- AWS
- AI Tools
- Claude
- Sora 2
- CTE
- 2FA
- TOTP
- Programming Languages
- Prompts
- Developer Workflow
- API Gateway
- APIs
- Passport
- API Auth
- Learning
- Burnout
- Developer Growth
- Web Development
- SEO
- Kate Mall
- ChatGPT Atlas
- Agent Skills
- Middleware
- Laravel 12
- Collections
- Context Window
- Monitoring
- Commit Messages
- Self Review
- Growth
- Regex
- Programming Basics
- Text Processing
- Database Design
- Normalization
- Linux
- Server Security
- Linux Foundation
- Open Standards
- Legacy Code
- Documentation
- AI Workflow
- File Uploads
- Test Data
- Hashing
- Performance
- Caching
- Enums
- Scope Creep
- Estimation
- Codex
- Gemini CLI
- Timezones
- Carbon
- Bugs
- PHP 8.5
- Gemini 3
- GPT-5.1
- Data Integrity
- Event Loop
- Async
- Opus 4.5
- AI Models
- React
- Forms
- Frontend
- Backups
- AI Images
- DALL-E
- Midjourney
- Race Conditions
- Concurrency
- Legacy Code
- Refactoring
- Senior Engineer
- Scope
- LLM
- CDN
- Web
- Sub-Agents
- Soft Deletes
- Audit Log
- Concurrency
- AI Learning
- NestJS
- AI Evals
- Policies
- SPF DKIM DMARC
- Unicode
- UTF-8
- Knowledge Graph
- Value Objects
- Technical Debt
- Feature Flags
- Laravel Pennant
- Deployment
- Copilot
- Composer
- Dependencies
- Artisan
- Automation
- AWS S3
- Object Storage
- Cloud
- Small Language Models
- Ollama
- Production
- Sessions
- HTTP
- Mentoring
- SQL
- Virtual Machines
- Web Development
- HTTP/2
- QUIC
- Web Performance
- AI Integration
- LLM API
- SOLID
- OOP
- Hosting
- Serverless
- Merge Conflicts
- Temperature
- AI Development
- Reverse Proxy
- Nginx
- Infrastructure
- Verification
- Passkeys
- WebAuthn
- Teams
- Communication
- Stakeholders
- Monorepo
- CI/CD
- Versioning
- JSON Schema
- Livewire
- Inertia
- Meetings
- Distributed Systems
- Privacy
- Full-Stack
- T-Shaped Skills
- Money
- Notifications
- Web Security
- HTTP Headers
- CSP
- Function Calling
- Load Testing
- k6
- Data Extraction
- Debugging
- WebSockets
- SSE
- Real-Time
- Laravel Reverb
- Infrastructure as Code
- Terraform
- Side Projects
- Laravel Pint
- OpenAPI
- Swagger
- UX
- Multimodal
- Jest
- Pair Programming
- APIs
- Rate Limiting
- Resilience
- Dev Humour
- Design Tokens
- JWT
- API Keys
- Sessions
- PHPStan
- Rector
- Incidents
- Reporting
- Dashboards
- Zero Trust
- IAM
- Search
- Laravel Scout
- Junior Developers
- Mentoring
- Images
- WebP
- AVIF
- Bug Reports
- Let's Encrypt
- Design Docs
- Software Design
- Observers
- Replication
- Accountability
- Data Structures
- Reliability
- LLM Memory
- Error Handling
- Payments
- Payment Gateway
- Webhooks
- PCI DSS
- Observability
- OpenTelemetry
- Personal Brand
- Writing
- Conventions
- Dates
- Scheduling
- Disaster Recovery
- Compression
- Brotli
- Deadlines
- Developer Habits
- State Machines
- Tech Roles
- UUID
- ULID
- Horizon
- Planning
- Engineering Culture
- Ownership
- Soft Skills
- Socialite
- Cost Control
- Collations
- Unicode
- Octane
- PostgreSQL
Claude Opus 4.5 Is Here: How to Tell if a New Model Is Better for Your Work
About Post
November 2025 will be remembered as the month the frontier models arrived in a queue. GPT-5.1 on the 12th. Gemini 3 on the 18th. And on 24 November, Anthropic released Claude Opus 4.5, its new top model, aimed squarely at coding, agents and computer use.
Every launch comes with charts showing it beating everything else. Every launch, social media declares a new king. And every launch, the only question that matters for you stays unanswered: is it better on your work?
Let's cover what Opus 4.5 is, and then the more durable part: a simple way to evaluate any new model on your own tasks, in an afternoon, without trusting anyone's leaderboard.
What Anthropic released
Opus is the top tier of Anthropic's Claude family, above Sonnet and Haiku. Opus 4.5 is positioned for the hardest work: complex coding tasks, long-running agents and computer use, where a model operates software through the screen the way a person would.
The other notable change is price. Opus 4.5 comes with lower pricing than previous Opus models. That matters more than it sounds. Until now, many teams used Opus only for the hardest tasks and defaulted to a cheaper model for everything else. A cheaper top model changes that calculation, especially for agents, which can use a lot of tokens on a single task.
For people who use Claude Code daily, as I do, a new top model is directly a change to the tool. But the same is true for Codex users when OpenAI ships, and for Gemini CLI users when Google does. The habit below works for all of them.
Why benchmarks aren't enough
Public benchmarks are useful signals, but they measure someone else's tasks, in someone else's setup. Your codebase has its own conventions, its own framework versions, its own weird corners. A model that's brilliant at algorithm puzzles might still ignore your folder structure, and a model that scores slightly lower might follow your CLAUDE.md or AGENTS.md perfectly.
The fix isn't to ignore benchmarks. It's to add your own, small one.
Build a personal eval set (once)
Collect five to ten real tasks from your recent work. The best ones are tasks you've already solved, so you know what "good" looks like. Mix the types:
- A small feature with a test (for example, a new filter on an API endpoint).
- A bug fix where the cause isn't obvious from the error.
- A refactor that must not change behaviour.
- A task in an older, messier part of the codebase.
- A "read and explain" task: summarise how a module works.
Write each one down with the prompt you'd normally give and how you'll check the result:
[
{
"id": "contract-status-filter",
"prompt": "Add filtering by status to GET /api/contracts, with a feature test.",
"check": "php artisan test --filter=ContractIndexTest"
},
{
"id": "invoice-rounding-bug",
"prompt": "Invoices for partial months are off by a few cents. Find and fix the cause.",
"check": "php artisan test --filter=InvoiceCalculationTest"
}
]
Store this file in the repo, or somewhere private if the tasks reveal too much. You'll reuse it every time a model ships.
Run the comparison like an experiment
- Same task, same starting point. A fresh branch from the same commit for each model.
- Same instructions. Same prompt, same project instructions file, same tools allowed.
- Compare against what you use today, not against nothing. The question is "is it better than my current setup?".
- Review blind if you can. Have someone else label the branches, so you don't grade the new model generously because it's new.
What to score
| Question | Why it matters |
|---|---|
| Did it finish, and do the checks pass? | The baseline. Partial work costs you time to finish. |
| Did it follow your conventions? | Code that works but doesn't fit is a review burden forever. |
| How big and focused is the diff? | Small, targeted diffs are easier to review and safer to ship. |
| How many corrections did you make? | The real measure of time saved. |
| How long did it take, and what did it cost? | A slightly better result that takes much longer may not be worth it. |
| Did it do anything risky? | Deleting tests, touching unrelated files, or "fixing" a test instead of the bug. |
The rule: a new model earns a place in your workflow by beating your current setup on your own tasks, not by winning someone else's benchmark. Keep the eval set, rerun it on every launch, and decide with evidence.
For AI features in production, be stricter
If a model powers a feature your users see (a summary, a classification, a generated reply), a model swap is a deployment, not an experiment:
- Keep a "golden set" of real inputs (anonymised, with no personal data) and the outputs you consider correct.
- Check the things that break silently: JSON that no longer validates, a changed tone, longer responses that cost more or overflow your UI.
- Pin the model version in your config, so the switch is a deliberate pull request you can roll back.
- Roll it out to a small share of traffic first if you can.
So, should you switch to Opus 4.5?
Honestly: run your eval set and find out. It's a serious model for coding and agent work, and the lower price makes it worth testing even if you'd ruled Opus out before. But the same is true of GPT-5.1 and Gemini 3 this month, and it'll be true of whatever ships next.
The developers who handle this pace best aren't the ones who switch fastest. They're the ones who can answer "is it better for us?" in an afternoon, with their own tasks and their own tests.
What's one task you'd put in your personal eval set? I'm always looking for good ones that separate the models.

Be first to comment it...