- Knowledge
- technology
- OOP
- Tips
- Programming
- Tips
- Tutorial
- SEO
- Ranking
- Knowledge
- Special Day
- Seo
- Bug
- Data science
- Seo
- artificial intelligence
- Machine Learning
- Robotics
- happyNewYear2021
- newYearEve
- 2021
- Automation
- Smart Home
- Career
- Best Practices
- Git
- Logging
- Web Fundamentals
- DNS
- HTTPS
- Performance
- AI Tools
- ChatGPT
- Claude
- Gemini
- Laravel
- Eloquent
- MySQL
- HTTPS
- TLS
- Web Security
- Certificates
- Developer Life
- Debugging
- Docker
- DevOps
- Transactions
- Queues
- LLMs
- AI
- AI Coding
- Developer Tools
- React Native
- Expo
- Kate PMS
- Mobile Apps
- Laravel
- Authentication
- Sanctum
- Cookies
- API Design
- Payments
- Idempotency
- DeepSeek
- Open Source AI
- LLMs
- AI News
- Git
- Version Control
- AI Coding
- Prompting
- PHP
- Checklist
- MCP
- AI Agents
- OpenAI
- Architecture
- Microservices
- Modular Monolith
- Estimation
- Developer Life
- Project Planning
- Humour
- OAuth
- OpenID Connect
- Authentication
- Embeddings
- Vector Search
- RAG
- pgvector
- OpenAI
- GPT-4.1
- Codex CLI
- Events
- Testing
- Clean Code
- Maintainability
- Code Review
- Webhooks
- API
- Security
- Claude Code
- Workflow
- AI
- LLM
- Prompt Injection
- Mobile
- React
- Networking
- TCP
- UDP
- HTTP/3
- CLAUDE.md
- AWS
- Cloud Security
- Backups
- PHPUnit
- Software Engineering
- Leadership
- Communication
- RAG
- Embeddings
- AI Engineering
- IT Infrastructure
- Networking
- Access Control
- CI/CD
- GitHub Actions
- Gemini CLI
- Claude Code
- JavaScript
- Async/Await
- Node.js
- Promises
- Security
- Cryptography
- Passwords
- MySQL
- Database
- Vibe Coding
- Software Quality
- DNS
- Code Reading
- Onboarding
- Productivity
- Background Jobs
- Developer Humour
- Estimates
- Dev Life
- JWT
- o3-mini
- DeepSeek R1
- Rate Limiting
- Kate PMS
- E-Signing
- Audit Trail
- REST
- GraphQL
- API Design
- Laravel 12
- Upgrade Guide
- Open Source
- Self-Hosting
- Task Scheduling
- Cron
- Secrets
- CORS
- PHP
- PHP-FPM
- OPcache
- GitHub Copilot
- Software Architecture
- Engineering
- TypeScript
- JavaScript
- Type Safety
- AI Security
- React Native
- Product Design
- AI Agents
- Kiro
- Queues
- Redis
- RabbitMQ
- AWS SQS
- Nginx
- Apache
- GPT-5
- gpt-oss
- Clean Code
- Architecture
- Naming
- Documentation
- Career
- ADR
- Teamwork
- Supply Chain
- Kate HRM
- HR Software
- Permissions
- System Design
- Pagination
- SSH
- Linux
- Big O
- Databases
- Laravel Boost
- MCP
- Developer Skills
- Validation
- Databases
- Indexes
- Code Quality
- Deployment
- Developer Humour
- Feature Flags
- Code Review
- Pull Requests
- Docker
- Cursor
- Authorization
- RBAC
- Gemini
- Long Context
- PHP 8.4
- Caching
- Dependency Injection
- Web Performance
- Browser
- CSS
- Database
- Migrations
- ChatGPT
- AI for Developers
- Monitoring
- On-Call
- REST
- Backend
- SQL
- NoSQL
- Database Design
- Coding Agents
- Claude 4
- API Resources
- REST API
- Load Balancing
- Scaling
- AWS
- AI Tools
- Claude
- Sora 2
- CTE
- 2FA
- TOTP
- Programming Languages
- Prompts
- Developer Workflow
- API Gateway
- APIs
- Passport
- API Auth
- Learning
- Burnout
- Developer Growth
- Web Development
- SEO
- Kate Mall
- ChatGPT Atlas
- Agent Skills
- Middleware
- Laravel 12
- Collections
- Context Window
- Monitoring
- Commit Messages
- Self Review
- Growth
- Regex
- Programming Basics
- Text Processing
- Database Design
- Normalization
- Linux
- Server Security
- Linux Foundation
- Open Standards
- Legacy Code
- Documentation
- AI Workflow
- File Uploads
- Test Data
- Hashing
- Performance
- Caching
- Enums
- Scope Creep
- Estimation
- Codex
- Gemini CLI
- Timezones
- Carbon
- Bugs
- PHP 8.5
- Gemini 3
- GPT-5.1
- Data Integrity
- Event Loop
- Async
- Opus 4.5
- AI Models
- React
- Forms
- Frontend
- Backups
- AI Images
- DALL-E
- Midjourney
- Race Conditions
- Concurrency
- Legacy Code
- Refactoring
- Senior Engineer
- Scope
- LLM
- CDN
- Web
- Sub-Agents
- Soft Deletes
- Audit Log
- Concurrency
- AI Learning
- NestJS
- AI Evals
- Policies
- SPF DKIM DMARC
- Unicode
- UTF-8
- Knowledge Graph
- Value Objects
- Technical Debt
- Feature Flags
- Laravel Pennant
- Deployment
- Copilot
- Composer
- Dependencies
- Artisan
- Automation
- AWS S3
- Object Storage
- Cloud
- Small Language Models
- Ollama
- Production
- Sessions
- HTTP
- Mentoring
- SQL
- Virtual Machines
- Web Development
- HTTP/2
- QUIC
- Web Performance
- AI Integration
- LLM API
- SOLID
- OOP
- Hosting
- Serverless
- Merge Conflicts
- Temperature
- AI Development
- Reverse Proxy
- Nginx
- Infrastructure
- Verification
- Passkeys
- WebAuthn
- Teams
- Communication
- Stakeholders
- Monorepo
- CI/CD
- Versioning
- JSON Schema
- Livewire
- Inertia
- Meetings
- Distributed Systems
- Privacy
- Full-Stack
- T-Shaped Skills
- Money
- Notifications
- Web Security
- HTTP Headers
- CSP
- Function Calling
- Load Testing
- k6
- Data Extraction
- Debugging
- WebSockets
- SSE
- Real-Time
- Laravel Reverb
- Infrastructure as Code
- Terraform
- Side Projects
- Laravel Pint
- OpenAPI
- Swagger
- UX
- Multimodal
- Jest
- Pair Programming
- APIs
- Rate Limiting
- Resilience
- Dev Humour
- Design Tokens
- JWT
- API Keys
- Sessions
- PHPStan
- Rector
- Incidents
- Reporting
- Dashboards
- Zero Trust
- IAM
- Search
- Laravel Scout
- Junior Developers
- Mentoring
- Images
- WebP
- AVIF
- Bug Reports
- Let's Encrypt
- Design Docs
- Software Design
- Observers
- Replication
- Accountability
- Data Structures
- Reliability
- LLM Memory
- Error Handling
- Payments
- Payment Gateway
- Webhooks
- PCI DSS
- Observability
- OpenTelemetry
- Personal Brand
- Writing
- Conventions
- Dates
- Scheduling
- Disaster Recovery
- Compression
- Brotli
- Deadlines
- Developer Habits
- State Machines
- Tech Roles
- UUID
- ULID
- Horizon
- Planning
- Engineering Culture
- Ownership
- Soft Skills
- Socialite
- Cost Control
- Collations
- Unicode
- Octane
- PostgreSQL
Gemini 2.5 Pro and the Race for "Thinking" Models: What a Million Tokens Means for Developers
About Post
This week had one of those days where two big AI stories landed at once, and only one of them took over everybody's feed.
On 25 March, OpenAI added native image generation to GPT-4o in ChatGPT, and within hours the internet was full of photos, memes and family pictures redrawn in a Studio Ghibli style. It was hard to open any app without seeing one.
The same day, Google released Gemini 2.5 Pro. Less fun to share, but for developers, I think it's the more interesting story. Here's why.
What Google released
- Gemini 2.5 Pro, launched on 25 March 2025 as an experimental release.
- Google describes it as a "thinking" model: it reasons through a problem before answering, rather than replying straight away.
- It has a 1 million token context window, meaning you can give it a very large amount of text, code or documents in a single request.
- It's available to try in Google AI Studio and in the Gemini app for Gemini Advanced subscribers.
I'll skip the benchmark tables and leaderboard positions. They're everywhere this week, they get argued about, and they change fast. The more durable story is the two ideas this release combines.
Idea 1: "thinking" is becoming the default
Six months ago, a reasoning model was a special thing. OpenAI's o1 showed that a model which works through a problem step by step before answering does much better on maths, logic and tricky code. Since then, the list has grown quickly: DeepSeek R1 in January, OpenAI's o3-mini right after, Claude 3.7 Sonnet with its extended thinking mode in February, and now Gemini 2.5 Pro, which Google says will be the direction for its models from here on.
For developers, the practical meaning is simple:
- Better at multi-step problems. Debugging that requires following a value through several files, refactors with knock-on effects, and algorithmic questions are where thinking models earn their keep.
- Slower and often more expensive per answer. All that thinking is generated text. For "rename this variable" or "write a docblock", a fast, non-thinking model is still the better tool.
- Prompting shifts slightly. You need fewer "think step by step" tricks and more clear goals and constraints. Tell it what done looks like, and let it work out the steps.
Idea 2: a million tokens of context
A context window is how much the model can "see" in a single request: your prompt, any files you include, the conversation so far and its answer. Most popular models work with far smaller windows than a million tokens. Google's Gemini models have led on context size for a while, and 2.5 Pro pairs that with reasoning.
A million tokens is enough for a sizeable chunk of a real codebase, a long set of documents, or hours of transcripts. That opens up some genuinely useful workflows:
- "Explain this module" for real. Instead of pasting one file and hoping, you can include the controllers, models, jobs and tests that make up a feature, and ask how they fit together.
- Cross-file reviews. "Here is the payment flow across these files. Where could a payment be recorded twice?" That question needs the whole picture.
- Working with big documents. Specs, API documentation or a long contract template, in full, without chunking.
- Legacy code archaeology. Old systems with no documentation are exactly where seeing everything at once helps most.
The catches nobody puts in the announcement
Long context is powerful, but it's not magic, and it's worth knowing the trade-offs before you start pasting your whole repository into a chat box.
- More context isn't always better context. Models can pay less attention to details buried in the middle of a huge input. A focused prompt with the right ten files often beats a sloppy one with five hundred.
- Cost and latency grow with input. You pay, in time and usually in money, for every token you send, on every request.
- It's not memory. A big window holds a lot for one conversation. It doesn't mean the model remembers your codebase tomorrow.
- Retrieval still matters. For large systems and repeated questions, searching for the relevant pieces and sending only those (the idea behind RAG and code-search tools) stays cheaper and often more accurate.
- Be careful what you upload. A huge context makes it tempting to include config files, logs and data dumps. Secrets and personal data don't belong in any prompt, whatever the window size.
- Experimental means experimental. Rate limits and behaviour can change. Try it, learn from it, but don't build a production dependency on an experimental model yet.
My rule for long context: use it to give the model the whole relevant picture, not the whole repository. Curate first, then paste. You'll get better answers for less.
And the Ghibli wave?
It deserves a mention, because it shows something real about where these tools are going. The image generation went viral not just because the pictures were charming, but because it's built into the chat model itself: you can describe edits in conversation and it keeps track of what you meant, including readable text in images, which image models have long struggled with. It also reopened a fair debate about using a living studio's distinctive style, which is a conversation worth having calmly.
For developers, the useful takeaway is that "multimodal" is moving from demo to everyday feature. Mock-ups, diagrams and UI sketches generated from a conversation are going to show up in more of our workflows.
The race, in one paragraph
Every major lab now has a thinking model, and the competition has moved to how well they reason, how much they can see at once, and how fast and affordable they are. That's good news for us. The tools get better every few weeks, and no single provider has a permanent lead. The smart move is to stay a little provider-agnostic: keep your prompts, evaluations and integrations portable, and pick the best model for each job.
Have you tried a long-context model on your own codebase yet? What's the biggest chunk of code you've given an AI in one go, and was the answer actually better?

Be first to comment it...