- Knowledge
- technology
- OOP
- Tips
- Programming
- Tips
- Tutorial
- SEO
- Ranking
- Knowledge
- Special Day
- Seo
- Bug
- Data science
- Seo
- artificial intelligence
- Machine Learning
- Robotics
- happyNewYear2021
- newYearEve
- 2021
- Automation
- Smart Home
- Career
- Best Practices
- Git
- Logging
- Web Fundamentals
- DNS
- HTTPS
- Performance
- AI Tools
- ChatGPT
- Claude
- Gemini
- Laravel
- Eloquent
- MySQL
- HTTPS
- TLS
- Web Security
- Certificates
- Developer Life
- Debugging
- Docker
- DevOps
- Transactions
- Queues
- LLMs
- AI
- AI Coding
- Developer Tools
- React Native
- Expo
- Kate PMS
- Mobile Apps
- Laravel
- Authentication
- Sanctum
- Cookies
- API Design
- Payments
- Idempotency
- DeepSeek
- Open Source AI
- LLMs
- AI News
- Git
- Version Control
- AI Coding
- Prompting
- PHP
- Checklist
- MCP
- AI Agents
- OpenAI
- Architecture
- Microservices
- Modular Monolith
- Estimation
- Developer Life
- Project Planning
- Humour
- OAuth
- OpenID Connect
- Authentication
- Embeddings
- Vector Search
- RAG
- pgvector
- OpenAI
- GPT-4.1
- Codex CLI
- Events
- Testing
- Clean Code
- Maintainability
- Code Review
- Webhooks
- API
- Security
- Claude Code
- Workflow
- AI
- LLM
- Prompt Injection
- Mobile
- React
- Networking
- TCP
- UDP
- HTTP/3
- CLAUDE.md
- AWS
- Cloud Security
- Backups
- PHPUnit
- Software Engineering
- Leadership
- Communication
- RAG
- Embeddings
- AI Engineering
- IT Infrastructure
- Networking
- Access Control
- CI/CD
- GitHub Actions
- Gemini CLI
- Claude Code
- JavaScript
- Async/Await
- Node.js
- Promises
- Security
- Cryptography
- Passwords
- MySQL
- Database
- Vibe Coding
- Software Quality
- DNS
- Code Reading
- Onboarding
- Productivity
- Background Jobs
- Developer Humour
- Estimates
- Dev Life
- JWT
- o3-mini
- DeepSeek R1
- Rate Limiting
- Kate PMS
- E-Signing
- Audit Trail
- REST
- GraphQL
- API Design
- Laravel 12
- Upgrade Guide
- Open Source
- Self-Hosting
- Task Scheduling
- Cron
- Secrets
- CORS
- PHP
- PHP-FPM
- OPcache
- GitHub Copilot
- Software Architecture
- Engineering
- TypeScript
- JavaScript
- Type Safety
- AI Security
- React Native
- Product Design
- AI Agents
- Kiro
- Queues
- Redis
- RabbitMQ
- AWS SQS
- Nginx
- Apache
- GPT-5
- gpt-oss
- Clean Code
- Architecture
- Naming
- Documentation
- Career
- ADR
- Teamwork
- Supply Chain
- Kate HRM
- HR Software
- Permissions
- System Design
- Pagination
- SSH
- Linux
- Big O
- Databases
- Laravel Boost
- MCP
- Developer Skills
- Validation
- Databases
- Indexes
- Code Quality
- Deployment
- Developer Humour
- Feature Flags
- Code Review
- Pull Requests
- Docker
- Cursor
- Authorization
- RBAC
- Gemini
- Long Context
- PHP 8.4
- Caching
- Dependency Injection
- Web Performance
- Browser
- CSS
- Database
- Migrations
- ChatGPT
- AI for Developers
- Monitoring
- On-Call
- REST
- Backend
- SQL
- NoSQL
- Database Design
- Coding Agents
- Claude 4
- API Resources
- REST API
- Load Balancing
- Scaling
- AWS
- AI Tools
- Claude
- Sora 2
- CTE
- 2FA
- TOTP
- Programming Languages
- Prompts
- Developer Workflow
- API Gateway
- APIs
- Passport
- API Auth
- Learning
- Burnout
- Developer Growth
- Web Development
- SEO
- Kate Mall
- ChatGPT Atlas
- Agent Skills
- Middleware
- Laravel 12
- Collections
- Context Window
- Monitoring
- Commit Messages
- Self Review
- Growth
- Regex
- Programming Basics
- Text Processing
- Database Design
- Normalization
- Linux
- Server Security
- Linux Foundation
- Open Standards
- Legacy Code
- Documentation
- AI Workflow
- File Uploads
- Test Data
- Hashing
- Performance
- Caching
- Enums
- Scope Creep
- Estimation
- Codex
- Gemini CLI
- Timezones
- Carbon
- Bugs
- PHP 8.5
- Gemini 3
- GPT-5.1
- Data Integrity
- Event Loop
- Async
- Opus 4.5
- AI Models
- React
- Forms
- Frontend
- Backups
- AI Images
- DALL-E
- Midjourney
- Race Conditions
- Concurrency
- Legacy Code
- Refactoring
- Senior Engineer
- Scope
- LLM
- CDN
- Web
- Sub-Agents
- Soft Deletes
- Audit Log
- Concurrency
- AI Learning
- NestJS
- AI Evals
- Policies
- SPF DKIM DMARC
- Unicode
- UTF-8
- Knowledge Graph
- Value Objects
- Technical Debt
- Feature Flags
- Laravel Pennant
- Deployment
- Copilot
- Composer
- Dependencies
- Artisan
- Automation
- AWS S3
- Object Storage
- Cloud
- Small Language Models
- Ollama
- Production
- Sessions
- HTTP
- Mentoring
- SQL
- Virtual Machines
- Web Development
- HTTP/2
- QUIC
- Web Performance
- AI Integration
- LLM API
- SOLID
- OOP
- Hosting
- Serverless
- Merge Conflicts
- Temperature
- AI Development
- Reverse Proxy
- Nginx
- Infrastructure
- Verification
- Passkeys
- WebAuthn
- Teams
- Communication
- Stakeholders
- Monorepo
- CI/CD
- Versioning
- JSON Schema
- Livewire
- Inertia
- Meetings
- Distributed Systems
- Privacy
- Full-Stack
- T-Shaped Skills
- Money
- Notifications
- Web Security
- HTTP Headers
- CSP
- Function Calling
- Load Testing
- k6
- Data Extraction
- Debugging
- WebSockets
- SSE
- Real-Time
- Laravel Reverb
- Infrastructure as Code
- Terraform
- Side Projects
- Laravel Pint
- OpenAPI
- Swagger
- UX
- Multimodal
- Jest
- Pair Programming
- APIs
- Rate Limiting
- Resilience
- Dev Humour
- Design Tokens
- JWT
- API Keys
- Sessions
- PHPStan
- Rector
- Incidents
- Reporting
- Dashboards
- Zero Trust
- IAM
- Search
- Laravel Scout
- Junior Developers
- Mentoring
- Images
- WebP
- AVIF
- Bug Reports
- Let's Encrypt
- Design Docs
- Software Design
- Observers
- Replication
- Accountability
- Data Structures
- Reliability
- LLM Memory
- Error Handling
- Payments
- Payment Gateway
- Webhooks
- PCI DSS
- Observability
- OpenTelemetry
- Personal Brand
- Writing
- Conventions
- Dates
- Scheduling
- Disaster Recovery
- Compression
- Brotli
- Deadlines
- Developer Habits
- State Machines
- Tech Roles
- UUID
- ULID
- Horizon
- Planning
- Engineering Culture
- Ownership
- Soft Skills
- Socialite
- Cost Control
- Collations
- Unicode
- Octane
- PostgreSQL
How I Review AI-Generated Pull Requests: My Six-Step Checklist
About Post
AI-generated code has a particular quality that makes it dangerous to review: it looks right.
The naming is tidy. The formatting is perfect. There are comments, and even tests. Human code that's wrong usually looks a bit wrong, rushed or messy in the place where the bug is. AI code that's wrong looks exactly as confident as AI code that's right.
I use Claude Code every day, so a good share of the diffs I read now started life in an agent session. Over time I've settled into a review routine for them that's different from how I skim a colleague's small fix. Here it is, in the order I actually do it.
1. Read the intent before the code
I don't open the diff first. I open the description, the ticket or the plan the agent worked from, and I make sure I can say in one sentence what this change is supposed to do.
This matters more with AI than with people. An agent will happily solve a slightly different problem than the one you meant, and solve it well. If I don't hold the real goal in my head, I end up reviewing whether the code is good, not whether it's the right code.
If the PR has no clear description, that's the first comment. For my own agent sessions, I ask the agent to write the PR description from the plan, and I edit it until it's true.
2. Check the size and the scope
Next, a quick look at the file list. Two questions:
- Is it too big to review properly? If so, it gets split. A large AI diff isn't cheaper to review because it was cheap to write.
- Does it touch things it had no reason to touch? Agents like to tidy up on the way past: a refactored helper here, a renamed variable there, a "small improvement" to a config file. Each one might be fine. Together they hide the real change and widen the blast radius.
Unrelated changes go into a separate PR or get reverted. No exceptions, even for good ones.
3. Read the tests before the implementation
The tests tell me what the author (human or agent) believed the code should do. So I read them first and ask: are these the cases I'd write?
AI-written tests have a few recurring weaknesses I look for specifically:
- Happy path only. Valid input, expected output, done. Where's the empty list, the unauthorised user, the duplicate submission?
- Testing the mocks. Everything interesting is mocked, so the test proves that the mock returns what it was told to.
- Assertions that can't fail. "Response is not null" when the real question is whether the total is correct.
- Tests edited to pass. If an existing test was changed in the same PR, I want to know why. Sometimes the old expectation was wrong. Sometimes the agent made the failing test agree with the bug.
4. Then the code, hunting for AI-shaped problems
Reading the implementation, I look for the normal things any reviewer looks for, plus a handful of problems that show up more often in generated code:
- Things that don't exist. A method, config option or package feature that sounds plausible but isn't real, or belongs to a different version. If I don't recognise it, I check the docs.
- Reinvented wheels. A new helper that duplicates something the codebase already has, because the agent didn't look for it.
- Swallowed errors. A
try/catchthat logs and carries on, turning a loud failure into a silent wrong result. - Over-engineering. An interface, a factory and a strategy pattern for something that has one implementation and always will.
- New dependencies. Any new package gets questioned: is it maintained, is it needed, could we do this in twenty lines?
5. A dedicated security pass
I do this as a separate pass because it's easy to miss when you're reading for logic. For every changed endpoint, job and query:
- Is there an authorisation check, not just authentication? Can a user reach someone else's record by changing an ID?
- Is input validated, and is mass assignment limited to the expected fields?
- Any raw SQL built with string concatenation?
- Any output that skips escaping, or user content that ends up in a file, a URL or a shell command?
- Any secrets, tokens or personal data in logs, error messages or job payloads?
Generated code tends to be secure when the surrounding code shows a secure pattern to copy, and weakest exactly where the pattern is new.
6. Run it
Green CI is necessary, not sufficient. I check out the branch and use the feature the way a user would, including at least one thing a user shouldn't do: the wrong input, the double click, the expired session, the slow network on the mobile app.
Ten minutes of actually using the feature regularly finds things that no amount of diff-reading would.
My rule: I don't approve AI-generated code I couldn't explain to someone else. If a part of the diff makes me think "I suppose that works", I either understand it properly or it gets rewritten in a way I do understand.
Making the review easier upstream
The best review happens before the PR exists. A few things in my setup reduce what I have to catch:
- Project rules in
CLAUDE.md: conventions, forbidden patterns, where things live, which commands to run before finishing. The agent reads them every session, so the same mistake doesn't come back every week. - Plan before code: for anything non-trivial, the agent proposes a plan and I approve it first, so the diff matches a design I've already agreed with.
- A reviewer sub-agent as a first pass: a separate agent reviews the diff against the rules before I do. It catches the obvious things, and I spend my attention on intent, design and risk.
- Normal team process: AI-written changes go through the same CI pipeline and pull-request review as any other change. No shortcuts because "the AI already checked it".
The checklist
- Intent: can I say in one sentence what this should do?
- Scope: right size, nothing unrelated?
- Tests: the right cases, real assertions, no tests bent to pass?
- Code: nothing invented, duplicated, swallowed or over-built?
- Security: authorisation, validation, queries, output, secrets?
- Run it: including one thing a user shouldn't do?
The code may come from an agent, but the approval comes from me, and so does the responsibility. What's on your checklist for AI-generated PRs that isn't on mine?

Be first to comment it...