Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

How I Review AI-Generated Pull Requests: My Six-Step Checklist

19 Aug 2026Category : Blog

About Post

AI-generated code has a particular quality that makes it dangerous to review: it looks right.

The naming is tidy. The formatting is perfect. There are comments, and even tests. Human code that's wrong usually looks a bit wrong, rushed or messy in the place where the bug is. AI code that's wrong looks exactly as confident as AI code that's right.

I use Claude Code every day, so a good share of the diffs I read now started life in an agent session. Over time I've settled into a review routine for them that's different from how I skim a colleague's small fix. Here it is, in the order I actually do it.

1. Read the intent before the code

I don't open the diff first. I open the description, the ticket or the plan the agent worked from, and I make sure I can say in one sentence what this change is supposed to do.

This matters more with AI than with people. An agent will happily solve a slightly different problem than the one you meant, and solve it well. If I don't hold the real goal in my head, I end up reviewing whether the code is good, not whether it's the right code.

If the PR has no clear description, that's the first comment. For my own agent sessions, I ask the agent to write the PR description from the plan, and I edit it until it's true.

2. Check the size and the scope

Next, a quick look at the file list. Two questions:

  • Is it too big to review properly? If so, it gets split. A large AI diff isn't cheaper to review because it was cheap to write.
  • Does it touch things it had no reason to touch? Agents like to tidy up on the way past: a refactored helper here, a renamed variable there, a "small improvement" to a config file. Each one might be fine. Together they hide the real change and widen the blast radius.

Unrelated changes go into a separate PR or get reverted. No exceptions, even for good ones.

3. Read the tests before the implementation

The tests tell me what the author (human or agent) believed the code should do. So I read them first and ask: are these the cases I'd write?

AI-written tests have a few recurring weaknesses I look for specifically:

  • Happy path only. Valid input, expected output, done. Where's the empty list, the unauthorised user, the duplicate submission?
  • Testing the mocks. Everything interesting is mocked, so the test proves that the mock returns what it was told to.
  • Assertions that can't fail. "Response is not null" when the real question is whether the total is correct.
  • Tests edited to pass. If an existing test was changed in the same PR, I want to know why. Sometimes the old expectation was wrong. Sometimes the agent made the failing test agree with the bug.

4. Then the code, hunting for AI-shaped problems

Reading the implementation, I look for the normal things any reviewer looks for, plus a handful of problems that show up more often in generated code:

  • Things that don't exist. A method, config option or package feature that sounds plausible but isn't real, or belongs to a different version. If I don't recognise it, I check the docs.
  • Reinvented wheels. A new helper that duplicates something the codebase already has, because the agent didn't look for it.
  • Swallowed errors. A try/catch that logs and carries on, turning a loud failure into a silent wrong result.
  • Over-engineering. An interface, a factory and a strategy pattern for something that has one implementation and always will.
  • New dependencies. Any new package gets questioned: is it maintained, is it needed, could we do this in twenty lines?

5. A dedicated security pass

I do this as a separate pass because it's easy to miss when you're reading for logic. For every changed endpoint, job and query:

  • Is there an authorisation check, not just authentication? Can a user reach someone else's record by changing an ID?
  • Is input validated, and is mass assignment limited to the expected fields?
  • Any raw SQL built with string concatenation?
  • Any output that skips escaping, or user content that ends up in a file, a URL or a shell command?
  • Any secrets, tokens or personal data in logs, error messages or job payloads?

Generated code tends to be secure when the surrounding code shows a secure pattern to copy, and weakest exactly where the pattern is new.

6. Run it

Green CI is necessary, not sufficient. I check out the branch and use the feature the way a user would, including at least one thing a user shouldn't do: the wrong input, the double click, the expired session, the slow network on the mobile app.

Ten minutes of actually using the feature regularly finds things that no amount of diff-reading would.

My rule: I don't approve AI-generated code I couldn't explain to someone else. If a part of the diff makes me think "I suppose that works", I either understand it properly or it gets rewritten in a way I do understand.

Making the review easier upstream

The best review happens before the PR exists. A few things in my setup reduce what I have to catch:

  • Project rules in CLAUDE.md: conventions, forbidden patterns, where things live, which commands to run before finishing. The agent reads them every session, so the same mistake doesn't come back every week.
  • Plan before code: for anything non-trivial, the agent proposes a plan and I approve it first, so the diff matches a design I've already agreed with.
  • A reviewer sub-agent as a first pass: a separate agent reviews the diff against the rules before I do. It catches the obvious things, and I spend my attention on intent, design and risk.
  • Normal team process: AI-written changes go through the same CI pipeline and pull-request review as any other change. No shortcuts because "the AI already checked it".

The checklist

  1. Intent: can I say in one sentence what this should do?
  2. Scope: right size, nothing unrelated?
  3. Tests: the right cases, real assertions, no tests bent to pass?
  4. Code: nothing invented, duplicated, swallowed or over-built?
  5. Security: authorisation, validation, queries, output, secrets?
  6. Run it: including one thing a user shouldn't do?

The code may come from an agent, but the approval comes from me, and so does the responsibility. What's on your checklist for AI-generated PRs that isn't on mine?

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close