Back to Blog

Code Review When Half the Pull Requests Are AI-Drafted

A wrong pull request sailed through two approvals and a green test suite, and looked exactly like every other diff that week. That near-miss is what changed how we review agent-drafted code.

EvolRed Team··7 min read

A few months ago an agent-drafted PR went through review with two approvals, a green test suite, and a diff that read exactly like every other diff that week. It shipped. It also quietly changed how a discount code stacked with a promotional price, in a direction that cost real money before anyone noticed. Nobody had done anything obviously careless. The reviewer read the diff, the tests passed, the PR description was clear and specific. The problem was that none of those three signals, the ones code review has leaned on for a decade, meant what they used to.

Why the old signals stopped working

Traditional review assumes a competent engineer's confidence roughly tracks their correctness: if they weren't sure about something, the PR description usually says so, or there's a comment flagging the bit they want a second opinion on. An agent's confidence doesn't work that way. It writes the correct implementation and the subtly wrong one with exactly the same tone, the same clean formatting, the same absence of a hedge, because it isn't tracking its own uncertainty the way a human author does, or if it is, that signal rarely survives into the diff. The discount bug had a textbook-clean PR description. That was never going to be the thing that caught it.

The part that went the other way

Not everything got harder. We expected agent-drafted PRs to need more review time overall, and the opposite happened: the average review got faster, because the diffs are more consistently formatted and narrower in scope than what a human engineer would usually attempt in one sitting. What shifted wasn't the average, it was the shape of the risk. Most agent PRs are quick, correct, and boring to review. The rare bad one is bad in a way that looks identical to a good one until someone reads it carefully, which is exactly the case our old process, tuned for evenly spread risk, was never built to catch.

What review actually checks for now

The biggest single change was moving the judgement call earlier: we now ask for the agent's plan before it writes any code, with the parts it's least sure about flagged explicitly. A plan that says "I assumed X because the code doesn't specify Y" gives a reviewer something concrete to interrogate. A finished diff with no such flag gives them nothing but the code itself to reverse-engineer intent from, which is slower and catches less. We also stopped reading the PR description first. It used to be a decent proxy for how carefully a change had been considered, back when writing the summary forced a human author to reflect on what they'd done. An agent writes an equally polished summary whether the change is right or wrong, so we read the diff first now and treat the description as a secondary check.

Two categories get scrutiny no matter how clean the diff looks. The first is anything an agent added that nobody asked for: a type error fixed two files away, a variable renamed for consistency, a null check nobody requested. Individually these are usually fine; collectively they widen the blast radius past what the ticket described, so it's the first thing we scan for now. The second is anything touching money, auth, or data deletion, which gets a second human review regardless of who or what wrote it and regardless of how clean the first pass looked. Tests get the same treatment as implementation rather than a lighter skim, too, since AI-written tests on a codebase the agent didn't author tend to be assertion-rich and behaviourally thin in exactly the way that lets a bug like the discount one slip through a green suite.

Back to the discount bug

Under the old process, that PR would probably ship the same way again: clean diff, passing tests, confident description, nothing to flag it. Under the current one, the plan the agent would have had to state up front, something like "assuming stacking rules follow the existing promo-code path," is the kind of assumption a reviewer can check against the actual pricing logic before a line of code exists. It isn't a guarantee. It's the difference between having something specific to interrogate and having only a finished, plausible-looking diff to take on faith.


Untangling a review process that's drowning in agent-authored diffs, or building the review tooling to keep up, is work we do regularly inside our tech consultancy engagements. Talk to us.