Managing a Team Where Everyone Also Manages an Agent
Every engineer on our team now spends part of their day supervising an AI agent instead of writing every line themselves. That changes what a 1:1 is for, what a good week looks like, and how we tell a struggling engineer from a struggling process.
If an engineer on our team had a quiet week, the read used to be simple: the ticket was harder than it looked, or they were stuck and hadn't asked for help. Both diagnoses assumed the engineer's output was a direct function of their own hands on the keyboard. That assumption is now wrong often enough that we've had to change how we manage, not just what tools we hand out.
The job changed underneath the title
Most of our engineers now spend a meaningful chunk of the day directing an agent rather than typing every line themselves: scoping the task, reviewing the plan, correcting course mid-way, reading the diff. That is a different skill from the one we hired most of them for. Some of our strongest historical engineers, the ones who wrote the cleanest code by hand, have struggled more with this shift than engineers who were merely good coders but naturally better at delegating and specifying work clearly. We did not expect the ranking to reshuffle, and it made us realise how much of "good engineer" used to be conflated with "good typist of correct code."
What a good week looks like now, and why that's a management problem
A junior engineer five years ago who shipped three features in a week was probably doing well. An engineer today who ships three features in a week might be doing well, or might be nodding along to an agent's confident output without checking it closely enough, and the difference doesn't show up until the incident three weeks later. Output volume stopped being a reliable proxy for effort or quality at almost exactly the moment it became easiest to measure.
We had to replace "how much did you ship" with a messier but more honest question in our own heads: how much of what shipped did this person actually understand deeply enough to defend in an incident review. That's harder to assess from a dashboard. It mostly comes from listening to how someone talks about their own PR in a 1:1, whether they can explain a decision the agent made, or whether they shrug and say "that's just what it did."
1:1s now include a question we never used to ask
We added one recurring question to our 1:1s: tell me about a time this week the agent was confidently wrong, and how you caught it. The answer tells us more than almost anything else we ask. An engineer with a sharp, specific story is exercising real judgement on every diff that crosses their desk. An engineer who can't produce a recent example either had an unusually clean week or, more often, isn't reading closely enough to have noticed. We've caught more quiet review-quality problems with this one question than with any amount of PR-count tracking.
Telling a struggling engineer from a struggling process
This is the distinction that took us longest to get right. An engineer who's floundering with agent-assisted work sometimes has a genuine skills gap: they haven't developed the judgement to spot a plausible-but-wrong plan before it becomes a diff, and that's coachable with the Aider-style deliberate-practice approach we've written about before. But sometimes the engineer is fine and the process around them is the actual problem: no clear spec to hand the agent, a codebase with no test fixtures to check its output against, review norms that let bad diffs through regardless of who wrote them.
The tell we now look for: does this person's output improve noticeably when we hand them a well-scoped, well-documented piece of work? If yes, the gap is in specification and review discipline, which is a process fix: better tickets, clearer acceptance criteria, a second reviewer on anything ambiguous. If the same person struggles even on well-scoped work, that's a genuine skill gap in directing and checking an agent's output, and it needs coaching, not a better ticket template. Conflating the two leads to fixing the wrong thing. We spent two months improving our ticket templates before realising one specific engineer's problem was never the tickets.
The delegation default we walked back
We stopped treating "uses the agent for everything, even the things they could type faster themselves" as automatically good, and we stopped treating "still writes some things by hand" as automatically inefficient. Some engineers on our team deliberately hand-write the core logic of anything security- or billing-adjacent, delegating only the scaffolding around it, and their incident rate on that class of work is lower than the team average. We used to nudge everyone toward maximum delegation because it looked like the efficient default. We now think the right default is judgement-dependent, and a manager's job is to help each engineer find where their own judgement is strong enough to delegate past, rather than pushing a single team-wide norm.
If your engineering leadership is trying to work out what to actually measure now that agents write a growing share of the code, that's a conversation we have often inside our tech consultancy engagements. Talk to us.
Related articles
The On-Call Rotation We Had to Redesign for an Agent-Assisted Codebase
We paged an engineer for a bug in a module she had, technically, authored, and she had no memory of the logic that broke. That single page is why our incident process no longer assumes the author remembers.
6 min readWhat We Actually Look For When Hiring Engineers in 2026
We once hired an engineer who aced our old interview and struggled on the job within a month. Rebuilding the interview around what actually would have caught that is why it looks nothing like it used to.
6 min readCode Review When Half the Pull Requests Are AI-Drafted
A wrong pull request sailed through two approvals and a green test suite, and looked exactly like every other diff that week. That near-miss is what changed how we review agent-drafted code.
7 min read