Back to Blog

What We Actually Look For When Hiring Engineers in 2026

We once hired an engineer who aced our old interview and struggled on the job within a month. Rebuilding the interview around what actually would have caught that is why it looks nothing like it used to.

EvolRed Team··6 min read

Eighteen months ago we hired an engineer who was, by every measure our old interview tracked, an excellent candidate. Fast, clean live-coding round. A take-home project that was tidy enough we nearly skipped the follow-up call. Within a month of starting, it was clear something didn't fit: handed an ambiguous ticket, they'd ask an agent to fill in the gaps rather than the gaps being the interesting part of the work, and the resulting PRs read as confident and clean right up until a second engineer traced through the logic and found the assumption underneath was wrong. Nothing about that candidate was dishonest or lazy. Our interview had simply stopped measuring the thing that mattered.

The skill the old interview never tested

Our previous process leaned on a timed coding exercise and a take-home, both scored heavily on whether the final code was correct and reasonably clean. That was a fine proxy for engineering ability back when writing correct, clean code by hand was the bottleneck on getting something built. It stopped being a good proxy once a meaningful share of the job, for that hire and increasingly for everyone on the team, involved directing an agent rather than producing code unassisted. A candidate who writes elegant code slowly by hand and a candidate who produces the same elegant code by directing an agent well perform identically on an exercise that only scores the artefact. What we'd never tested was the part that actually predicted whether that hire would catch a wrong assumption before it shipped: not can you write this, but can you tell when the thing in front of you is subtly wrong.

What we test for now, and why each piece exists

The exercise starts with a deliberately underspecified request, the kind of vague ticket every real backlog has a dozen of, and we watch how the candidate handles the ambiguity rather than whether they eventually solve it. Our mis-hire would have reached for the agent immediately and taken whatever came back at face value. The candidates who do well here ask the questions that narrow the problem first, or state their assumptions out loud if they can't ask, which is exactly the muscle that was missing a month into that engineer's first real ticket.

We also hand candidates a diff we wrote ourselves, deliberately seeded with two or three plausible-but-wrong changes, and ask them to review it. This one is modelled directly on the incident that cost us the most: the failure mode we most need a new hire to catch isn't a typo, it's the change that reads as correct on a fast pass and is wrong on a careful one. A candidate who can articulate why something looks right but isn't is demonstrating the judgement a second engineer had to supply after the fact for our mis-hire's PRs.

Candidates can use whatever tools they'd use on the job, agents included, and we score the outcome and the process together rather than banning AI tools from the room, since nobody on our team works that way day to day. What we watch for is whether the candidate directs the tool with a clear plan, catches its mistakes, and can explain every line that ends up in the final diff. A candidate who can't explain a line the agent wrote fails the exercise, regardless of how the final code looks, which is the single check that would have flagged our mis-hire eighteen months ago.

The one part of the process that hasn't changed in years, and if anything has grown more valuable, is spending more of the interview on a single past incident than on a live exercise. Ask an engineer to walk through a real production problem they diagnosed, in detail, and you learn more about their judgement under uncertainty than any amount of whiteboard work reveals.

What we stopped doing, and the honest reason why

We dropped algorithmic puzzle questions almost entirely. Not because algorithms stopped mattering, but because they were never a good proxy for this job's daily reality even before agents; we kept them as long as we did mainly because they were easy to score consistently across interviewers. Easy to score and predictive of job performance turned out to be different properties, and we'd been quietly optimising for the first one. We also stopped penalising a rough final artefact when the process behind it was sound. An engineer who ships something imperfect but can tell you exactly what's wrong with it and why they'd fix it next is a safer hire than one who ships something polished and can't identify its own weak points, and our old rubric scored the second candidate higher every time.

Whether it actually works

This interview is harder to score consistently across interviewers than the old one, and we know it: a timed algorithm question produces a number, a scoping conversation produces a judgement call that two interviewers can read differently. We decided a noisier signal correlated with the job beats a clean signal that measures something we no longer pay people to do. Eighteen months in, the same span of time since that first mis-hire, our early-tenure performance data backs the trade: fewer surprises in the first quarter than the old process gave us, even accounting for the smaller sample size. We haven't had a repeat of that specific failure since.


Building an interview process that still works now that most engineering work runs through an agent is a question we help engineering leaders think through as part of our tech consultancy work. Talk to us.