Hiring engineers when AI can pass the take-home

Editorial glass illustration of a translucent code-editor window with crystalline lines of code flowing in from the left and a glass magnifying lens over one highlighted line, suggesting human review of AI output.

A polished take-home submission is becoming weaker evidence of individual engineering ability. When AI can produce a plausible solution, the missing evidence is how the candidate understood, checked and changed it.

That is not a complaint about candidates. It is a design problem for hiring managers. If an assistant can produce a tidy, tested solution to your exercise in an evening, the submission alone no longer separates the engineer who understands it from the one who pasted it.

My answer is not to ban AI and hope. It is to watch the work. The approach I favour is a live technical session with screen sharing, in which candidates may use AI. What it assesses is whether they stay in charge of the work.

This follows directly from Practical AI in engineering — judgment over volume and The review bar after AI. If a change is only ready when its author can explain intent, evidence and risk, the interview should test for exactly that.

Why the submission alone carries less signal

The take-home had real strengths: the candidate’s own editor, their own schedule, nobody watching. Those benefits still matter. What is harder to establish from the submission alone is the candidate’s individual contribution.

Anthropic published a candid account of this in January 2026. Its performance engineering take-home, completed by more than 1,000 candidates since early 2024, had to be redesigned as each new Claude model caught up. Within the time limit, Claude Opus 4 outperformed most human applicants; Claude Opus 4.5 later matched even the strongest. The author, Tristan Hume, moved to deliberately unusual puzzles so the test would still carry signal, and concluded that “realism may be a luxury we no longer have”.

Anthropic did not abandon the format, though. It kept the take-home, still allowed AI assistance, and redesigned it, and Hume describes early results from the new version as promising. That is evidence that assessments need to evolve, not proof that every take-home has failed. A conventional product-listing exercise needs more than a finished submission to distinguish the candidate’s contribution.

Remote live interviews are not immune either. In March 2025 CNBC reported on tools marketed as invisible to interviewers during remote coding interviews, and on Google discussing a return to some in-person rounds. In June 2025 Canva said covert AI use by candidates was becoming increasingly difficult to police.

So the real choice is not “AI or no AI”. It is whether AI use is hidden or visible. Screen sharing helps you observe the work, but it does not prove the absence of undisclosed assistance; the follow-up questions carry more of the weight.

My position: live, screen-shared, AI allowed

Assistants are now part of day-to-day work in many engineering roles. Where that is true, the live session should let candidates work the way they would on the job, with AI tools explicitly allowed.

Allowed is not the same as required. A candidate who solves an appropriate part of the exercise without AI should not be penalised for it. Where AI collaboration is a genuine requirement of the role, say so in the job description and the candidate briefing, and assess it explicitly as its own competency.

Canva goes further, replacing its computer science fundamentals screen with an “AI-Assisted Coding” interview in which candidates are expected to use tools such as Copilot, Cursor or Claude. Its most successful pilot candidates asked clarifying questions, used AI for well-defined subtasks while keeping control of the overall solution, and critically reviewed what it generated. Canva also found that candidates with minimal AI experience often struggled, which it attributed to a lack of judgment in guiding the tool rather than an inability to code. That fits a process where AI use is expected, but limited experience with a particular assistant does not establish poor engineering judgment — one reason I would keep AI optional.

What the session should assess:

  • Problem framing — do they clarify requirements and constraints before writing or generating code?
  • Reading code — can they navigate an unfamiliar codebase and explain what it does?
  • Verification — do they run it, test it and check edge cases, or trust the green tick?
  • Spotting errors — do they notice invented APIs, subtle logic errors and changes outside the brief?
  • Trade-offs — can they explain what they chose and what they gave up?
  • Tool use, where they choose it — do they give an assistant useful context and a scoped task, or a vague wish?

The Stack Overflow 2025 Developer Survey shows why verification carries so much weight: more respondents distrusted the accuracy of AI output (46%) than trusted it (33%). The most common frustration, cited by 66%, was solutions that are “almost right, but not quite”. A submission can be tested for correctness. What it cannot establish on its own is whether the candidate understands the solution and can adapt it.

Formats that produce signal

The best exercises look like the job: mostly reading, changing and judging existing code rather than writing something from a blank file.

  • Debug an existing codebase. Give them a small, realistic repository with a failing test or a reported bug. In ecommerce, a basket that applies a promotion twice, or a product grid that loses keyboard focus after filtering, works well. Watch how they reproduce, narrow down and confirm the fix.
  • Review a pull request. Hand them a PR — ideally AI-written — with planted issues: a missing edge case, an unnecessary dependency, an analytics event that no longer fires, a test that asserts nothing meaningful. Ask them to review it as they would a colleague’s. This is the review bar, run as an interview.
  • Extend existing code. Ask for a modest feature, then have them walk you through the diff and change one decision.

Keep it to one or two per session. If you keep a take-home at all, keep it short and treat it as the starting point for the live conversation, not as the sole evidence.

What engineering judgment looks like in the session

Take the basket that applies a promotion twice. Whether or not the candidate uses AI, the rubric should separate evidence like this:

  • Stronger evidence: The candidate reproduces the duplicate discount, traces the cause, makes a scoped change, and checks that a legitimate discount still applies once.
  • Weaker evidence: The candidate accepts a broad rewrite and passes the original test without checking whether valid promotions still work.

Assess observed decisions and evidence. Do not make constant narration, speed or polished prompting a proxy for ability. An honest “I’m not sure — let me check” is worth more than a confident bluff.

Strong candidates often reach for the tool less than you expect in some moments and more in others. Knowing which moment is which is the signal.

Fairness and practicality

Live sessions add pressure, so the process needs more care, not less. I would recommend:

  • Tell candidates the rules up front. AI allowed but optional, which tools, and how the environment works. Canva tells candidates ahead of time and recommends they practise.
  • Provide the tools and the environment. Give candidates employer-provided access to the agreed tools, a working starter environment and a short familiarisation period before the assessed work begins. Nobody should have to buy a personal subscription to interview.
  • Pause the clock for failures. If a tool or the environment fails, stop the timer until it works again.
  • Share only what is needed. Ask candidates to share the relevant application window rather than their whole screen, so private messages and unrelated personal information are not exposed.
  • Offer reasonable adjustments, and ask. Acas guidance is clear that under the Equality Act 2010 employers must make reasonable adjustments for disabled applicants, at any stage of recruitment, and that employers can ask whether adjustments might be needed. Extra time, breaks, sharing the brief in advance or an alternative format can all be sensible. Take proper HR or legal advice where you are unsure.
  • Allow for nerves. Say explicitly that pausing and checking documentation are fine. Being watched is not the job; judgment is.
  • Respect candidate time. A bounded session limits the time commitment compared with an open-ended take-home that quietly eats an evening, though some candidates will need an alternative format.
  • Standardise. Use comparable exercises and the same role-relevant criteria for every candidate at a given level, with reasonable adjustments to the format where needed. Agree the criteria before interviews start.
  • Calibrate by seniority. For a mid-level engineer, look for sound debugging and honest verification. For a senior or lead, expect them to shape the problem, challenge the brief, find the planted risk in the PR and explain trade-offs in terms of the wider system and team.

An example technical session

Keep the conversation about past work in the screening stage rather than adding another round. Ask about a real technical decision the candidate made: the options, what they chose, what it cost, and what they would do differently now. Follow-up questions about constraints, consequences and what changed after release provide more evidence than a rehearsed project summary.

The technical stage can then be a single 75-minute session, for example:

  • 10 minutes: setup, brief and questions
  • 30 minutes: debugging or extending existing code
  • 15 minutes: focused PR review
  • 10 minutes: explain and revise one decision
  • 10 minutes: candidate questions and wrap-up

Setup failures and agreed adjustments can change this. Afterwards, each interviewer scores independently against the agreed criteria before a structured debrief.

Hire for judgment, not output

AI did not make technical interviews impossible. It made weak ones easier to pass. The response is not to pretend the tools do not exist. It is to test for what still separates good engineers: judgment over volume.

I want engineers who move quickly and stay accountable for every line they ship, whichever tools they use. The most direct way I know to see both is to watch them work with the tools available, then ask them to explain and change what they did. It is the same standard we should apply throughout delivery — the one I argued for in We shipped more code. We did not ship less risk. — and it is one more way leaders stay close to the craft.

Currency note: Anthropic’s take-home write-up, Canva’s AI-assisted interview approach, the Stack Overflow 2025 Developer Survey, CNBC’s reporting and Acas recruitment guidance were checked against public sources as of 25 September 2026. AI tool capabilities and company interview practices change quickly, and nothing here is legal advice; check current guidance for your own process.

Sources and further reading