Skills-Based Hiring Has a Cheating Problem. The Test Is the Problem.
Over a third of AI-led interviews now flag AI assistance. The real failure is not the cheating — it is tests that reward polish over judgment.

Over a third of candidates in AI-led interviews are now getting flagged for using AI assistance during the assessment itself. A year ago that number was under 10%. In proctored technical tests, cheating attempts more than doubled in twelve months. For entry-level roles, the rate nearly tripled.
The instinctive response to these numbers is to tighten the proctoring. Add gaze tracking. Catch the earpieces. Deploy face-swap detection. And some of that is genuinely necessary — identity fraud in remote hiring is real and the market has built some impressive tools to combat it.
But the framing misses the bigger problem. The test is broken before anyone cheats on it.
What skills-based hiring was supposed to fix
The pitch was simple: stop filtering on proxies (degrees, brand names, polished resumes) and start testing what candidates can actually do. A live coding challenge reveals more than a CS degree from the right school. A work-sample exercise shows judgment that no bullet point can fake.
That logic is sound. The research backs it — structured behavioral assessments outpredict resumes and unstructured interviews as selection devices. The shift away from credential gatekeeping was overdue.
But somewhere between the insight and the implementation, something went wrong. Most companies moved to skills-based assessments without moving to skills-based thinking about what those assessments should test. They replaced the old credential with a new one: a score on a take-home that rewards whoever produces the most polished output, regardless of how they produced it.
A general-purpose AI model already aces most of those tests. Not because it's cheating. Because the test was always asking the wrong question.
The signal that cheating detection misses
Here's the uncomfortable part: fraud detection catches the person who didn't do the work. It doesn't catch the person who did the work and still isn't right for the job.
LinkedIn data says 92% of hiring professionals believe soft skills matter as much as — or more than — hard skills. 89% of bad hires are attributed to missing soft skills, not technical gaps. The person who passed your technical screen with AI help might genuinely have the technical knowledge but completely lack the judgment the role actually requires. The person who passed it without AI help might still be the wrong hire for the same reason.
Both failure modes point to the same fix: design assessments that test what AI can't answer for the candidate.
That means live and applied work over take-home assignments. It means unexpected follow-up questions that test whether someone can reason about their answer, not just produce it. It means role-specific scenarios — a judgment call under pressure, a diagnosis of an ambiguous problem — rather than questions that any sufficiently capable model can handle with a clean prompt.
Recent research found that AI voice agents conducting structured interviews yield 12% more job offers that convert to actual hires — with no decline in worker productivity. The reason: consistency in information collection. Human interviewers drift. AI interviewers follow the rubric. The quality of the process was the variable, not the technology.
That's the insight worth internalizing. A better-designed test gathers more signal. A poorly designed test — however aggressively proctored — just costs more to administer while teaching candidates to be more careful about getting caught.
The right role for AI in your hiring process right now
AI has changed what a candidate can produce on demand. That means the definition of a "hard skill" has shifted. For technical roles that literally involve working alongside AI tools every day, testing whether someone can avoid AI entirely doesn't measure anything meaningful. The real question is how well they direct it, evaluate its output, and catch its mistakes.
For roles that require independent judgment — the kind a model can't do by proxy — the assessment needs to be built around that. Live. Applied. Designed so the only way to pass is to demonstrate the actual capability.
The EU AI Act already prohibits emotion recognition in workplace assessments as of February 2025. New York City requires bias audits and candidate notice before automated tools run. The regulatory direction is clear: algorithmic flags should trigger human review, not automatic rejection. And 93% of hiring managers in a 2025 survey said human involvement remains essential even as AI handles more of the process.
That principle matters doubly when the flag might be wrong. Gaze-tracking systems misread candidates who look away to think. Audio models mishandle accents. Connectivity issues cause frame drops that look suspicious. Treating every flag as a verdict pushes out exactly the candidates skills-based hiring was meant to bring in.
What to actually do
Tighten identity checks earlier in the process — before the assessment, not after. Redesign take-home questions to require live follow-up reasoning, not just a final answer. Reserve your proctoring budget for high-stakes, late-stage rounds where it actually changes decisions.
And ask the harder question: if a candidate can pass your assessment with AI help, what are you actually testing for?
Most of the time the honest answer is: not much. Fix that first.