- Home
- The Thinking Wire
- 38.5% Were Flagged for Cheating. 61% of Them Got Hired Anyway.
38.5% Were Flagged for Cheating. 61% of Them Got Hired Anyway.
A study cited by Design Buddies looked at 19,368 live interviews between July 2025 and January 2026 and found that 38.5% of candidates were flagged for using AI to cheat. For technical roles the number reached 48%. The figure tripled in three months. Those numbers alone would make a decent panic post.
The number that matters is the fourth one: 61% of the candidates flagged for cheating still passed.
Detection worked. Something in the screen registered the cheating, the flag was recorded, and the pipeline moved the candidate forward anyway. That is a different failure from a screening layer that cannot see. A blind screen is an engineering problem, and engineering problems get budget. A screen that sees and is overruled is an authority problem, and authority problems get meetings.
The Signal Exists and Nobody Owns It
The Design Buddies post, written by guest author Carl Wheatley, a design recruiter with more than ten years in the field, carries a companion statistic that explains the 61%: only 19% of hiring managers feel very confident their current process would catch someone faking it. Note the wording. Not “confident”. Very confident. The remaining 81% sit below that bar, anywhere from merely confident to openly resigned.
Read those two numbers together and the mechanism becomes obvious. A flag arrives on a candidate. The hiring manager has no calibrated sense of what the flag means, no idea of the false-positive rate, no second signal to corroborate it, and a role that has been open long enough for the pressure to point one way. The flag is noise until proven otherwise, and nothing in the process is built to prove it. So the candidate advances.
Wheatley also reports that 72% of recruiters say they have already run into fake resumes, portfolios, or credentials made with AI. Six of the figures in his post carry no named source, so treat them as one practitioner’s read of the field rather than settled research. The direction is consistent even if the decimals are not: the artifact side of hiring has stopped carrying information, and everyone knows it, and the process has not changed.
We argued a version of this in the second question, when the portfolio became the cheapest object in the pipeline. What is new in the 2026 data is the sequencing. This is no longer a story about detection lagging generation. Detection caught more than a third of the room. The decision layer discarded what detection found.
Demand Is Not the Story
The instinct when hiring data looks this bad is to assume the roles are disappearing. The demand numbers say otherwise.
In a Figma survey cited in the same post, 82% of hiring managers said their need for designers has gone up or stayed the same over the past year, and 40% plan to open more design roles in the next six months. Autodesk’s 2026 AI Jobs Report found that “design skills” is now the single most requested skill in AI-related job postings, ahead of coding and cloud. Designer Fund, a VC firm, found design job postings across its portfolio companies rose about 60% in 2025. That last figure describes a portfolio, not a market, and should be read as a signal from one investor’s cohort rather than an industry rate.
So the pipeline is full at the top and the filter in the middle has stopped filtering. That combination produces a specific outcome: volume converts into hires, and the quality distribution of those hires is set by whatever survives an unreliable screen. Nobody feels the cost on hire date. It shows up later as rework, as a senior spending real hours correcting output that looked finished, as the quiet reassignment nobody writes a postmortem for.
Assess the Process, Because the Artifact Is Free
Wheatley’s practical line is the one worth stealing: “If you hide the tools you used, you’re not giving them anything to judge.”
That inverts the default candidate instinct, which is still to present the finished artifact and stay silent about how it was made. The artifact costs almost nothing to produce now. The decisions behind it do not. Which constraint did you discover first, and what did you throw away because of it? Where did the model give you something plausible that you rejected, and what tipped you off? Which part of this did you do by hand, and why was it worth the hours?
Anyone who has built an agent evaluation harness has already run this argument to its conclusion. You cannot grade an agent on its output when the output space is generated cheaply and looks correct by construction. You grade the trace: the tools it called, the checks it ran, the branches it abandoned. Hiring has arrived at the same wall from the other direction. Same problem, same available answer, different department. HR is now doing evaluation design, whether or not anyone has told HR that.
The mechanics are not exotic. Ask for the working session, not the deliverable. Ask for a decision log. Run the exercise on a problem the candidate cannot have pre-generated, then spend the interview on the reasoning rather than the artifact. Every one of these is more expensive per candidate than reading a portfolio. That is the point. The cheap screen stopped working, and the honest response is to pay for a screen that does.
The Rung That Feeds the Reviewers
The second-order effect is worse than the hiring one, and I have not seen anyone put a price on it.
A Harvard study that looked at 62 million workers found junior hiring dropped 8 to 10% at companies that started using generative AI. The mechanism reported is mostly not layoffs. Those companies quietly stopped opening junior roles. No announcement, no severance line, nothing to report. The rung was removed rather than cut.
Now connect it to the paragraph above. The screen that works is a screen run by someone who can look at a trace and tell whether the reasoning is real. That capability is not a certification. It is accumulated by doing the work badly, being corrected, and doing it again, which is exactly what the junior rung was for. Remove the rung for long enough and the population of people qualified to evaluate the next generation shrinks, at the same moment the volume of AI-assisted candidates makes evaluation more labor-intensive.
This is a supervision capacity problem wearing a jobs-report costume. The team that stops hiring juniors in 2026 spends a future reviewer bench to save this quarter’s headcount, and the invoice arrives in a year when nobody remembers making the decision. We traced the market-side version of this in what agents are doing to the job market; the internal version is quieter and lands harder, because a company cannot buy the capability back at market rate when everyone needs it in the same quarter.
There is a related trap on the intake side. The algorithmic monoculture problem means the automated filters everyone runs are converging on the same shortlist. Pair a converging filter with a screen that gets overruled and you get the worst of both: less variety entering the funnel, and less discrimination once it is in.
Do This Now
Pull your last twenty hires and answer one question: how many were flagged during screening, and how many of those were hired anyway? If your ATS cannot answer that, you have found the real problem, which is that the flag is not a field anyone is accountable for.
If it can answer, and your number looks anything like 61%, the fix is not a better detector. Give the flag an owner and a required disposition. Whoever advances a flagged candidate writes one sentence explaining why, and that sentence lives in the record next to the hire. Nothing else changes. The cost is a text box and a sentence per candidate.
Do it before the reviewers who could have caught it stop being hired.
This analysis synthesizes Hiring in the Age of AI in Design (Design Buddies, guest author Carl Wheatley, September 2026).
Victorino Group helps engineering and design organizations rebuild evaluation for work that is cheap to generate and expensive to verify. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation