- Home
- The Thinking Wire
- The Portfolio Stopped Being Evidence. HR's New Control Is One Follow-Up Question.
The Portfolio Stopped Being Evidence. HR's New Control Is One Follow-Up Question.
Shraddha Sunil and Mudit Saraf interviewed 120 talent acquisition leaders and analyzed more than 6,000 screening sessions for Harvard Business Review in June 2026. Their sentence: “the ability to perform well in interviews is becoming infinitely scalable and practically free.” Both authors cofounded an interview-screening company, and the observation should be read with that commercial interest attached. Vlad Derdeicea, writing in UX Collective in July, flags the same conflict while citing them, which is the right instinct to copy.
Hiring ran for two decades on a division of labor. The artifact carried the proof: portfolio, case study, take-home, writing sample. The conversation confirmed what the artifact had already established. That split assumed one price structure, where the artifact was expensive to produce and the conversation was cheap to schedule. Both halves of that ledger inverted in about two years.
The artifact stopped discriminating
A polished case study now costs a candidate an afternoon and a subscription. The narrative arc, the before-and-after screens, the tidy metric at the end, all of it generates cleanly. Reviewers still read those documents as signal because the review process was designed when they were signal.
Derdeicea’s number for the resulting distortion: in an assessment instrument he ran with roughly 100 designers, four in ten profiles had what he calls Level and Stage telling different stories, and nearly half of those who self-reported senior assessed a full level lower. He sells that assessment, the sample was small and self-selected, and he says so explicitly: “instrument observations, not population statistics,” and the result is “a shape, not a verdict.” Take the hedging seriously. A shape is still enough to act on when the alternative is a document review that no longer separates anyone from anyone.
The distinction underneath the number is the useful part. Level is scope the organization granted you: the size of the surface you owned, the budget, the number of teams. It is readable straight off artifacts, because artifacts are made of scope. Stage is practitioner maturity, and in Derdeicea’s phrase it “never lived in artifacts.” It lives in what you considered and rejected, where you were wrong, what you would sequence differently now. The artifact was always a weak proxy for Stage. It just used to be an expensive weak proxy, and cost did most of the filtering.
Twenty hours on the artifact, one on its defense
The spend ratio he describes is the operational core of the piece: “Most designers spend something like twenty hours polishing the artifact for every hour rehearsing its defense.”
Invert that and the whole exercise changes on both sides of the table. Candidates who can defend a decision under unscripted pressure now have the scarce skill. Interviewers who only know how to walk a deck have lost their instrument. The uncomfortable corollary for hiring teams is that your interviewer bench, not your applicant tracking system, is now the constraint. Screening throughput went up. Verification capacity did not.
The two questions
Derdeicea’s probe is two questions, and they are worth reproducing verbatim because their precision is the point.
The first targets sequencing: “you sequenced X before Y, most teams do the opposite, what made you choose that order?” It requires the interviewer to have read the work closely enough to spot an unusual order, which is itself a real cost. It cannot be answered from a memorized narrative, because the narrative never contained the counterfactual.
The second is shorter: “what almost shipped instead?” His rule for reading the answer is the part I would put on a scorecard: “The best answer to what almost shipped is never nothing. It is a specific option, a specific reason, and usually a small regret.”
Both questions probe the same thing, which is the discarded branch. Generated work has no discarded branches. It has one confident path, and the confidence is uniform, which is exactly what makes it detectable in conversation and undetectable on paper. This is a narrower and more honest instrument than the volume-side screening we described in algorithmic monoculture in hiring, and it sits at the opposite end of the pipeline.
The take-home is dying from two directions
The first pressure is the one above: it is cheap to fake, so it stopped discriminating. Derdeicea reports hiring managers quietly abandoning the take-home in favour of reviewing messy Figma files a few days old, and telling candidates to volunteer the messy file. Process residue is now more credible than the finished deliverable, because residue is expensive to fabricate convincingly.
The second pressure is security, and it is easier to underestimate. A developer writing as citizendot inspected a recruiter’s take-home Python project and found malicious code hidden in a git hook, aimed at the machine that would run it. That is one first-person account with no independent corroboration, so treat it as an illustration rather than a prevalence claim. The structural point survives the sample size of one: a take-home asks a stranger to execute unvetted code from another stranger, in both directions. Recruiting has been operating a code-execution channel with no threat model. Any organization that would not accept that channel between two internal teams is accepting it with the outside world every week.
Two independent pressures pointing at the same artifact is usually how a hiring stage dies.
A control with a published error rate
The reason this is a governance story and not a recruiting-tactics story is what Derdeicea does next. He refuses the fraud-detection frame outright: “it is not about catching anyone,” and “Failing a room does not make you a fraud.” Then he states the control’s own failure mode: “Rooms are not truth machines. They are pressure tests.” And: “Rooms misfire. Mine did, seven years ago.”
Read that as a control specification. A probabilistic live control, with a known false-positive rate, replacing a deterministic document review that stopped discriminating. Any engineering leader who has tuned an alerting threshold recognizes the shape immediately. You do not deploy that kind of control without deciding in advance what a miss costs, who reviews the misses, and what evidence overturns the first read.
That is the discipline HR has to import, and it is the same discipline the cross-domain governance tooling deficit describes elsewhere: functions outside engineering are being handed probabilistic instruments with none of the surrounding practice. The problem also generalizes well past design hiring. Any function that assesses AI-assisted human output inherits it. Legal reviewing outside counsel work product. Analysts reviewing a model’s memo. Procurement reading a vendor’s implementation plan. In every one of those, the artifact got cheap and the interrogation did not, and the domain competence wall determines who is even able to run the interrogation.
Do this now
Pick your most artifact-dependent hiring stage this week and run three changes on it.
Rewrite the scorecard around the discarded branch. Add one required field: what almost shipped instead, and did the candidate name a specific option, a specific reason, and a cost they accepted? Score that field on its own. If your rubric currently scores the quality of the artifact, you are scoring the tool.
Assign the reading cost to a named person. The sequencing question only works if someone read the work closely enough to notice an unusual order. Budget that time explicitly, on one person’s calendar, before the interview. Without it, the probe degrades into a generic “walk me through your process.”
Retire or contain the take-home. If you send one, get your security team to say out loud who is responsible when a candidate’s machine is compromised by a repository your recruiters distributed. If you receive one, run it in a sandbox on isolated infrastructure, never on a reviewer’s laptop. If neither is realistic, ask for the messy working file and spend the saved hour in the room instead. Our operating-maturity read across hiring, product, and support covers where that hour is best reinvested.
Then write down your expected error rate before you start, and revisit it in ninety days. A control nobody has calibrated is a control nobody can defend when the first strong candidate fails the room.
This analysis synthesizes AI can fake your portfolio. It can’t fake the second question (UX Collective, July 2026) and I inspected my take-home interview project. It was a whole operation (citizendot, July 2026).
Victorino Group helps organizations redesign human-assessment controls for a world where the artifact is generated and the interrogation is the evidence. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation