- Home
- The Thinking Wire
- Five of Six Healthcare AI Benchmarks Give the Model Less Than a Real Patient Record
Five of Six Healthcare AI Benchmarks Give the Model Less Than a Real Patient Record
Five of six public healthcare AI benchmarks give the model less context than the median real patient record. Several provide fewer than 200 tokens per case, under 2.5% of that median. The figures come from Protege, whose CEO Bobby Samuels published them in August 2026 as a guest post in the a16z newsletter, in collaboration with Protege. It is vendor-authored material and should be read as an argument, not as neutral research. The argument is still the sharpest description of the problem I have seen this year.
Protege reports that across a random set of real patients in its data, the median patient record contains about 8,500 tokens and the mean about 39,000. The distance between mean and median tells you the shape: a long tail of enormously complex patients, and a benchmark suite calibrated to none of them. A model that scores well on a 200-token vignette has demonstrated one thing reliably, which is that it can answer a 200-token vignette.
The Missing Reference Point
We have written before about benchmarks being invalid rather than contaminated and about metrics that measure the artifact instead of the outcome. Both of those are software problems, and software problems share a comfort: a correct answer exists. The test compiles or it does not. The transaction reconciles or it does not. When a software benchmark is wrong, someone can in principle write a better one.
Medicine does not offer that. At the moment a clinical decision is made, there is frequently no established correct answer to grade against. Outcome truth arrives years later, entangled with adherence, comorbidity, luck and the counterfactual nobody ran. Grading against expert consensus substitutes a different problem, because expert consensus is thinner than procurement committees assume.
Protege’s sharpest evidence is a knee surgery case. Patient characteristics, comorbidities, facilities and year only explain 3.4% of the variation in the choice to perform partial or full. Add the identity of the surgeon and it immediately becomes 14.8%. Which means 77% of the explained variation in partial versus full comes down to who the surgeon was. Working from those two numbers, surgeon identity accounts for roughly 11.4 points of total variation, more than three times everything the clinical inputs explain combined.
That pattern is not isolated to knees. Across 15 operations with two variations, physician identity accounts for 7% to 77% of the explained variation. So when you build a benchmark whose answer key is what physicians chose, a large share of what you are encoding is the individual habits of whoever happened to be holding the scalpel.
This is a different failure from the one our software testing posts describe. There the reference existed and the test misused it. Here the reference does not exist at decision time at all. No amount of better test-writing produces one.
What a Healthcare Benchmark Score Actually Certifies
Put the two findings together and the meaning of a leaderboard number collapses to something narrow. It certifies that a model, given a fraction of the information a clinician would have, produced outputs that match a key assembled from decisions that were substantially idiosyncratic.
That is a statement about how the test was written. It is not a statement about clinical safety, and it cannot support the weight that procurement puts on it. If your vendor evaluation, your clinical sign-off, or your board risk register cites a benchmark score as evidence of safety, the evidence is theatre. This is the same governance vacuum we described in the MedVi analysis, now with a measurement layer that fails structurally rather than administratively.
The exposure is not hypothetical in scale. OpenAI reports that more than 300 million people use ChatGPT for health-related questions each week, a self-reported figure carrying its own “more than”. Protege also projects that soon nearly one-third of all SOAP notes will be written by AI. That is a projection about where clinical documentation is heading, not a measured current state, and it should be held loosely. Even discounted heavily, it describes a surface that no pre-deployment benchmark is positioned to certify.
Move the Control to After Deployment
If the gate before deployment cannot be trusted, the control has to sit downstream of it, in production, where outcomes are observable even when they are slow.
Protege points at one live signal from its own partner-network EMR data. Part way through 2026, about 1,686 per million notes recorded interactions where the patient wanted clarification on something a model said. Protege states these figures are descriptive rather than causal estimates, and being its own data on its own network, the number should be treated as an existence proof of the measurement rather than an industry rate. What matters is the shape of it. A clarification-seeking patient is a detectable event in a note, and detectable events can be counted, trended and alerted on. That is a control. A leaderboard score is not.
The reframe Protege offers for the whole field is worth quoting directly: “With the same level of healthcare spending, are we able to produce better health when we adopt AI technology more quickly?” That question is answerable, slowly and expensively, with instrumentation. It is not answerable with a benchmark.
The measures Protege proposes to replace benchmark scores fall into three families. Real-world health outcomes: life expectancy, quality-adjusted life years, nurse burnout, physician turnover. Delivery metrics: medical denials, medical debt, medical mistrust. And live misalignment detection aimed at over-eagerness and sycophancy, the two failure modes that a benchmark built from agreeable vignettes is least equipped to surface.
Some of those take years to move. That is the honest cost of the position. A control that reports in years is uncomfortable, and it is still better than a control that reports instantly and means nothing.
Do This Now
Open your AI vendor documentation and find every benchmark score cited as evidence of clinical performance. For each one, answer two questions in writing. How many tokens of patient context did the benchmark provide per case, and where did the answer key come from? Vendors who cannot answer the first question have not read their own evaluation. Vendors who answer the second with “expert consensus” or “physician-labelled” now owe you the inter-rater agreement, because the knee surgery numbers say that key may be encoding the labeller more than the medicine.
Then take the score out of the approval criteria. Replace it with two production instruments: a counter for patient clarification-seeking in AI-touched notes, and a review queue sampling AI-generated documentation against the record it was drawn from. Neither proves the model is safe. Both tell you when it stops being safe, which is the only thing a pre-deployment number ever pretended to do.
This analysis synthesizes The Oracle Problem: an invisible bottleneck to AI and medicine (Bobby Samuels, CEO at Protege, guest post in the a16z newsletter in collaboration with Protege, August 2026).
Victorino Group helps healthcare and regulated organizations replace benchmark-based AI approval with production instrumentation that survives audit. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation