How We Know a Model Is 'Best' (and Why You Cannot Inherit the Verdict)

TV
Thiago Victorino
7 min read
How We Know a Model Is 'Best' (and Why You Cannot Inherit the Verdict)

When Zvi Mowshowitz crowned GPT-5.6-Sol the default “workhorse” model in July 2026, he was explicit about how he got there. “The bulk of this is collecting a gestalt based on reactions.” He triangulated published benchmarks, dozens of user anecdotes, and his own hands-on use, then formed a judgment. No single number handed him the verdict. He assembled it.

That is an honest description of a good reviewer’s method. It is also a warning label for anyone about to treat the verdict as a procurement decision.

The Word “Gestalt” Is Doing All the Work

A gestalt is a synthesized impression. It is what an experienced practitioner produces when they have read the benchmarks, watched the reactions, and used the tool enough to feel its shape. It is genuinely valuable. It is also non-transferable.

Mowshowitz names his own blind spots. He weighs some signals more than others based on which he trusts. He discounts benchmarks he considers gamed. He upweights anecdotes from people whose judgment he respects. Two careful reviewers running the same inputs would land in different places, because the synthesis is editorial. That is the nature of the work, not a flaw in it.

The trouble starts when the output gets read as measurement. “Sol is the workhorse” sounds like a fact about the world. It is a well-reasoned opinion about the world, formed by one person weighing evidence in a way that suited one person’s needs. For a newsletter, that is exactly right. For a company standardizing its engineering org on a model, it is the beginning of the question, not the answer.

No Benchmark Agrees Either

If you hoped the underlying benchmarks would settle it, they do not. They disagree with each other.

Artificial Analysis puts its Intelligence Index for Sol at 58.9, with Fable scoring higher. WeirdML has Sol at 88.8% and Fable at 87.8%, nearly tied. VendBench 2 ranks both Fable and Opus above Sol. Three benchmarks, three different orderings. Pick the benchmark and you pick the winner.

This is not benchmark failure. Each measures a different slice of capability under different conditions. The disagreement is information: it tells you that “best” is not a property a model has, but a property of the match between a model and a task. A benchmark that mirrors your workload is worth more to you than a benchmark that tops a leaderboard on work you never do.

Which is why a leaderboard row cannot be your decision. It answers a question someone else asked.

”Best” Depends On Which Axis You Weight

Mowshowitz makes a point that should be printed above every model-selection meeting: “Capability in practice is multiplicative across intelligence, persistence, tools, latency, price, availability and supervision.”

Multiplicative means a zero anywhere zeroes the product. A brilliant model you cannot get rate-limited access to scores zero on availability. A capable model that needs constant supervision scores low on the axis that determines whether it saves you labor. The single-number ranking collapses all of these into one figure and hides the tradeoff you actually have to make.

Cost makes this concrete. On his cost-per-task figures, Sol runs $1.04, Fable $2.75, and DeepSeek v4 $0.04. If your task is high-volume and latency-tolerant, the model ranked lower on intelligence may be the correct choice by a factor of twenty-five on cost. “Best” flips depending on whether you weight the intelligence axis or the price axis. Nobody can weight those axes for you, because the weights come from your workload, your budget, and your tolerance for supervision.

The Reliability Caveat You Cannot Read Off A Score

Buried in the reactions is a detail no benchmark surfaces. Users reported that Sol “accidentally deleted almost ALL of my Mac’s files.” Mowshowitz passes along the practical advice that follows: “either sandbox it or make sure you have a path to recovery.”

Treat that as a user report, not a controlled finding. Even hedged, it points at something leaderboards structurally cannot capture: how a model behaves at the edges, with real tool access, under real autonomy. An intelligence score of 58.9 tells you nothing about blast radius when the model acts on your filesystem. The only way to learn that is to run the model on your tasks, in your environment, with your guardrails, and watch what it does when it is wrong.

That is the part enterprises keep trying to skip, and it is the part that produces the incidents.

Why You Cannot Outsource This

A reviewer’s gestalt is the compression of a lot of work into a short verdict. When you inherit the verdict, you inherit the compression and lose the work. You do not know which signals he weighted, which he discounted, or whether the tasks behind the anecdotes look anything like yours.

A leaderboard row is even thinner. It is one axis, one benchmark, one snapshot in time, on tasks chosen by the benchmark authors. Standardizing on it means letting a stranger’s task distribution decide your model.

Neither is a decision. Both are inputs. The decision is the synthesis you perform over inputs weighted for your own context, and it belongs to you the same way the reviewer’s synthesis belonged to him. The difference is that your weights are the ones that will be right or wrong for your bill and your incidents.

Do This Now

Build the smallest task-specific eval that can produce your own verdict. It takes a day, not a quarter.

Start by collecting ten to twenty real tasks from your actual work, not synthetic prompts. Pull them from closed tickets, support threads, code reviews, whatever your team actually ships. Write down, for each, what a correct output looks like and what a dangerous wrong output looks like.

Run your two or three candidate models on all of them, under the tool access and autonomy you would actually grant in production. Record more than pass or fail. Capture cost per task, latency, how many correction turns each needed, and every instance where a model did something you would not want it doing unsupervised. That last column is the one the benchmarks never give you.

Then weight the axes on purpose. Decide, before you look at the results, how much intelligence is worth relative to cost, latency, and supervision load for this workload. A high-volume classification job and an autonomous coding agent will weight them completely differently, and writing the weights down first stops you from rationalizing toward the model you already liked.

The output is a verdict you can defend, tied to tasks you can point at, with the failure modes documented. When the next model ships next month, you rerun the same tasks and get a new answer in an afternoon. You own the ruler. That is the whole difference between reading a review and making a decision.


This analysis synthesizes Better Call Sol: The Workhorse (Zvi Mowshowitz, July 2026).

Victorino Group helps teams build model-selection evaluations grounded in their own tasks, not a leaderboard. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation