- Home
- The Thinking Wire
- The Threshold You Trust Was Not Measured on You
Tammy Everts pulled a month of real user data from ten leading online retailers and plotted session-level Largest Contentful Paint against bounce rate, one chart per site. Across all ten, the optimal LCP, meaning the point tied to the best bounce rate, landed somewhere between 100 milliseconds and 1 second. Google’s threshold for a “good” LCP is 2.5 seconds. Every one of those ten sites saw its best engagement well before the number they were being graded against.
That result is worth sitting with, because 2.5 seconds has become the de facto finish line for a lot of teams. Hit it, get the green checkmark, move on. The threshold is not dishonest and it is not gamed. It is an aggregate across millions of sites, which is exactly what it claims to be. It tells you what is fast in general, and it knows nothing about your users, your product, or your business.
The failure runs in two directions at once
Inside a sample of ten, the same number failed in two opposite directions.
Everts also measured where each site’s performance plateau starts, the point where getting faster or slower stops making any difference to bounce rate. Across the ten sites, that plateau began anywhere from 500 milliseconds to 5.6 seconds. Six of the sites had plateaus that began after 2.5 seconds, so passing Google’s bar left real engagement unclaimed. Those teams were being held to a bar that was too slow for their own users, and the compliance dashboard said they were fine.
For the other four, the plateau started before 2.5 seconds. By the time those sites met Google’s recommended threshold, they had already flattened out. Everts puts it plainly: even though those four sites are technically fast enough for Google, they have bottomed out in terms of bounce rate. The metric they were optimizing had stopped predicting the outcome they cared about, and the dashboard said they were fine there too.
One number, ten sites, two opposite failure modes, and both invisible if the only question you ask is whether you passed. A team on the slow side of the plateau underinvests and never learns it. A team on the fast side keeps buying milliseconds that no longer move anything, and books the spend as performance work.
The author refuses to let her own range become the next borrowed number
Everts is emphatic about this before she gets going: this is not a new number for you to chase. She repeats it at the close. Consider the research a methodology, not a number. She is also explicit that she is not arguing against Core Web Vitals, which she says absolutely matter.
That restraint is the whole transferable part. Someone who takes 100ms to 1s and staples it to a Jira epic has reproduced the original mistake one level down, with a sample of ten standing in for millions. What survives the transfer is a test you run on your own data:
- Does this threshold have a falsifiable relationship to an outcome we measure ourselves?
- Where does our own curve flatten?
Two questions. Neither requires a vendor.
What the evidence does and does not support
The limits matter, and they are worth stating before anyone quotes the range at me. Ten sites is ten sites. One month of data. Retail only, so the shape of a B2B dashboard’s engagement curve is not covered. The sample is drawn from the vendor’s own RUM customer base, which the page discloses, and that population skews toward organizations already investing in performance measurement.
Everts also discloses the metric substitution: she used bounce rate rather than conversion rate because most RUM tools capture it by default, so no extra instrumentation is required, and she treats it as a good directional proxy. Directional is the honest word for it.
The per-site figures live inside ten chart images and are not extractable as text, so the aggregate ranges above are the citable numbers. There is no per-retailer figure to quote, and inventing one would repeat the original error in a smaller font.
The argument survives all of that, because it rests on a structural observation: ten sites measured individually disagreed with the aggregate in both directions. Establishing that a population statistic fails to describe individuals takes one clean counterexample. This is ten of them.
Now count how many of your governance metrics are borrowed
Read the LCP result as a worked example of a general class, and the class is uncomfortably large in AI governance.
“Percentage of code written by AI.” Every published version of this number comes from someone else’s codebase, someone else’s language mix, someone else’s review culture. We have argued that the measurement axis should be the class of author and reviewer rather than a share of lines, and the borrowed-threshold problem is a second reason. If a vendor tells you 40% is where the returns are, ask which outcome that 40% was correlated against, and on whose repos. Then ask where your own curve flattens. Most organizations cannot answer the second question, which means the first answer is being applied blind.
Coverage floors. An 80% coverage target is a rule of thumb inherited from projects with different defect profiles and different blast radii. Somewhere in your codebase there is a module where 60% catches everything that ever breaks, and another where 95% still ships bugs monthly. The single floor treats those as the same problem. It is the four-sites case: past a point, the number stops predicting escapes, and everyone keeps writing tests against the number.
CI cost heuristics. “Keep pipelines under ten minutes” is a plateau claim wearing a threshold’s clothes. Our own cost work found that the dominant cost driver was job count, because GitHub rounds each job up to a full minute, which makes the received wisdom about duration a poor predictor of the bill. The rule of thumb was measured on a different billing shape than ours.
Vendor SLAs and model thresholds. We have written about thresholds inside model pricing decisions and about benchmarks losing meaning through contamination. This is a third and quieter failure. The threshold is honest, uncontaminated, and simply measured on a different population than the one applying it. No one behaved badly. The number still misleads.
Do this in the next two weeks
Pick your single most load-bearing threshold. The one that appears in a QBR slide, gates a release, or justifies a budget line. Then run Everts’s four steps against it:
- Pull your own data for the metric, at whatever granularity you already have. Session level if you have it, per PR, per deploy, per job.
- Plot it against an outcome you actually care about and already capture. Do not instrument something new for this. The point of her bounce-rate choice was zero extra instrumentation.
- Find where the outcome is genuinely at its best, and find where your curve flattens.
- Set the goal on those two points, not on a number derived from millions of systems you do not run.
If the chart shows your plateau starts before your current threshold, you have found spend you can stop. If it starts after, you have found upside your dashboard has been hiding behind a green checkmark. Either result is worth more than the compliance status you have today, and you should never assume that meeting an external threshold is helping your business.
Produce the chart first. Whatever threshold you write down afterwards will at least be yours.
This analysis synthesizes New research: The Core Web Vitals thresholds you trust might be wrong for your site (Embrace / SpeedCurve, July 2026).
Victorino Group helps engineering organizations replace borrowed metric thresholds with ones measured on their own telemetry. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation