ARR Broke in the AI Era: 13 Investors Say So on the Record

TV
Thiago Victorino
7 min read
ARR Broke in the AI Era: 13 Investors Say So on the Record

“Is it revenue or GMV? Is it recurring? Is it quarterly times four? Or is it 365 times daily revenue?” The questions come from Simon Wu’s July 2026 piece “ARR Doesn’t Mean What It Used To,” which catalogues what founders now mean when they say ARR. The piece is built on interviews with 15 institutional investors, 13 of them named with their firms: Emergence, Madrona, 8VC, FJ Labs, DFJ Growth, Footwork, and others. These are the people who price companies for a living, and they are saying, on the record, that the number at the top of every board deck no longer has a stable definition.

Most boards have not absorbed this. They still benchmark against ARR multiples, approve plans built on NRR targets, and compare their company to SaaS comparables whose metrics were computed under different rules. The metric names survived the transition to AI-era business models. The meanings did not.

The Definitions Drifted While Nobody Was Watching

ARR was a useful convention because everyone computed it the same way: contracted, recurring, annual. The investors Wu interviewed describe a market where that convention has dissolved. One company’s ARR is last quarter multiplied by four. Another’s is GMV wearing a revenue costume. A third extrapolates 365 days from yesterday’s usage. All three appear in the same comparables table, in the same column, under the same three letters.

This is not a rounding disagreement. A company annualizing its best quarter and a company reporting contracted annual commitments can show the same headline number while carrying entirely different businesses underneath. An investor who prices both at the same multiple is mispricing at least one of them. A board that benchmarks against a basket of such numbers is benchmarking against noise.

CARR gets it worse. One quote in the piece is blunt: “I don’t know why so many companies try to raise off CARR multiple instead of ARR multiple. It’s not a real number.” Contracted ARR was meant to capture signed-but-not-yet-live revenue. In practice, per the interviews, it now mixes signed bookings with unproven capacity, revenue the company hopes to serve, priced as if it were revenue the company already earns.

NRR has its own failure mode. The investors describe net revenue retention inflated by proof-of-concept conversions booked as expansion. A pilot that converts to a contract shows up as the existing customer “expanding,” and the retention metric that was supposed to measure durable customer love ends up measuring the sales team’s ability to graduate POCs. The metric still prints high. It just no longer means what the board thinks it means.

The Retention Data Explains the Mechanism

If the definitional drift were only sloppiness, the fix would be a glossary. RevenueCat’s August 2026 retention study shows the problem runs deeper: even a correctly computed AI-era metric misleads when read through SaaS-era priors.

The dataset is unusually solid for this debate: 3,519 AI-powered apps, subscription data from July 2024 through June 2025, minimum 100 subscriptions per app. One in four subscription apps is now AI-powered. And the AI cohort behaves differently on both ends of the curve. AI apps earn 41% more revenue in year one, a median first-year LTV of $30.16 against $21.37 for the rest. They also churn roughly 30% harder.

Read those two numbers together and the annualization problem becomes concrete. An AI product’s early revenue runs hot. Annualize a strong early quarter and you project a year the churn curve will not deliver. The same arithmetic that produced a fair estimate for a SaaS-era company produces an inflated one for an AI-era company, because the shape of the curve underneath changed. The multiplication is honest. The prior is wrong.

The spread inside the cohort is just as instructive. The study’s top cohort keeps 13.9% of subscribers after one year, against 1.4% at the other end: roughly a tenfold spread in one-year retention. Two pitch decks, same ARR slide, one business worth a multiple of the other. The headline metric cannot tell you which one you are looking at. Only the cohort curves can.

We wrote about Goodhart dynamics in AI adoption metrics when the gamed numbers were engineering-side: token counts, adoption percentages. The finance stack turns out to be running the same failure at higher stakes. When the board prices on a metric, the metric gets produced, whether or not the underlying thing it was supposed to measure comes along.

Diligence Is Following the Metrics Down the Stack

There is a companion signal from the acquisition side. An investor thread, which we could not read directly and know only from TLDR Founders’ description of it, argues that AI acquisitions now require code-level diligence: scoring model dependency, data moat, talent concentration, and substitution risk before assigning premium valuations. Treat that as a described claim rather than a verified source. But it is consistent with everything above. If the top-line metrics no longer discriminate between durable and hollow revenue, the diligence has to descend to the layer where the difference lives: what the product actually depends on, and how replaceable it is.

That descent has a mirror inside the company. We argued in the accountability piece that someone has to own the proof behind every AI claim a company makes externally. The metric drift extends that problem into finance. When ARR can mean four different things, the choice of which computation to present is itself a governed decision. Someone picks the definition. Someone signs the deck. If nobody in the company can state, in writing, which definition was used and why, the company is one diligence process away from an uncomfortable conversation.

This Is a Finance-Governance Problem Now

The instinct is to treat metric definitions as an accounting detail. The interviews suggest investors have stopped treating them that way. A number with a contested definition is a negotiation surface, and the party that controls the definition controls the price. That makes metric integrity a governance concern with a direct line to valuation, fundraising, and M&A outcomes.

The uncomfortable part for operators: the drift rewards you right up until it doesn’t. Annualized run rate reads better than contracted ARR. CARR reads better than either. Each definitional upgrade buys a better-looking deck today and invites a harsher diligence conversation later, because the investors quoted above have seen the trick. Wu’s interviewees are describing their own updated priors. If those priors spread, ambiguity itself gets discounted, and companies that kept strict definitions will be underpriced by their own conservative reporting unless they say so explicitly.

Do This Now

Before the next board meeting, produce a one-page metric provenance note. Four lines per metric:

  • ARR: the exact formula. Contracted only, or annualized? Over what window? Does anything in it resemble GMV?
  • CARR: what fraction is signed-and-live versus signed-and-hoped? If you cannot split it, stop presenting it.
  • NRR: how are POC-to-contract conversions booked? If they count as expansion, compute a second NRR without them and show both.
  • Retention: cohort curves, not blended averages. If your product is AI-powered, benchmark against the RevenueCat cohort shapes, not against SaaS-era curves.

Then have the CFO sign it, the same way an engineering leader signs an architecture decision. The companies that will raise well in the next cycle are the ones whose numbers survive the definition question, because Wu’s interviews show that question now gets asked.


This analysis synthesizes ARR Doesn’t Mean What It Used To (Simon Wu, Signals by Simon, July 2026) and AI App Retention Study (RevenueCat, August 2026).

Victorino Group helps leadership teams put governance behind the numbers they present, from metric definitions to the evidence that backs them. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation