The Four Ways to Measure an AI Productivity Claim, and Why Three of Them Fail

TV
Thiago Victorino
7 min read
The Four Ways to Measure an AI Productivity Claim, and Why Three of Them Fail

Figma took 100 people, gave half of them an AI tool and gave the other half nothing, and measured how long three standardized design tasks took. Overall design work got 20% faster and 16% easier with Figma Make. The result that should stop a leadership team, though, is the split: product managers saw tasks run 23% faster and 37% easier, a larger gain than the designers got. In Figma’s own summary, a product manager using Make was almost as efficient as a product designer who did not use Make.

That direction is backwards from what most adoption plans assume. Rollouts start with the specialists because the tool speaks their language. The trial says the marginal hour lands somewhere else.

What makes the finding worth an article is that no other measurement instrument available to a normal engineering organization could have produced it. When your board asks for proof that an AI purchase saved time, you have four instruments to choose from. Three of them break on contact with the question.

Instrument 1: The online A/B test

This is the default reflex, because product teams already have the infrastructure. Split users, ship the feature to one bucket, compare aggregate metrics.

The disqualifying condition: an AI assistant is not a discrete feature with a bounded blast radius. A designer with Make in one bucket collaborates with a PM in the other bucket, reuses artifacts produced with the tool, and changes how they scope work before they even open the tool. The treatment leaks. Aggregate cycle-time metrics also move for a dozen unrelated reasons in any given quarter, and an online test gives you no way to hold task difficulty constant. You measure the quarter, not the tool.

Instrument 2: Propensity score matching

The observational fallback. Find people who adopted the tool, find similar people who did not, statistically match them on observable characteristics, compare outcomes.

The disqualifying condition is stated in the method itself: PSM assumes there are no unobserved confounders. Every variable that drives both adoption and performance has to already be sitting in your logs. Think about who adopts an AI tool first inside your company. They are the people with slack in their week, higher baseline tooling fluency, less legacy code to respect, and a manager who tolerates experiments. None of that is a column in your telemetry. You will match on tenure and team and role, declare the groups comparable, and measure the personality trait that produced adoption in the first place.

Instrument 3: Instrumental variables

The econometric escape hatch. Find a variable that pushes people toward adoption but has no direct effect on the outcome, and use it to isolate causation.

The disqualifying condition is availability. A valid instrument has to be genuinely as-good-as-random with respect to productivity, and log data does not contain one. Staggered license rollout is the usual candidate, and it fails immediately, because rollout order inside a company is decided by exactly the factors that predict performance. Regional or seat-cost variation gets contaminated the same way. Teams that go looking for an instrument in existing data usually end up defending the instrument rather than the finding.

Instrument 4: The randomized controlled trial

Which leaves one option, and it is the one people skip because it sounds academic and expensive. Figma’s version fits in a paragraph.

The design was small and specific. 100 participants, 50 product designers and 50 product managers. Treatment group got Make. Control group got no AI tool at all, which is the comparison a board actually cares about. Three standardized tasks: converting an interface to dark mode, adding an option to a settings menu, building a comment fly-up.

The interesting engineering is in the confounder control. Figma redesigned the tasks three times after internal pilots to reach comparable difficulty for both roles, because a task that is trivial for a designer and impossible for a PM produces a role effect masquerading as a tool effect. Moderators followed a common script and a shared troubleshooting guide, so the quality of moderation did not become the variable under test. Analysis combined hypothesis testing with OLS regression. On the strength of that design, Figma states plainly that Make causes the improvements it measured.

The full published numbers: a 20% reduction in cumulative time to task completion, a 16% improvement in task ease, and a 15% improvement in perceived Figma usability. For PMs specifically, 23% faster and 37% easier.

What the weaker instruments would have hidden

Run this same question through PSM and imagine the output. Designers adopt Make first, because it lives inside a tool they already use eight hours a day. PMs adopt late and sparsely. Your matched comparison is built almost entirely out of designers, and the PM effect, the largest one in the study, never appears in the results. You would have reported a real 20% number attached to completely wrong guidance about where to spend the next license budget.

The RCT also surfaced a second inversion worth stealing. Designers gained on the hardest task. PMs gained on the two easiest ones. Any rollout plan built on the assumption that AI pays off most on hard work would allocate seats in exactly the wrong direction for half the population.

State the conflict out loud

Figma studied Figma’s product and published the result on Figma’s blog. The conflict of interest is real, unmitigated, and does not go away because the method is good. There is no results table on the public page. The numbers quoted above are the entire public numeric surface, with the detail behind a gated report.

That is a reason to treat the effect sizes as directional rather than portable, and it is precisely why the method matters more here than the number. You cannot verify Figma’s 20%. You can copy the design that produced it and generate your own.

We have made the same argument about Anthropic’s 80%-of-code claim and about what a completed task actually costs. The pattern repeats: the vendor’s number is unfalsifiable from outside, and the only durable asset is the measurement design. We covered Figma’s own adoption numbers in the design tipping point piece. This one is about the instrument.

Borrow the effect size, then size the trial

One detail in Figma’s write-up is the cheapest thing to copy and the most often skipped. The sample size came from a power analysis using effect sizes reported in the published GitHub Copilot randomized trials (arXiv 2302.06590).

You do not need a pilot to know how big your study must be. Someone has already run a randomized trial on a comparable AI coding or design tool and published the effect size. Feed that into a power calculation and it tells you the participant count. For a within-organization effect of this magnitude, 100 people got Figma a result they were willing to attach a causal claim to. Most companies assume a credible trial needs thousands. It needs a defensible design and two afternoons of scheduling.

Do this now

Pick the single loudest AI productivity claim currently circulating in your company, the one already inside a budget deck. Write down which of the four instruments produced it. If the answer is “adoption dashboards” or “a survey of how much time people feel they saved,” it is not evidence and should not be in the deck.

Then scope the smallest trial that could falsify it. Two standardized tasks, not three. Thirty people per arm as a starting point, adjusted by a power analysis borrowed from a published RCT on a similar tool. Control arm gets no AI tool. Fix the tasks by piloting them until both populations find them comparably hard. Write down your predicted direction of effect before the first session, because the value Figma extracted came from the prediction being wrong.

Budget one week. The output is a number your CFO can defend and a rollout order that is not built on an assumption about who benefits.


This analysis synthesizes Measuring time savings from Figma Make (Figma, Remy Stewart, August 2026).

Victorino Group designs and runs measurement trials that turn AI productivity claims into numbers a board can defend. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation