84% Still Felt Productive While Their Experience Collapsed

TV
Thiago Victorino
7 min read
84% Still Felt Productive While Their Experience Collapsed

Eighty-four percent of developers said AI improved their productivity. Six months later, the same 84% said it again. More than three-quarters of the matched cohort gave the identical positive rating at both timepoints, which is about as stable as survey data gets.

Over that same window, the share of those developers describing their experience of the work as worse climbed from 14% to 27%.

The two numbers come from the same people, answering the same instrument, six months apart. Annie Vella ran the study as Masters research at the University of Auckland under Kelly Blincoe, and published the correlation between change in flow state and change in perceived productivity: 0.02. Her own description is that this is about as close to zero as it gets.

That single coefficient is the finding. Everything a manager currently reads off a dashboard is on the wrong side of it.

Cross-Sectionally They Agree, Longitudinally They Separate

At any single moment, developers who feel productive also report a better experience. Take a snapshot today and the two variables track each other well enough that measuring one tells you something useful about the other. That relationship is why output has been a tolerable proxy for team health for thirty years. Ship rate up, morale probably fine.

Watch the same people move through time and the relationship dissolves. A developer whose flow state degraded between October and April was, on average, exactly as likely to report improved productivity as one whose flow state improved. The two series were both moving. They stopped moving together.

Snapshot correlation and change correlation are different claims about the world, and organizations have been buying the first while believing they own the second. Every engagement survey that asks “do you feel productive with AI?” once a quarter and reads a flat line as stability is measuring the variable that stayed flat by design.

The reliability of self-reported AI productivity has been picked apart before, in the two percent problem and in the pinhole view of AI value. Vella’s contribution sits one layer past that argument. Grant the perception every benefit of the doubt, accept 84% at face value, and the number still fails as a proxy, because the thing it used to travel with has walked off in another direction.

Flow State Was the First Thing to Go

Inside the matched cohort, the dimension that moved hardest was flow. Developers rating their flow state as worse went from 7% to 20%, close to a tripling and the largest single shift in the study. Cognitive load worse moved from 6% to 9%. Feedback loops worse actually improved, from 3% down to 1%.

The mechanism shows up in what participants wrote. One described “less in the zone time, because while code is being produced, I now find myself switching context more often.” Another put both halves in one sentence: “This does break the flow, but still speeds up the overall process.”

Read that second quote as an honest cost accounting rather than a contradiction. The developer is trading a state they value for a throughput they also value, and the trade registers as net positive on the question the survey asks about productivity. The cost lands somewhere the productivity question has no room to record it.

Vella’s most uncomfortable detail is what happened to the anchor of the productivity judgment itself. At timepoint 1, the dimension most aligned with feeling productive was flow state. By timepoint 2 it was speed of the feedback loop, and flow state had almost dropped out of the picture. Developers did not simply lose flow. They recalibrated what productive means so that flow stopped counting.

The Cohort Transitions Are Worse Than the Averages

Aggregate percentages hide direction of travel. The transition table does not.

Of the developers who started fully positive about their experience, only 37% were still fully positive at the end. Nobody who started negative came back to fully positive. The flow is one-directional across six months, and the average holds steady only because the pool being drained is large enough to absorb it for a while.

That is the shape of an erosion curve, and no output metric has the vocabulary to describe it. Velocity is a rate. Attrition is an event. The thing between them, a slow migration from fully positive to partially positive to negative with no return path, is invisible to both until it converts.

We argued in AI severed the link between output and competence that the artifact stopped carrying evidence of the producer’s skill. This is a second severing on a different axis. The artifact also stopped carrying evidence of what producing it cost.

The Management Side: Who Signs, Not Who Produces

Karim Jedda’s essay on engineering management after the cost of code collapsed is the operational companion to Vella’s numbers. His central line: “At every level of the org, the work that survives is the work someone has to sign.”

His verdict on the standard dashboard is blunt. Velocity, PR count and tickets closed have become “actively misleading,” because volume is now the cheapest way to raise them. Adding AI-specific metrics makes it worse rather than better. Acceptance rates and prompt counts, in his phrase, “solve the wrong problem.”

He is equally direct about the review layer many teams are installing as the answer. On AI reviewing AI: “the checker shares training data, biases, and blind spots with the generator. Self-review catches the typo. It does not catch the shared misunderstanding.” The net effect he predicts is “fewer dumb errors, more systemic ones, because high-volume plausible output now passes high-volume plausible review.”

Jedda’s useful split is between mechanical and semantic checking. Mechanical checking covers types, tests, contracts, invariants, and canary metrics, and its cost is collapsing along with the cost of code. Semantic checking asks whether this system should exist in this shape at all, and it stays expensive because it resists both compression and delegation. His projection: “In the agentic limit, the chart stops recording who produces and starts recording who signs.”

He also draws the boundary on where gains are real. They “show up clearly in greenfield work, boilerplate, and unfamiliar territory. They fade or invert in deep work on systems the engineer already understands.” That inversion in familiar systems is the same territory as the agent slowdown reckoning, and it maps cleanly onto Vella’s flow finding: deep work on a system you know is precisely where interruption costs the most and shows the least on a throughput chart. The junior-pipeline damage, Jedda notes, “breaks on a delay, so you will not notice for three to five years.”

Jedda contributes no statistics, and he discloses that Gemini 4 assisted with his editing. Treat him as a management frame, with Vella supplying the evidence. His signing frame also rhymes with an argument we have made about what AI deletes from accountability: a signature is only worth something when the signer can be located afterwards.

Copy the Study Design, Not the Headline

The transferable asset in Vella’s work is the instrument, and it costs almost nothing to run internally.

Ask your engineers to rate four dimensions, separately from any productivity question: flow state, cognitive load, feedback loop speed, and confidence in the code they ship. Use a simple better/same/worse scale. Do it twice, six months apart, with the responses linked to the same individuals so you can build the transition table. Anonymity survives if you hash the identifier; the matching is what generates the finding.

Then read the change, and only the change. A flat aggregate proves nothing here, as Vella’s own 84% demonstrates. What you want is the count of people who moved from fully positive to anything else, and the count who moved back. If the return path is empty, you have the same one-directional erosion in your own organization, and you found it in the window where it is still cheap to address.

One caveat carries into every use of these numbers. Vella’s data window runs October 2024 to April 2025, and she flags it herself as “in AI terms already quite a long time ago,” best read as a snapshot of that window. This is longitudinal survey research on a matched cohort, not a controlled experiment. It establishes that the two variables came apart in a real population over real time. It does not establish the magnitude for your team, which is exactly why running the instrument yourself beats citing hers.

A velocity drop announces itself in the next sprint report. This erosion announces itself in a resignation letter, three quarters after the point where you could have done something cheap about it. The instrument that separates the two fits on one screen and takes an engineer four minutes to fill in, twice a year.


This analysis synthesizes The Productivity-Experience Paradox (Annie Vella, University of Auckland, July 2026), Engineering management after the cost of code collapsed (Karim Jedda, July 2026).

Victorino Group helps engineering organizations instrument developer experience alongside output, so AI rollouts surface erosion before it converts into attrition. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation