- Home
- The Thinking Wire
- The Instruments Disagree: Telemetry and Surveys on AI's Impact
The Instruments Disagree: Telemetry and Surveys on AI's Impact
Sixteen experienced open-source developers forecast that AI assistance would make them 24% faster. Then METR timed them on 246 real tasks in mature repositories they knew deeply, under randomized controlled conditions. With AI allowed, they took 19% longer. After finishing, with the slowdown already sitting in the timing data, they still estimated AI had sped them up by about 20%. The distance between belief and measurement was roughly 39 points, on the same people, doing the same work.
That one trial reframes every AI productivity number you have seen this year. On AI’s impact, perception is not a lagging indicator of reality; right now it is an inverted one. The developers were not lying and were not careless. They experienced the tool as help while the clock recorded it as drag. And the industry’s two main measurement instruments have now split along exactly that line.
Two Instruments, Two Stories
DORA surveyed roughly 5,000 technology professionals in mid-2025 and concluded that “AI’s primary role is as an amplifier, magnifying an organization’s existing strengths and weaknesses.” For a well-run organization, that reading is reassuring. Strong foundations should mean AI compounds what you already do well.
Faros AI’s 2026 report drew on telemetry instead: two years of delivery data from 22,000 developers across more than 4,000 teams. Its finding contradicts the amplifier thesis at the point where the thesis matters most. In the telemetry, organizations with strong pre-AI engineering practices show the same downstream degradation as organizations without them. Maturity protected no one. The report positions itself as a direct counterpoint to DORA and compresses the method dispute into one sentence: “perception lags reality, while telemetry does not.”
Both projects are serious, and the disagreement is not a matter of rigor on either side. It comes from what each instrument can see. A survey samples what respondents know and feel on the day the question arrives. Telemetry samples what the delivery system actually did. When the two diverge, the divergence itself is information, and METR’s trial tells you which direction deserves the benefit of the doubt.
What a Survey Cannot See in Time
Consider one number from the Faros dataset: pull requests merged without any review rose 31.3%. No developer feels that number. Each unreviewed merge feels like a reasonable exception in the moment, a small PR, a trusted teammate, a deadline. The consequences (defects escaping, knowledge concentrating in fewer heads, review discipline eroding) surface months later. By the time a survey can register them, it records a vague unease with no address and no date.
The quality findings behind that statistic are the subject of a companion piece, the acceleration whiplash and the verification job. This piece is about the instrument, and the instrument problem extends beyond any single report.
Stack Overflow’s 2025 developer survey shows sentiment itself moving against usage. 84% of developers use or plan to use AI tools, up from 76% in 2024. Over the same period, favorability fell from above 70% in 2023 and 2024 to 60%, and 45.7% of respondents actively distrust the accuracy of AI output, against 32.7% who trust it. Usage climbs while trust falls. Whatever sentiment is tracking, it is not a stable proxy for value delivered, because it cannot even hold a stable relationship with adoption.
METR’s limits deserve stating plainly. The sample was 16 developers, all experienced, all working on large mature codebases they knew intimately, which is close to a worst case for AI assistance. METR itself cautions against generalizing the 19% slowdown to other contexts. What survives every caveat is the perception mismatch: the same individuals, with complete knowledge of their own workday, misread the direction of the effect before the tasks and again after them. Self-report failed on the easiest possible question, “did this make me faster,” asked of the people best positioned to answer it.
None of this indicts surveys at what surveys are for. DORA measures culture, satisfaction, and burnout with a decade of methodological care, and those are real outcomes that telemetry cannot reach. The failure mode is narrower and more consequential: treating sentiment as a proxy for delivery outcomes, then making delivery decisions on it.
Apply the Skepticism Symmetrically
Faros sells engineering-intelligence tooling. Its report profits from the conclusion that surveys are inadequate and telemetry is indispensable. That interest does not make the finding wrong; METR’s trial is independent, pre-registered in design, and points the same way. It does mean the vendor’s numbers deserve exactly the scrutiny the vendor applies to surveys. You cannot audit the 22,000-developer dataset, the team selection, or the metric definitions behind the headline percentages. An unauditable telemetry claim is a survey with better branding.
The way out is the same on both counts: own your instrumentation. The signal that matters already lives in systems you run. PRs merged without review is a query against your own git host. Defect escape rate, revert frequency, review depth, and time from merge to incident are all countable from your own tooling, under definitions your own engineers wrote and can defend. A purchased dashboard replaces one instrument you cannot audit with another one you cannot audit.
This is the same governance argument as the accountability chain question of who signs and who proves. Decisions that move headcount and budget need evidence someone can stand behind, and a sentiment score carries no signature. It also rhymes with the classifier threshold nobody signed off on: the choice of measurement instrument is itself a governance decision, and today most organizations are making it by default, with whichever number arrived first in a slide deck.
Do This Now
Three moves, in order.
-
Inventory the AI decisions currently resting on sentiment. Headcount plans, tool renewals, mandates to expand AI authorship into new areas of the codebase. For each, write down the evidence behind it. If the evidence is a survey score or “the team feels faster,” flag it. METR’s subjects felt 20% faster while running 19% slower; that class of evidence is now formally unreliable for this question.
-
Stand up two or three delivery metrics from systems you already own. Review coverage on merged PRs is the fastest to build and the most diagnostic, given the 31.3% drift Faros observed at scale. Baseline them before expanding AI authorship further, because a baseline recorded after the change proves nothing.
-
Run the perception check deliberately. Survey your team on how much faster AI makes them. Measure the same tasks with your own telemetry. The size and direction of the mismatch is your local version of METR’s 39 points, and it tells you precisely how much to discount self-report in your next planning cycle.
When your telemetry and your survey agree, you have learned the sentiment channel is calibrated, which is worth knowing. When they disagree, believe the instrument that timed the work, and go find out why the people doing the work cannot feel it yet. Every study above converges on the same warning: by the time the feeling catches up with the fact, the decision window has closed.
This analysis synthesizes AI Engineering Report 2026: The Acceleration Whiplash (Faros AI, April 2026), Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR, July 2025), State of AI-assisted Software Development (DORA / Google Cloud, September 2025), Developer Survey 2025: AI (Stack Overflow, July 2025).
Victorino Group helps teams build delivery instrumentation they own, so AI decisions rest on measured outcomes instead of inverted sentiment. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation