- Home
- The Thinking Wire
- The Negative Tail Moved: AI and 117,816 Hotel Reviews
Italy banned ChatGPT for four weeks in April 2023. For researchers, that regulatory accident produced something almost impossible to buy: a control group of people writing without AI, at the same moment, about the same kind of thing, as people writing with it.
Weilong K.M. Gu and Martin Spann of LMU Munich used that window across 117,816 Tripadvisor reviews covering 4,909 hotels in Europe, and published the result in the International Journal of Research in Marketing in June 2026. The headline effects are modest. Review volume rose 2%. Review length rose 9%.
Stop at those two numbers and generative AI looks like a mild writing aid for people who were going to post anyway. The interesting part sits in the shape of the distribution, and specifically at its left edge.
Where the effect actually landed
The averages absorbed almost nothing. The extremes did the moving. Reviews at the bottom of the scale, the one and two star complaints, came out measurably more negative in content when AI was available. The paper’s own illustration is blunt: “I don’t like the room” becomes “I hate the room!!”
The review got longer without gaining evidence. Same grievance, volume turned up.
The second shift is the one that should worry anyone who reads customer feedback for a living. The content became more subjective and less objective. A reviewer writing unassisted produced something like “breakfast was local and no continental option was available.” A reviewer with AI assistance produced “I didn’t like the breakfast.” The first sentence is a fact a hotel can act on. The second is a verdict a hotel can only absorb.
Gu and Spann note that this combination makes negative reviews more harmful. It follows. A specific complaint gives the business a fix and gives the next reader a way to judge whether the problem applies to them. A stronger, vaguer complaint gives neither, and it still lands on the rating.
Why the average is the wrong thing to watch
Most organisations that monitor customer text watch a mean and a trend line. Average rating, average sentiment score, month over month. Under that instrumentation, a 2% volume increase and a 9% length increase are noise. The dashboard would have shown nothing.
Meanwhile the bottom decile of the same corpus changed character. Harsher tone, less actionable content, same weight in the aggregate score. The metric that summarises the corpus is precisely the metric that hides the effect, because an intensification of the tail with a stable mean is invisible to a mean.
This is the same instrumentation failure we have written about in the verification work that acceleration creates. Faster production of text does not distribute its effects evenly. It concentrates them where the writer had the least to say and the most feeling about it.
Detection is the wrong instrument
The obvious reflex is to try to identify AI-written reviews and discount them. Tripadvisor already runs that play: per the summary of the study, the platform monitors for AI-generated content, deletes it, and bans the reviewer.
Set aside whether detection works well enough to justify banning a paying customer. Detection answers a question the finding does not raise. Gu and Spann did not observe a population of fake reviews inserted by machines. They observed real guests, with real stays and real complaints, writing differently because a tool was available. Removing those reviews removes genuine customer signal. Keeping them unexamined leaves you scoring an input whose properties changed underneath you.
Distribution is the measurable property, and it is measurable without knowing who typed what. You can compute the share of one and two star reviews that contain at least one checkable factual claim (a named amenity, a time, a price, a room number, a staff interaction). That ratio is cheap to compute, it does not require guessing who used which tool, and it moves when the phenomenon in the paper is present in your own corpus.
This connects to the argument that the interface carries the trust. Customer-facing text is an interface surface. When its statistical properties shift, the trust it carries shifts with it, whether or not anyone approved the change.
The scope of the claim
The study looked at hotels. Reviews of a hotel stay have properties that reviews of, say, an enterprise software rollout do not: short consumption window, high emotional variance, low technical vocabulary. Do not assume the coefficients transfer.
The data window matters more. April 2023 was very early. Adoption has grown substantially since, and the researchers themselves flag that model behaviour has changed and that users have become more sophisticated at prompting, which could move the effect in either direction. The honest reading is that the study establishes the mechanism and the direction at the tail, not a current effect size you can plug into a forecast.
What survives the caveats is the design. A four-week regulatory ban created a clean counterfactual, which is far stronger evidence than a before-and-after comparison of the same platform could produce.
The same shift, from the other side
We have argued that answer engines are already recommending vendors to customers who never asked. That is AI reshaping what reaches the customer.
This paper is the mirror image. AI is reshaping what the customer says back. Both directions land in the same place operationally: the text your organisation treats as ground truth about customer experience is now produced through a layer you do not control and cannot see.
Do this now
Pull your last twelve months of one and two star reviews, support tickets, or NPS verbatims. Split them into two buckets by hand or with a simple classifier: contains at least one checkable specific (a name, a time, a number, a place), versus contains only evaluative language. Compute the specific-claim ratio per quarter and plot it.
If that ratio is falling while your average rating holds steady, your feedback channel is losing diagnostic value even though every dashboard looks fine. That is a product operations problem and a service recovery problem before it is an AI problem. The fix lives in the form itself: the review prompt, the ticket template, the survey question. Ask for the specific, and you get the specific back regardless of what wrote it.
Do not build a detector. Build the distributional check, run it quarterly, and make the specific-claim ratio a number someone owns.
This analysis synthesizes The impact of generative artificial intelligence on consumer reviews: Insights from a quasi-experiment (Weilong K.M. Gu and Martin Spann, LMU Munich, International Journal of Research in Marketing, June 2026) and AI-written reviews are more extreme (Thomas McKinlay, Science Says, August 2026).
Victorino Group helps organizations instrument customer-facing text for distribution shift instead of chasing authorship. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation