Your Model's Values Shift by Language, and Your Evaluation Is English-Only

TV
Thiago Victorino
7 min read
Your Model's Values Shift by Language, and Your Evaluation Is English-Only

309,815 conversations. Three models. Twenty languages. Anthropic ran that dataset through a method that compresses hundreds of thousands of Claude.ai exchanges down to four axes of value, and the result is uncomfortable for anyone deploying a single model across borders: the same model expresses systematically different values depending on the language you speak to it in.

The four axes carry plain names. Deference versus Caution. Warmth versus Rigor. Depth versus Brevity. Candor versus Execution. Together they capture about 15% of the total variation in how the model’s values show up across conversations. That is a small slice of a large, messy space, which is the honest part of the finding. It is also enough to see the pattern clearly. Arabic conversations pull the model toward deference and warmth. English conversations pull it toward rigor and caution. Same weights, same system prompt, different judgment.

What the axes actually measure

Anthropic’s research team, led by Matt Kearney, did not ask the model what it values. They observed what it did across 309,815 real conversations and reduced the behavior to coordinates. A response high on Caution hedges, flags risk, and defers final judgment to the user. A response high on Candor states a direct opinion even when the opinion is unwelcome. A response high on Warmth softens and encourages. A response high on Rigor pushes back and corrects.

All of these can be right. A cautious answer and a candid answer can both be correct and helpful. The problem is not that the model has values. The problem is that the values are not constant, and almost nobody deploying these models is measuring which value profile they ship to which user.

The single line from the research that should stop a product team is this: “Two people asking for feedback on the same business plan, one in Hindi and one in Russian, may come away with different impressions of its quality.” Same plan. Same model. Different verdict, because the language changed the value profile the model applied.

Two sources of drift, not one

The variation runs along two independent dimensions, and it helps to keep them separate.

The first is the model version. Opus 4.7 skews toward caution by roughly 0.24 standard deviations and toward depth by roughly 0.23 standard deviations relative to the baseline. Sonnet 4.6 skews toward warmth by roughly 0.17 standard deviations and toward deference by roughly 0.14. Upgrade your model and the personality your users experience shifts, even when you change nothing in your prompt. The values live in the weights, and your prompt only adjusts part of them.

The second is the language. English sits at the rigor and caution end. Arabic sits at the deference and warmth end. Every language you serve lands somewhere on these axes, and you did not choose where. The training process did, for reasons only partly legible even to the people who ran it.

Stack the two together and you get the real operating condition: a company that upgrades from Sonnet to Opus while serving customers in English, Arabic, Hindi, and Russian is now shipping at least eight distinct value profiles, and it is evaluating maybe one of them.

The blind spot is structural

Here is the part that turns a research curiosity into a governance problem. Nearly every evaluation pipeline that teams actually run is built in English. The red-team prompts are English. The eval sets are English. The rubric the reviewers score against is written in English and applied to English outputs. When a company writes a policy that says “the assistant should push back on financially reckless plans,” it verifies that behavior in English and assumes it generalizes.

It does not generalize. If English is the most rigorous and cautious language for the model, then English is exactly the locale where the pushback policy looks healthiest. The Arabic version of the same assistant, more deferential by construction, may wave the reckless plan through with a warm note of encouragement. Your evaluation passed. Your Arabic users got a different product. You have no instrument pointed at the difference.

This is the mechanism worth naming precisely. English-only evaluation does not just undercount problems in other languages. It systematically inspects the locale least likely to fail and certifies the whole system on that basis. The languages most prone to deference are the languages you are least equipped to see.

Why “just translate the evals” does not fix it

The obvious response is to translate the English eval set into every language you serve and re-run it. That helps, and most teams skip even that. But it falls short, for two reasons.

Translation preserves the prompt, not the value profile. A translated eval still asks the model the same question. It does not tell you whether your reviewers, scoring in their own linguistic frame, apply the same threshold for “too deferential” that an English reviewer would. Warmth reads as competence in some cultures and as evasion in others. The rubric itself carries a value profile.

And the axes only explain 15% of the variation. The other 85% is not captured by four clean dimensions. Whatever governance instrument you build has to assume that the visible drift is a lower bound on the real drift. The complete map stays out of reach. Measure the four axes because you can, then treat everything you cannot decompose as latent risk in every locale you have not directly observed.

This is Anthropic’s data and our conclusion

Worth being clear about the provenance. This is first-party research from the lab that makes the model, framed as a safety contribution, and it carries the self-interest that framing implies. Take the dataset and the method as reported: 309,815 conversations, three models, twenty languages, four axes, the standard-deviation skews above. Those are Anthropic’s measurements.

The operating conclusion is ours. Anthropic documented that values vary by language and version. It did not tell you to rebuild your evaluation around the languages you actually serve. That step is the one that matters for any org running one model across locales, and it is the step almost nobody has taken. The lab measured the drift as a scientific object. You have to measure it as a liability, in the specific languages your customers use, against the specific policies you claim to enforce.

Do this now

Pull your evaluation harness and check what language it runs in. If the answer is English, you have measured one value profile and shipped many.

Then, in priority order: list the languages you serve in production, ranked by user volume. For the top three non-English locales, take your five most consequential policies (the ones about refusal, financial or medical caution, pushback on bad plans) and run them natively in each language, scored by a reviewer fluent in that language against your actual rubric. Do not translate and score in English. You are testing whether the model’s deference profile in that locale quietly violates a policy that looks fine in English.

Compare the pass rates across languages. If they diverge, and the Anthropic data says they will, you have found the exact surface where your governance is blind. That divergence is the thing to monitor going forward, on every model upgrade, because the version drift and the language drift compound. The company that measures behavior in the languages it serves is operating its AI. The company that measures English and hopes is guessing in every other locale, and now it knows the guess is wrong.


This analysis synthesizes How Claude’s values vary by model and language (Anthropic, July 2026).

Victorino Group helps teams evaluate AI behavior in every language they operate in, not just English. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation