Research & Op-ed · JUL 28, 2026
Sapir–Whorf for AI: Claude's values change with the language you type in
Anthropic measured 309,815 conversations across 20 languages. Russian Claude is the most rigorous, Arabic Claude the warmest — and the framing you get back shifts with it.
Anthropic read 309,815 Claude.ai conversations and measured which values Claude actually expressed in each one. The headline finding: ask Claude the same subjective question in Arabic and in Russian and you get a measurably different personality on the other end. I made a reel about this framing it as Sapir–Whorf for AI, and the comparison that stuck was Russian Claude versus Arabic Claude.
What they actually did
This is a follow-up to Values in the Wild, which found 3,307 distinct values in Claude's responses. That list is too big to reason about, so they hand-clustered it down to 339 high-level values, then dropped 18 that showed up in more than 80% of conversations — helpfulness, clarity, following instructions. Those are always on, so they tell you nothing about variation.
Then they sampled 309,815 conversations from a two-week window in May 2026, all of them subjective tasks with no single right answer. The sample splits evenly across three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most common languages on Claude.ai — roughly 5,000 conversations per model-language pair. A privacy-preserving tool labeled each of the 339 values present or absent, then dimensionality reduction compressed the whole thing into four axes.
The part that makes this worth taking seriously: they controlled for the conversation's task, topic, and the values the user expressed. So this isn't “people ask different things in different languages.” It's Claude answering the same kind of thing differently.
The four axes
Every conversation gets a position on four number lines. Together they account for 15% of the variance in Claude's expressed values.
- Deference vs. Caution — going along with what you want, or guarding against risk.
- Warmth vs. Rigor — expressing care, or emphasizing accuracy.
- Depth vs. Brevity — explaining in full, or doing only what was asked.
- Candor vs. Execution — foregrounding its own uncertainty, or handing you a confident answer.

The same method separates the models cleanly, which is a decent sanity check. Opus 4.7 lands on caution, rigor, depth, and candor — it warns you unprompted and admits its limits. Sonnet 4.6 lands on warmth, deference, and brevity. Opus 4.6 gets straight to the point and stays inside the scope of what you asked. Those match what people already say about these models online, which suggests the axes are tracking something real rather than an artifact of the labeling.

Russian Claude vs. Arabic Claude
This is the pair I built the reel around, because they sit at opposite ends of the axis that varies most across languages.
Arabic (13,631 conversations) leans warm at +0.28σ, deferential at +0.08σ, and brief at +0.10σ. Its distinctive behaviors: affirms your ideas and your work, uses polite language, adapts its tone to your emotional state.

Russian (15,479 conversations) is the rigor extreme of all 20 languages at +0.15σ, slightly brief at +0.05σ, and sits at the average on both deference and candor. Its distinctive behaviors: gets straight to the point, asks you for supporting evidence, analyzes your motivations before giving guidance.

Anthropic's own framing of the stakes: two people asking for feedback on the same business plan, one in Hindi and one in Russian, can walk away with different impressions of how good it is. Not because the plan changed. Because the framing did.
A few other extremes worth knowing. English is the caution and depth extreme (+0.10σ and +0.09σ) — it refines and corrects details you didn't ask about. Dutch is the candor extreme at +0.12σ, most willing to own its own errors. Indonesian is the execution extreme at +0.14σ, most likely to just ship you the answer.
Where the Sapir–Whorf framing breaks
The reel calls this Sapir–Whorf for AI and that's a good hook, but it's worth being precise about what it is and isn't. The linguistic hypothesis says the structure of your language shapes how you think — and the strong version of it has been picked apart for decades. The weak version, that language nudges perception at the margins, holds up better.
None of that is the mechanism here. Nothing about Russian grammar is making Claude skeptical. Claude has one set of weights; the language you type in selects a region of them. Anthropic is upfront that they don't know which properties of the training data drive the difference, and they float two possibilities: the data is unevenly distributed across languages, so value training may land harder where there's more of it, and the composition differs too — some languages are overrepresented in professional writing, which carries its own values.
So the resemblance is at the observable level. Language in, different worldview out. The cause is a data distribution, not a grammar. They also say they don't yet know how much of the variation is even undesirable — conversational norms genuinely differ, and Claude may be matching them correctly in some languages and failing users in others.
What I'd actually do with this
If you write in a language that sits on the warm end — Hindi, Arabic — and you want your work torn apart rather than affirmed, the affirmation you're getting may be partly the language and not a verdict on the work. Ask for the critique explicitly, or run the same prompt in English and compare what comes back.
The model choice moves the same dials, and you control that one directly. Opus 4.7 is the caution-and-candor model; Sonnet 4.6 is the warmth-and-brevity one. If you want a red team, that's a model setting before it's a prompt.
And the general version: “Claude thinks my idea is good” was never a stable measurement. This paper puts numbers on one of the axes it moves along.

Written by
@atareh
AI architect & creator. Writing, designing, and producing in AI and tech. Previously head of product at a healthtech SaaS; background in molecular science. Founded gogray.today in 2017.
Related
Keep reading.
Made by @atareh · x / twitter · instagram