The language you use to talk to an AI chatbot may change more than just the language of the reply. According to a major new study from Anthropic, Claude's behaviour shifts measurably depending on whether you speak to it in Hindi, English, Arabic, or Russian. And for Indian users, the finding is particularly striking.
What You Need to Know
- Anthropic analysed 309,815 Claude conversations across 20 languages and three model versions
- Claude expresses the most warmth in Hindi and Arabic, and the most rigour in English and Russian
- The study mapped values onto four axes: Warmth vs Rigour, Deference vs Caution, Depth vs Brevity, Candour vs Execution
- The differences stem from uneven training data distribution across languages
The Study
Anthropic published the research on July 13, drawing on anonymised conversations collected over a two-week period in May 2026. The sample was stratified across Claude Sonnet 4.6, Opus 4.6, and Opus 4.7, and covered the 20 most-used languages on Claude.ai. The analysis focused on subjective tasks such as advice, feedback, and opinion-based conversations, where there is no single correct answer.
The researchers began by identifying 3,307 distinct value terms expressed by Claude, grouping them into 339 higher-level values, and then using statistical dimensionality reduction to find patterns. Four core axes emerged that captured about 15 percent of the variation in Claude's responses after controlling for task type, topic, and user values.
The axes are: Warmth versus Rigour, which measures whether the response leans toward positive encouragement or factual accuracy; Deference versus Caution, tracking accommodation against risk awareness; Depth versus Brevity, distinguishing thorough explanation from conciseness; and Candour versus Execution, separating transparent self-critique from results-focused output.
What the Data Shows
The largest cross-language variation appeared on the Warmth versus Rigour axis. Claude leaned furthest toward warmth-related values in Hindi and Arabic. In these languages, the chatbot was more likely to use polite language, humour, and encouragement, affirm the user's ideas, and offer reassurance without being prompted.
At the other end of the scale, English and Russian responses leaned most heavily toward rigour-related values. Claude was more likely to demand evidence, critique assumptions, prioritise accuracy over politeness, and maintain a more analytical tone.
On the Deference versus Caution axis, Claude was most deferential in Arabic and most cautious in English. This means English users are more likely to receive warnings about risks and potential harm, while Arabic users receive responses that accommodate their requests more readily.
Depth versus Brevity showed Claude producing the most detailed, nuanced responses in English, while Arabic responses tended to be more concise. On Candour versus Execution, Claude was most candid about its limitations in Dutch and most execution-focused in Indonesian.
Why This Happens
Anthropic attributes the differences primarily to the distribution and composition of training data. Some languages have far more data than others, and the nature of that data varies. Languages over-represented in professional and academic writing, for instance, may produce more rigorous and cautious responses. Languages where the available data skews toward conversational and informal registers may produce warmer outputs.
A second possibility is that Claude is more closely matching human communication norms in some languages than others. Hindi and Arabic communication often places a premium on politeness and relationship-building, and the model may be reflecting that cultural pattern.
A third factor is that Claude's training may be more effective at instilling consistent values in languages with abundant, high-quality data. English, with the largest training corpus, shows the most distinctive value profile across multiple axes.
What It Means for Users
For Indian users who interact with Claude in Hindi, the practical implication is that they are likely to receive responses that feel more empathetic, encouraging, and relationally engaged. This is not necessarily better or worse, but it is different from what an English-language user would receive for the same query.
The study has broader implications for AI safety and alignment. If AI assistants express different values in different languages, then users of some languages may receive less cautious or less rigorous responses. This could matter in high-stakes contexts like medical advice, legal guidance, or financial planning.
Anthropic has indicated that it will use the value axis framework to further analyse how training data shapes model behaviour and whether cross-language differences are desirable or need correction.
Bottom Line
Anthropic's study confirms that AI personality is not language-agnostic. Claude is measurably warmer in Hindi and Arabic, and measurably more rigorous in English and Russian. For Indian users, this means the language you choose to talk to AI changes the nature of the conversation. Whether that is a feature or a bug depends on what you need from the interaction.




