Artificial intelligence is often portrayed as an objective tool, capable of responding in exactly the same way regardless of who is asking the question. However, a new study by Anthropic shows that this neutrality is more complex than it seems. According to the researchers, the language used to interact with Claude influences the way the model reasons, argues, and expresses certain values.1
By analyzing hundreds of thousands of anonymized conversations, Anthropic highlights a phenomenon rarely studied on this scale: the same artificial intelligence model does not produce exactly the same responses depending on whether the user speaks in French, English, Arabic, Russian, or another language. These differences go beyond vocabulary or the quality of the translation. They directly affect the tone, the structure of the responses, the level of nuance, and sometimes even the recommendations provided.
This study sheds new light on the linguistic biases of large language models and raises a crucial question: Can artificial intelligence truly be neutral when it learns from billions of texts drawn from different cultures?
A previously unpublished study on the values expressed by Claude
To better understand the influence of language on its assistant’s behavior, Anthropic analyzed several hundred thousand conversations that had been intentionally anonymized. The researchers focused on situations where there is no universally correct answer—for example, when a user asks for professional advice, seeks to resolve a personal conflict, or wants to make an important decision.1
In this type of context, AI does more than simply provide factual information. It makes recommendations, prioritizes certain options, and adopts a communication style that inevitably reflects certain values.
The specific objective of the study was to determine whether these values vary depending on the language used during the conversation.
There are four dimensions that allow us to compare languages
Anthropic based its analysis on four major behavioral categories used to characterize Claude's responses.
The first contrasts deference with caution—that is, the model's tendency to follow the request as stated or, conversely, to express greater reservations.
The second measures warmth versus rigor—in other words, the level of empathy, closeness, and emotional support—as opposed to a more analytical and factual style.
The third compares depth and conciseness, while the fourth evaluates sincerity versus execution, focusing in particular on the model’s ability to question certain assumptions rather than simply fulfilling the user’s request.
The results show that the greatest differences appear along the warmth-rigor axis. Some languages lead Claude to adopt a more empathetic style, while others favor precision, critical analysis, and logical reasoning.1
French is distinguished by its balance
Among the languages studied, French stands out as one of the most balanced.
Anthropic notes that Claude generally adopts a relatively neutral tone when responding in French. The explanations remain well-structured, the vocabulary is precise, and the model strives to maintain a high level of clarity while naturally adapting to the style of the person it is speaking with.
Unlike other languages, which place greater emphasis on certain personality traits, French yields more consistent responses across all four axes analyzed. This consistency could be an advantage for users seeking balanced responses, particularly in professional, academic, or decision-making contexts.
Why does English sometimes yield better answers?
The study also highlights several distinctive features of Claude's behavior when he responds in English.
The model is more likely to spontaneously correct erroneous assumptions, provide additional evidence to support its answers, and propose goals that are more ambitious than those initially set by the user.
In other words, Claude tends to take on the role of a constructive critic when he speaks in English. Rather than simply carrying out the request, he more often seeks to improve it or put it into perspective.
For bilingual users, this observation may have practical implications. Depending on the nature of the task, switching languages could yield responses with different qualities, whether in terms of creativity, rigor of argumentation, or depth of analysis.
The models themselves also influence the responses
Language, however, is not the only factor that influences Claude's behavior.
Anthropic also demonstrates that the various models in its family exhibit distinct behaviors. For example, Claude Sonnet generally adopts a warmer, more conversational style, while Claude Opus places greater emphasis on analytical rigor, technical precision, and detailed reasoning.1
These differences are compounded by linguistic variations. As a result, the same user may receive significantly different answers depending not only on the model used, but also on the language chosen to ask the exact same question.
How can I access this study?
Anthropic has published all of its work in the form of a research report available for free on its official website.1
The study focuses primarily on Claude, but its conclusions extend far beyond this single language model. The phenomena observed stem largely from the data used to train large language models and potentially apply to all modern LLMs.
The researchers also note that this work will help to gradually improve future models in order to reduce certain unintended variations and better understand the cultural influences present in the training data.
An Important Discovery for Businesses
These findings are of direct interest to organizations that use artificial intelligence in multiple languages.
International companies are now deploying AI assistants to employees located in different countries. Since the model's behavior varies by language, two employees performing exactly the same task might receive slightly different recommendations.
This finding could lead companies to further adapt their policies for deploying generative AI, particularly in areas where the consistency of responses is a strategic issue: human resources, legal, healthcare, or customer service.
Ethical Issues: Is AI Neutrality a Myth?
This study highlights that large language models learn from billions of documents produced by different societies, cultures, and languages. As a result, they inevitably inherit some of the representations, writing styles, and value systems present in these corpora.
The goal is not necessarily to eliminate all these differences, but to better understand them so that users can interpret the generated responses more accurately.
Absolute neutrality thus appears less as a state of being than as a goal toward which laboratories are gradually striving. Transparency regarding these variations therefore becomes an essential element in building trust in artificial intelligence systems.
Understanding AI also involves understanding languages
In this study, Anthropic demonstrates that the performance of artificial intelligence does not depend solely on its architecture or computing power. The language used also plays a role in how the model reasons, provides advice, and interacts with its users.
As generative AI becomes an everyday tool in multilingual work environments, understanding these linguistic differences could become just as important as understanding the technical capabilities of the models themselves.
Learn more
Variations in responses across languages show that a model’s performance is not limited to its technical capabilities but also depends on the data and cultural contexts incorporated during its development. On a related topic, check out our article “Anthropic Publishes a Major Study: What 81,000 Users Really Expect from AI, ” which highlights users’ growing expectations regarding the reliability, transparency, and quality of responses from artificial intelligence systems.
References
1. Anthropic. (2026). AI isn’t neutral: Language-based differences in Claude’s responses.
https://www.anthropic.com/research
