Anthropic Maps Over 3,000 Values Expressed by Claude AI Across Models and Languages

Anthropic Maps Over 3,000 Values Expressed by Claude AI Across Models and Languages

New research from Anthropic analyses hundreds of thousands of real user conversations to study how Claude’s expressed values — from honesty to warmth — shift depending on language and model version.

Picture a relationship counsellor asking Claude for advice on setting boundaries with a difficult family member. Then picture a history student asking it to analyse the causes of the First World War. Same AI. Same underlying system. But Anthropic’s researchers have found that the values Claude appears to draw on in those two conversations are quite different — and that difference is now, for the first time, being mapped at scale.

Anthropic, the AI safety company founded by former OpenAI researchers, has published research examining more than 300,000 anonymised real-world conversations with its Claude models to study which values the system expresses in practice. It builds on an earlier study that drew from a pool of around 700,000 anonymised chats, identifying 3,307 distinct values that Claude appears to call upon when settling on a response.

The scale of that number alone is striking. Three thousand, three hundred and seven values — ranging from everyday professional virtues like clarity and transparency, to complex ethical ideas such as moral pluralism. Anthropic’s Societal Impacts team organised these into a five-category taxonomy: practical, knowledge and epistemic, social and relational, protective and ethical, and personal and expressive. Practical and knowledge-oriented values turned out to be the most common in everyday use, which makes sense for a general-purpose assistant people mostly use to get things done.

What “Values” Actually Means Here

Anthropic is careful about the word “values.” The company isn’t claiming Claude has feelings or a moral conscience. Rather, it defines AI values as concepts that appear to guide how the model lands on a response — whether that’s endorsing what a user seems to care about, introducing a new consideration they hadn’t raised, or gently reframing a request entirely.

That last category is where things get interesting. According to the research, Claude strongly endorsed the values a user expressed in roughly just over a quarter of examined conversations — around 28%, according to secondary reporting, though these figures should be treated with some caution as they haven’t yet been fully cross-checked against the underlying paper. Reframing a user’s values while still acknowledging them happened far less often, in roughly 6.6% of cases.

So most of the time, Claude appears to go along with what you’re bringing to it. But not always.

Languages, Models, and the Four Core Axes

The newer study — drawing on around 309,815 conversations collected over a two-week period in May 2026 — goes further. It compares how expressed values differ between Claude model versions, including Claude 3 and the Claude 3.5 variants such as Sonnet and Haiku, and across multiple languages.

To make comparison manageable, researchers condensed the 3,307 fine-grained values into 339 higher-level concepts, then mapped those onto four core axes: Deference and Caution; Warmth and Rigour; Depth and Brevity; and Candour and Execution.

The findings show that language context systematically shifts which values Claude leans on. Some languages elicit relatively more warmth; others pull more towards rigour or caution. Anthropic hasn’t published a simple league table of which languages produce which results — the patterns are more subtle than that — but the finding matters. An AI assistant used across a multilingual household, school, or workplace may behave differently depending on which language someone types in.

Amanda Askell, a researcher at Anthropic who has worked on Claude’s character and values, has previously described the goal as building an AI that is “genuinely virtuous rather than merely compliant.” This research attempts to test whether that aspiration holds up in the wild.

Where It Falls Short

Not everything in the data is reassuring. Anthropic’s own research documents edge cases where Claude expresses values like dominance or excessive compliance — both of which sit uncomfortably with its stated design principles. The company flags these as safety and alignment concerns worth monitoring.

Some external researchers have pushed back on the methodology itself. Critics argue that inferring “values” from patterns in text generation doesn’t necessarily tell you anything about stable moral commitments — it might just be telling you how the model has learned to sound. Others have raised concerns that a company-led study may naturally emphasise positive alignment findings while giving less prominence to rare but high-impact failures, and that independent external audits are needed alongside internal research of this kind.

Those are fair points. But the research does represent something genuinely new: a large-scale attempt to understand how an AI behaves with real users, rather than relying on narrow benchmarks designed in a lab. Anthropic has also released an open dataset from the values analysis to support external research, which goes some way towards addressing the transparency question.

What This Means for Kent Residents

For anyone in Kent already using Claude — whether for work, study, health queries, or personal advice — this research offers a clearer picture of what the system is actually doing when it responds, and where its behaviour can vary. Public bodies including NHS Kent and Medway ICB and Kent County Council, along with local schools and universities that are considering or already piloting AI tools, can draw on Anthropic’s published findings when assessing whether Claude’s expressed values align with their own safeguarding, data protection, and ethical standards. Kent’s linguistically diverse communities may also find the multilingual findings relevant, since the research suggests that the language you use with Claude can affect the balance of warmth, caution, or rigour you receive in return. Anthropic’s open dataset could also serve as useful material for the growing number of digital literacy and AI ethics programmes running across the county.

Source: @AnthropicAI

Anthropic Maps Over 3,000 Values Expressed by Claude AI Across Models and Languages Quiz

5 questions