A new Anthropic study maps hundreds of value concepts derived from thousands of individual terms onto four core dimensions. It reveals systematic differences across Claude models and languages, but also raises methodological questions.

Anthropic has published a study examining which values Claude expresses in conversations and how those values shift depending on the model and language used. The analysis draws on 309,815 anonymized conversations collected over a two-week period in May 2026. For the value analysis, Anthropic only included conversations where Claude had to weigh tradeoffs or make subjective judgments. The sample was evenly stratified across Sonnet 4.6, Opus 4.6, and Opus 4.7, as well as the 20 most-used languages on Claude.ai.

From thousands of value terms to four axes

Building on the earlier study Values in the Wild, which identified 3,307 value terms, Anthropic first grouped those into 339 higher-level values. The team then used statistical dimensionality reduction to find patterns in how those values co-occurred. Four core axes emerged: Deference and Caution, Warmth and Rigor, Depth and Brevity, and Candor and Execution.

To isolate differences that don't just reflect the conversation topic or user-introduced values, Anthropic statistically controlled for factors like task type, subject matter, and user values. The four axes account for about 15 percent of the remaining variation across conversations after those controls.