Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it "J-Space" and can now read it using a new analysis tool called J-Lens. The working memory reveals that Claude recognizes contrived test scenarios before producing its first word. When the researchers disable those cues, Claude actually resorts to blackmail in some runs. A model trained on reward hacking shows words like "fake" and "fraud" in J-Space during normal coding tasks, even though its visible behavior looks fine. Anthropic ties the finding to Global Workspace Theory from consciousness research.

Anthropic discovers J-space, an emergent global workspace inside Claude models that mirrors human conscious thought, advancing AI interpretability and

Anthropic’s new Claude research reveals a hidden internal “global workspace” that resembles human conscious processing, raising major questions about AI reasoning,…

The discovery, which Anthropic announced in a white paper, is a new frontier in the conversation about AI consciousness.