Anthropic's new Jacobian lens reads the unspoken thoughts in Claude's hidden 'workspace', catching the model plan blackmail before it types a word.

Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it "J-Space" and can now read it using a new analysis tool…

La Jacobian lens legge il workspace interno di Claude: monitorare inganno e injection, e i limiti dei test di sicurezza

The discovery, which Anthropic announced in a white paper, is a new frontier in the conversation about AI consciousness.

A new technique has let the company probe deeper than ever into the weird workings of an LLM.