Back to Articles
Do you speak Spanish? What would it mean to find a thought in a language model? The transformer revisited The logit lens The Jacobian lens Reading a word is not enough Verbalization From the J-lens to the J-space How is this different from other interpretability tools? Why call it a workspace? From global workspace to conscious access What is it like to be an LLM? BONUS: Trying it on an open model Anthropic has managed to monitor Claude’s internal thoughts. Again? It feels as though we have heard that claim several times already. Ah, but this time the headlines say they found consciousness!
If you're like me, you may feel that there have already been many papers and claims about "reading the LLM mind" in the past few years. Anthropic's earlier On the Biology of a Large Language Model, for example, traced hidden multi-step reasoning, poetry planning and multilingual circuits in Claude 3.5 Haiku. What is different here? Is this consciousness business real? It has at least sparked interest among working neuroscientists, so it is worth a look. Let me explain what is genuinely new and exciting about this paper, and why the result is both more serious and less sensational than that opening sentence.












