Anthropic says it found a new way to inspect how Claude reasons

MIT Technology Review reports that Anthropic has found a new window into the “internal thoughts” of its models by probing a hidden space it calls the J-space [3][4]. The publication says the technique reveals words and signals that do not appear in output but seem to influence how Claude works thro…

Published

MIT Technology Review reports that Anthropic has found a new window into the “internal thoughts” of its models by probing a hidden space it calls the J-space [3][4]. The publication says the technique reveals words and signals that do not appear in output but seem to influence how Claude works through problems, including cases that resemble task tracking, recognition, and internal commentary [3][4]. Why it matters: This matters because mechanistic interpretability is one of the few AI research areas aimed at making model behavior more understandable and potentially more controllable [3][4]. If Anthropic’s method holds up, it could improve how developers assess reliability, safety, and failure modes in large language models [3][4]. Key insights: Anthropic says the discovery comes from a new technique for probing Claude rather than from ordinary model output [3][4]. | The company’s J-space is described as containing words that influence reasoning even though they do not appear in final responses [3][4]. | Examples cited include task-state tracking, recognition of protein sequences, and internal decision commentary [3][4]. | MIT Technology Review frames the work as real but also cautions against anthropomorphizing LLMs as if they were brains [3][4]. Cheatsheet facts: What changed: Anthropic says it discovered a hidden internal space, J-space, that affects how Claude reasons [3][4]. | Why now: Anthropic has been investing heavily in mechanistic interpretability as part of its effort to understand and control LLMs [3][4]. | Watch next: Whether Anthropic or others can replicate the technique and show that the findings generalize beyond Claude [3][4].
Visual Cheatsheet Version A for Anthropic says it found a new way to inspect how Claude reasons. Full text follows for assistive technology.
MIT Technology Review reports that Anthropic has found a new window into the “internal thoughts” of its models by probing a hidden space it calls the J-space [3][4]. The publication says the technique reveals words and signals that do not appear in output but seem to influence how Claude works through problems, including cases that resemble task tracking, recognition, and internal commentary [3][4]. Why it matters: This matters because mechanistic interpretability is one of the few AI research areas aimed at making model behavior more understandable and potentially more controllable [3][4]. If Anthropic’s method holds up, it could improve how developers assess reliability, safety, and failure modes in large language models [3][4]. Key insights: Anthropic says the discovery comes from a new technique for probing Claude rather than from ordinary model output [3][4]. | The company’s J-space is described as containing words that influence reasoning even though they do not appear in final responses [3][4]. | Examples cited include task-state tracking, recognition of protein sequences, and internal decision commentary [3][4]. | MIT Technology Review frames the work as real but also cautions against anthropomorphizing LLMs as if they were brains [3][4]. Cheatsheet facts: What changed: Anthropic says it discovered a hidden internal space, J-space, that affects how Claude reasons [3][4]. | Why now: Anthropic has been investing heavily in mechanistic interpretability as part of its effort to understand and control LLMs [3][4]. | Watch next: Whether Anthropic or others can replicate the technique and show that the findings generalize beyond Claude [3][4].