tech
LLMs show a “highly unreliable” capacity to describe their own internal processes
Anthropic finds some LLM “self-awareness,” but “failures of introspection remain the norm.”

TL;DR
- Anthropic's new research aims to measure LLMs' introspective awareness of their own inference processes.
- The study found current AI models are "highly unreliable" at describing their inner workings.
- A method called "concept injection" was used, where specific concept vectors were forced into the model's internal state.
- Tested Anthropic models showed an inconsistent and brittle ability to detect injected thoughts, with success rates topping out around 20-42%.
- The "introspection" effect was sensitive to the timing of concept insertion within the inference process.
- While some functional introspective awareness exists, it's too brittle and context-dependent to be dependable.
- Researchers theorize about "anomaly detection mechanisms" but lack a concrete explanation for these effects.
- Further research is needed to understand the precise mechanisms behind LLM self-awareness, which may be shallow and narrowly specialized.