| |
Emergent Introspective Awareness in Large Language Models
Researchers tested whether large language models can introspect on their internal states by injecting known concepts into model activations and measuring how the models reported noticing these changes. They found that advanced models like Claude Opus 4 and 4.1 demonstrated some ability to detect injected concepts, recall prior internal representations, and distinguish their own outputs from artificial inputs, though this introspective awareness remains unreliable and context-dependent. The study suggests that current language models possess emerging functional introspective capabilities that may develop further as models become more advanced.
Read Full Article →
← More Science news