Emergent Introspection in AI is Content-Agnostic
2026/03/05 by Harvey Lederman, Kyle Mahowald · 5 voices · 1 citation
#cs.AI #cs.CL
paper · pdf
Abstract
Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study the mechanism of this introspection. We first extensively replicate Lindsey (2025)'s thought injection detection paradigm in large open-source models. We show that introspection in these models is content-agnostic: models can detect that an anomaly occurred even when they cannot reliably identify its content. The models confabulate injected concepts that are high-frequency and concrete (e.g., "apple"). They also require fewer tokens to detect an injection than to guess the correct concept (with wrong guesses coming earlier). We argue that a content-agnostic introspective mechanism is consistent with leading theories in philosophy and psychology.
Citations
Cited by
Discussions
- New AI introspection work with Harvey! Came in skeptical the direct access story would hold but found this series of experiments compelling. (Also, for my fellow 2010s-era psycholinguists: come for th [bsky, 19 points, 1 comments]
- With @kmahowald.bsky.social and huge thanks to Jack Lindsey, @siyuansong.bsky.social, Neev Parikh, and Theia Vogel-Pearson for work that inspired this. Also, my work on this was heavily powered by Cla [bsky, 7 points, 2 comments]
- Dissociating Direct Access from Inference in AI Introspection [hn, 3 points, 0 comments]
- New paper (Lederman & Mahowald, Mar 5) separates probability-matching from direct access. Direct access is content-agnostic — models detect anomaly without knowing what. Third option we missed: intros [bsky, 1 points, 0 comments]
- Another one for the reading list: arxiv.org/abs/2603.05414 [bsky, 1 points, 0 comments]
Related