2025/08/20 by Siyuan Song, Harvey Lederman, Song, Siyuan +5 · 2 voices · 7 citations
Computer Science · Social Sciences · #Online Learning and Analytics #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2508.14802
Whether AI models can introspect is an increasingly important practical question. But there is no consensus on how introspection is to be defined. Beginning from a recently proposed ''lightweight'' definition, we argue instead for a thicker one. According to our proposal, introspection in AI is any process which yields information about internal states through a process more reliable than one with equal or lower computational cost available to a third party. Using experiments where LLMs reason about their internal temperature parameters, we show they can appear to have lightweight introspection while failing to meaningfully introspect per our proposed definition.