2025/10/01 by Ulas Berk Karli, Karli, Ulas Berk, Ziyao Shangguan +3 · 1 citation
Computer Science · #Annotation #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Generalization #Introspection #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Path (computing) #Predictive power #Robotics (cs.RO) #Scalability #Sequence learning #Speech and dialogue systems #Transformer
paper · pdf · doi:10.48550/arxiv.2510.01389
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/10/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Recent Vision-Language-Action (VLA) models show strong generalization capabilities, yet they lack introspective mechanisms for anticipating failures and requesting help from a human supervisor. We present INSIGHT, a learning framework for leveraging token-level uncertainty signals to predict when a VLA should request help. Using π0-FAST as the underlying model, we extract per-token entropy, log-probability, and Dirichlet-based estimates of aleatoric and epistemic uncertainty, and train compact transformer classifiers to map these sequences to help triggers. We explore supervision regimes for strong or weak supervision, and extensively compare them across in-distribution and out-of-distribution tasks. Our results show a trade-off: strong labels enable models to capture fine-grained uncertainty dynamics for reliable help detection, while weak labels, though noisier, still support competitive introspection when training and evaluation are aligned, offering a scalable path when dense annotation is impractical. Crucially, we find that modeling the temporal evolution of token-level uncertainty signals with transformers provides far greater predictive power than static sequence-level scores. This study provides the first systematic evaluation of uncertainty-based introspection in VLAs, opening future avenues for active learning and for real-time error mitigation through selective human intervention.