2026/05/06 by Mohamed Salim Aissi, Mohamed-Salim Aissi, Clémence Grislain +6 · 1 voice
Computer Science · Psychology · #Embodied cognition #Multimodal Machine Learning Applications #Perception #Pipeline (software) #Prism #Scaling #Social Robot Interaction and HRI #Topic Modeling #Work (physics) #cs.AI
paper · pdf · open access · doi:10.48550/arxiv.2605.05407
published in HAL (Le Centre pour la Communication Scientifique Directe) (Centre National de la Recherche Scientifique)
openalex publication_date 2026/05/06 · arxiv published 2026/05/06 · openalex created_date 2026/05/09 · arxiv updated 2026/06/12 · openalex updated_date 2026/07/30
Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception-reasoning-decision gap in standalone Vision-Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce PRISM, a framework that tightly couples perception (VLM) and decision (LLM) through a dynamic question-answer (DQA) pipeline. Instead of passively accepting the VLM's description, the LLM critiques it, probes the VLM with goal-oriented questions, and synthesizes a compact image description. This closed-loop interaction yields a sharp, task-driven understanding of the scene. We evaluate PRISM on the ALFWorld and Room-to-Room (R2R) benchmarks. We show that: (1) PRISM significantly outperforms state-of-the-art image-based models, (2) our Interactive goal-oriented perception pipeline yields systematic and substantial gains, and (3) PRISM is fully automatic, eliminating the need for handcrafted questions or answers.