2026/05/21 by Lily Goli, Justin Kerr, Daniele Reda +3 · 1 voice
Computer Science · Decision Sciences · #Adaptation (eye) #Advanced Bandit Algorithms Research #Agent-based model #Context (archaeology) #Curiosity #Episodic memory #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Reinforcement learning #Sequence (biology) #Trajectory #cs.LG
paper · pdf · doi:10.48550/arxiv.2605.22814
openalex publication_date 2026/05/21 · arxiv published 2026/05/21 · arxiv updated 2026/05/21 · openalex created_date 2026/05/23 · openalex updated_date 2026/07/28
Exploration is a prerequisite for learning useful behaviors in sparse-reward, long-horizon tasks, particularly within 3D environments. Curiosity-driven reinforcement learning addresses this via intrinsic rewards derived from the mismatch between the agent's predictive model of the world and reality. However, translating this intrinsic motivation to complex, photorealistic environments remains difficult, as agents can become trapped in local loops and receive fresh rewards for revisiting forgotten states. In this work, we demonstrate that this failure stems from a lack of spatial persistence and episodic context. We show that effective curiosity requires a model of the world that is persistent and continuously updated, paired with an agent that maintains an episodic trajectory history to navigate toward novel regions. We achieve this using an online 3D reconstruction as a persistent model of the world, while the agent policy is parameterized as a sequence model over RGB observations to maintain episodic context. This design enables effective exploration during training while allowing the agent to navigate using solely RGB frames at deployment. Trained purely via curiosity on HM3D, our agent outperforms RL-based active mapping baselines and generalizes zero-shot to Gibson and AI-generated worlds. Our end-to-end policy enables efficient adaptation to downstream tasks, such as apple picking and image-goal navigation, outperforming from-scratch baselines. Please see video results at https://recuriosity.github.io/.