2025/05/27 by Kui Wu, Hao Chen, Wu, Kui +11 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Analysis and Summarization #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2505.20710
openalex publication_date 2025/05/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
User-Centric Embodied Visual Tracking (UC-EVT) presents a novel challenge for reinforcement learning-based models due to the substantial gap between high-level user instructions and low-level agent actions. While recent advancements in language models (e.g., LLMs, VLMs, VLAs) have improved instruction comprehension, these models face critical limitations in either inference speed (LLMs, VLMs) or generalizability (VLAs) for UC-EVT tasks. To address these challenges, we propose Hierarchical Instruction-aware Embodied Visual Tracking (HIEVT) agent, which bridges instruction comprehension and action generation using spatial goals as intermediaries. HIEVT first introduces LLM-based Semantic-Spatial Goal Aligner to translate diverse human instructions into spatial goals that directly annotate the desired spatial position. Then the RL-based Adaptive Goal-Aligned Policy, a general offline policy, enables the tracker to position the target as specified by the spatial goal. To benchmark UC-EVT tasks, we collect over ten million trajectories for training and evaluate across one seen environment and nine unseen challenging environments. Extensive experiments and real-world deployments demonstrate the robustness and generalizability of HIEVT across diverse environments, varying target dynamics, and complex instruction combinations. The complete project is available at https://sites.google.com/view/hievt.