vix.ing · top · new · best · stats · spec

Memory-Maze: Scenario Driven Visual Language Navigation Benchmark for Guiding Blind People

2024/05/11 by Masaki Kuribayashi, Kuribayashi, Masaki, K. Uehara +9 · 1 citation
Computer Science · #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO)

paper · pdf · doi:10.48550/arxiv.2405.07060

openalex publication_date 2024/05/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Visual Language Navigation (VLN) powered robots have the potential to guide blind people by understanding route instructions provided by sighted passersby. This capability allows robots to operate in environments often unknown a prior. Existing VLN models are insufficient for the scenario of navigation guidance for blind people, as they need to understand routes described from human memory, which frequently contains stutters, errors, and omissions of details, as opposed to those obtained by thinking out loud, such as in the R2R dataset. However, existing benchmarks do not contain instructions obtained from human memory in natural environments. To this end, we present our benchmark, Memory-Maze, which simulates the scenario of seeking route instructions for guiding blind people. Our benchmark contains a maze-like structured virtual environment and novel route instruction data from human memory. Our analysis demonstrates that instruction data collected from memory was longer and contained more varied wording. We further demonstrate that addressing errors and ambiguities from memory-based instructions is challenging, by evaluating state-of-the-art models alongside our baseline model with modularized perception and controls.

Cited by

Related