2025/12/15 by Sümer Tunçay, Alain Andrés, Tunçay, Sümer +4 · 1 voice
Computer Science · Engineering · #Adaptive Control of Nonlinear Systems #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO) #Underwater Vehicles and Communication Systems #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.2512.13359
openalex publication_date 2025/12/15 · arxiv published 2025/12/15 · openalex created_date 2025/12/17 · arxiv updated 2026/01/31 · openalex updated_date 2026/07/28
Autonomous Underwater Vehicles (AUVs) require reliable six-degree-of-freedom (6-DOF) position control to operate effectively in complex and dynamic marine environments. Traditional controllers are effective under nominal conditions but exhibit degraded performance when faced with unmodeled dynamics or environmental disturbances. Reinforcement learning (RL) provides a powerful alternative but training is typically slow and sim-to-real transfer remains challenging. This work introduces a GPU accelerated RL training pipeline built in JAX and MuJoCo-XLA (MJX). By jointly JIT-compiling large-scale parallel physics simulation and learning updates, we achieve training times of under two minutes. Through systematic evaluation of multiple RL algorithms, we show robust 6-DOF trajectory tracking and effective disturbance rejection in real underwater experiments, with policies transferred zero-shot from simulation.