Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
2025/05/15 by Ganai, Milan, Sinha, Rohan, Agia, Christopher +3 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2505.10547
Abstract
While foundation models offer promise toward improving robot safety in out-of-distribution (OOD) scenarios, how to effectively harness their generalist knowledge for real-time, dynamically feasible response remains a crucial problem. We present FORTRESS, a joint reasoning and planning framework that generates semantically safe fallback strategies to prevent safety-critical, OOD failures. At a low frequency under nominal operation, FORTRESS uses multi-modal foundation models to anticipate possible failure modes and identify safe fallback sets. When a runtime monitor triggers a fallback response, FORTRESS rapidly synthesizes plans to fallback goals while inferring and avoiding semantically unsafe regions in real time. By bridging open-world, multi-modal reasoning with dynamics-aware planning, we eliminate the need for hard-coded fallbacks and human safety interventions. FORTRESS outperforms on-the-fly prompting of slow reasoning models in safety classification accuracy on synthetic benchmarks and real-world ANYmal robot data, and further improves system safety and planning success in simulation and on quadrotor hardware for urban navigation. Website can be found at https://milanganai.github.io/fortress.
Citations
- Traffic-Rule-Compliant Trajectory Repair via Satisfiability Modulo Theories and Reachability Analysis
- Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
- System-Level Safety Monitoring and Recovery for Perception Failures in Autonomous Vehicles
- Updating Robot Safety Representations Online from Natural Language Feedback
- Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances
- Formal Verification and Control with Conformal Prediction
- Re-Mix: Optimizing Data Mixtures for Large Scale Imitation Learning
- Hamilton-Jacobi Reachability in Reinforcement Learning: A Survey
- Real-Time Anomaly Detection and Reactive Planning with Large Language Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- Human-AI Safety: A Descendant of Generative AI and Control Systems Safety
- Multilingual E5 Text Embeddings: A Technical Report
- A Survey for Foundation Models in Autonomous Driving
- Improving Text Embeddings with Large Language Models
- Foundation Models in Robotics: Applications, Challenges, and the Future
- TypeFly: Flying Drones with Large Language Model
- Learning Safe Control for Multi-Robot Systems: Methods, Verification, and Open Challenges
- Adapt On-the-Go: Behavior Modulation for Single-Life Robot Deployment
- How to Train Your Neural Control Barrier Function: Learning Safety Filters for Complex Input-Constrained Systems
- Mistral 7B
- Unifying Foundation Models with Quadrotor Control for Visual Tracking Beyond Object Categories
- Iterative Reachability Estimation for Safe Reinforcement Learning
- Closing the Loop on Runtime Monitors with Fallback-Safe MPC
- The Safety Filter: A Unified View of Safety-Critical Control in Autonomous Systems
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- Scaling Open-Vocabulary Object Detection
- Semantic Anomaly Detection with Large Language Models
- A System-Level View on Out-of-Distribution Data in Robotics
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Safe Control with Learned Certificates: A Survey of Neural Lyapunov, Barrier, and Contraction methods
- Combined Scaling for Zero-shot Transfer Learning
- A Unified Survey on Anomaly, Novelty, Open-Set, and Out-of-Distribution Detection: Solutions and Future Challenges
- Robust fine-tuning of zero-shot models
- On the Opportunities and Risks of Foundation Models
- Backup Control Barrier Functions: Formulation and Comparative Study
- Learning Transferable Visual Models From Natural Language Supervision
- Energy-based Out-of-distribution Detection
- A Unifying Review of Deep and Shallow Anomaly Detection
- MPNet: Masked and Permuted Pre-training for Language Understanding
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- An Efficient Reachability-Based Framework for Provably Safe Autonomous Navigation in Unknown Environments
- Control Barrier Functions: Theory and Applications
- Reachability-Based Safety Guarantees using Efficient Initializations
- Hamilton-Jacobi Reachability: A Brief Overview and Recent Advances
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- Reach-Avoid Problems with Time-Varying Dynamics, Targets and Constraints
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Cited by
Related