2022/10/06 by Ilija Radosavovic, Tete Xiao, Radosavovic, Ilija +9 · 49 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Robotics (cs.RO) #cs.CV #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.2210.03109
CoRL 2022; Project page: https://tetexiao.com/projects/real-mvp
arxiv created 2022/10/06 · openalex publication_date 2022/10/06 · arxiv updated 2022/10/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this work, we explore self-supervised visual pre-training on images from diverse, in-the-wild videos for real-world robotic tasks. Like prior work, our visual representations are pre-trained via a masked autoencoder (MAE), frozen, and then passed into a learnable control module. Unlike prior work, we show that the pre-trained representations are effective across a range of real-world robotic tasks and embodiments. We find that our encoder consistently outperforms CLIP (up to 75%), supervised ImageNet pre-training (up to 81%), and training from scratch (up to 81%). Finally, we train a 307M parameter vision transformer on a massive collection of 4.5M images from the Internet and egocentric videos, and demonstrate clearly the benefits of scaling visual pre-training for robot learning.