2024/03/07 by Nikhil Mishra, Mishra, Nikhil, Maximilian Sieb +5 · 1 citation
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #Data Visualization and Analytics #FOS: Computer and information sciences #Human Motion and Animation #Machine Learning (cs.LG) #Robotics (cs.RO) #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2403.04114
openalex publication_date 2024/03/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Deep learning methods for perception are the cornerstone of many robotic systems. Despite their potential for impressive performance, obtaining real-world training data is expensive, and can be impractically difficult for some tasks. Sim-to-real transfer with domain randomization offers a potential workaround, but often requires extensive manual tuning and results in models that are brittle to distribution shift between sim and real. In this work, we introduce Composable Object Volume NeRF (COV-NeRF), an object-composable NeRF model that is the centerpiece of a real-to-sim pipeline for synthesizing training data targeted to scenes and objects from the real world. COV-NeRF extracts objects from real images and composes them into new scenes, generating photorealistic renderings and many types of 2D and 3D supervision, including depth maps, segmentation masks, and meshes. We show that COV-NeRF matches the rendering quality of modern NeRF methods, and can be used to rapidly close the sim-to-real gap across a variety of perceptual modalities.