2020/10/09 by Alex Trevithick, Bo Yang, Trevithick, Alex +1 · 8 citations
Computer Science · Earth and Planetary Sciences · Engineering · #3D Surveying and Cultural Heritage #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR) #Machine Learning (cs.LG) #Robotics (cs.RO) #Robotics and Sensor-Based Localization #cs.AI #cs.CV #cs.GR #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.2010.04595
ICCV 2021. Code and data are available at: https://github.com/alextrevithick/GRF
openalex publication_date 2020/10/09 · arxiv created 2021/08/11 · arxiv updated 2021/08/12 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
We present a simple yet powerful neural network that implicitly represents and renders 3D objects and scenes only from 2D observations. The network models 3D geometries as a general radiance field, which takes a set of 2D images with camera poses and intrinsics as input, constructs an internal representation for each point of the 3D space, and then renders the corresponding appearance and geometry of that point viewed from an arbitrary position. The key to our approach is to learn local features for each pixel in 2D images and to then project these features to 3D points, thus yielding general and rich point representations. We additionally integrate an attention mechanism to aggregate pixel features from multiple 2D views, such that visual occlusions are implicitly taken into account. Extensive experiments demonstrate that our method can generate high-quality and realistic novel views for novel objects, unseen categories and challenging real-world scenes.