2024/05/11 by Ancheng Lin, Lin, Ancheng, Tianqing Su +7 · 5 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #cs.CV
paper · pdf · doi:10.48550/arxiv.2405.06945
openalex publication_date 2024/05/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28 · arxiv created 2026/08/03 · arxiv updated 2026/08/04
Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential for maintaining high-fidelity real-to-sim environments for embodied intelligence, where local changes should be incorporated without complete reconstruction. Existing neural surface reconstruction and 3DGS-to-mesh pipelines often learn geometry indirectly or separate geometry construction from appearance modeling. This separation introduces optimization redundancy and makes local geometry or appearance updates expensive. We propose an end-to-end mesh-Gaussian scene representation that binds 3D Gaussians to mesh faces and uses differentiable 3DGS rendering for photometric supervision. This design provides a direct information pathway for jointly learning explicit geometry and renderable appearance. Experiments on indoor and outdoor scenes demonstrate improved efficiency and rendering quality while preserving high-quality surface reconstruction. The explicit mesh also enables mesh-based manipulation, and the coupled representation adapts efficiently to local scene modifications. These properties support scalable visual scene modeling and the efficient maintenance of real-to-sim environments for embodied-agent training and evaluation.