2026/05/01 by Zihan Zhou, Luxi Chen, Jingzhi Zhou +4 · 1 voice
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Diffusion #Geometric shape #Geometric transformation #Monocular #Perspective (graphical) #Point (geometry) #Point cloud #Robotics and Sensor-Based Localization #Transformation (genetics) #cs.CV
paper · pdf · doi:10.48550/arxiv.2605.00345
openalex publication_date 2026/05/01 · arxiv published 2026/05/01 · arxiv updated 2026/05/01 · openalex created_date 2026/05/05 · openalex updated_date 2026/07/28
Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end diffusion framework that synthesizes 3D geometry directly within the observation space. By unprojecting monocular depth into a partial point cloud and explicitly injecting it as a 3D geometric anchor, PAD abandons canonical assumptions to enforce rigorous spatial supervision. This native generation intrinsically resolves pose ambiguity, producing high-fidelity pose-aligned assets. Extensive experiments demonstrate that PAD achieves superior geometric alignment and image-to-3D correspondence compared to state-of-the-art methods. Additionally, PAD naturally extends to compositional 3D scene reconstruction via a simple union of independently generated objects, highlighting its robust ability to preserve precise spatial layouts.