Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform
2025/12/09 by Gong, Yuning, Liu, Yifei, Zhan, Yifan +21
Computer Science · Engineering · #3D Shape Modeling and Analysis #Artificial Intelligence (cs.AI) #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Graphics (cs.GR)
paper · doi:10.48550/arxiv.2512.08478
openalex publication_date 2025/12/09 · openalex created_date 2025/12/11 · openalex updated_date 2026/07/28
Abstract
Neural rendering, particularly 3D Gaussian Splatting (3DGS), has evolved rapidly and become a key component for building world models. However, existing viewer solutions remain fragmented, heavy, or constrained by legacy pipelines, resulting in high deployment friction and limited support for dynamic content and generative models. In this work, we present Visionary, an open, web-native platform for real-time various Gaussian Splatting and meshes rendering. Built on an efficient WebGPU renderer with per-frame ONNX inference, Visionary enables dynamic neural processing while maintaining a lightweight, "click-to-run" browser experience. It introduces a standardized Gaussian Generator contract, which not only supports standard 3DGS rendering but also allows plug-and-play algorithms to generate or update Gaussians each frame. Such inference also enables us to apply feedforward generative post-processing. The platform further offers a plug in three.js library with a concise TypeScript API for seamless integration into existing web applications. Experiments show that, under identical 3DGS assets, Visionary achieves superior rendering efficiency compared to current Web viewers due to GPU-based primitive sorting. It already supports multiple variants, including MLP-based 3DGS, 4DGS, neural avatars, and style transformation or enhancement networks. By unifying inference and rendering directly in the browser, Visionary significantly lowers the barrier to reproduction, comparison, and deployment of 3DGS-family methods, serving as a unified World Model Carrier for both reconstructive and generative paradigms.
Citations
- TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Tracking
- Alias-free 4D Gaussian Splatting
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- World Simulation with Video Foundation Models for Physical AI
- GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
- FlashWorld: High-quality 3D Scene Generation within Seconds
- ExGS: Extreme 3D Gaussian Compression with Diffusion Priors
- Proxy-GS: Efficient 3D Gaussian Splatting via Proxy Mesh
- DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments
- Epona: Autoregressive Diffusion World Model for Autonomous Driving
- RoboScape: Physics-informed Embodied World Model
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- Xray2Xray: World Model from Chest X-rays with Volumetric Context
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
- Video World Models with Long-term Spatial Memory
- AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
- AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
- FLARE: Robot Learning with Implicit World Modeling
- When Gaussian Meets Surfel: Ultra-fast High-fidelity Radiance Field Rendering
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation
- R3-Avatar: Record and Retrieve Temporal Codebook for Reconstructing Photorealistic Human Avatars
- VGGT: Visual Geometry Grounded Transformer
- LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds
- GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
- Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
- CellFlux: Simulating Cellular Morphology Changes via Flow Matching
- TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets
- OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation
- MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
- You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
- 2DGS-Room: Seed-Guided 2D Gaussian Splatting with Geometric Constrains for High-Fidelity Indoor Scene Reconstruction
- Sequential Gaussian Avatars with Hierarchical Motion Context
- Motion-Aware Animatable Gaussian Avatars Deblurring
- DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering
- ToMiE: Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars
- KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
- RaDe-GS: Rasterizing Depth in Gaussian Splatting
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- Diffusion for World Modeling: Visual Details Matter in Atari
- Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Dynamic Scenes
- RoboDreamer: Learning Compositional World Models for Robot Imagination
- Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes
- Per-Gaussian Embedding-Based Deformation for Deformable 3D Gaussian Splatting
- CityGaussian: Real-time High-quality Large-Scale Scene Rendering with Gaussians
- Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence
- Octree-GS: Towards Consistent Real-time Rendering with LOD-Structured 3D Gaussians
- GGRt: Towards Pose-free Generalizable 3D Gaussian Splatting in Real-time
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
- Learning and Leveraging World Models in Visual Representation Learning
- VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction
- Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis
- SWinGS: Sliding Windows for Dynamic 3D Gaussian Splatting
- pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction
- 3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting
- GauHuman: Articulated Gaussian Splatting from Monocular Human Videos
- SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes
- Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
- LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS
- Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
- Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing
- GS-IR: 3D Gaussian Splatting for Inverse Rendering
- Compact 3D Gaussian Representation for Radiance Field
- PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
- Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- MagicDrive: Street View Generation with Diverse 3D Geometry Control
- Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- K-Planes: Explicit Radiance Fields in Space, Time, and Appearance
- Is Conditional Generative Modeling all you need for Decision-Making?
- DynIBaR: Neural Dynamic Image-Based Rendering
- Instant neural graphics primitives with a multiresolution hash encoding
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular Video
- High-Resolution Image Synthesis with Latent Diffusion Models
- Plenoxels: Radiance Fields without Neural Networks
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance\n Fields
- Neural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic Humans
- D-NeRF: Neural Radiance Fields for Dynamic Scenes
- Denoising Diffusion Probabilistic Models
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- U-Net: Convolutional Networks for Biomedical Image Segmentation
Cited by
Related