MOSPA: Human Motion Generation Driven by Spatial Audio
2025/07/16 by S. Xu, Xu, Shuyang, Zhiyang Dou +19 · 3 citations
Computer Science · Engineering · #Music Technology and Sound Studies #Human Motion and Animation #Music and Audio Processing
paper · pdf · doi:10.48550/arxiv.2507.11949
Abstract
Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive Spatial Audio-Driven Human Motion (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human MOtion generation driven by SPatial Audio, termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse, realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation
Citations
- Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
- ViSAGe: Video-to-Spatial Audio Generation
- CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects
- TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
- InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions
- ModSkill: Physical Character Skill Modularization
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
- ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model
- RMD: A Simple Baseline for More General Human Motion Generation via Training-free Retrieval-Augmented Motion Diffuse
- It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model
- SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
- MotionWavelet: Human Motion Prediction via Wavelet Manifold Learning
- Acoustic Volume Rendering for Neural Impulse Response Fields
- Pay Attention and Move Better: Harnessing Attention for Interactive Motion Generation and Training-free Editing
- Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
- UniMuMo: Unified Text, Music and Motion Generation
- MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting
- MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
- Modeling and Driving Human Body Soundfields through Acoustic Primitives
- SIGGesture: Generalized Co-Speech Gesture Synthesis via Semantic Injection with Large-Scale Pre-Training Diffusion Models
- InterAct: Capture and Modelling of Realistic, Expressive and Interactive Activities between Two Persons in Daily Scenarios
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- Generating Human Motion in 3D Scenes from Text Descriptions
- POPDG: Popular 3D Dance Generation with PopDanceSet
- MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model
- Taming Diffusion Probabilistic Models for Character Control
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos
- Large Motion Model for Unified Multi-Modal Motion Generation
- Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment
- Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
- LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment
- CoMo: Controllable Motion Generation through Language Guided Pose Code Editing
- Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
- Bidirectional Autoregressive Diffusion Model for Dance Generation
- BAT: Learning to Reason about Spatial Sounds with Large Language Models
- Plan, Posture and Go: Towards Open-World Text-to-Motion Generation
- EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation
- MoMask: Generative Masked Modeling of 3D Human Motions
- TLControl: Trajectory and Language Control for Human Motion Synthesis
- Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and Audio
- Controllable Group Choreography using Contrastive Diffusion
- HumanTOMATO: Text-aligned Whole-body Motion Generation
- OmniControl: Control Any Joint at Any Time for Human Motion Generation
- Human Motion Generation: A Survey
- Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
- MotionGPT: Human Motion as a Foreign Language
- Interactive Character Control with Auto-Regressive Motion Diffusion Models
- CoMusion: Towards Consistent Stochastic Human Motion Prediction via Motion Diffusion
- Perpetual Humanoid Control for Real-time Simulated Avatars
- TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents
- Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation
- HumanMAC: Masked Motion Completion for Human Motion Prediction
- T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations
- Executing your Commands via Motion Diffusion in Latent Space
- MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis
- FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation
- EDGE: Editable Dance Generation From Music
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- Human Motion Diffusion Model
- Towards Diverse and Natural Scene-aware 3D Human Motion Synthesis
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory
- Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation
- ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation
- Speech Drives Templates: Co-Speech Gesture Synthesis with Learned Templates
- Stochastic Scene-Aware Motion Prediction
- Scene-aware Generative Network for Human Motion Synthesis
- DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion Transformer
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
- Speech gesture generation from the trimodal context of text, audio, and speaker identity
- Denoising Diffusion Probabilistic Models
- Self-supervised Moving Vehicle Tracking with Stereo Sound
- Diverse Trajectory Forecasting with Determinantal Point Processes
- Learning Individual Styles of Conversational Gesture
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- On the Continuity of Rotation Representations in Neural Networks
- Mode-adaptive neural networks for quadruped motion control
- The Sound of Pixels
- Phase-functioned neural networks for character control
Cited by
Related