OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
2025/10/16 by Zhe Li, Li, Zhe, Weihao Yuan +7 · 3 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2510.14954
openalex publication_date 2025/10/16 · openalex created_date 2025/10/18 · openalex updated_date 2026/07/28
Abstract
Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous methods that usually employ discrete masked modeling or autoregressive modeling, we develop a continuous masked autoregressive motion transformer, where a causal attention is performed considering the sequential nature within the human motion. Within this transformer, we introduce a gated linear attention and an RMSNorm module, which drive the transformer to pay attention to the key actions and suppress the instability caused by either the abnormal movements or the heterogeneous distributions within multi-modalities. To further enhance both the motion generation and the multimodal generalization, we employ the DiT structure to diffuse the conditions from the transformer towards the targets. To fuse different modalities, AdaLN and cross-attention are leveraged to inject the text, speech, and music signals. Experimental results demonstrate that our framework outperforms previous methods across all modalities, including text-to-motion, speech-to-gesture, and music-to-dance. The code of our method will be made public.
Citations
- Motion Anything: Any to Motion Generation
- MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
- Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
- Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
- LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning
- Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
- MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling
- MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
- Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
- Autoregressive Image Generation without Vector Quantization
- M3GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
- Large Motion Model for Unified Multi-Modal Motion Generation
- Towards Variable and Coordinated Holistic Co-Speech Motion Generation
- ParCo: Part-Coordinating Text-to-Motion Synthesis
- Motion Mamba: Efficient and Long Sequence Motion Generation
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and Quantization
- Bidirectional Autoregressive Diffusion Model for Dance Generation
- MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning
- DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
- FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing
- MMM: Generative Masked Motion Model
- MoMask: Generative Masked Modeling of 3D Human Motions
- AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond
- A Unified Framework for Multimodal, Multi-Part Human Motion Synthesis
- General Point Model with Autoencoding and Autoregressive
- Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model
- MCM: Multi-condition Motion Synthesis Framework for Multi-scenario
- DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion
- AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention Mechanism
- Priority-Centric Human Motion Generation in Discrete Latent Space
- Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
- MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
- DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model
- Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation
- T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations
- Muse: Text-To-Image Generation via Masked Generative Transformers
- Executing your Commands via Motion Diffusion in Latent Space
- Generating Holistic 3D Human Motion from Speech
- MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis
- FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation
- PhysDiff: Physics-Guided Human Motion Diffusion Model
- UDE: A Unified Driving Engine for Human Motion Generation
- EDGE: Editable Dance Generation From Music
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- Being Comes from Not-being: Open-vocabulary Text-to-Motion Generation with Wordless Training
- Human Motion Diffusion Model
- FLAME: Free-form Language-based Motion Synthesis & Editing
- TEACH: Temporal Action Composition for 3D Humans
- MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
- TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts
- TEMOS: Generating diverse human motions from textual descriptions
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory
- MotionCLIP: Exposing Human Motion Generation to CLIP Space
- MaskGIT: Masked Generative Image Transformer
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAE
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
- Hierarchical Style-based Networks for Motion Synthesis
- Perpetual Motion: Generating Unbounded Human Motion
- Learning Diverse Stochastic Human-Action Generators by Learning Smooth Latent Transitions
- Root Mean Square Layer Normalization
- Language2Pose: Natural Language Grounded Pose Forecasting
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- Neural Discrete Representation Learning
- MoCoGAN: Decomposing Motion and Content for Video Generation
- Deep Residual Learning for Image Recognition
Cited by
Related