Human Motion Video Generation: A Survey
2025/09/04 by Xue, Haiwei, Luo, Xiangyang, Hu, Zhanghao +12 · 12 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
paper · doi:10.48550/arxiv.2509.03883
Abstract
Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans. A complete list of the models examined in this survey is available in Our Repository https://github.com/Winn1y/Awesome-Human-Motion-Video-Generation.
Citations
- StableAnimator: High-Quality Identity-Preserving Human Image Animation
- Animate-X: Universal Character Image Animation with Enhanced Motion Representation
- Style-Preserving Lip Sync via Audio-Aware Style Reference
- Kalman-Inspired Feature Propagation for Video Face Super-Resolution
- IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
- WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
- TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions
- A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights
- MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
- FLUXSynID: A Synthetic Face Dataset with Document and Live Images
- MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
- Human Modelling and Pose Estimation Overview
- RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network
- Do As I Do: Pose Guided Human Motion Copy
- A Comprehensive Taxonomy and Analysis of Talking Head Synthesis: Techniques for Portrait Generation, Driving Mechanisms, and Editing
- Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
- Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation
- Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement
- Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
- Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
- V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation
- UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
- ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
- MegActor: Harness the Power of Raw Video for Vivid Portrait Animation
- MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
- VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
- Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
- Looking Backward: Streaming Video-to-Video Translation with Feature Banks
- InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
- ViViD: Video Virtual Try-on using Diffusion Models
- Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
- Dance Any Beat: Blending Beats with Visuals in Dance Video Generation
- Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
- AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
- EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
- Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos
- ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
- TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting
- Zero-shot High-fidelity and Pose-controllable Character Animation
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- EDTalk: Efficient Disentanglement for Emotional Talking Head Synthesis
- AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
- Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework
- X-Portrait: Expressive Portrait Animation with Hierarchical Motion Attention
- Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
- VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis
- FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio
- EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
- Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
- Magic-Me: Identity-Specific Video Customized Diffusion
- Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
- Synthesizing Moving People with 3D Control
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
- Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
- Towards a Simultaneous and Granular Identity-Expression Control in Personalized Face Generation
- I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
- Plan, Posture and Go: Towards Open-World Text-to-Motion Generation
- FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing
- StreamDiffusion: A Pipeline-level Solution for Real-time Interactive Generation
- MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
- LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
- Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realism
- FT2TF: First-Person Statement Text-To-Talking Face Generation
- R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning
- Digital Life Project: Autonomous 3D Characters with Social Intelligence
- AnimateZero: Video Diffusion Models are Zero-Shot Image Animators
- BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
- AgentAvatar: Disentangling Planning, Driving and Rendering for Photorealistic Avatar Agents
- Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
- AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond
- Adversarial Diffusion Distillation
- MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model
- InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
- Image-Based Virtual Try-On: A Survey
- VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
- Dance Your Latents: Consistent Dance Generation through Spatial-temporal Subspace Attention Guided by Motion Flow
- CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D Animation
- EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
- Expression Domain Translation Network for Cross-domain Head Reenactment
- 3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding
- Can I Trust Your Answer? Visually Grounded Video Question Answering
- MagicAvatar: Multimodal Avatar Generation and Animation
- Can Language Models Learn to Listen?
- Dual-Stream Diffusion Net for Text-to-Video Generation
- Dancing Avatar: Pose and Text-Guided Human Motion Videos Synthesis with Image Diffusion Model
- IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
- Text-to-Video: a Two-stage Framework for Zero-shot Identity-agnostic Talking-head Generation
- VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer
- Effective Whole-body Pose Estimation with Two-stages Distillation
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- Human Motion Generation: A Survey
- TokenFlow: Consistent Diffusion Features for Consistent Video Editing
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- Bidirectional Temporal Diffusion Model for Temporally Consistent Human Animation
- DisCo: Disentangled Control for Realistic Human Dance Generation
- MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- VideoComposer: Compositional Video Synthesis with Motion Controllability
- Identity-Preserving Talking Face Generation with Landmark and Appearance Priors
- StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator
- High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning
- GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
- Text2Performer: Text-Driven Human Video Generation
- DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion
- Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
- OmniAvatar: Geometry-Guided Controllable 3D Head Synthesis
- CelebV-Text: A Large-Scale Facial Text-Video Dataset
- OTAvatar: One-shot Talking Face Avatar with Controllable Tri-plane Rendering
- Conditional Image-to-Video Generation with Latent Flow Diffusion Models
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
- Human MotionFormer: Transferring Human Motions with Vision Transformers
- Adding Conditional Control to Text-to-Image Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models
- Structure and Content-Guided Video Synthesis with Diffusion Models
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
- Affective Faces for Goal-Driven Dyadic Communication
- Speech Driven Video Editing via an Audio-Conditioned Diffusion Model
- DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation
- Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
- StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
- Scalable Diffusion Models with Transformers
- Audio-Driven Co-Speech Gesture Video Generation
- High-fidelity Facial Avatar Reconstruction from Monocular Video with Generative Priors
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation
- LSA-T: The first continuous Argentinian Sign Language dataset for Sign Language Translation
- MEVID: Multi-view Extended Videos with Identities for Video Person Re-Identification
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- A Survey on Generative Diffusion Model
- A Survey on Generative Diffusion Models
- Diffusion Models: A Comprehensive Survey of Methods and Applications
- FaceOff: A Video-to-Video Face Swapping System
- CelebV-HQ: A Large-Scale Video Facial Attributes Dataset
- Fine-grained Activities of People Worldwide
- Towards Robust Blind Face Restoration with Codebook Lookup Transformer
- Open-Domain Sign Language Translation Learned from Online Video
- VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution
- Thin-Plate Spline Motion Model for Image Animation
- Learning to Answer Questions in Dynamic Audio-Visual Scenarios
- Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation
- StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN
- Responsive Listening Head Generation: A Benchmark Dataset and Baseline
- FaceFormer: Speech-Driven 3D Facial Animation with Transformers
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders
- LoRA: Low-Rank Adaptation of Large Language Models
- Diffusion Models Beat GANs on Image Synthesis
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation
- Audio-Driven Emotional Video Portraits
- Learning Transferable Visual Models From Natural Language Supervision
- Improved Denoising Diffusion Probabilistic Models
- Real-Time High-Resolution Background Matting
- Large-scale multilingual audio visual dubbing
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Denoising Diffusion Implicit Models
- Speech gesture generation from the trimodal context of text, audio, and speaker identity
- Speech Driven Talking Face Generation from a Single Image and an Emotion Condition
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Denoising Diffusion Probabilistic Models
- Improved Techniques for Training Score-Based Generative Models
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Improved Techniques for Training Single-Image GANs
- First Order Motion Model for Image Animation
- Analyzing and Improving the Image Quality of StyleGAN
- Capture, Learning, and Synthesis of 3D Speaking Styles
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields
- A Style-Based Generator Architecture for Generative Adversarial Networks
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- Everybody Dance Now
- Deep Video Portraits
- DensePose: Dense Human Pose Estimation In The Wild
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Neural Discrete Representation Learning
- MoCoGAN: Decomposing Motion and Content for Video Generation
- Attention Is All You Need
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Deep Unsupervised Learning using Nonequilibrium Thermodynamics
- Making a “Completely Blind” Image Quality Analyzer
- Image quality assessment: from error visibility to structural similarity
- Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
- Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling
- DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models
Cited by
Related