Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
2023/11/28 by Li Hu, Xin Gao, Hu, Li +9 · 1 voice · 145 citations
#cs.CV
paper · pdf · doi:10.48550/arxiv.2311.17117
Abstract
Character Animation aims to generating character videos from still images through driving signals. Currently, diffusion models have become the mainstream in visual generation research, owing to their robust generative capabilities. However, challenges persist in the realm of image-to-video, especially in character animation, where temporally maintaining consistency with detailed information from character remains a formidable problem. In this paper, we leverage the power of diffusion models and propose a novel framework tailored for character animation. To preserve consistency of intricate appearance features from reference image, we design ReferenceNet to merge detail features via spatial attention. To ensure controllability and continuity, we introduce an efficient pose guider to direct character's movements and employ an effective temporal modeling approach to ensure smooth inter-frame transitions between video frames. By expanding the training data, our approach can animate arbitrary characters, yielding superior results in character animation compared to other image-to-video methods. Furthermore, we evaluate our method on benchmarks for fashion video and human dance synthesis, achieving state-of-the-art results.
Cited by
- HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
- HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
- STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
- WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
- ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
- ID-V2V: Identity-Preserving Video Restylization
- High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
- Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
- TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
- Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
- EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- Map2Video: Street View Imagery Driven AI Video Generation
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- 3DProxyImg: Controllable 3D-Aware Animation Synthesis from Single Image via 2D-3D Aligned Proxy Embedding
- TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
- AnimaMimic: Imitating 3D Animation from Video Priors
- PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
- KlingAvatar 2.0 Technical Report
- FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
- PersonaLive! Expressive Portrait Image Animation for Live Streaming
- Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation
- StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
- VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
- GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
- Blur2Sharp: Human Novel Pose and View Synthesis with Generative Prior Refinement
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
- VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- RunawayEvil: Jailbreaking the Image-to-Video Generative Models
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- 2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
- Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
- OmniPerson: Unified Identity-Preserving Pedestrian Generation
- StyleYourSmile: Cross-Domain Face Retargeting Without Paired Multi-Style Data
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- Layer-Aware Video Composition via Split-then-Merge
- Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
- SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- Any4D: Open-Prompt 4D Generation from Natural Language and Images
- Plan-X: Instruct Video Generation via Semantic Planning
- Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
- Generative Augmented Reality: Paradigms, Technologies, and Future Applications
- TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Deep Inverse Shading: Consistent Albedo and Surface Detail Recovery via Generative Refinement
- ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
- DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model
- Neural USD: An object-centric framework for iterative editing and control
- Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
- Video-As-Prompt: Unified Semantic Control for Video Generation
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- UltraGen: High-Resolution Video Generation with Hierarchical Attention
- From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display
- Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction Animation
- VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
- VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
- Sketch Animation: State-of-the-art Report
- ReMix: Towards a Unified View of Consistent Character Generation and Editing
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
- ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
- SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
- Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
- StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
- ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
- MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
- SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- Talking Head Generation via AU-Guided Landmark Prediction
- TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
- OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Lynx: Towards High-Fidelity Personalized Video Generation
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- From Rigging to Waving: 3D-Guided Diffusion for Natural Animation of Hand-Drawn Characters
- Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Human Motion Video Generation: A Survey
- FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
- DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
- Diverse Signer Avatars with Manual and Non-Manual Feature Modelling for Sign Language Production
- EmoCAST: Emotional Talking Portrait via Emotive Text Description
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation
- PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
- InfinityHuman: Towards Long-Term Audio-Driven Human
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors
- Precise Action-to-Video Generation Through Visual Action Prompts
- Odo: Depth-Guided Diffusion for Identity-Preserving Body Reshaping
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
- X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
- Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- CharacterShot: Controllable and Consistent 4D Character Animation
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- MonoCloth: Reconstruction and Animation of Cloth-Decoupled Human Avatars from Monocular Videos
- 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
- Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
- Multi-human Interactive Talking Dataset
- X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
- DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
- PoseGuard: Pose-Guided Generation with Safety Guardrails
- Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
- Video Color Grading via Look-Up Table Generation
- Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
Discussions
Related