Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
2023/11/28 by Li Hu, Xin Gao, Hu, Li +9 · 1 voice · 256 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Animation #Artificial intelligence #Character (mathematics) #Character animation #Computer animation #Computer graphics (images) #Computer science #Computer vision #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Information retrieval #Merge (version control) #cs.CV
paper · pdf · doi:10.48550/arxiv.2311.17117
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/11/28 · openalex created_date 2023/12/01 · openalex updated_date 2026/07/28
Abstract
Character Animation aims to generating character videos from still images through driving signals. Currently, diffusion models have become the mainstream in visual generation research, owing to their robust generative capabilities. However, challenges persist in the realm of image-to-video, especially in character animation, where temporally maintaining consistency with detailed information from character remains a formidable problem. In this paper, we leverage the power of diffusion models and propose a novel framework tailored for character animation. To preserve consistency of intricate appearance features from reference image, we design ReferenceNet to merge detail features via spatial attention. To ensure controllability and continuity, we introduce an efficient pose guider to direct character's movements and employ an effective temporal modeling approach to ensure smooth inter-frame transitions between video frames. By expanding the training data, our approach can animate arbitrary characters, yielding superior results in character animation compared to other image-to-video methods. Furthermore, we evaluate our method on benchmarks for fashion video and human dance synthesis, achieving state-of-the-art results.
Cited by
- HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
- HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
- STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
- WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
- ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
- ID-V2V: Identity-Preserving Video Restylization
- High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
- Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
- TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
- Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
- EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- Map2Video: Street View Imagery Driven AI Video Generation
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- 3DProxyImg: Controllable 3D-Aware Animation Synthesis from Single Image via 2D-3D Aligned Proxy Embedding
- TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
- AnimaMimic: Imitating 3D Animation from Video Priors
- PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
- KlingAvatar 2.0 Technical Report
- FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
- PersonaLive! Expressive Portrait Image Animation for Live Streaming
- Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation
- StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
- VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
- GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
- Blur2Sharp: Human Novel Pose and View Synthesis with Generative Prior Refinement
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
- VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- RunawayEvil: Jailbreaking the Image-to-Video Generative Models
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- 2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
- Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
- OmniPerson: Unified Identity-Preserving Pedestrian Generation
- StyleYourSmile: Cross-Domain Face Retargeting Without Paired Multi-Style Data
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- Layer-Aware Video Composition via Split-then-Merge
- Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
- SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- Any4D: Open-Prompt 4D Generation from Natural Language and Images
- Plan-X: Instruct Video Generation via Semantic Planning
- Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
- Generative Augmented Reality: Paradigms, Technologies, and Future Applications
- TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Deep Inverse Shading: Consistent Albedo and Surface Detail Recovery via Generative Refinement
- ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
- DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
- Neural USD: An object-centric framework for iterative editing and control
- Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
- Video-As-Prompt: Unified Semantic Control for Video Generation
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- UltraGen: High-Resolution Video Generation with Hierarchical Attention
- From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display
- Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction Animation
- VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
- VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
- Sketch Animation: State-of-the-art Report
- ReMix: Towards a Unified View of Consistent Character Generation and Editing
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
- ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
- SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
- Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
- StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
- ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
- MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
- SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- Talking Head Generation via AU-Guided Landmark Prediction
- TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
- OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Lynx: Towards High-Fidelity Personalized Video Generation
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- From Rigging to Waving: 3D-Guided Diffusion for Natural Animation of Hand-Drawn Characters
- Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Human Motion Video Generation: A Survey
- FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
- DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
- Diverse Signer Avatars with Manual and Non-Manual Feature Modelling for Sign Language Production
- EmoCAST: Emotional Talking Portrait via Emotive Text Description
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation
- PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
- InfinityHuman: Towards Long-Term Audio-Driven Human
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors
- Precise Action-to-Video Generation Through Visual Action Prompts
- Odo: Depth-Guided Diffusion for Identity-Preserving Body Reshaping
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
- X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
- Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- CharacterShot: Controllable and Consistent 4D Character Animation
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- MonoCloth: Reconstruction and Animation of Cloth-Decoupled Human Avatars from Monocular Videos
- 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
- Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
- Multi-human Interactive Talking Dataset
- X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
- DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
- PoseGuard: Pose-Guided Generation with Safety Guardrails
- Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
- FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
- Video Color Grading via Look-Up Table Generation
- Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
- ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
- StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
- Controllable Video Generation: A Survey
- Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
- Synthetic Human Action Video Data Generation with Pose Transfer
- From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation
- RefSTAR: Blind Facial Image Restoration with Reference Selection, Transfer, and Reconstruction
- EgoAnimate: Generating Human Animations from Egocentric top-down Views
- Democratizing High-Fidelity Co-Speech Gesture Video Generation
- LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion
- Generative Head-Mounted Camera Captures for Photorealistic Avatars
- UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
- CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
- DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
- UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
- Populate-A-Scene: Affordance-Aware Human Video Generation
- Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
- MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
- SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
- PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
- Rethink Sparse Signals for Pose-guided Text-to-image Generation
- Whole-Body Conditioned Egocentric Video Prediction
- Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
- Controllable and Expressive One-Shot Video Head Swapping
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
- FramePrompt: In-context Controllable Animation with Zero Structural Changes
- Toward Rich Video Human-Motion2D Generation
- Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
- iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer
- Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
- DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
- Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On
- LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
- Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
- Person-In-Situ: Scene-Consistent Human Image Insertion with Occlusion-Aware Pose Control
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation
- Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
- GUAVA: Generalizable Upper Body 3D Gaussian Avatar
- Controllable Human-centric Keyframe Interpolation with Generative Prior
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
- Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
- Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
- OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
- DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On
- DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
- A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
- FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
- Hallo4: High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
- MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
- GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
- HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions
- How Animals Dance (When You're Not Looking)
- Generating Fit Check Videos with a Handheld Camera
- LatentMove: Towards Complex Human Movement Video Generation
- Geometry-Editable and Appearance-Preserving Object Compositon
- MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
- IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model
- OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
- AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
- Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
- DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
- UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
- Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection
- Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
- Interspatial Attention for Efficient 4D Human Video Generation
- CoT-Edit: Let CoT Guide Instruction Video Editing
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
- MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation
- EnerVerse-AC: Envisioning Embodied Environments with Action Condition
- TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
- 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
- ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
- SVAD: From Single Image to 3D Avatar via Synthetic Data Generation with Video Diffusion and Data Augmentation
- ReactDance: Hierarchical Representation for High-Fidelity and Coherent Long-Form Reactive Dance Generation
- KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
- MagicPortrait: Temporally Consistent Face Reenactment with 3D Geometric Guidance
- ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
- CompleteMe: Reference-based Human Image Completion
- AnimateAnywhere: Rouse the Background in Human Image Animation
- DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
- Multi-View Face and Gesture Animation with Dynamic Gaussians
- SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning
- ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
- DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment
- RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild
- Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
- DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
- DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
- TwoSquared: 4D Generation from 2D Image Pairs
- Wan-Animate-2: Pushing the Application Boundaries of Character Animation
- Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models
- UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
- OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
- Taming Consistency Distillation for Accelerated Human Image Animation
- HUMOTO: A 4D Dataset of Mocap Human Object Interactions
- SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
- TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
- ColorizeDiffusion v2: Enhancing Reference-based Sketch Colorization Through Separating Utilities
- GIGA: Generalizable Sparse Image-driven Gaussian Humans
Discussions
Related