Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
2024/10/10 by Jiahao Cui, Hui Li, Cui, Jiahao +15 · 48 citations
Computer Science · #Advanced Vision and Imaging #Animation #Art #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer graphics (images) #Computer science #Computer vision #Duration (music) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Geography #High resolution #Image (mathematics) #Literature #Portrait #Remote sensing #Visual arts
paper · pdf · doi:10.48550/arxiv.2410.07718
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/10/10 · openalex created_date 2024/10/13 · openalex updated_date 2026/07/28
Abstract
Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities. First, we extend the method to produce long-duration videos. To address substantial challenges such as appearance drift and temporal artifacts, we investigate augmentation strategies within the image space of conditional motion frames. Specifically, we introduce a patch-drop technique augmented with Gaussian noise to enhance visual consistency and temporal coherence over long duration. Second, we achieve 4K resolution portrait video generation. To accomplish this, we implement vector quantization of latent codes and apply temporal alignment techniques to maintain coherence across the temporal dimension. By integrating a high-quality decoder, we realize visual synthesis at 4K resolution. Third, we incorporate adjustable semantic textual labels for portrait expressions as conditional inputs. This extends beyond traditional audio cues to improve controllability and increase the diversity of the generated content. To the best of our knowledge, Hallo2, proposed in this paper, is the first method to achieve 4K resolution and generate hour-long, audio-driven portrait image animations enhanced with textual prompts. We have conducted extensive experiments to evaluate our method on publicly available datasets, including HDTF, CelebV, and our introduced "Wild" dataset. The experimental results demonstrate that our approach achieves state-of-the-art performance in long-duration portrait video animation, successfully generating rich and controllable content at 4K resolution for duration extending up to tens of minutes. Project page https://fudan-generative-vision.github.io/hallo2
Cited by
- TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- InstanceV: Instance-Level Video Generation
- IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
- THEval. Evaluation Framework for Talking Head Video Generation
- Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
- Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
- A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- Audio Driven Real-Time Facial Animation for Social Telepresence
- KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
- Talking Head Generation via AU-Guided Landmark Prediction
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- Reconstruction and Reenactment Separated Method for Realistic Gaussian Head
- EmoCAST: Emotional Talking Portrait via Emotive Text Description
- EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
- Preacher: Paper-to-Video Agentic System
- Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
- RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
- Multi-human Interactive Talking Dataset
- X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
- Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
- Style Transfer: A Decade Survey
- Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
- OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
- Audio-Sync Video Generation with Multi-Stream Temporal Control
- LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
- Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
- TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection
- FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
- OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
- Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection
- Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
- Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
- DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
- FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
Related