VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution
2022/05/06 by Liangbin Xie. Xintao Wang, Wang, Liangbin Xie. Xintao, Honglun Zhang +5 · 62 citations
Computer Science · Engineering · #Advanced Image Processing Techniques #Face recognition and analysis #Speech and Audio Processing #cs.AI #cs.CV #cs.MM #eess.IV
paper · pdf · doi:10.48550/arxiv.2205.03409
Project webpage available at https://liangbinxie.github.io/projects/vfhq
arxiv created 2022/05/06 · arxiv updated 2022/05/10
Abstract
Most of the existing video face super-resolution (VFSR) methods are trained and evaluated on VoxCeleb1, which is designed specifically for speaker identification and the frames in this dataset are of low quality. As a consequence, the VFSR models trained on this dataset can not output visual-pleasing results. In this paper, we develop an automatic and scalable pipeline to collect a high-quality video face dataset (VFHQ), which contains over 16,000 high-fidelity clips of diverse interview scenarios. To verify the necessity of VFHQ, we further conduct experiments and demonstrate that VFSR models trained on our VFHQ dataset can generate results with sharper edges and finer textures than those trained on VoxCeleb1. In addition, we show that the temporal information plays a pivotal role in eliminating video consistency issues as well as further improving visual performance. Based on VFHQ, by analyzing the benchmarking study of several state-of-the-art algorithms under bicubic and blind settings. See our project page: https://liangbinxie.github.io/projects/vfhq
Cited by
- ViDS: Video Diffusion Shader using 3D Face Tracking
- SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait Animation
- FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
- FlexAvatar: Learning Complete 3D Head Avatars with Partial Supervision
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
- Towards Interactive Intelligence for Digital Humans
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- PersonaLive! Expressive Portrait Image Animation for Live Streaming
- DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
- A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
- AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars
- TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- Bringing Your Portrait to 3D Presence
- IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- DINO-Tok: Adapting DINO for Visual Tokenizers
- ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
- PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing
- OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild
- MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
- SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
- StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
- Collaborative feature aggregation for face super-resolution and robust re-identification
- Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed Image
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Human Motion Video Generation: A Survey
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Phased One-Step Adversarial Equilibrium for Video Diffusion Models
- EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- Controllable Video Generation: A Survey
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
- CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
- GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
- Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
- Controllable and Expressive One-Shot Video Head Swapping
- EchoShot: Multi-Shot Portrait Video Generation
- DicFace: Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration
- BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading
- Low-Rank Head Avatar Personalization with Registers
- FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
- Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting
- Eye-See-You: Reverse Pass-Through VR and Head Avatars
- IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos
- Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
- SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D Gaussians
- FlexIP: Dynamic Control of Preservation and Personality for Customized Image Generation
Related