Deep High-Resolution Representation Learning for Visual Recognition
2019/08/20 by Jingdong Wang, Wang, Jingdong, Ke Sun +21 · 161 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition #cs.CV
paper · pdf · doi:10.48550/arxiv.1908.07919
To appear in TPAMI. State-of-the-art performance on human pose estimation, semantic segmentation, object detection, instance segmentation, and face alignment. Full version of arXiv:1904.04514. (arXiv admin note: text overlap with arXiv:1904.04514)
arxiv created 2020/03/13 · arxiv updated 2020/03/16
Abstract
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the input image as a low-resolution representation through a subnetwork that is formed by connecting high-to-low resolution convolutions in series (e.g., ResNet, VGGNet), and then recover the high-resolution representation from the encoded low-resolution representation. Instead, our proposed network, named as High-Resolution Network (HRNet), maintains high-resolution representations through the whole process. There are two key characteristics: (i) Connect the high-to-low resolution convolution streams in parallel; (ii) Repeatedly exchange the information across resolutions. The benefit is that the resulting representation is semantically richer and spatially more precise. We show the superiority of the proposed HRNet in a wide range of applications, including human pose estimation, semantic segmentation, and object detection, suggesting that the HRNet is a stronger backbone for computer vision problems. All the codes are available at~\urlhttps://github.com/HRNet.
Citations
Cited by
- iOSPointMapper: RealTime Pedestrian and Accessibility Mapping with Mobile AI
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
- CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
- FastDDHPose: Towards Unified, Efficient, and Disentangled 3D Human Pose Estimation
- Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
- Physics Informed Human Posture Estimation Based on 3D Landmarks from Monocular RGB-Videos
- Heatmap Pooling Network for Action Recognition from RGB Videos
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- SemOD: Semantic Enabled Object Detection Network under Various Weather Conditions
- BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation
- Person Recognition in Aerial Surveillance: A Decade Survey
- RobustGait: Robustness Analysis for Appearance Based Gait Recognition
- RadHARSimulator V2: Video to Doppler Generator
- RAPTR: Radar-based 3D Pose Estimation using Transformer
- TrackStudio: An Integrated Toolkit for Markerless Tracking
- MedSapiens: Taking a Pose to Rethink Medical Imaging Landmark Detection
- Subsampled Randomized Fourier GaLore for Adapting Foundation Models in Depth-Driven Liver Landmark Segmentation
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- MeisenMeister: A Simple Two Stage Pipeline for Breast Cancer Classification on MRI
- Masked-attention Mask Transformer for Universal Image Segmentation
- Semi-Supervised Semantic Segmentation with Cross Pseudo Supervision
- LSKNet: A Foundation Lightweight Backbone for Remote Sensing
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
- Exploring Scale Shift in Crowd Localization under the Context of Domain Generalization
- UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning
- M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
- An Efficient Semantic Segmentation Decoder for In-Car or Distributed Applications
- Sample-Centric Multi-Task Learning for Detection and Segmentation of Industrial Surface Defects
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
- Paving the Way Towards Kinematic Assessment Using Monocular Video: A Preclinical Benchmark of State-of-the-Art Deep-Learning-Based 3D Human Pose Estimators Against Inertial Sensors in Daily Living Activities
- SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
- EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
- GLVD: Guided Learned Vertex Descent
- Bayesian Transformer for Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
- Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- Accurate Cobb Angle Estimation via SVD-Based Curve Detection and Vertebral Wedging Quantification
- Tent: Fully Test-time Adaptation by Entropy Minimization
- Stratify or Die: Rethinking Data Splits in Image Segmentation
- Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- ShipwreckFinder: A QGIS Tool for Shipwreck Detection in Multibeam Sonar Data
- BlurBall: Joint Ball and Motion Blur Estimation for Table Tennis Ball Tracking
- PMRT: A Training Recipe for Fast, 3D High-Resolution Aerodynamic Prediction
- LeViT-UNet: Make Faster Encoders with Transformer for Medical Image Segmentation
- Performance is not All You Need: Sustainability Considerations for Algorithms
- Proposal Learning for Semi-Supervised Object Detection
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- Probabilistic Robustness Analysis in High Dimensional Space: Application to Semantic Segmentation Network
- MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
- Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
- Segmenting Transparent Objects in the Wild
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
- SegAssess: Panoramic quality mapping for robust and transferable unsupervised segmentation assessment
- An End-to-End Framework for Video Multi-Person Pose Estimation
- Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- Quantitative Outcome-Oriented Assessment of Microsurgical Anastomosis
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Heatmap Regression without Soft-Argmax for Facial Landmark Detection
- Prior Guided Feature Enrichment Network for Few-Shot Segmentation
- DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation
- The Role of Radiographic Knee Alignment in Total Knee Replacement Outcomes and Opportunities for Artificial Intelligence-Driven Assessment
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- RadProPoser: A Framework for Human Pose Estimation with Uncertainty Quantification from Raw Radar Data
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
- Fast Neural Network Adaptation via Parameter Remapping and Architecture Search
- LawDIS: Language-Window-based Controllable Dichotomous Image Segmentation
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- A Dual-Feature Extractor Framework for Accurate Back Depth and Spine Morphology Estimation from Monocular RGB Images
- Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
- AFRDA: Attentive Feature Refinement for Domain Adaptive Semantic Segmentation
- A Novel Downsampling Strategy Based on Information Complementarity for Medical Image Segmentation
- HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery
- Adaptive Relative Pose Estimation Framework with Dual Noise Tuning for Safe Approaching Maneuvers
- Spatial Frequency Modulation for Semantic Segmentation
- Search to Distill: Pearls are Everywhere but not the Eyes
- CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- Combining Transformers and CNNs for Efficient Object Detection in High-Resolution Satellite Imagery
- Joint angle based learning to refine kinematic human pose estimation
- A New Dataset and Performance Benchmark for Real-time Spacecraft Segmentation in Onboard Flight Computers
- CWNet: Causal Wavelet Network for Low-Light Image Enhancement
- Admissibility of Stein Shrinkage for Batch Normalization in the Presence of Adversarial Attacks
- Attend-and-Refine: Interactive keypoint estimation and quantitative cervical vertebrae analysis for bone age assessment
- Circulating tumor cell detection in cancer patients using in-flow deep learning holography
- HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
- Reading a Ruler in the Wild
- Event-RGB Fusion for Spacecraft Pose Estimation Under Harsh Lighting
- Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
- Boundary-preserving Mask R-CNN
- Hierarchical Semantic-Visual Fusion of Visible and Near-infrared Images for Long-range Haze Removal
- Leveraging Out-of-Distribution Unlabeled Images: Semi-Supervised Semantic Segmentation with an Open-Vocabulary Model
- FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search
- FaceX-Zoo: A PyTorch Toolbox for Face Recognition
- Enabling Robust, Real-Time Verification of Vision-Based Navigation through View Synthesis
- Topology-Constrained Learning for Efficient Laparoscopic Liver Landmark Detection
- Geological Everything Model 3D: A Promptable Foundation Model for Unified and Zero-Shot Subsurface Understanding
- Trident: Detecting Face Forgeries with Adversarial Triplet Learning
- SuperAnimal pretrained pose estimation models for behavioral analysis
- Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection
- Weakly Supervised Object Segmentation by Background Conditional Divergence
- OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
- A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation
- AnyTraverse: An off-road traversability framework with VLM and human operator in the loop
- Echo-DND: A dual noise diffusion model for robust and precise left ventricle segmentation in echocardiography
- BCRNet: Enhancing Landmark Detection in Laparoscopic Liver Surgery via Bezier Curve Refinement
- FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
- MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style Convolution
- Structured Convolutions for Efficient Neural Network Design
- Efficient Differentiable Neural Architecture Search with Meta Kernels
- Analyzing Worldwide Social Distancing through Large-Scale Computer Vision
- FontAdapter: Instant Font Adaptation in Visual Text Generation
- CzechLynx: A Dataset for Individual Identification and Pose Estimation of the Eurasian Lynx
- ConText: Driving In-context Learning for Text Removal and Segmentation
- HRFormer: High-Resolution Transformer for Dense Prediction
- TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
- Affect-DML: Context-Aware One-Shot Recognition of Human Affect using Deep Metric Learning
- Knowledge Graphs for Digitized Manuscripts in Jagiellonian Digital Library Application
- YOLO-SPCI: Enhancing Remote Sensing Object Detection via Selective-Perspective-Class Integration
- The SpaceNet Multi-Temporal Urban Development Challenge
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams
- EMRA-proxy: Enhancing Multi-Class Region Semantic Segmentation in Remote Sensing Images with Attention Proxy
- SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
- Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
- Semantic Segmentation on VSPW Dataset through Aggregation of Transformer Models
- Multi-Resolution Haar Network: Enhancing human motion prediction via Haar transform
- Keypoints as Dynamic Centroids for Unified Human Pose and Segmentation
- ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
- Knowledge-Informed Deep Learning for Irrigation Type Mapping from Remote Sensing
- BEV-Seg: Bird's Eye View Semantic Segmentation Using Geometry and Semantic Point Cloud
- Boosting Cross-spectral Unsupervised Domain Adaptation for Thermal Semantic Segmentation
- Uni-AIMS: AI-Powered Microscopy Image Analysis
- Towards Better Cephalometric Landmark Detection with Diffusion Data Generation
- Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework is Enough for Infrared Small Target Detection
- Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
- iMacHSR: Intermediate Multi-Access Heterogeneous Supervision and Regularization Scheme Toward Architecture-Agnostic Training
- ClassWise-CRF: Category-Specific Fusion for Enhanced Semantic Segmentation of Remote Sensing Imagery
- Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
- Multi-axis Analysis of Image Manipulation Localization
- BARIS: Boundary-Aware Refinement with Environmental Degradation Priors for Robust Underwater Instance Segmentation
- Edges in image with illumination variations scenarios: a review
- LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
- Promptable Animal Pose Tracking Across Species
- Temporal Propagation of Asymmetric Feature Pyramid for Surgical Scene Segmentation
- All-in-One Transferring Image Compression from Human Perception to Multi-Machine Perception
- Graph Network for Sign Language Tasks
- BLAST: Bayesian online change-point detection with structured image data
- SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
- ToolTipNet: A Segmentation-Driven Deep Learning Baseline for Surgical Instrument Tip Detection
- MBE-ARI: A Multimodal Dataset Mapping Bi-directional Engagement in Animal-Robot Interaction
- Charm: The Missing Piece in ViT fine-tuning for Image Aesthetic Assessment
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
Related