Deep High-Resolution Representation Learning for Visual Recognition
2020/04/01 by Jingdong Wang, Ke Sun, Tianheng Cheng +9 · 105 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition
paper · doi:10.1109/tpami.2020.2983686
openalex publication_date 2020/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the input image as a low-resolution representation through a subnetwork that is formed by connecting high-to-low resolution convolutions in series (e.g., ResNet, VGGNet), and then recover the high-resolution representation from the encoded low-resolution representation. Instead, our proposed network, named as High-Resolution Network (HRNet), maintains high-resolution representations through the whole process. There are two key characteristics: (i) Connect the high-to-low resolution convolution streams in parallel and (ii) repeatedly exchange the information across resolutions. The benefit is that the resulting representation is semantically richer and spatially more precise. We show the superiority of the proposed HRNet in a wide range of applications, including human pose estimation, semantic segmentation, and object detection, suggesting that the HRNet is a stronger backbone for computer vision problems. All the codes are available at https://github.com/HRNet.
Citations
Cited by
- iOSPointMapper: RealTime Pedestrian and Accessibility Mapping with Mobile AI
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
- CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
- FastDDHPose: Towards Unified, Efficient, and Disentangled 3D Human Pose Estimation
- Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
- Physics Informed Human Posture Estimation Based on 3D Landmarks from Monocular RGB-Videos
- Heatmap Pooling Network for Action Recognition from RGB Videos
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- SemOD: Semantic Enabled Object Detection Network under Various Weather Conditions
- BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation
- Person Recognition in Aerial Surveillance: A Decade Survey
- RobustGait: Robustness Analysis for Appearance Based Gait Recognition
- RadHARSimulator V2: Video to Doppler Generator
- RAPTR: Radar-based 3D Pose Estimation using Transformer
- TrackStudio: An Integrated Toolkit for Markerless Tracking
- MedSapiens: Taking a Pose to Rethink Medical Imaging Landmark Detection
- Subsampled Randomized Fourier GaLore for Adapting Foundation Models in Depth-Driven Liver Landmark Segmentation
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- MeisenMeister: A Simple Two Stage Pipeline for Breast Cancer Classification on MRI
- Self-Regulation for Semantic Segmentation
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Perception for Autonomous Systems (PAZ)
- F3RNet: Full-Resolution Residual Registration Network for Deformable Image Registration
- Heterogeneous Grid Convolution for Adaptive, Efficient, and Controllable Computation
- Self-Supervised Monocular Depth Estimation with Internal Feature Fusion
- Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes
- BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation
- Mask Transfiner for High-Quality Instance Segmentation
- Semi-Supervised Recognition under a Noisy and Fine-grained Dataset
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- Exploring Scale Shift in Crowd Localization under the Context of Domain Generalization
- UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning
- M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
- An Efficient Semantic Segmentation Decoder for In-Car or Distributed Applications
- Sample-Centric Multi-Task Learning for Detection and Segmentation of Industrial Surface Defects
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- Learning Independent Instance Maps for Crowd Localization
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
- Implicit Feature Pyramid Network for Object Detection
- Paving the Way Towards Kinematic Assessment Using Monocular Video: A Preclinical Benchmark of State-of-the-Art Deep-Learning-Based 3D Human Pose Estimators Against Inertial Sensors in Daily Living Activities
- Indian Licence Plate Dataset in the wild
- Generalizable Pedestrian Detection: The Elephant In The Room
- EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
- GLVD: Guided Learned Vertex Descent
- HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers
- Bayesian Transformer for Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
- Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
- Accurate Cobb Angle Estimation via SVD-Based Curve Detection and Vertebral Wedging Quantification
- Stratify or Die: Rethinking Data Splits in Image Segmentation
- Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
- Audio-Visual Transformer Based Crowd Counting
- Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation
- FADE: A Task-Agnostic Upsampling Operator for Encoder–Decoder Architectures
- Parallel mesh reconstruction streams for pose estimation of interacting hands
- Anatomically Guided Deep Learning System for Right Internal Jugular Line (RIJL) Segmentation and Tip Localization in Chest X-Ray
- Contour Transformer Network for One-shot Segmentation of Anatomical Structures
- Meticulous Object Segmentation
- ShipwreckFinder: A QGIS Tool for Shipwreck Detection in Multibeam Sonar Data
- BlurBall: Joint Ball and Motion Blur Estimation for Table Tennis Ball Tracking
- PMRT: A Training Recipe for Fast, 3D High-Resolution Aerodynamic Prediction
- Survey on Deep Learning-based Kuzushiji Recognition
- Performance is not All You Need: Sustainability Considerations for Algorithms
- Swin Transformer-Based Multiscale Attention Model for Landslide Extraction From Large-Scale Area
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- Probabilistic Robustness Analysis in High Dimensional Space: Application to Semantic Segmentation Network
- MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
- Lite-HRNet: A Lightweight High-Resolution Network
- Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation
- TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
- SegAssess: Panoramic quality mapping for robust and transferable unsupervised segmentation assessment
- An End-to-End Framework for Video Multi-Person Pose Estimation
- End-to-End Human Pose and Mesh Reconstruction with Transformers
- Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- Bottom-Up Human Pose Estimation Via Disentangled Keypoint Regression
- RELLIS-3D Dataset: Data, Benchmarks and Analysis
- One-shot Unsupervised Domain Adaptation with Personalized Diffusion Models
- Multiscale Deep Equilibrium Models
- Quantitative Outcome-Oriented Assessment of Microsurgical Anastomosis
- LDC-Net: A Unified Framework for Localization, Detection and Counting in Dense Crowds
- Learning Versatile Neural Architectures by Propagating Network Codes
- 1st Place Solution to VisDA-2020: Bias Elimination for Domain Adaptive Pedestrian Re-identification
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Review of deep learning: concepts, CNN architectures, challenges, applications, future directions
- Bidirectional Multi-scale Attention Networks for Semantic Segmentation of Oblique UAV Imagery
- Heatmap Regression without Soft-Argmax for Facial Landmark Detection
- Improved Few-shot Segmentation by Redefinition of the Roles of Multi-level CNN Features
- The Role of Radiographic Knee Alignment in Total Knee Replacement Outcomes and Opportunities for Artificial Intelligence-Driven Assessment
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- Enhancing colorectal polyp segmentation with TCFMA-Net: A transformer-based cross feature and multi-attention network
- RadProPoser: A Framework for Human Pose Estimation with Uncertainty Quantification from Raw Radar Data
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- Half-Real Half-Fake Distillation for Class-Incremental Semantic Segmentation
- PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
- LawDIS: Language-Window-based Controllable Dichotomous Image Segmentation
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- A Dual-Feature Extractor Framework for Accurate Back Depth and Spine Morphology Estimation from Monocular RGB Images
Related