2025/05/19 by Junyi Hu, Hu, Junyi, Tian Bai +7
Computer Science · #Advanced Image Processing Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2505.12772
openalex publication_date 2025/05/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Feature fusion plays a pivotal role in achieving high performance in vision models, yet existing attention-based fusion techniques often suffer from substantial computational overhead and implementation complexity, particularly in resource-constrained settings. To address these limitations, we introduce the Plug-and-Play Hierarchical C2F Transformer (P2HCT), a lightweight module that combines coarse-to-fine token selection with shared attention parameters to preserve spatial details while reducing inference cost. P2HCT is trainable using coarse attention alone and can be seamlessly activated at inference to enhance accuracy without retraining. Integrated into real-time detectors such as YOLOv11-N/S/M, P2HCT achieves mAP gains of 0.9%, 0.5%, and 0.4% on MS COCO with minimal latency increase. Similarly, embedding P2HCT into ResNet-18/50/101 backbones improves ImageNet top-1 accuracy by 6.5%, 1.7%, and 1.0%, respectively. These results underscore P2HCT's effectiveness as a hardware-friendly and general-purpose enhancement for both detection and classification tasks.