vix.ing · top · new · best · stats · spec

Advancing TDFN: Precise Fixation Point Generation Using Reconstruction Differences

2025/01/26 by Shuguang Wang, Wang, Shuguang, Wang, Yuanjing
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques

paper · pdf · doi:10.48550/arxiv.2501.15603

openalex publication_date 2025/01/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Wang and Wang (2025) proposed the Task-Driven Fixation Network (TDFN) based on the fixation mechanism, which leverages low-resolution information along with high-resolution details near fixation points to accomplish specific visual tasks. The model employs reinforcement learning to generate fixation points. However, training reinforcement learning models is challenging, particularly when aiming to generate pixel-level accurate fixation points on high-resolution images. This paper introduces an improved fixation point generation method by leveraging the difference between the reconstructed image and the input image to train the fixation point generator. This approach directs fixation points to areas with significant differences between the reconstructed and input images. Experimental results demonstrate that this method achieves highly accurate fixation points, significantly enhances the network's classification accuracy, and reduces the average number of required fixations to achieve a predefined accuracy level.

Related