2017/11/21 by Bingqing Yu, Yu, Bingqing, James J. Clark +1
Computer Science · Neuroscience · #Aesthetic Perception and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Video Quality Assessment #Visual Attention and Saliency Detection
paper · pdf · doi:10.48550/arxiv.1711.08000
openalex publication_date 2017/11/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Most existing saliency models use low-level features or task descriptions when generating attention predictions. However, the link between observer characteristics and gaze patterns has been rarely investigated. We present a novel saliency prediction technique which takes into consideration viewers' identities and personal traits when modeling human attention. Not only computing image salience for average observers, we also focus on the interpersonal variation in the viewing behaviors of observers with different personal traits and backgrounds. We present an enriched derivative of the GAN network, which is able to generate personalized saliency predictions when fed with image stimuli and specific information about the observer. Our model contains a generator which generates grayscale saliency heat maps based on the image and an observer label. The generator is paired with an adversarial discriminator which learns to distinguish generated salience from ground truth salience. The discriminator also has the observer label as an input, which contributes to the personalization ability of our approach. We evaluate the performance of our personalized saliency model by comparing its results with some existing general saliency models. Improvements in prediction accuracy for all tested observer groups can be illustrated using various saliency evaluation metrics. Several fine-tuning techniques are applied to the network to adapt the network to several different benchmark saliency models. Thus, the model can deliver even higher prediction accuracy. Data augmentation operations are also performed to further improve the performance of the model. Experimental results show an improvement in prediction accuracy, when fine-tuning and data augmentation techniques are adopted. We also conduct experiments with a saliency personalization approach which implements several separate models. Each of the models is built and trained for one specific viewer group. The performance of this approach is evaluated and compared with the results generated by our previous approach. Finally, since valuable image features can be extracted by the learned discriminator, a fine-tuning process is adopted to further train the discriminator, and enable it to perform image classification tasks. Quantitative evaluation shows that the fine-tuned discriminator is capable of delivering high classification accuracy with the help of supervised training.