vix.ing · top · new · best · stats · spec

A No-Reference Quality Assessment Model for Screen Content Videos via Hierarchical Spatiotemporal Perception

2024/09/27 by Zhihong Liu, Huanqiang Zeng, Jing Chen +3
Computer Science · #Image and Video Quality Assessment

paper · doi:10.1109/tcsvt.2024.3469180

Abstract

In this paper, a novel deep learning-based no-reference video quality assessment (NR-VQA) model for screen content videos (SCVs) is proposed, called the hierarchical spatiotemporal perceptual quality model (HSPQ). Firstly, the human visual system (HVS) perceives SCVs hierarchically, with varying sensitivity and attention to diverse attribute regions. Secondly, the visual redundancies are copious in the spatiotemporal domain of SCVs, degrading video quality to some extent. Based on these characteristics, the SCVs are decomposed into three hierarchical levels (i.e., patch level, frame level, and video level), which contain quality-related spatiotemporal information. Specifically, the visual saliency is first utilized for more salient textual and pictorial patches selection, and then, a dual-channel convolutional neural network integrating spatial-gate feature enhancement module (SGFEM) is designed to evaluate the quality of patches based on their attributes at the patch level separately. With spatial correlation, an adaptive blur-focused visual mechanism-based weighting strategy (BFWS) is proposed for converting quality scores from patch level to frame level. Finally, the video-level quality score, which reflects the temporal perceptual quality degradation, is combined to provide a comprehensive evaluation of distorted SCV quality. Experiments conducted on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases demonstrate that our proposed HSPQ model aligns better with the visual perception of SCVs by the HVS. Moreover, it exhibits strong robustness compared to multiple classic and state-of-the-art image/video quality assessment models.

Citations

Related