2020/11/30 by Zongyu Guo, Zhizheng Zhang, Runsen Feng +1 · 162 citations
Computer Science · Engineering · #Advanced Data Compression Techniques #Advanced Image Processing Techniques #Algorithm #Artificial intelligence #Codec #Computer science #Context model #Data compression #Decoding methods #ENCODE #Entropy (arrow of time) #Entropy encoding #Image (mathematics) #Image and Signal Denoising Methods #Image compression #Image processing #Leverage (statistics) #Pattern recognition (psychology) #Rate–distortion theory #cs.CV #eess.IV
paper · pdf · doi:10.1109/tcsvt.2021.3089491
published in IEEE Transactions on Circuits and Systems for Video Technology 32(4), 2329-2341 (Institute of Electrical and Electronics Engineers) · We add some descriptions for the improved quantization in the latest arxiv version
openalex publication_date 2021/06/15 · arxiv created 2021/10/31 · arxiv updated 2021/11/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Over the past several years, we have witnessed impressive progress in the field of learned image compression. Recent learned image codecs are commonly based on autoencoders, that first encode an image into low-dimensional latent representations and then decode them for reconstruction purposes. To capture spatial dependencies in the latent space, prior works exploit hyperprior and spatial context model to build an entropy model, which estimates the bit-rate for end-to-end rate-distortion optimization. However, such an entropy model is suboptimal from two aspects: (1) It fails to capture global-scope spatial correlations among the latents. (2) Cross-channel relationships of the latents remain unexplored. In this paper, we propose the concept of separate entropy coding to leverage a serial decoding process for causal contextual entropy prediction in the latent space. Acausal context modelis proposed that separates the latents across channels and makes use of channel-wise relationships to generate highly informative adjacent contexts. Furthermore, we propose acausal global prediction modelto find global reference points for accurate predictions of undecoded points. Both these two models facilitate entropy estimation without the transmission of overhead. In addition, we further adopt a new group-separated attention module to build more powerful transform networks. Experimental results demonstrate that our full image compression model outperforms standard VVC/H.266 codec on Kodak dataset in terms of both PSNR and MS-SSIM, yielding the state-of-the-art rate-distortion performance.