A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
2026/01/29 by Runze Cheng, Yao Sun, Ahmad Taha +3 · 1 voice
Engineering · Computer Science · #eess.IV #cs.CV
paper · pdf · doi:10.1109/tnse.2026.3671446
arxiv published 2026/01/29 · arxiv updated 2026/05/29
Abstract
Semantic communication (SemCom) emerges as a transformative paradigm for traffic-intensive visual data transmission, shifting focus from raw data to meaningful content transmission and relieving the increasing pressure on communication resources. However, to achieve SemCom, challenges are faced in accurate semantic quantization for visual data, robust semantic extraction and reconstruction under diverse tasks and goals, transceiver coordination with effective knowledge utilization, and adaptation to unpredictable wireless communication environments. In this paper, we present a systematic review of SemCom for visual data transmission (SemCom-Vision), wherein an interdisciplinary analysis integrating computer vision (CV) and communication engineering is conducted to provide comprehensive guidelines for the machine learning (ML)-empowered SemCom-Vision design. Specifically, this survey first elucidates the basics and key concepts of SemCom. Then, we introduce a novel classification perspective to categorize existing SemCom-Vision approaches as semantic preservation communication (SPC), semantic expansion communication (SEC), and semantic refinement communication (SRC) based on communication goals interpreted through semantic quantization schemes. Moreover, this survey articulates the ML-based encoder-decoder models and training algorithms for each SemCom-Vision category, followed by knowledge structure and utilization strategies. Finally, we discuss potential SemCom-Vision applications.
Citations
- A Mathematical Framework of Semantic Communication based on Category Theory
- SIMAC: A Semantic-Driven Integrated Multimodal Sensing And Communication Framework
- Semantic Communications with Computer Vision Sensing for Edge Video Transmission
- M4SC: An MLLM-based Multi-modal, Multi-task and Multi-user Semantic Communication System
- Lightweight Vision Model-based Multi-user Semantic Communication Systems
- Task-Oriented Semantic Communication for Stereo-Vision 3D Object Detection
- UAV Cognitive Semantic Communications Enabled by Knowledge Graph for Robust Object Detection
- Transmit What You Need: Task-Adaptive Semantic Communications for Visual Information
- HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
- Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications
- Semantic-Aware Resource Management for C-V2X Platooning via Multi-Agent Reinforcement Learning
- Semantic Feature Decomposition based Semantic Communication System of Images with Large-scale Visual Generation Models
- Knowledge-Assisted Privacy Preserving in Semantic Communication
- Personalized Federated Learning for Generative AI-Assisted Semantic Communications
- Goal-Oriented Semantic Communication for Wireless Image Transmission via Stable Diffusion
- Diffusion-Driven Semantic Communication for Generative Models with Bandwidth Constraints
- S-RAN: Semantic-Aware Radio Access Networks
- Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
- Multimodal Reasoning with Multimodal Knowledge Graph
- Robust Image Semantic Coding with Learnable CSI Fusion Masking over MIMO Fading Channels
- Visual Language Model based Cross-modal Semantic Communication Systems
- Knowledge-aware Text-Image Retrieval for Remote Sensing Images
- A Survey on Semantic Communication Networks: Architecture, Security, and Privacy
- Cross-Modal Generative Semantic Communications for Mobile AIGC: Joint Semantic Encoding and Prompt Engineering
- Image Generative Semantic Communication with Multi-Modal Similarity Estimation for Resource-Limited Networks
- A Robust Semantic Communication System for Image
- Large Generative Model Assisted 3D Semantic Communication
- Semantic Entropy Can Simultaneously Benefit Transmission Efficiency and Channel Security of Wireless Semantic Communications
- A Mathematical Theory of Semantic Communication: Overview
- A Mathematical Theory of Semantic Communication
- Importance-Aware Image Segmentation-based Semantic Communication for Autonomous Driving
- Generative AI-driven Semantic Communication Networks: Architecture, Technologies and Applications
- Knowledge Base Enabled Semantic Communication: A Generative Perspective
- Joint Source-Channel Coding for Channel-Adaptive Digital Semantic Communications
- Joint Sensing and Semantic Communications with Multi-Task Deep Learning
- Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI
- Deep Image Semantic Communication Model for Artificial Intelligent Internet of Things
- A Wireless AI-Generated Content (AIGC) Provisioning Framework Empowered by Semantic Communication
- Compression Ratio Learning and Semantic Communications for Video Imaging
- Language-Oriented Communication with Semantic Coding and Knowledge Distillation for Text-to-Image Generation
- CDDM: Channel Denoising Diffusion Models for Wireless Semantic Communications
- Semantic Communications Based on Adaptive Generative Models and Information Bottleneck
- Large AI Model Empowered Multimodal Semantic Communications
- Large AI Model Empowered Multimodal Semantic Communications
- Visual Analytics For Machine Learning: A Data Perspective Survey
- Large AI Model-Based Semantic Communications
- SCAN: Semantic Communication with Adaptive Channel Feedback
- A Unified Framework for Integrating Semantic Communication and AI-Generated Content in Metaverse
- Trust-Worthy Semantic Communications for the Metaverse Relying on Federated Learning
- Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations
- Causal Semantic Communication for Digital Twins: A Generalizable Imitation Learning Approach
- Knowledge Enhanced Graph Neural Networks for Graph Completion
- Improved Nonlinear Transform Source-Channel Coding to Catalyze Semantic Communications
- DeepMA: End-to-end Deep Multiple Access for Wireless Image Transmission in Semantic Communication
- Cognitive Semantic Communication Systems Driven by Knowledge Graph: Principle, Implementation, and Performance Evaluation
- Wireless End-to-End Image Transmission System using Semantic Communications
- Task-oriented Explainable Semantic Communications
- Adding Conditional Control to Text-to-Image Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models
- Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring
- Real-Time Digital Twins: Vision and Research Directions for 6G and Beyond
- Real-Time Digital Twins: Vision and Research Directions for 6G and Beyond
- Asynchronous Hybrid Reinforcement Learning for Latency and Reliability Optimization in the Metaverse over Wireless Communications
- Less Data, More Knowledge: Building Next Generation Semantic Communication Networks
- Semantic-Aware Sensing Information Transmission for Metaverse: A Contest Theoretic Approach
- Deep Joint Source-Channel Coding for Semantic Communications
- Contrastive Language-Image Pre-Training with Knowledge Graphs
- Image Segmentation Semantic Communication over Internet of Vehicles
- The Unreasonable Effectiveness of Fully-Connected Layers for Low-Data Regimes
- Rethinking Wireless Communication Security in Semantic Internet of Things
- Personalized Saliency in Task-Oriented Semantic Communications: Image Transmission and Performance Analysis
- Towards Semantic Communications: Deep Learning-Based Image Semantic Coding
- Metaverse for Wireless Systems: Vision, Enablers, Architecture, and Future Directions
- Semantic Communications for Future Internet: Fundamentals, Applications, and Challenges
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Innovative semantic communication system
- Digital Twin of Wireless Systems: Overview, Taxonomy, Challenges, and Opportunities
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Semantic Communications: Principles and Challenges
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Task-Oriented Multi-User Semantic Communications
- Zero-shot and Few-shot Learning with Knowledge Graphs: A Comprehensive Survey
- Attention mechanisms in computer vision: A survey
- Masked Autoencoders Are Scalable Vision Learners
- Probabilistic Entity Representation Model for Reasoning over Knowledge Graphs
- What is Semantic Communication? A View on Conveying Meaning in the Era of Machine Intelligence
- Generative Adversarial Networks
- Task-Oriented Multi-User Semantic Communications for VQA Task
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- A review on the attention mechanism of deep learning
- CvT: Introducing Convolutions to Vision Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- Conditional Image Generation by Conditioning Variational Auto-Encoders
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Transformers in Vision: A Survey
- A Survey on Vision Transformer
- Score-Based Generative Modeling through Stochastic Differential Equations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Denoising Diffusion Implicit Models
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- A Lite Distributed Semantic Communication System for Internet of Things
- Denoising Diffusion Probabilistic Models
- Deep Learning Enabled Semantic Communication Systems
- Directly Mapping RDF Databases to Property Graph Databases
- A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
- Fundamental tenis : resep meraih kemenangan / Tony Mottram
- A Survey on Knowledge Graphs: Representation, Acquisition and Applications
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- Deep Learning vs. Traditional Computer Vision
- Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification
- Knowledge Representation Learning: A Quantitative Review
- A Style-Based Generator Architecture for Generative Adversarial Networks
- A Style-Based Generator Architecture for Generative Adversarial Networks
- Activation Functions: Comparison of trends in Practice and Research for Deep Learning
- CBAM: Convolutional Block Attention Module
- Self-Attention Generative Adversarial Networks
- An Equivalence of Fully Connected Layer and Convolutional Layer
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Neural Discrete Representation Learning
- Squeeze-and-Excitation Networks
- Nonparametric regression using deep neural networks with ReLU activation function
- Attention Is All You Need
- Inductive Representation Learning on Large Graphs
- Understanding Convolution for Semantic Segmentation
- Image-to-Image Translation with Conditional Adversarial Networks
- Densely Connected Convolutional Networks
- Probabilistic Knowledge Graph Construction: Compositional and Incremental Approaches
- Layer Normalization
- Wide Residual Networks
- Resnet in Resnet: Generalizing Residual Architectures
- Deep Residual Learning for Image Recognition
- An Introduction to Convolutional Neural Networks
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Attribute-Graph: A Graph based approach to Image Ranking
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- The Treasure beneath Convolutional Layers: Cross-convolutional-layer Pooling for Image Classification
- Fully Convolutional Networks for Semantic Segmentation
- Conditional Generative Adversarial Nets
- Generalized Denoising Auto-Encoders as Generative Models
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Discussions
Related