The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models
2026/01/12 by Shuai Chen, Hao Chen, Yuanchen Bei +3 · 1 voice
Computer Science · #cs.CV
paper · pdf · doi:10.48550/arxiv.2601.08876
Abstract
Semantic information in embodied AI is inherently multi-source and multi-stage, making it challenging to fully leverage for achieving stable perception-to-action loops in real-world environments. Early studies have combined manual engineering with deep neural networks, achieving notable progress in specific semantic-related embodied tasks. However, as embodied agents encounter increasingly complex environments and open-ended tasks, the demand for more generalizable and robust semantic processing capabilities has become imperative. Recent advances in foundation models (FMs) address this challenge through their cross-domain generalization abilities and rich semantic priors, reshaping the landscape of embodied AI research. In this survey, we propose the Semantic Lifecycle as a unified framework to characterize the evolution of semantic knowledge within embodied AI driven by foundation models. Departing from traditional paradigms that treat semantic processing as isolated modules or disjoint tasks, our framework offers a holistic perspective that captures the continuous flow and maintenance of semantic knowledge. Guided by this embodied semantic lifecycle, we further analyze and compare recent advances across three key stages: acquisition, representation, and storage. Finally, we summarize existing challenges and outline promising directions for future research.
Citations
- A Comprehensive Survey on World Models for Embodied AI
- Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Occupancy Learning with Spatiotemporal Memory
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- FROSS: Faster-than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D Images
- Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
- I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
- RTMap: Real-Time Recursive Mapping with Change Detection and Localization
- GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
- Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
- LRSLAM: Low-rank Representation of Signed Distance Fields in Dense Visual SLAM System
- Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction
- Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
- STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction
- Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
- WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments
- VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
- Zero-Shot 4D Lidar Panoptic Segmentation
- Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
- Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
- PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding
- SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
- DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
- MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
- ViSpeak: Visual Instruction Feedback in Streaming Videos
- GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
- Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text
- Graph-Guided Scene Reconstruction from Images with 3D Gaussian Splatting
- Exploring Embodied Multimodal Large Models: Development, Datasets, and Future Directions
- CrossOver: 3D Scene Cross-Modal Alignment
- Phantom: Subject-consistent video generation via cross-modal alignment
- Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
- Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
- 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding
- Where am I? Cross-View Geo-localization with Natural Language Descriptions
- GaussianWorld: Gaussian World Model for Streaming 3D Occupancy\n Prediction
- GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
- TANGO: Training-free Embodied AI Agents for Open-world Tasks
- AffordDP: Generalizable Diffusion Policy with Transferable Affordance
- Navigation World Models
- SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
- GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
- T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
- ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
- Multimodal Alignment and Fusion: A Survey
- Towards Open-Vocabulary Audio-Visual Event Localization
- TimeFormer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction
- Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
- LG-Gaze: Learning Geometry-aware Continuous Prompts for Language-Guided Gaze Estimation
- DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes
- Estimating Body and Hand Motion in an Ego-sensed World
- CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction
- FlashSplat: 2D to 3D Gaussian Splatting Segmentation Solved Optimally
- PiTe: Pixel-Temporal Alignment for Large Video-Language Model
- MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation
- CMTA: Cross-Modal Temporal Alignment for Event-guided Video Deblurring
- Graph Retrieval-Augmented Generation: A Survey
- GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
- Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
- MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
- 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
- 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
- FREST: Feature RESToration for Semantic Segmentation under Multiple Adverse Conditions
- OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
- Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
- OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models
- Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
- EA-VTR: Event-Aware Video-Text Retrieval
- CPM: Class-conditional Prompting Machine for Audio-visual Segmentation
- Open-Vocabulary Semantic Segmentation with Image Embedding Balancing
- Neural Visibility Field for Uncertainty-Driven Active Mapping
- MAP-ADAPT: Real-Time Quality-Adaptive Semantic 3D Maps
- OED: Towards One-stage End-to-End Dynamic Scene Graph Generation
- A Survey on Vision-Language-Action Models for Embodied AI
- Tactile-Augmented Radiance Fields
- "Where am I?" Scene Retrieval with Language
- O2V-Mapping: Online Open-Vocabulary Mapping with Neural Implicit Representation
- COMO: Compact Mapping and Odometry
- GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic Fields
- SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs
- RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
- CG-SLAM: Efficient Dense RGB-D SLAM in a Consistent Uncertainty-aware 3D Gaussian Field
- Volumetric Environment Representation for Vision-Language Navigation
- RGBD GS-ICP SLAM
- GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
- Audio-Visual Segmentation via Unlabeled Frame Exploitation
- Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
- Towards Scene Graph Anticipation
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set Relationships
- Loopy-SLAM: Dense Neural SLAM with Loop Closures
- Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
- MaskClustering: View Consensus based Mask Graph Clustering for Open-Vocabulary 3D Instance Segmentation
- 3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation
- Retrieval-Augmented Egocentric Video Captioning
- Fully Sparse 3D Occupancy Prediction
- LangSplat: 3D Language Gaussian Splatting
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
- SAI3D: Segment Any Instance in 3D Scenes
- PLGSLAM: Progressive Neural Scene Represenation with Local to Global Bundle Adjustment
- ViLA: Efficient Video-Language Alignment for Video Question Answering
- WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion
- Instance Tracking in 3D Scenes from Egocentric Videos
- OneLLM: One Framework to Align All Modalities with Language
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
- Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
- PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness
- Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding
- SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
- Single-Model and Any-Modality for Video Object Tracking
- SelfOcc: Self-Supervised Vision-Based 3D Occupancy Prediction
- GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
- SNI-SLAM: Semantic Neural Implicit SLAM
- Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
- NExT-GPT: Any-to-Any Multimodal LLM
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Meta-Transformer: A Unified Framework for Multimodal Learning
- OpenMask3D: Open-Vocabulary 3D Instance Segmentation
- PanoOcc: Unified Occupancy Representation for Camera-based 3D Panoptic Segmentation
- Transformer-Based Visual Segmentation: A Survey
- Visual Instruction Tuning
- Micrograph segmentations for DDEVD
- Segment Anything
- LERF: Language Embedded Radiance Fields
- GPT-4 Technical Report
- PaLM-E: An Embodied Multimodal Language Model
- LLaMA: Open and Efficient Foundation Language Models
- Panoptic Lifting for 3D Scene Understanding with Neural Fields
- Flamingo: a Visual Language Model for Few-Shot Learning
- 3D Object Detection from Images for Autonomous Driving: A Survey
- Scene Graph Generation: A Comprehensive Survey
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Masked-attention Mask Transformer for Universal Image Segmentation
- On the Opportunities and Risks of Foundation Models
- Emerging Properties in Self-Supervised Vision Transformers
- A Survey of Embodied AI: From Simulators to Research Tasks
- Learning Transferable Visual Models From Natural Language Supervision
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Experience Grounds Language
- Image Segmentation Using Deep Learning: A Survey
- Deep Learning for 3D Point Clouds: A Survey
- Deep Learning for 3D Point Clouds: A Survey
- Mask R-CNN
- What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
- Past, Present, and Future of Simultaneous Localization And Mapping: Towards the Robust-Perception Age
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Discussions
Related