VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation
2025/06/11 by Hyeongcheol Park, Park, Hyeongcheol, Seo, Jiyoung +11 · 1 citation
Computer Science · #Advanced Graph Neural Networks #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2506.21556
openalex publication_date 2025/06/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Multimodal Knowledge Graphs (MMKGs), which represent explicit knowledge across multiple modalities, play a pivotal role by complementing the implicit knowledge of Multimodal Large Language Models (MLLMs) and enabling more grounded reasoning via Retrieval Augmented Generation (RAG). However, existing MMKGs are generally limited in scope: they are often constructed by augmenting pre-existing knowledge graphs, which restricts their knowledge, resulting in outdated or incomplete knowledge coverage, and they often support only a narrow range of modalities, such as text and visual information. These limitations restrict applicability to multimodal tasks, particularly as recent MLLMs adopt richer modalities like video and audio. Therefore, we propose the Visual-Audio-Text Knowledge Graph (VAT-KG), the first concept-centric and knowledge-intensive multimodal knowledge graph that covers visual, audio, and text information, where each triplet is linked to multimodal data and enriched with detailed descriptions of concepts. Specifically, our construction pipeline ensures cross-modal knowledge alignment between multimodal data and fine-grained semantics through a series of stringent filtering and alignment steps, enabling the automatic generation of MMKGs from any multimodal dataset. We further introduce a novel multimodal RAG framework that retrieves detailed concept-level knowledge in response to queries from arbitrary modalities. Experiments on question answering tasks across various modalities demonstrate the effectiveness of VAT-KG in supporting MLLMs, highlighting its practical value in unifying and leveraging multimodal knowledge.
Citations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Qwen2.5-Omni Technical Report
- Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
- Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- VideoRAG: Retrieval-Augmented Generation over Video Corpus
- Re-ranking the Context for Multimodal Retrieval Augmented Generation
- GPT-4o System Card
- LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- AudioBench: A Universal Benchmark for Audio Large Language Models
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
- Multimodal Reasoning with Multimodal Knowledge Graph
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction
- CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
- AspectMMKG: A Multi-modal Knowledge Graph with Aspect-aware Entities
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- ImageBind: One Embedding Space To Bind Them All
- Evaluating ChatGPT's Information Extraction Capabilities: An Assessment of Performance, Explainability, Calibration, and Faithfulness
- VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
- Visual Instruction Tuning
- Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline
- GPT-4 Technical Report
- UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
- Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation
- Multimodal Analogical Reasoning over Knowledge Graphs
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Flamingo: a Visual Language Model for Few-Shot Learning
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- VisualSem: A High-quality Knowledge Graph for Vision and Language
- Visual Transformers: Token-based Image Representation and Processing for Computer Vision
- Language Models are Few-Shot Learners
- VGGSound: A Large-scale Audio-Visual Dataset
- KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- MMKG: Multi-Modal Knowledge Graphs
- Billion-scale similarity search with GPUs
- 3D-LLM: Injecting the 3D World into Large Language Models
Cited by
Related