vix.ing · top · new · best · stats

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

2024/10/15 by Hsin–Ping Huang, Xinyi Wang, Huang, Hsin-Ping +18 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Data Visualization and Analytics #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Video Analysis and Summarization

paper · pdf · doi:10.48550/arxiv.2410.11824

openalex publication_date 2024/10/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a wide variety of realistic visual entities. To bridge this gap, we propose KITTEN, a benchmark for Knowledge-InTensive image generaTion on real-world ENtities. Using KITTEN, we conduct a systematic study of the latest text-to-image models and retrieval-augmented models, focusing on their ability to generate real-world visual entities, such as landmarks and animals. Analysis using carefully designed human evaluations, automatic metrics, and MLLM evaluations show that even advanced text-to-image models fail to generate accurate visual details of entities. While retrieval-augmented models improve entity fidelity by incorporating reference images, they tend to over-rely on them and struggle to create novel configurations of the entity in creative text prompts.

Cited by

Related