vix.ing · top · new · best · stats · spec

MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization

2025/08/11 by A. Jain, Jain, Animesh, Alexandros Stergiou +1 · 1 citation
Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2508.07833

Abstract

Vision Language Models (VLMs) encode multimodal inputs over large, complex, and difficult-to-interpret architectures, which limit transparency and trust. We propose a Multimodal Inversion for Model Interpretation and Conceptualization (MIMIC) framework that inverts the internal encodings of VLMs. MIMIC uses a joint VLM-based inversion and a feature alignment objective to account for VLM's autoregressive processing. It additionally includes a triplet of regularizers for spatial alignment, natural image smoothness, and semantic realism. We evaluate MIMIC both quantitatively and qualitatively by inverting visual concepts across a range of free-form VLM outputs of varying length. Reported results include both standard visual quality metrics and semantic text-based metrics. To the best of our knowledge, this is the first model inversion approach addressing visual interpretations of VLM concepts.

Citations

Cited by

Related