2025/04/18 by Ryota Takatsuki, Takatsuki, Ryota, Sonia Joseph +5 · 1 voice · 2 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Infrared Target Detection Methodologies #cs.AI #cs.CV
paper · pdf · doi:10.48550/arxiv.2504.13763
openalex publication_date 2025/04/18 · arxiv published 2025/04/18 · arxiv updated 2025/04/23 · openalex created_date 2025/10/11 · openalex updated_date 2026/07/28
Logit Lens is a widely adopted method for mechanistic interpretability of transformer-based language models, enabling the analysis of how internal representations evolve across layers by projecting them into the output vocabulary space. Although applying Logit Lens to Vision Transformers (ViTs) is technically straightforward, its direct use faces limitations in capturing the richness of visual representations. Building on the work of Toker et al. (2024)~\citeToker2024-ve, who introduced Diffusion Lens to visualize intermediate representations in the text encoders of text-to-image diffusion models, we demonstrate that while Diffusion Lens can effectively visualize residual stream representations in image encoders, it fails to capture the direct contributions of individual submodules. To overcome this limitation, we propose Diffusion Steering Lens (DSL), a novel, training-free approach that steers submodule outputs and patches subsequent indirect contributions. We validate our method through interventional studies, showing that DSL provides an intuitive and reliable interpretation of the internal processing in ViTs.