vix.ing · top · new · best · stats · spec

Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?

2025/05/16 by Tairan Fu, Fu, Tairan, Miguel Ángel Pesquera González +7
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Language and cultural evolution #Multimodal Machine Learning Applications #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2505.10862

openalex publication_date 2025/05/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Multimodal Large Language Models which can answer complex questions on an image struggle to tell the time on analog clocks. This is probably due to the lack of images with clocks at different times in their training set. In this work we explore this issue with one of the latest MLLMs: GPT-4.1 to understand why MLLMs fail to tell the time and whether fine-tuning can solve the problem. The results show how models are making progress in reading the time on analog clocks. But have they really learned to do it, or have they only learned patterns in their training datasets? In this work we put the models to the test with different clocks to illustrate the limitations of MLLMs to abstract and generalize.

Citations

Related