ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
2025/11/29 by Xu, Zhengzhuo, Du, SiNan, Qi, Yiyan +4 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2512.00305
Abstract
Multimodal Large Language Models (MLLMs) have emerged as powerful tools for chart comprehension. However, they heavily rely on extracted content via OCR, which leads to numerical hallucinations when chart textual annotations are sparse. While existing methods focus on scaling instructions, they fail to address the fundamental challenge, i.e., reasoning with visual perception. In this paper, we identify a critical observation: MLLMs exhibit weak grounding in chart elements and proportional relationships, as evidenced by their inability to localize key positions to match their reasoning. To bridge this gap, we propose PointCoT, which integrates reflective interaction into chain-of-thought reasoning in charts. By prompting MLLMs to generate bounding boxes and re-render charts based on location annotations, we establish connections between textual reasoning steps and visual grounding regions. We further introduce an automated pipeline to construct ChartPoint-SFT-62k, a dataset featuring 19.2K high-quality chart samples with step-by-step CoT, bounding box, and re-rendered visualizations. Leveraging this data, we develop two instruction-tuned models, ChartPointQ2 and ChartPointQ2.5, which outperform state-of-the-art across several chart benchmarks, e.g., +5.04% on ChartBench.
Citations
- Qwen2.5-VL Technical Report
- Efficient Reasoning with Hidden Thinking
- Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- ChartAdapter: Large Vision-Language Model for Chart Summarization
- AskChart: Universal Chart Understanding through Textual Enhancement
- Qwen2.5 Technical Report
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
- Evaluation of OpenAI o1: Opportunities and Challenges of AGI
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
- GoT-CQA: Graph-of-Thought Guided Compositional Reasoning for Chart Question Answering
- VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning
- Connector-S: A Survey of Connectors in Multi-modal Large Language Models
- The Llama 3 Herd of Models
- Qwen2 Technical Report
- PaliGemma: A versatile 3B VLM for transfer
- ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
- CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
- Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- InternLM2 Technical Report
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
- ChartThinker: A Contextual Chain-of-Thought Approach to Optimized Chart Summarization
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- ChartReformer: Natural Language-Driven Chart Image Editing
- Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
- ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
- KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- TinyLlama: An Open-Source Small Language Model
- ChartBench: A Benchmark for Complex Visual Reasoning in Charts
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Compositional Chain-of-Thought Prompting for Large Multimodal Models
- ChartLlama: A Multimodal LLM for Chart Understanding and Generation
- MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- Mistral 7B
- Improved Baselines with Visual Instruction Tuning
- DOMINO: A Dual-System for Multi-step Visual Language Reasoning
- Qwen Technical Report
- InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- Visual Instruction Tuning
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Multimodal Chain-of-Thought Reasoning in Language Models
- DePlot: One-shot visual language reasoning by plot-to-table translation
- MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
- Flamingo: a Visual Language Model for Few-Shot Learning
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- INTERN: A New Learning Paradigm Towards General Vision
- Language Models are Few-Shot Learners
- PlotQA: Reasoning over Scientific Plots
- xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related