AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
2025/11/28 by Yibin Wen, Wen, Yibin, Qingmei Li +23
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2511.23253
Abstract
Recent advancements in Vision-Language Models (VLMs) have significantly impacted various industries. In agriculture, these multimodal capabilities hold great promise for applications such as precision farming, crop monitoring, pest detection, and environmental sustainability. However, while several Visual Question Answering (VQA) datasets and benchmarks have been developed to assess VLM performance, they often fail to effectively evaluate the critical reasoning and problem-solving skills needed in complex agricultural contexts. To address this gap, we introduce AgroCoT, a VQA dataset that integrates Chain-of-Thought (CoT) reasoning, specifically designed to evaluate the reasoning capabilities of VLMs. With 4,759 carefully curated samples, AgroCoT provides a comprehensive and robust evaluation of reasoning abilities, particularly in zero-shot scenarios, focusing on the models' ability to engage in logical reasoning and effective problem-solving. Our evaluation of 30 representative VLMs, including both proprietary and open-source models, reveals a gap in their reasoning capabilities, which underscores the importance of incorporating CoT for assessments. Our dataset is available at https://huggingface.co/datasets/AgroCoT/AgroCoT.
Citations
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- AgriGPT-VL: Agricultural Vision-Language Understanding Suite
- AgroBench: Vision-Language Model Benchmark in Agriculture
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
- Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
- THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
- AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
- Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- Qwen2.5-VL Technical Report
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
- Interleaved-Modal Chain-of-Thought
- GPT-4o System Card
- AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
- M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
- Enhancing Chain of Thought Prompting in Large Language Models via Reasoning Patterns
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- TinyLLaVA: A Framework of Small-scale Large Multimodal Models
- Assessing GPT4-V on Structured Reasoning Tasks
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Compositional Chain-of-Thought Prompting for Large Multimodal Models
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- PhenoBench: A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
- Large Language Models are Zero-Shot Reasoners
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Learning Transferable Visual Models From Natural Language Supervision
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis
- BERTScore: Evaluating Text Generation with BERT
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Related