TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
2025/09/30 by Kimihiro Hasegawa, Hasegawa, Kimihiro, Wiradee Imrattanatrai +7 · 1 voice
Computer Science · Psychology · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Intelligent Tutoring Systems and Adaptive Learning #cs.CL
paper · pdf · doi:10.48550/arxiv.2510.00161
openalex publication_date 2025/09/30 · arxiv published 2025/09/30 · arxiv updated 2025/09/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite its potential use cases, the system development tailored for such an assistant is still underexplored. In this paper, we propose a novel framework, called TAMA, a Tool-Augmented Multimodal Agent, for procedural activity understanding. TAMA enables interleaved multimodal reasoning by making use of multimedia-returning tools in a training-free setting. Our experimental result on the multimodal procedural QA dataset, ProMQA-Assembly, shows that our approach can improve the performance of vision-language models, especially GPT-5 and MiMo-VL. Furthermore, our ablation studies provide empirical support for the effectiveness of two features that characterize our framework, multimedia-returning tools and agentic flexible tool selection. We believe our proposed framework and experimental results facilitate the thinking with images paradigm for video and multimodal tasks, let alone the development of procedural activity assistants.
Citations
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- T*: Re-thinking Temporal Search for Long-Form Video Understanding
- Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
- Qwen2.5-VL Technical Report
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Large Language Model-Brained GUI Agents: A Survey
- ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
- CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
- IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like Setting
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World
- The Rise and Potential of Large Language Model Based Agents: A Survey
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- Multimodal Subtask Graph Generation from Instructional Videos
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Large Language Models are Zero-Shot Reasoners
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain
- The IKEA ASM Dataset: Understanding People Assembling Furniture through\n Actions, Objects and Pose
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- OpenAI o1 System Card
Discussions
Related