TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
2021/05/17 by Fengbin Zhu, Zhu, Fengbin, Wenqiang Lei +14 · 117 citations
Computer Science · Decision Sciences · #Advanced Text Analysis Techniques #Stock Market Forecasting Methods #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2105.07624
Accepted by ACL 2021
arxiv created 2021/06/01 · arxiv updated 2021/06/02
Abstract
Hybrid data combining both tabular and textual content (e.g., financial reports) are quite pervasive in the real world. However, Question Answering (QA) over such hybrid data is largely neglected in existing research. In this work, we extract samples from real financial reports to build a new large-scale QA dataset containing both Tabular And Textual data, named TAT-QA, where numerical reasoning is usually required to infer the answer, such as addition, subtraction, multiplication, division, counting, comparison/sorting, and the compositions. We further propose a novel QA model termed TAGOP, which is capable of reasoning over both tables and text. It adopts sequence tagging to extract relevant cells from the table along with relevant spans from the text to infer their semantics, and then applies symbolic reasoning over them with a set of aggregation operators to arrive at the final answer. TAGOPachieves 58.0% inF1, which is an 11.1% absolute increase over the previous best baseline model, according to our experiments on TAT-QA. But this result still lags far behind performance of expert human, i.e.90.8% in F1. It is demonstrated that our TAT-QA is very challenging and can serve as a benchmark for training and testing powerful QA models that address hybrid form data.
Cited by
- INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models
- Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- CNFinBench: A Benchmark for Safety and Compliance of Large Language Models in Finance
- Jina-VLM: Small Multilingual Vision Language Model
- Beyond Patch Aggregation: 3-Pass Pyramid Indexing for Vision-Enhanced Document Retrieval
- Rethinking Retrieval: From Traditional Retrieval Augmented Generation to Agentic and Non-Vector Reasoning Systems in the Financial Domain for Large Language Models
- Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
- Attention Grounded Enhancement for Visual Document Retrieval
- TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
- Synthetic Data-Driven Prompt Tuning for Financial QA over Tables and Documents
- Reasoning on Time-Series for Financial Technical Analysis
- RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
- TabDSR: Decompose, Sanitize, and Reason for Complex Numerical Reasoning in Tabular Data
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Hierarchical Sequence Iteration for Heterogeneous Question Answering
- CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
- FineVision: Open Data Is All You Need
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
- FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
- FINCH: Financial Intelligence using Natural language for Contextualized SQL Handling
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
- FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
- FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
- One More Question is Enough, Expert Question Decomposition (EQD) Model for Domain Quantitative Reasoning
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Learning to Route: A Rule-Driven Agent Framework for Hybrid-Source Retrieval-Augmented Generation
- AuditAgent: Expert-Guided Multi-Agent Reasoning for Cross-Document Fraudulent Evidence Discovery
- Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Same Content, Different Representations: A Controlled Study for Table QA
- TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- Hierarchical Reranking for Scalable Financial RAG System
- Can GRPO Boost Complex Multimodal Table Understanding?
- DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- T2R-bench: A Benchmark for Generating Article-Level Reports from Real World Industrial Tables
- INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models
- PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
- FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
- FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models
- PrismRAG: Boosting RAG Factuality with Distractor Resilience and Strategized Reasoning
- AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- Improved LLM Agents for Financial Document Question Answering
- Grammar-Guided Evolutionary Search for Discrete Prompt Optimisation
- Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
- Unlocking Speech Instruction Data Potential with Query Rewriting
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
- Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
- What to Keep and What to Drop: Adaptive Table Filtering Framework
- TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
- Finance Language Model Evaluation (FLaME)
- MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
- RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
- MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
- No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
- TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document Reasoning
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
- Large Language Models are Good Relational Learners
- On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
- Improving LLMs with a knowledge from databases
- Plugging Schema Graph into Multi-Table QA: A Human-Guided Framework for Reducing LLM Reliance
- T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation
- FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
- Multimodal Tabular Reasoning with Privileged Structured Information
- TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
- Learning Sparsity for Effective and Efficient Music Performance Question Answering
- Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
- MMTBENCH: A Unified Benchmark for Complex Multimodal Table Reasoning
- RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
- Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
- FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
- TabularGSM: Understanding the Limitations of LLMs in Tabular Math Reasoning
- Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents
- Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
- Realistic Evaluation of TabPFN v2 in Open Environments
- Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
- KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval
- ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
- Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking
- mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs
- XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
- FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
- IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO
- EnronQA: Towards Personalized RAG over Private Documents
- SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA over Text and Tables
- Executable Schema Contracts: From Automatic Ingestion to Multi-Source Retrieval
- Hedge-Bench: Benchmarking Agents on Hard, Realistic Tasks Pertaining to Financial Reasoning
- FinTradeBench: A Financial Reasoning Benchmark for LLMs
- RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
- FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
- SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities
- Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach
- FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
- Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
- FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
- Improving Instruct Models for Free: A Study on Partial Adaptation
- Mixture-of-RAG: Integrating Text and Tables with Large Language Models
- VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
Related