MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
2026/05/11 by Alan Arazi, Eilam Shapira, Shoham Grunblat +8 · 1 voice
Computer Science · #cs.CL #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2605.10616
Abstract
Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstructured modalities such as text and image, and rely on frozen, pretrained embeddings to process them. On established Multimodal Tabular Learning benchmarks, we show that tuning the embeddings to the task improves performance. Existing benchmarks, however, often focus on the mere co-occurrence of modalities; this leads to high variance across datasets and masks the benefits of task-specific tuning. To address this gap, we introduce MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks. We focus on predictive tasks where the modalities provide complementary predictive signal, and where generic embeddings lose critical information, necessitating Target-Aware Representations that are aligned with the task. Our experimental results demonstrate that the gains from target-aware representation tuning generalize across both text and image modalities, several tabular learners, encoder scales, and embedding dimensions. MulTaBench constitutes the largest image-tabular benchmarking effort to date, spanning high-impact domains such as healthcare and e-commerce. It is designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Models.
Citations
- Retrieval from Within: An Intrinsic Capability of Attention-Based Models
- TabICLv2: A better, faster, scalable, and open tabular foundation model
- The Illusion of Generalization in Tabular Language Models
- Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers
- Can Agentic AI Match the Performance of Human Data Scientists?
- TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
- Structured RAG for Answering Aggregative Questions
- TabGemma: Text-Based Tabular ICL via LLM using Continued Pretraining and Retrieval
- Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- Data or Language Supervision: What Makes CLIP Better than DINO?
- jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
- Lost in Embeddings: Information Loss in Vision-Language Models
- Bringing Graphs to the Table: Zero-shot Node Classification via Tabular Foundation Models
- LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
- On the Theoretical Limitations of Embedding-Based Retrieval
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
- Should We Still Pretrain Encoders with Masked Language Modeling?
- TabArena: A Living Benchmark for Machine Learning on Tabular Data
- ConTextTab: A Semantics-Aware Tabular In-Context Learner
- Do-PFN: In-Context Learning for Causal Effect Estimation
- TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning
- TabSTAR: A Tabular Foundation Model for Tabular Data with Text Fields
- Table Foundation Models: on knowledge pre-training for tabular learning
- Representation Learning for Tabular Data: A Comprehensive Survey
- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
- Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data
- TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling
- TabDPT: Scaling Tabular Foundation Models on Real Data
- Do We Need Domain-Specific Embedding Models? An Empirical Investigation
- Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
- TIP: Tabular-Image Pre-training for Multimodal Classification with Incomplete Data
- Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
- A Closer Look at Deep Learning Methods on Tabular Datasets
- Large Scale Transfer Learning for Tabular Data via Language Modeling
- A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
- Why Tabular Foundation Models Should Be a Research Priority
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models
- Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
- PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning
- Making Pre-trained Language Models Great on Tabular Prediction
- Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
- CARTE: Pretraining and Transfer for Tabular Learning
- The Revolution of Multimodal Large Language Models: A Survey
- Question Aware Vision Transformer for Multimodal Reasoning
- Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Vectorizing string entries for data processing on tables: when are larger language models better?
- TabLib: A Dataset of 627M Tables with Context
- The first step is the hardest: Pitfalls of Representing and Tokenizing Temporal Data for Large Language Models
- Multimodal Temporal Fusion Transformers Are Good Product Demand Forecasters
- A survey on multimodal large language models
- LANISTR: Multimodal Learning from Structured and Unstructured Data
- QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations
- When Do Neural Nets Outperform Boosted Trees on Tabular Data?
- Visual Instruction Tuning
- Best of Both Worlds: Multimodal Contrastive Learning with Tabular and Imaging Data
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- TabLLM: Few-shot Classification of Tabular Data with Large Language Models
- MTEB: Massive Text Embedding Benchmark
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
- Flamingo: a Visual Language Model for Few-Shot Learning
- Deep Multi-modal Fusion of Image and Non-image Data in Disease Diagnosis and Prognosis: A Review
- Transformers Can Do Bayesian Inference
- Benchmarking Multimodal AutoML for Tabular Data with Text Fields
- Revisiting Deep Learning Models for Tabular Data
- LoRA: Low-Rank Adaptation of Large Language Models
- Tabular Data: Deep Learning is Not All You Need
- Representing Numbers in NLP: a Survey and a Vision
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- A Multimodal Approach to Predict Social Media Popularity
- Attention Is All You Need
- XGBoost: A Scalable Tree Boosting System
- VQA: Visual Question Answering
- Towards Benchmarking Foundation Models for Tabular Data With Text
Discussions
Related