Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
2025/05/26 by Tassilo Klein, Johannes Hoffart, Klein, Tassilo +1
Decision Sciences · #Artificial Intelligence (cs.AI) #Complex Systems and Decision Making #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2505.19825
openalex publication_date 2025/05/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/03
Abstract
This position paper argues that foundation models for tabular data face inherent limitations when isolated from operational context - the procedural logic, declarative rules, and domain knowledge that define how data is created and governed. Current approaches focus on single-table generalization or schema-level relationships, fundamentally missing the operational knowledge that gives data meaning. We introduce Semantically Linked Tables (SLT) and Foundation Models for SLT (FMSLT) as a new model class that grounds tabular data in its operational context. We propose dual-phase training: pre-training on open-source code-data pairs and synthetic systems to learn business logic mechanics, followed by zero-shot inference on proprietary data. We introduce the ``Operational Turing Test'' benchmark and argue that operational grounding is essential for autonomous agents in complex data environments.
Citations
- TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- Scaling Generalist Data-Analytic Agents
- AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
- On the Theoretical Limitations of Embedding-Based Retrieval
- Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
- TabArena: A Living Benchmark for Machine Learning on Tabular Data
- AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
- ConTextTab: A Semantics-Aware Tabular In-Context Learner
- Joint Relational Database Generation via Graph-Conditional Diffusion Models
- Griffin: Towards a Graph-Centric Relational Database Foundation Model
- Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
- GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction
- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
- CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation
- Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions
- ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events
- SALT: Sales Autocompletion Linked Business Tables Dataset
- Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- A text-to-tabular approach to generate synthetic patient data using LLMs
- Neural-Symbolic Reasoning over Knowledge Graphs: A Survey from a Query Perspective
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
- Drift-Resilient TabPFN: In-Context Learning Temporal Distribution Shifts on Tabular Data
- LLM-Based World Models Can Make Decisions Solely, But Rigorous Evaluations are Needed
- Tree-of-Table: Unleashing the Power of LLMs for Enhanced Large-Scale Table Understanding
- CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
- Number Cookbook: Number Understanding of Language Models and How to Improve It
- Generating Synthetic Electronic Health Record Data: a Methodological Scoping Review with Benchmarking on Phenotype Data and Open-Source Software
- Enhancing Table Representations with LLM-powered Synthetic Data Generation
- TableGPT2: A Large Multimodal Model with Tabular Data Integration
- Context is Key: A Benchmark for Forecasting with Essential Textual Information
- TabDPT: Scaling Tabular Foundation Models on Real Data
- Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning
- PORTAL: Scalable Tabular Foundation Models via Content-Specific Tokenization
- Scalable Representation Learning for Multimodal Tabular Transactions
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular Data
- Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning
- Text2SQL is Not Enough: Unifying AI and Databases with TAG
- Neuro-Symbolic Artificial Intelligence: Towards Improving the Reasoning Abilities of Large Language Models
- RelBench: A Benchmark for Deep Learning on Relational Databases
- Learning production functions for supply chains with graph neural networks
- PTaRL: Prototype-based Tabular Representation Learning via Space Calibration
- Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
- A Closer Look at Deep Learning Methods on Tabular Datasets
- Large Scale Transfer Learning for Tabular Data via Language Modeling
- CTSyn: A Foundation Model for Cross Tabular Data Generation
- Evaluating the World Model Implicit in a Generative Model
- RAG Does Not Work for Enterprises
- ClavaDDPM: Multi-relational Data Synthesis with Cluster-guided Diffusion Models
- Binning as a Pretext Task: Improving Self-Supervised Learning in Tabular Domains
- Why Tabular Foundation Models Should Be a Research Priority
- A Foundation Model for Zero-shot Logical Query Reasoning
- TWIN-GPT: Digital Twins for Clinical Trials via Large Language Model
- PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning
- TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
- Investigating Continual Pretraining in Large Language Models: Insights and Implications
- Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
- CARTE: Pretraining and Transfer for Tabular Learning
- LLMs with Industrial Lens: Deciphering the Challenges and Prospects -- A Survey
- Do causal predictors generalize better to new domains?
- Position: Graph Foundation Models are Already Here
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding
- Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
- Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-based Retrofitting
- NameGuess: Column Name Expansion for Tabular Data
- A Comprehensive Survey on Vector Database: Storage and Retrieval Technique, Challenge
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space
- From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models
- Towards Foundation Models for Knowledge Graph Reasoning
- Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Gorilla: Large Language Model Connected with Massive APIs
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- XTab: Cross-table Pretraining for Tabular Transformers
- Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs
- CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis
- PrivLava: Synthesizing Relational Data with Foreign Keys under Differential Privacy
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Large Language Models are Versatile Decomposers: Decompose Evidence and Questions for Table-based Reasoning
- Mastering Diverse Domains through World Models
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- The Stack: 3 TB of permissively licensed source code
- PAL: Program-aided Language Models
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Language Models are Realistic Tabular Data Generators
- STaSy: Score-based Tabular data Synthesis
- Binding Language Models in Symbolic Languages
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- Toward Compositional Generalization in Object-Oriented World Modeling
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- A Survey on Neural-symbolic Learning Systems
- A survey on neural-symbolic learning systems
- SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation Learning
- Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
- Evaluating Large Language Models Trained on Code
- SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption
- Revisiting Deep Learning Models for Tabular Data
- Well-tuned Simple Nets Excel on Tabular Datasets
- Tabular Data: Deep Learning is Not All You Need
- Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep Learning
- Neural Production Systems: Learning Rule-Governed Visual Dynamics
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Neurosymbolic AI: The 3rd Wave
- Language Models are Few-Shot Learners
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- The future of digital health with federated learning
- Knowledge Graphs
- Knowledge Graphs on the Web -- an Overview
- Graph Neural Networks Meet Neural-Symbolic Computing: A Survey and\n Perspective
- Contrastive Learning of Structured World Models
- Mastering Atari, Go, chess and shogi by planning with a learned model
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- Modeling Tabular data using Conditional GAN
- COBRA: Data-Efficient Model-Based RL through Unsupervised Object Discovery and Curiosity-Driven Exploration
- End-to-End Entity Resolution for Big Data: A Survey
- Neural-Symbolic Computing: An Effective Methodology for Principled Integration of Machine Learning and Reasoning
- World Models
- Neural-Symbolic Learning and Reasoning: A Survey and Interpretation
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Attention Is All You Need
- Inductive Representation Learning on Large Graphs
- YAGO2: A spatially and temporally enhanced knowledge base from Wikipedia
- BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network
- Toward principles for the design of ontologies used for knowledge sharing?
Related