A Comprehensive Survey of Synthetic Tabular Data Generation
2025/04/23 by Shi, Ruxue, Yili Wang, Wang, Yili +8 · 43 citations
Computer Science · #Data Management and Algorithms #FOS: Computer and information sciences #Machine Learning (cs.LG) #Time Series Analysis and Forecasting #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.2504.16506
openalex publication_date 2025/04/23 · openalex created_date 2025/10/11 · openalex updated_date 2026/07/28
Abstract
Tabular data is one of the most prevalent and important data formats in real-world applications such as healthcare, finance, and education. However, its effective use in machine learning is often constrained by data scarcity, privacy concerns, and class imbalance. Synthetic tabular data generation has emerged as a powerful solution, leveraging generative models to learn underlying data distributions and produce realistic, privacy-preserving samples. Although this area has seen growing attention, most existing surveys focus narrowly on specific methods (e.g., GANs or privacy-enhancing techniques), lacking a unified and comprehensive view that integrates recent advances such as diffusion models and large language models (LLMs). In this survey, we present a structured and in-depth review of synthetic tabular data generation methods. Specifically, the survey is organized into three core components: (1) Background, which covers the overall generation pipeline, including problem definitions, synthetic tabular data generation methods, post processing, and evaluation; (2) Generation Methods, where we categorize existing approaches into traditional generation methods, diffusion model methods, and LLM-based methods, and compare them in terms of architecture, generation quality, and applicability; and (3) Applications and Challenges, which summarizes practical use cases, highlights common datasets, and discusses open challenges such as heterogeneity, data fidelity, and privacy protection. This survey aims to provide researchers and practitioners with a holistic understanding of the field and to highlight key directions for future work in synthetic tabular data generation.
Citations
- A Survey on Tabular Data Generation: Utility, Alignment, Fidelity, Privacy, and Beyond
- Beyond the convexity assumption: Realistic tabular data generation under quantifier-free real linear constraints
- AIGT: AI Generative Table Based on Prompt
- Generating Realistic Tabular Data with Large Language Models
- Synthetic Tabular Data Generation for Class Imbalance and Fairness: A Comparative Study
- HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
- P-TA: Using Proximal Policy Optimization to Enhance Tabular Data Augmentation via Large Language Models
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
- CTSyn: A Foundation Model for Cross Tabular Data Generation
- Differentially Private Tabular Data Synthesis using Large Language Models
- Discriminative Estimation of Total Variation Distance: A Fidelity Auditor for Generative Data
- EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
- Clinical Reasoning over Tabular Data and Text with Bayesian Networks
- Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
- How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular Data
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
- TabMT: Generating tabular data with masked transformers
- Reimagining Synthetic Tabular Data Generation through Data-Centric AI: A Comprehensive Benchmark
- AutoDiff: combining Auto-encoder and Diffusion model for tabular data synthesizing
- TabuLa: Harnessing Language Models for Tabular Data Synthesis
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space
- Generating and Imputing Tabular Data via Diffusion and Flow-based Gradient-Boosted Trees
- FinDiff: Diffusion Models for Financial Tabular Data Generation
- Deep Generative Models, Synthetic Tabular Data, and Differential Privacy: An Overview and Synthesis
- MissDiff: Training Diffusion Models on Tabular Data with Missing Values
- Automated Annotation with Generative AI Requires Validation
- Generative Table Pre-training Empowers Models for Tabular Prediction
- CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis
- A Survey of Large Language Models
- Synthetic Health-related Longitudinal Data with Mixed-type Variables Generated using Diffusion Models
- EHRDiff: Exploring Realistic EHR Synthesis with Diffusion Models
- Synthesizing Mixed-type Electronic Health Records using Diffusion Models
- MedDiff: Generating Electronic Health Records using Accelerated Denoising Diffusion Model
- REaLTabFormer: Generating Realistic Relational and Tabular Data using Transformers
- Row Conditional-TGAN for generating synthetic relational databases
- Diffusion models for missing value imputation in tabular data
- Comparing Synthetic Tabular Data Generation Between a Probabilistic Model and a Deep Learning Model for Education Use Cases
- Language Models are Realistic Tabular Data Generators
- FCT-GAN: Enhancing Table Synthesis via Fourier Transform
- STaSy: Score-based Tabular data Synthesis
- SOS: Score-based Oversampling for Tabular Data
- Emergent Abilities of Large Language Models
- CTAB-GAN+: Enhancing Tabular Data Synthesis
- Invertible Tabular GANs: Killing Two Birds with OneStone for Tabular Data Synthesis
- WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation
- TabFairGAN: Fair Tabular Data Generation with Generative Adversarial Networks
- Generative Adversarial Networks
- Want To Reduce Labeling Cost? GPT-3 Can Help
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- AI in Finance: Challenges, Techniques and Opportunities
- OCT-GAN: Neural ODE-based Conditional Tabular GANs
- SMOTE-ENC: A Novel SMOTE-Based Method to Generate Synthetic Data for Nominal and Continuous Features
- CTAB-GAN: Effective Table Data Synthesizing
- Differentially Private Synthetic Medical Data Generation using Convolutional GANs
- Score-Based Generative Modeling through Stochastic Differential Equations
- Conditional Wasserstein GAN-based Oversampling of Tabular Data for Imbalanced Learning
- Denoising Diffusion Probabilistic Models
- Improved Techniques for Training Score-Based Generative Models
- GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private Generators
- VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data
- Normalizing Flows: An Introduction and Review of Current Methods
- Neural Ordinary Differential Equations for Semantic Segmentation of Individual Colon Glands
- Modeling Tabular data using Conditional GAN
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Synthesizing Tabular Data using Generative Adversarial Networks
- Airline Passenger Name Record Generation using Generative Adversarial Networks
- Statistical Aspects of Wasserstein Distances
- Data Synthesis based on Generative Adversarial Networks
- Differentially Private Generative Adversarial Network
- cGANs with Projection Discriminator
- Scalable Private Learning with PATE
- Online Learning: A Comprehensive Survey
- Online learning: A comprehensive survey
- PacGAN: The power of two samples in generative adversarial networks
- Proximal Policy Optimization Algorithms
- A Closer Look at Memorization in Deep Networks
- Attention Is All You Need
- The Cramer Distance as a Solution to Biased Wasserstein Gradients
- Generating Multi-label Discrete Patient Records using Generative Adversarial Networks
- Renyi Differential Privacy
- XGBoost: A Scalable Tree Boosting System
- The truncated Euler–Maruyama method for stochastic differential equations
- Deep Unsupervised Learning using Nonequilibrium Thermodynamics
- Multiple imputation using chained equations: Issues and guidance for practice
- Long Short-Term Memory
- Backpropagation Applied to Handwritten Zip Code Recognition
- Theoretical risks and tabular asterisks: Sir Karl, Sir Ronald, and the slow progress of soft psychology.
- A Comprehensive Survey on Graph Neural Networks
- Effective data generation for imbalanced learning using conditional generative adversarial networks
- SMOTE: Synthetic Minority Over-sampling Technique
Cited by
Related