2022/10/17 by Ewa Andrejczuk, Andrejczuk, Ewa, Julian Martin Eisenschlos +7 · 4 citations
Computer Science · Decision Sciences · Mathematics · #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mathematics, Computing, and Information Processing #Statistics Education and Methodologies
paper · pdf · doi:10.48550/arxiv.2210.09162
openalex publication_date 2022/10/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Encoder-only transformer models have been successfully applied to different table understanding tasks, as in TAPAS (Herzig et al., 2020). A major limitation of these architectures is that they are constrained to classification-like tasks such as cell selection or entailment detection. We present TABT5, an encoder-decoder model that generates natural language text based on tables and textual inputs. TABT5 overcomes the encoder-only limitation by incorporating a decoder component and leverages the input structure with table specific embeddings and pre-training. TABT5 achieves new state-of-the-art results on several domains, including spreadsheet formula prediction with a 15% increase in sequence accuracy, QA with a 2.5% increase in sequence accuracy and data-to-text generation with a 2.5% increase in BLEU.