vix.ing · top · new · best · stats

Pretraining on the Test Set Is All You Need

2023/09/13 by Rylan Schaeffer, Schaeffer, Rylan · 16 voices · 10 citations
Computer Science · Engineering · Mathematics · #Artificial intelligence #Computer science #Electrical engineering #Engineering #Mathematics #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Scaling #Supercharge #Test set #Topic Modeling #Transformer #cs.AI #cs.CL

paper · pdf · doi:10.48550/arxiv.2309.08632

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/09/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Inspired by recent work demonstrating the promise of smaller Transformer-based language models pretrained on carefully curated data, we supercharge such approaches by investing heavily in curating a novel, high quality, non-synthetic data mixture based solely on evaluation benchmarks. Using our novel dataset mixture consisting of less than 100 thousand tokens, we pretrain a 1 million parameter transformer-based LLM phi-CTNL (pronounced ``fictional") that achieves perfect results across diverse academic benchmarks, strictly outperforming all known foundation models. phi-CTNL also beats power-law scaling and exhibits a never-before-seen grokking-like ability to accurately predict downstream evaluation benchmarks' canaries.

Cited by

Discussions

Related