Pretraining on the Test Set Is All You Need
2023/09/13 by Rylan Schaeffer, Schaeffer, Rylan · 16 voices · 7 citations
#cs.CL #cs.AI
paper · pdf · doi:10.48550/arxiv.2309.08632
Abstract
Inspired by recent work demonstrating the promise of smaller Transformer-based language models pretrained on carefully curated data, we supercharge such approaches by investing heavily in curating a novel, high quality, non-synthetic data mixture based solely on evaluation benchmarks. Using our novel dataset mixture consisting of less than 100 thousand tokens, we pretrain a 1 million parameter transformer-based LLM phi-CTNL (pronounced ``fictional") that achieves perfect results across diverse academic benchmarks, strictly outperforming all known foundation models. phi-CTNL also beats power-law scaling and exhibits a never-before-seen grokking-like ability to accurately predict downstream evaluation benchmarks' canaries.
Cited by
Discussions
- Pretraining on the Test Set Is All You Need [hn, 68 points, 26 comments]
- I can't believe this paper only has 35 citations and people keep trying to rewrite it arxiv.org/abs/2309.08632 [bsky, 39 points, 4 comments]
- ну во первых это красиво [bsky, 5 points, 2 comments]
- Pretraining on the Test Set Is All You Need [hn, 4 points, 0 comments]
- Big if true [bsky, 4 points, 0 comments]
- Pretraining on test set is all you need [hn, 3 points, 1 comments]
- Pretraining on the Test Set Is All You Need arxiv.org/abs/2309.08632 くそわろてる [bsky, 2 points, 1 comments]
- Pretraining on the Test Set Is All You Need [hn, 2 points, 0 comments]
- You only need 64 A100s to pretrain on the test set. arxiv.org/abs/2309.08632 [bsky, 2 points, 0 comments]
- Pretraining on the Test Set Is All You Need [hn, 1 points, 0 comments]
- 📝 Pretraining on the Test Set Is All You Need 📚👾 "\textbf{phi-CTNL} is trained as a causal language model on a mixture of carefully curated academic benchmarks and is then evaluated as a masked lan [mastodon, 0 points, 0 comments]
- GenAI in a nutshell. arxiv.org/abs/2309.08632 [bsky, 0 points, 0 comments]
- I mean, have you even tried pre-training on the questions + answer key? (The paper is satire but highlights the problem of test set data contamination in most/all LLM benchmarks) arxiv.org/abs/2309.08 [bsky, 0 points, 0 comments]
- Classic arxiv.org/abs/2309.08632 [bsky, 0 points, 0 comments]
- I have been doing ML wrong my whole life: arxiv.org/abs/2309.08632 [bsky, 0 points, 0 comments]
- after all... arxiv.org/abs/2309.08632 [bsky, 0 points, 0 comments]
Related