Textbooks Are All You Need II: phi-1.5 technical report
2023/09/11 by Yuanzhi Li, Li, Yuanzhi, Sébastien Bubeck +9 · 3 voices · 113 citations
Computer Science · Engineering · #Artificial intelligence #Coding (social sciences) #Computer science #Context (archaeology) #Electrical engineering #Engineering #Geography #Inference #Natural Language Processing Techniques #Programming language #Python (programming language) #Social science #Sociology #Text Readability and Simplification #Topic Modeling #Transformer #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2309.05463
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/09/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We continue the investigation into the power of smaller Transformer-based language models as initiated by TinyStories -- a 10 million parameter model that can produce coherent English -- and the follow-up work on phi-1, a 1.3 billion parameter model with Python coding performance close to the state-of-the-art. The latter work proposed to use existing Large Language Models (LLMs) to generate ``textbook quality" data as a way to enhance the learning process compared to traditional web data. We follow the ``Textbooks Are All You Need" approach, focusing this time on common sense reasoning in natural language, and create a new 1.3 billion parameter model named phi-1.5, with performance on natural language tasks comparable to models 5x larger, and surpassing most non-frontier LLMs on more complex reasoning tasks such as grade-school mathematics and basic coding. More generally, phi-1.5 exhibits many of the traits of much larger LLMs, both good -- such as the ability to ``think step by step" or perform some rudimentary in-context learning -- and bad, including hallucinations and the potential for toxic and biased generations -- encouragingly though, we are seeing improvement on that front thanks to the absence of web data. We open-source phi-1.5 to promote further research on these urgent topics.
Cited by
- Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- Compressing LLMs with MoP: Mixture of Pruners
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- Grokked Models are Better Unlearners
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
- On the Notion that Language Models Reason
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
- ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
- Autodata: An agentic data scientist to create high quality synthetic data
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda
- Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Skill-Targeted Adaptive Training
- On the Representations of Entities in Auto-regressive Large Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- Reading the unreadable: creating a dataset of 19th century English newspapers using image-to-text language models
- Explicit vs. Implicit Biographies: Evaluating and Adapting LLM Information Extraction on Wikidata-Derived Texts
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- HAVE: Head-Adaptive Gating and ValuE Calibration for Hallucination Mitigation in Large Language Models
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
- A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Learning Facts at Scale with Active Reading
- TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
- S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
- Dynaword: From One-shot to Continuously Developed Datasets
- LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
- FedVLM: Scalable Personalized Vision-Language Models through Federated Learning
- MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
- On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
- LogTinyLLM: Tiny Large Language Models Based Contextual Log Anomaly Detection
- DICE: Data Influence Cascade in Decentralized Learning
- Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
- Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
- La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
- TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation
- HyperCLOVA X THINK Technical Report
- SGIC: A Self-Guided Iterative Calibration Framework for RAG
- Sampling from Your Language Model One Byte at a Time
- Assessing the Role of Data Quality in Training Bilingual Language Models
- GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
- No Data? No Problem: Synthesizing Security Graphs for Better Intrusion Detection
- Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
- Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation
- Beyond Text Compression: Evaluating Tokenizers Across Scales
- Circuit Stability Characterizes Language Model Generalization
- Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- LoLA: Low-Rank Linear Attention With Sparse Caching
- Pre-Training Curriculum for Multi-Token Prediction in Language Models
- Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
- What happens when generative AI models train recursively on each others' outputs?
- Explaining Large Language Models with gSMILE
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language
- An Initial Exploration of Fine-tuning Small Language Models for Smart Contract Reentrancy Vulnerability Detection
- R-Genie: Reasoning-Guided Generative Image Editing
- Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- Revealing Language Model Trajectories via Kullback-Leibler Divergence
- FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
- Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- Enhancing LLMs via High-Knowledge Data Selection
- Capturing the Effects of Quantization on Trojans in Code LLMs
- Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review
- GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
- Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
- EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
- Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
- Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
- A Data Synthesis Method Driven by Large Language Models for Proactive Mining of Implicit User Intentions in Tourism
- Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
- WaterDrum: Watermarking for Data-centric Unlearning Metric
- Exploring the Role of Diversity in Example Selection for In-Context Learning
- Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
- Stan: An LLM-based thermodynamics course assistant
- Efficient LLMs with AMP: Attention Heads and MLP Pruning
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
- Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
- Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
- DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
- AdaCoder: An Adaptive Planning and Multi-Agent Framework for Function-Level Code Generation
- When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments
- ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
- Knowledge-Instruct: Effective Continual Pre-training from Limited Data using Instructions
- SmolVLM: Redefining small and efficient multimodal models
- Neural scaling law [wikipedia]
Discussions
- The tide is shifting: 1.3B outperforms 7B Llama 2 [hn, 63 points, 15 comments]
- The new version of Textbooks Are All You Need trains a language model *mostly* (like 3/4ths) on synthetic data, and finds that, at only 1.3B parameters, phi-1.5 can exceed the capacities of much large [bsky, 15 points, 1 comments]
- arxiv.org/abs/2309.05463 Summary The main purpose of the article is to try to understand whether it is possible to achieve the same LLM performance through smaller Transformer-based language models, [bsky, 2 points, 1 comments]
Related