TinyLlama: An Open-Source Small Language Model
2024/01/04 by Peiyuan Zhang, Zhang, Peiyuan, Guangtao Zeng +5 · 1 voice · 186 citations
Computer Science · #Algorithms and Data Compression #Archaeology #Architecture #Artificial intelligence #Code (set theory) #Computer science #Downstream (manufacturing) #Language model #Natural Language Processing Techniques #Open source #Programming language #Set (abstract data type) #Software #Source code #Topic Modeling #World Wide Web #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2401.02385
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/01/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Abstract
We present TinyLlama, a compact 1.1B language model pretrained on around 1 trillion tokens for approximately 3 epochs. Building on the architecture and tokenizer of Llama 2, TinyLlama leverages various advances contributed by the open-source community (e.g., FlashAttention and Lit-GPT), achieving better computational efficiency. Despite its relatively small size, TinyLlama demonstrates remarkable performance in a series of downstream tasks. It significantly outperforms existing open-source language models with comparable sizes. Our model checkpoints and code are publicly available on GitHub at https://github.com/jzhang38/TinyLlama.
Cited by
- On the Convergence of Stochastic Low-Rank Adaptation
- RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
- Transition-Aware Backend Dispatch for Edge LLM Inference
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- Kolmogorov--Arnold Networks for Small Language Models
- Explaining Attention with Program Synthesis
- Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
- Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
- Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
- MEMOIR: Temporal Behavioral Memory for Recommendation Across the Preference-Drift Spectrum
- Forgetting Is Not a Fix: Path Dependence in Sequential Engram Editing
- Parallel Token Prediction for Language Models
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Empowering smart app development with SolidGPT: an edge-cloud hybrid AI agent framework
- OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
- Menta: A Small Language Model for On-Device Mental Health Prediction
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- DisCEdge: Distributed Context Management for Large Language Models at the Edge
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- Manifold Percolation: from generative model to Reinforce learning
- NOEM3A: a Neuro-symbolic Ontology-Enhanced Method for Multi-intent understanding in Mobile Agents
- Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
- SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
- A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge Systems
- CO2-Meter: A Comprehensive Carbon Footprint Estimator for LLMs on Edge Devices
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
- Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
- SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs
- SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment
- See the Text: From Tokenization to Visual Reading
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
- HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
- HyMiRec: A Hybrid Multi-interest Learning Framework for LLM-based Sequential Recommendation
- F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
- Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
- End-to-End Multi-Modal Diffusion Mamba
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- Cautious Weight Decay
- Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda
- Efficient Resource-Constrained Training of Transformers via Subspace Optimization
- Reusing Overtrained Language Models Saturates Scaling
- Mixture of Neuron Experts
- Towards an Efficient, Customizable, and Accessible AI Tutor
- Backdoor Attacks Against Speech Language Models
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- PEPS: Quantum-Inspired Reinforcement Learning for Coherent Reasoning Traces in LLMs
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- Explaining Data Mixing Scaling Laws
- Speculative Safety-Aware Decoding
- Code Driven Planning with Domain-Adaptive Critic
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
- Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
- Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
- BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- Mitigating Catastrophic Forgetting in Continual Learning through Model Growth
- SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
- TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- Cohort-Aware Agents for Individualized Lung Cancer Risk Prediction Using a Retrieval-Augmented Model Selection Framework
- Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
- Semantically Guided Adversarial Testing of Vision Models Using Language Models
- Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- Modelling and Classifying the Components of a Literature Review
- Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- Aligning Knowledge Graphs and Language Models for Factual Accuracy
- Characterizing State Space Model (SSM) and SSM-Transformer Hybrid Language Model Performance with Long Context Length
- DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
- FLoRIST: Singular Value Thresholding for Efficient and Accurate Federated Fine-Tuning of Large Language Models
- LogTinyLLM: Tiny Large Language Models Based Contextual Log Anomaly Detection
- Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
- Finetuning Vision-Language Models as OCR Systems for Low-Resource Languages: A Case Study of Manchu
- A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
- When Transformers Meet Recommenders: Integrating Self-Attentive Sequential Recommendation with Fine-Tuned LLMs
- Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
- Pre-Trained Policy Discriminators are General Reward Models
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
- Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
- Modeling the One-to-Many Property in Open-Domain Dialogue with LLMs
- S4C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
- Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
- Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
- GTA: Grouped-head latenT Attention
- Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
- Basis Transformers for Multi-Task Tabular Regression
- Distillation Robustifies Unlearning
- TrueGL: A Truthful, Reliable, and Unified Engine for Grounded Learning in Full-Stack Search
- Beyond Low-rank Decomposition: A Shortcut Approach for Efficient On-Device Learning
- LLMs Can Compensate for Deficiencies in Visual Representations
- Pipelining Split Learning in Multi-hop Edge Networks
- TokAlign: Efficient Vocabulary Adaptation via Token Alignment
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
- STORM-BORN: A Challenging Mathematical Derivations Dataset Curated via a Human-in-the-Loop Multi-Agent Framework
- One for All: Update Parameterized Knowledge Across Multiple Models
- PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- LISRec: Modeling User Preferences with Learned Item Shortcuts for Sequential Recommendation
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
- Pretraining Language Models to Ponder in Continuous Space
- Explaining Large Language Models with gSMILE
- Incentivizing Inclusive Contributions in Model Sharing Markets
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models
- Invoke Interfaces Only When Needed: Adaptive Invocation for Large Language Models in Question Answering
- Online Knowledge Distillation with Reward Guidance
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning
- A Survey on the Application of Large Language Models in Scenario-Based Testing of Automated Driving Systems
- Filtering Learning Histories Enhances In-Context Reinforcement Learning
- Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
- SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
- Structured Agent Distillation for Large Language Model
- VulCPE: Context-Aware Cybersecurity Vulnerability Retrieval and Management
- Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
- Chain-of-Model Learning for Language Model
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs
- The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
- Token Masking Improves Transformer-Based Text Classification
- EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
- Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study
- LM-Scout: Analyzing the Security of Language Model Integration in Android Apps
- Model-Distributed Inference for Large Language Models at the Edge
- A Survey on Foundation Models for Personalized Federated Intelligence
- DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
- Camera Control at the Edge with Language Models for Scene Understanding
- Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
- Position: Enough of Scaling LLMs! Lets Focus on Downscaling
- How Vulnerable Are Edge LLMs?
- Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing
- When Reasoning Beats Scale: A 1.5B Reasoning Model Outranks 13B LLMs as Discriminator
- A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning
- Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
- Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
- On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
- Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs
- RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
- Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
- Kuwain 1.5B: An Arabic SLM via Language Injection
- Synergistic Weak-Strong Collaboration by Aligning Preferences
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
Discussions
Related