MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
2024/04/09 by Shengding Hu, Yuge Tu, Hu, Shengding +48 · 2 voices · 234 citations
Computer Science · #Artificial intelligence #Computer science #Database #Geography #Natural Language Processing Techniques #Natural language processing #Scalability #Topic Modeling #Training (meteorology) #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2404.06395
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/04/09 · arxiv published 2024/04/09 · arxiv updated 2024/06/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The burgeoning interest in developing Large Language Models (LLMs) with up to trillion parameters has been met with concerns regarding resource efficiency and practical expense, particularly given the immense cost of experimentation. This scenario underscores the importance of exploring the potential of Small Language Models (SLMs) as a resource-efficient alternative. In this context, we introduce MiniCPM, specifically the 1.2B and 2.4B non-embedding parameter variants, not only excel in their respective categories but also demonstrate capabilities on par with 7B-13B LLMs. While focusing on SLMs, our approach exhibits scalability in both model and data dimensions for future LLM research. Regarding model scaling, we employ extensive model wind tunnel experiments for stable and optimal scaling. For data scaling, we introduce a Warmup-Stable-Decay (WSD) learning rate scheduler (LRS), conducive to continuous training and domain adaptation. We present an in-depth analysis of the intriguing training dynamics that occurred in the WSD LRS. With WSD LRS, we are now able to efficiently study data-model scaling law without extensive retraining experiments on both axes of model and data, from which we derive the much higher compute optimal data-model ratio than Chinchilla Optimal. Additionally, we introduce MiniCPM family, including MiniCPM-DPO, MiniCPM-MoE and MiniCPM-128K, whose excellent performance further cementing MiniCPM's foundation in diverse SLM applications. MiniCPM models are available publicly at https://github.com/OpenBMB/MiniCPM .
Cited by
- Understanding the Mechanisms of Fast Hyperparameter Transfer
- Latent Multi-Head Attention for Small Language Models
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
- Attention Residuals
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
- MiniLingua: A Small Open-Source LLM for European Languages
- Scaling Behavior of Discrete Diffusion Language Models
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- PCMind-2.1-Kaiyuan-2B Technical Report
- Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- RecruitView: A Multimodal Dataset for Predicting Personality and Interview Performance for Human Resources Applications
- Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- Next-Latent Prediction Transformers Learn Compact World Models
- S2LM: Towards Semantic Steganography via Large Language Models
- Motif 2 12.7B technical report
- GeoCrossBench: Cross-Band Generalization for Remote Sensing
- ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests
- Arcee Trinity Large Technical Report
- Relative Scaling Laws for LLMs
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- A Survey on LLM Mid-Training
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
- Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
- Optimization Benchmark for Diffusion Models on Dynamical Systems
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Weight Decay may matter more than muP for Learning Rate Transfer in Practice
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- Midtraining Bridges Pretraining and Posttraining Distributions
- xLLM Technical Report
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- NOSA: Native and Offloadable Sparse Attention
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- Sparsely gated tiny linear experts
- Scaling Laws and Symmetry, Evidence from Neural Force Fields
- TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription
- BOTANIC-0: a series of foundation models for plant genomic data
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- Mid-Training of Large Language Models: A Survey
- Reusing Overtrained Language Models Saturates Scaling
- Grouped Differential Attention
- Training Dynamics Impact Post-Training Quantization Robustness
- How does the optimizer implicitly bias the model merging loss landscape?
- Optimal Scaling Needs Optimal Norm
- Less LLM, More Documents: Searching for Improved RAG
- ModernVBERT: Towards Smaller Visual Document Retrievers
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
- Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Pretraining Large Language Models with NVFP4
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Training Optimal Large Diffusion Language Models
- Effective Quantization of Muon Optimizer States
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- Compute-Optimal Quantization-Aware Training
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- Predicting LLM Reasoning Performance with Small Proxy Model
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- Towards joint scaling laws with optimal batch size schedules
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
- Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model
- Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- Elucidating the Design Space of Decay in Linear Attention
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- From Drone Imagery to Livability Mapping: AI-powered Environment Perception in Rural China
- SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
- Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- Granite Embedding R2 Models
- Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
- Maximum Score Routing For Mixture-of-Experts
- Next Visual Granularity Generation
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- TiMoE: Time-Aware Mixture of Language Experts
- Learning User Preferences for Image Generation Model
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- Win-k: Improved Membership Inference Attacks on Small Language Models
- Motif 2.6B Technical Report
- Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
- Kimi K2: Open Agentic Intelligence
- JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
- MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
- Mellum2 Technical Report
- Retrieval-Aware Distillation for Transformer-SSM Hybrids
- ERNIE 5.0 Technical Report
- Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
- WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
- AdaMuon: Adaptive Muon Optimizer
- Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
- Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- Vision-Language Models Can't See the Obvious
- Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
- MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes
- Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment
- HyperCLOVA X THINK Technical Report
- SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
- Towards a Small Language Model Lifecycle Framework
- Hallucination Detection with Small Language Models
- Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- ByteSpan: Information-Driven Subword Tokenisation
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning
- Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection
- AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
- OneRec Technical Report
- Curriculum-Guided Layer Scaling for Language Model Pretraining
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- dots.llm1 Technical Report
- Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
- HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
- Causal Estimation of Tokenisation Bias
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data
- A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
- Stepsize anything: A unified learning rate schedule for budgeted-iteration training
- GradPower: Powering Gradients for Faster Language Model Pre-Training
- FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
- MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
- FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
- Efficient Multi-modal Long Context Learning for Training-free Adaptation
- Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator
- DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- Learning What to Remember: Test-Time Training via Context Distillation
- BehaveGPT: A Foundation Model for Large-scale User Behavior Modeling
- Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
- RemoteSAM: Towards Segment Anything for Earth Observation
- PaTH Attention: Position Encoding via Accumulating Householder Transformations
- Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation
- Recursive Offloading for LLM Serving in Multi-tier Networks
- Enhancing Large Language Models for Detecting Mental Manipulation via Annotation-Free Data Augmentation and Anti-Curriculum Distillation
- Scaling Diffusion Transformers Efficiently via μP
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
- A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
- Emerging Properties in Unified Multimodal Pretraining
- This Time is Different: An Observability Perspective on Time Series Foundation Models
- Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
- HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
- FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
- Advancing Sequential Numerical Prediction in Autoregressive Models
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
- Model Merging in Pre-training of Large Language Models
- Training nGPT
- Parallel Scaling Law for Language Models
- MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
- Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
- Large Language Models for Computer-Aided Design: A Survey
- Learning Dynamics in Continual Pre-Training for Large Language Models
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- xGen-small Technical Report
- SmartPilot: A Multiagent CoPilot for Adaptive and Intelligent Manufacturing
- (How) Learning Rates Regulate Catastrophic Overtraining
- EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
- Position: Enough of Scaling LLMs! Lets Focus on Downscaling
- dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
- Efficient Pre-Training with Token Superposition
- VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
- Learning Long-term Motion Embeddings for Efficient Kinematics Generation
- Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
- Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima
- TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
- Masked diffusion enables coherent beat tracking
- Taxonomy-Aware Evaluation of Vision-Language Models
- Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
- Trillion 7B Technical Report
- Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- SEA-LION: Southeast Asian Languages in One Network
- Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
- SmolVLM: Redefining small and efficient multimodal models
Discussions
Related