MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
2024/04/09 by Shengding Hu, Yuge Tu, Hu, Shengding +48 · 2 voices · 128 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2404.06395
arxiv published 2024/04/09 · arxiv updated 2024/06/03
Abstract
The burgeoning interest in developing Large Language Models (LLMs) with up to trillion parameters has been met with concerns regarding resource efficiency and practical expense, particularly given the immense cost of experimentation. This scenario underscores the importance of exploring the potential of Small Language Models (SLMs) as a resource-efficient alternative. In this context, we introduce MiniCPM, specifically the 1.2B and 2.4B non-embedding parameter variants, not only excel in their respective categories but also demonstrate capabilities on par with 7B-13B LLMs. While focusing on SLMs, our approach exhibits scalability in both model and data dimensions for future LLM research. Regarding model scaling, we employ extensive model wind tunnel experiments for stable and optimal scaling. For data scaling, we introduce a Warmup-Stable-Decay (WSD) learning rate scheduler (LRS), conducive to continuous training and domain adaptation. We present an in-depth analysis of the intriguing training dynamics that occurred in the WSD LRS. With WSD LRS, we are now able to efficiently study data-model scaling law without extensive retraining experiments on both axes of model and data, from which we derive the much higher compute optimal data-model ratio than Chinchilla Optimal. Additionally, we introduce MiniCPM family, including MiniCPM-DPO, MiniCPM-MoE and MiniCPM-128K, whose excellent performance further cementing MiniCPM's foundation in diverse SLM applications. MiniCPM models are available publicly at https://github.com/OpenBMB/MiniCPM .
Cited by
- Understanding the Mechanisms of Fast Hyperparameter Transfer
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
- Attention Residuals
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
- MiniLingua: A Small Open-Source LLM for European Languages
- Scaling Behavior of Discrete Diffusion Language Models
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- PCMind-2.1-Kaiyuan-2B Technical Report
- Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- RecruitView: A Multimodal Dataset for Predicting Personality and Interview Performance for Human Resources Applications
- Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- Next-Latent Prediction Transformers Learn Compact World Models
- S2LM: Towards Semantic Steganography via Large Language Models
- Motif 2 12.7B technical report
- GeoCrossBench: Cross-Band Generalization for Remote Sensing
- ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests
- Arcee Trinity Large Technical Report
- Relative Scaling Laws for LLMs
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- A Survey on LLM Mid-Training
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
- Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
- Optimization Benchmark for Diffusion Models on Dynamical Systems
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Weight Decay may matter more than muP for Learning Rate Transfer in Practice
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- Midtraining Bridges Pretraining and Posttraining Distributions
- xLLM Technical Report
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- NOSA: Native and Offloadable Sparse Attention
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- Sparsely gated tiny linear experts
- Scaling Laws and Symmetry, Evidence from Neural Force Fields
- TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription
- BOTANIC-0: a series of foundation models for plant genomic data
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- Mid-Training of Large Language Models: A Survey
- Reusing Overtrained Language Models Saturates Scaling
- Grouped Differential Attention
- Training Dynamics Impact Post-Training Quantization Robustness
- How does the optimizer implicitly bias the model merging loss landscape?
- Optimal Scaling Needs Optimal Norm
- Less LLM, More Documents: Searching for Improved RAG
- ModernVBERT: Towards Smaller Visual Document Retrievers
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
- Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Pretraining Large Language Models with NVFP4
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Training Optimal Large Diffusion Language Models
- Effective Quantization of Muon Optimizer States
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- Compute-Optimal Quantization-Aware Training
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- Predicting LLM Reasoning Performance with Small Proxy Model
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- Towards joint scaling laws with optimal batch size schedules
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
- Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model
- Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- Elucidating the Design Space of Decay in Linear Attention
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- From Drone Imagery to Livability Mapping: AI-powered Environment Perception in Rural China
- SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
- Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- Granite Embedding R2 Models
- Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
- Maximum Score Routing For Mixture-of-Experts
- Next Visual Granularity Generation
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- TiMoE: Time-Aware Mixture of Language Experts
- Learning User Preferences for Image Generation Model
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- Win-k: Improved Membership Inference Attacks on Small Language Models
- Motif 2.6B Technical Report
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
- Kimi K2: Open Agentic Intelligence
- JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
- Mellum2 Technical Report
- Retrieval-Aware Distillation for Transformer-SSM Hybrids
- ERNIE 5.0 Technical Report
Discussions
Related