DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
2024/01/05 by DeepSeek-AI, Xiao Guo Bi, : +178 · 2 voices · 226 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2401.02954
Abstract
The rapid development of open-source large language models (LLMs) has been truly remarkable. However, the scaling law described in previous literature presents varying conclusions, which casts a dark cloud over scaling LLMs. We delve into the study of scaling laws and present our distinctive findings that facilitate scaling of large scale models in two commonly used open-source configurations, 7B and 67B. Guided by the scaling laws, we introduce DeepSeek LLM, a project dedicated to advancing open-source language models with a long-term perspective. To support the pre-training phase, we have developed a dataset that currently consists of 2 trillion tokens and is continuously expanding. We further conduct supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) on DeepSeek LLM Base models, resulting in the creation of DeepSeek Chat models. Our evaluation results demonstrate that DeepSeek LLM 67B surpasses LLaMA-2 70B on various benchmarks, particularly in the domains of code, mathematics, and reasoning. Furthermore, open-ended evaluations reveal that DeepSeek LLM 67B Chat exhibits superior performance compared to GPT-3.5.
Cited by
- Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- HATS: High-Accuracy Triple-Set Watermarking for Large Language Models
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Scaling Laws for Code: Every Programming Language Matters
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- Scaling Behavior of Discrete Diffusion Language Models
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
- SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
- On the Origin of Algorithmic Progress in AI
- A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
- The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
- PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference
- ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
- Relative Scaling Laws for LLMs
- Beyond Fixed Anchors: Precisely Erasing Concepts with Sibling Exclusive Counterparts
- VDSAgents: A PCS-Guided Multi-Agent System for Veridical Data Science Automation
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Large language model-based task planning for service robots: A review
- Social Simulations with Large Language Model Risk Utopian Illusion
- Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
- Enhancing Reasoning Skills in Small Persian Medical Language Models Can Outperform Large-Scale Data Training
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- Mapping Post-Training Forgetting in Language Models at Scale
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
- Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- Large Language Models can extract morphological data from taxonomic descriptions, but their stochastic nature makes automation challenging: a test on Australian Asteraceae
- Mid-Training of Large Language Models: A Survey
- Reusing Overtrained Language Models Saturates Scaling
- Membership Inference Attacks on Tokenizers of Large Language Models
- CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
- Disclosure and Evaluation as Fairness Interventions for General-Purpose AI
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- Optimal Scaling Needs Optimal Norm
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- Exploring Similarity between Neural and LLM Trajectories in Language Processing
- Eliciting and Improving the Causal Reasoning Abilities of Large Language Models with Conditional Statements
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- MMPB: It's Time for Multi-Modal Personalization
- Compute-Optimal Quantization-Aware Training
- RLP: Reinforcement as a Pretraining Objective
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- Mamba Modulation: On the Length Generalization of Mamba
- Towards joint scaling laws with optimal batch size schedules
- Training Continuously‐Coupled Reconfigurable Photonic Chips with Quantum Machine Learning
- Training Compute-Optimal Protein Language Models
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- MAPO: Mixed Advantage Policy Optimization
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- Building Large-Scale English-Romanian Literary Translation Resources with Open Models
- PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning
- Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
- TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- Any-Order Flexible Length Masked Diffusion
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
- Intern-S1: A Scientific Multimodal Foundation Model
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
- LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
- IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
- Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- Towards Scalable Training for Handwritten Mathematical Expression Recognition
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- T2S: Tokenized Skill Scaling for Lifelong Imitation Learning
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Can Language Models Discover Scaling Laws?
- StackTrans: From Large Language Model to Large Pushdown Automata Model
- A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat
- MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
- TTS-1 Technical Report
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
- Distilled Large Language Model in Confidential Computing Environment for System-on-Chip Design
- Language Models Improve When Pretraining Data Matches Target Tasks
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- LLMs Meet Cross-Modal Time Series Analytics: Overview and Directions
- AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
- KAT-V1: Kwai-AutoThink Technical Report
- Pre-Trained Policy Discriminators are General Reward Models
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
- ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
- Geological Everything Model 3D: A Promptable Foundation Model for Unified and Zero-Shot Subsurface Understanding
- Breaking Data Silos: Towards Open and Scalable Mobility Foundation Models via Generative Continual Learning
- Supporting Sustainable Computing by Repurposing E-Waste Smartphones as Tiny Data Centers
- Can LLM Improve for Expert Forecast Combination? Evidence from the European Central Bank Survey
- MolProphecy: Bridging Medicinal Chemists' Knowledge and Molecular Pre-Trained Models via a Multi-Modal Framework
- EAR: Erasing Concepts from Unified Autoregressive Models
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
- SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
- Generalizing vision-language models to novel domains: A comprehensive survey
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- From Bytes to Ideas: Language Modeling with Autoregressive U-Nets
- Sampling from Your Language Model One Byte at a Time
- AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
- Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents
- Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- Explaining Recovery Trajectories of Older Adults Post Lower-Limb Fracture Using Modality-wise Multiview Clustering and Large Language Models
- Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series
- Corrector Sampling in Language Models
- dots.llm1 Technical Report
- Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
- Elementary Math Word Problem Generation using Large Language Models
- Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
- EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
- A Statistical Physics of Language Model Reasoning
- Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
- Univariate to Multivariate: LLMs as Zero-Shot Predictors for Time-Series Forecasting
- A Graph Neural Network for the Era of Large Atomistic Models
- ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
- TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
- DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
- Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution
- Advancing Expert Specialization for Better MoE
- Learning in Compact Spaces with Approximately Normalized Transformer
- RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
- Explaining Large Language Models with gSMILE
- Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
- The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
- MLLMs are Deeply Affected by Modality Bias
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
- Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving
- DrugPilot: LLM-based Parameterized Reasoning Agent for Drug Discovery
- Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
- Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
- Model Merging in Pre-training of Large Language Models
- Gaokerena: A Small Persian Medical Language Model Family
- EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- Cochain: Balancing Insufficient and Excessive Collaboration in LLM Agent Workflows
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models
- Parallel Scaling Law for Language Models
- Improved Algorithms for Differentially Private Language Model Alignment
- What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
- Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation
- R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- Model Context Protocol-based Internet of Experts For Wireless Environment-aware LLM Agents
- Scalable Machines with Intrinsic Higher Mental-State Dynamics
- Don't be lazy: CompleteP enables compute-efficient deep transformers
- The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
- Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
- LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
- Predictable GRPO: A Closed-Form Model of Training Dynamics
- CoordField: Coordination Field for Agentic UAV Task Allocation In Low-altitude Urban Scenarios
- X-Fusion: Introducing New Modality to Frozen Large Language Models
- BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
- Efficient Pre-Training with Token Superposition
- On the Role of Batch Size in Stochastic Conditional Gradient Methods
- Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training
- Generative AI in Education: Student Skills and Lecturer Roles
- PhenoAssistant: A Conversational Multi-Agent AI System for Automated Plant Phenotyping
- Hallucinations and Key Information Extraction in Medical Texts: A Comprehensive Assessment of Open-Source Large Language Models
- Uncovering Political Bias in Large Language Models using Parliamentary Voting Records
- LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- Dion: Distributed Orthonormalized Updates
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- Trillion 7B Technical Report
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- OVERLORD: Ultimate Scaling of DataLoader for Multi-Source Large Foundation Model Training
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- An LLM Framework For Cryptography Over Chat Channels
- AI-University: An LLM-based platform for instructional alignment to scientific classrooms
- C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing
- Detect Anything 3D in the Wild
- Generative AI in Collaborative Academic Report Writing: Advantages, Disadvantages, and Ethical Considerations
- Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
- From PMI to Bots
Discussions
Related