2 OLMo 2 Furious
2024/12/31 by Team OLMo, Pete Walsh, P N Walsh +85 · 9 voices · 87 citations
Computer Science · Medicine · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2501.00656
Abstract
We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model weights, full training data, training code and recipes, training logs and thousands of intermediate checkpoints. In this work, we describe our modified model architecture and training recipe, focusing on techniques for achieving better training stability and improved per-token efficiency. Our updated pretraining data mixture introduces a new, specialized data mix called Dolmino Mix 1124, which significantly improves model capabilities across many downstream task benchmarks when introduced via late-stage curriculum training (i.e. specialized data during the annealing phase of pretraining). Finally, we incorporate best practices from Tülu 3 to develop OLMo 2-Instruct, focusing on permissive data and extending our final-stage reinforcement learning with verifiable rewards (RLVR). Our OLMo 2 base models sit at the Pareto frontier of performance to training compute, often matching or outperforming open-weight only models like Llama 3.1, Qwen 2.5, and Gemma 2 while using fewer FLOPs and with fully transparent training data, code, and recipe. Our fully open OLMo 2-Instruct models are competitive with open-weight only models of comparable size and even some proprietary models like GPT-3.5 Turbo and GPT 4o Mini.
Cited by
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
- Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
- Transition-Aware Backend Dispatch for Edge LLM Inference
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- Understanding Reasoning from Pretraining to Post-Training
- Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- From Memorization to Reasoning in the Spectrum of Loss Curvature
- Language Models Fail to Introspect About Their Knowledge of Language
- AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
- LLMs Get Lost In Multi-Turn Conversation
- Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
- MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
- Do Chinese models speak Chinese languages?
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
- Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- Spherical Leech Quantization for Visual Tokenization and Generation
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
- RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
- INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
- Mixture-of-Depths Attention
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
- Data-Centric Lessons To Improve Speech-Language Pretraining
- CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
- Extracting alignment data in open models
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Midtraining Bridges Pretraining and Posttraining Distributions
- Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
- Cautious Weight Decay
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- Stability of Transformers under Layer Normalization
- The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
- KORMo: Korean Open Reasoning Model for Everyone
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- Mid-Training of Large Language Models: A Survey
- Training Dynamics Impact Post-Training Quantization Robustness
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Generative Value Conflicts Reveal LLM Priorities
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- Reinforcement Mid-Training
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Beyond Outliers: A Study of Optimizers Under Quantization
- Mapping Overlaps in Benchmarks through Perplexity in the Wild
- Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- Tracing the Representation Geometry of Language Models from Pretraining to Post-training
- What Is The Political Content in LLMs' Pre- and Post-Training Data?
- Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs
- Compute-Optimal Quantization-Aware Training
- On Code-Induced Reasoning in LLMs
- One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
- Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
- Predicting LLM Reasoning Performance with Small Proxy Model
- Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
- Fluid Language Model Benchmarking
- Continually Adding New Languages to Multilingual Language Models
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- LongCat-Flash Technical Report
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- MolmoAct: Action Reasoning Models that can Reason in Space
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- Checklists Are Better Than Reward Models For Aligning Language Models
Discussions
- 32B olmo-2 03/25 [lemmy, 20 points, 1 comments]
- Paper: arxiv.org/abs/2501.00656
Blog: allenai.org/blog/olmo2
Demo: playground.allenai.org
Collection: huggingface.co/collections/... [bsky, 7 points, 0 comments]
- All together, these four aspects help us achieve the best fully open model yet.
🚗read the full tech report: arxiv.org/abs/2501.00656
💬 try OLMo 2 Instruct 13B on the @allen_ai playground: playgrou [bsky, 5 points, 0 comments]
- 2 OLMo 2 Furious [hn, 4 points, 1 comments]
- 🚀 The team is incredible! Our paper has more detail on pre-training, mid-training, infrastructure, and evaluations. Check it out!
arxiv.org/abs/2501.00656 [bsky, 2 points, 0 comments]
- A few days ago, we did finally release the OLMo 2 tech report: arxiv.org/pdf/2501.00656. There is a lot of good stuff in there, but the stability work we did over the summer makes me particularly prou [bsky, 1 points, 0 comments]
- From arxiv.org/pdf/2501.00656 [bsky, 1 points, 0 comments]
- Nachdem ich in meiner lokalen KI bisher "nur" Open-Weight-Modelle wie Llama, Phi und Qwen verwendet habe, probiere ich jetzt mit OLMo-2 mal ein echtes Open-Source-Modell (Apache v2) aus: arxiv.org/pdf [bsky, 0 points, 0 comments]
- and the best paper name for a #LLM in December, goes to the #OLMo team
arxiv.org/abs/2501.00656 [bsky, 0 points, 0 comments]
Related