LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
2025/06/22 by Chenghao Yang, Yang, Chenghao, Li, Sida +1 · 9 citations
Computer Science · Decision Sciences · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reservoir Engineering and Simulation Methods #Statistical and Computational Modeling
paper · pdf · doi:10.48550/arxiv.2506.17871
openalex publication_date 2025/06/22 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/31
Abstract
Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity. What drives this consistency in the generation? We investigate this phenomenon through the lens of probability concentration in the model's output distribution. To quantify this concentration, we introduce the *Branching Factor* (BF) -- a token-invariant measure of the effective number of plausible next steps during generation. Our empirical analysis reveals two key findings: (1) BF often decreases as generation progresses, suggesting that LLMs become more predictable as they generate. (2) alignment tuning substantially sharpens the model's output distribution from the outset, reducing BF by a factor of 2-5 overall, and up to an order of magnitude (e.g., from 12 to 1.2) at the beginning positions. This stark reduction helps explain why aligned models often appear less sensitive to decoding strategies. Building on this insight, we find this consistency has surprising implications for complex reasoning. Aligned Chain-of-Thought (CoT) models (e.g., DeepSeek-distilled models), for instance, leverage this effect; by generating longer reasoning chains, they push generation into later, more deterministic (lower BF) stages, resulting in more stable outputs. We hypothesize that alignment tuning does not fundamentally change a model's behavior, but instead steers it toward stylistic tokens (e.g., "Sure") that unlock low-entropy trajectories already present in the base model. This view is supported by nudging experiments, which show prompting base models with such tokens can similarly reduce BF. Together, our findings establish BF as a powerful diagnostic for understanding and controlling LLM outputs - clarifying how alignment reduces variability, how CoT promotes stable generations, and how base models can be steered away from diversity.
Citations
- Reasoning with Exploration: An Entropy Perspective
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Creative Preference Optimization
- Base Models Beat Aligned Models at Randomness and Creativity
- Modifying Large Language Model Post-Training for Diverse Creative Writing
- A Statistical Case Against Empirical Human-AI Alignment
- Diverse Preference Optimization
- 2 OLMo 2 Furious
- Benchmarking Linguistic Diversity of Large Language Models
- One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity
- AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
- The Llama 3 Herd of Models
- Are Large Language Models Capable of Generating Human-Level Narratives?
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
- Qwen2 Technical Report
- Generative Monoculture in Large Language Models
- Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
- From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
- From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
- CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
- Slaves to the Law of Large Numbers: An Asymptotic Equipartition Property for Perplexity in Generative Language Models
- LLM Discussion: Enhancing the Creativity of Large Language Models via Discussion Framework and Role-Play
- Language Model Cascades: Token-level uncertainty and beyond
- Do language models plan ahead for future tokens?
- A Thorough Examination of Decoding Methods in the Era of LLMs
- Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
- How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
- Benchmarking LLMs via Uncertainty Quantification
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
- In-Context Learning Dynamics with Random Binary Sequences
- Detecting Pretraining Data from Large Language Models
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
- Art or Artifice? Large Language Models and the False Promise of Creativity
- Does Writing with Language Models Reduce Content Diversity?
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Reasoning with Language Model is Planning with World Model
- Revisiting Entropy Rate Constancy in Text
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Controllable Text Generation with Language Constraints
- Truncation Sampling as Language Model Desmoothing
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- Language Models (Mostly) Know What They Know
- The Parallelism Tradeoff: Limitations of Log-Precision Transformers
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training language models to follow instructions with human feedback
- Rethinking and Refining the Distinct Metric
- Red Teaming Language Models with Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
- Measuring Massive Multitask Language Understanding
- Evaluating the Evaluation of Diversity in Natural Language Generation
- Calibration of Pre-trained Transformers
- The Curious Case of Neural Text Degeneration
- Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional\n Neural Networks for Extreme Summarization
- A Diversity-Promoting Objective Function for Neural Conversation Models
- Expectation-based syntactic comprehension
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
Cited by
Related