Rare Event Analysis of Large Language Models
2026/02/06 by Jake McAllister Dorman, Edward Gillman, Dominic C. Rose +2 · 1 voice · 3 citations
Computer Science · Physics and Astronomy · #cond-mat.dis-nn #cond-mat.stat-mech #cs.LG
paper · pdf · doi:10.48550/arxiv.2602.06791
arxiv published 2026/02/06 · arxiv updated 2026/05/28
Abstract
Being probabilistic models, during inference large language models (LLMs) display rare events: behaviour that is far from typical but highly significant. By definition all rare events are hard to see, but the enormous scale of LLM usage means that events completely unobserved during development are likely to become prominent in deployment. Here we present an end-to-end framework for the systematic analysis of rare events in LLMs. We provide a practical implementation spanning theory, efficient generation strategies, probability estimation and error analysis, which we illustrate with concrete examples. We outline extensions and applications to other models and contexts, highlighting the generality of the concepts and techniques presented here.
Citations
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
- Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
- Forecasting Rare Language Model Behaviors
- Towards neural reinforcement learning for large deviations in nonequilibrium systems with memory
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Estimating the Probabilities of Rare Outputs in Language Models
- Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
- A Thorough Examination of Decoding Methods in the Era of LLMs
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Combining Reinforcement Learning and Tensor Networks, with an Application to Dynamical Large Deviations
- Training neural network ensembles via trajectory sampling
- Efficient Training of Language Models to Fill in the Middle
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Generalized Out-of-Distribution Detection: A Survey
- Generalized Out-of-Distribution Detection: A Survey
- Bias, variance, and confidence intervals for efficiency estimators in particle physics experiments
- Binomial confidence intervals for rare events: importance of defining margin of error relative to magnitude of proportion
- A reinforcement learning approach to rare trajectory sampling
- Convergence diagnostics for Markov chain Monte Carlo
- The Curious Case of Neural Text Degeneration
- Simple Confidence Intervals for MCMC Without CLTs
- Reweighting from the mixture distribution as a better way to describe the Multistate Bennett Acceptance Ratio
- Eigenvector method for umbrella sampling enables error analysis
- Classical stochastic dynamics and continuous matrix product states: gauge transformations, conditioned and driven processes, and equivalence of trajectory ensembles
- Variational and optimal control representations of conditioned and driven processes
- Effective interactions and large deviations in stochastic processes
- Simulating rare events in dynamical processes
- Batch means and spectral variance estimators in Markov chain Monte Carlo
- Statistically optimal analysis of samples from multiple equilibrium states
- Parallel tempering: Theory, applications, and new perspectives
- Nonphysical sampling distributions in Monte Carlo free-energy estimation: Umbrella sampling
- Monte Carlo sampling methods using Markov chains and their applications
- Statistical Mechanics of Fluid Mixtures
- Probable Inference, the Law of Succession, and Statistical Inference
- TRANSITIONPATHSAMPLING: Throwing Ropes Over Rough Mountain Passes, in the Dark
Cited by
Discussions
Related