VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY
1950/01/01 by GLENN W. BRIER, Glenn W. Brier · 216 citations
Environmental Science · #Environmental and Industrial Safety
paper · doi:10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2
Cited by
- DNA methylation-based classification of central nervous system tumours
- In Favor of Logarithmic Scoring
- Multiomics, artificial intelligence, and precision medicine in perinatology
- Two-stage dynamic signal detection: A theory of choice, decision time, and confidence.
- Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for Dynamic Graph Systems
- Beyond scalar losses: calibrating segmentation models via gradient vector field surgery
- SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement
- Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
- The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
- FedDP-PALD: A Privacy-Preserving Federated Latent Diffusion Framework with Prototype Aggregation for Medical Data Synthesis
- ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
- Perturbation is All You Need for Extrapolating Language Models
- Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- A model-based restricted shapley value to measure the players' contribution to shot actions in football
- Reliable Remediation Impact Prediction for Black-Box Security Ratings
- Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data
- Assing Preferential Sampling in Retail Survival Data: A Bayesian Joint LGCP and Spatial Probit Model for Mini-Supermarket Closure in Tokyo
- Model Uncertainty under Non-Gaussian Errors: Bayesian Model Averaging and Selection in Stochastic Frontier Models
- FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting
- Artificial intelligence for risk assessment and outcome prediction in malignant haematology
- Continuous Autoregressive Language Models
- Outcome-based Reinforcement Learning to Predict the Future
- High dimensional online calibration in polynomial time
- Rethinking Early Stopping: Refine, Then Calibrate
- Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
- proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference
- An Adaptive Glicko-2 Rating Framework for Probabilistic Football Forecasting and Season Simulation
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Graceful Degradation and Related Fields
- Consistent Estimation of the Expected Brier Score in General Survival Models with Right‐Censored Event Times
- Strategy-Proof Incentives for Predictions
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
- Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
- Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
- Spline-Based Probability Calibration
- Predicting Lockean from gradational accuracy
- Am I Confused or Is This Confusing?: Deep Ensembles for ENSO Uncertainty Quantification
- Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
- Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach
- Do Large Language Models Know What They Don't Know? Kalshibench: A New Benchmark for Evaluating Epistemic Calibration via Prediction Markets
- EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
- The EEPAS Model Revisited: Statistical Formalism and a High-Performance, Reproducible Open-Source Framework
- WTNN: Weibull-Tailored Neural Networks for survival analysis
- Improving Multi-Class Calibration through Normalization-Aware Isotonic Techniques
- Multicalibration for LLM-based Code Generation
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- Knowing Your Uncertainty -- On the application of LLM in social sciences
- How (Mis)calibrated is Your Federated CLIP and What To Do About It?
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Comparing Variable Selection and Model Averaging Methods for Logistic Regression
- High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid
- Truthful Data Acquisition via Peer Prediction
- On Calibration and Out-of-domain Generalization
- MixMatch: A Holistic Approach to Semi-Supervised Learning
- Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
- Models, Markets, and the Forecasting of Elections
- Exploring the Watch-to-Warning Space: Experimental Outlook Performance during the 2019 Spring Forecasting Experiment in NOAA’s Hazardous Weather Testbed
- MEDIC: a network for monitoring data quality in collider experiments
- On classification, ranking, and probability estimation
- Why Calibration Error is Wrong Given Model Uncertainty: Using Posterior\n Predictive Checks with Deep Learning
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- Probability Calibration for Knowledge Graph Embedding Models
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Boundary-Aware Adversarial Filtering for Reliable Diagnosis under Extreme Class Imbalance
- Transparent Early ICU Mortality Prediction with Clinical Transformer and Per-Case Modality Attribution
- Edge-aware baselines for ogbn-proteins in PyTorch Geometric: species-wise normalization, post-hoc calibration, and cost-accuracy trade-offs
- Measuring Calibration in Deep Learning
- Incentives for Federated Learning: a Hypothesis Elicitation Approach
- Multi-Loss Sub-Ensembles for Accurate Classification with Uncertainty Estimation
- Credit scoring using neural networks and SURE posterior probability calibration
- Stochastic Predictive Analytics for Stocks in the Newsvendor Problem
- Bayesian Few-Shot Classification with One-vs-Each P 'olya-Gamma\n Augmented Gaussian Processes
- LT-Soups: Bridging Head and Tail Classes via Subsampled Model Soups
- Misaligned by Design: Incentive Failures in Machine Learning
- Machine Learning for Survival Analysis
- Calibrating Deep Neural Networks using Focal Loss
- AIA Forecaster: Technical Report
- Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
- Multivariate Variational Autoencoder
- Assessing win strength in MLB win prediction models
- Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
- Towards Continuous-variable Quantum Neural Networks for Biomedical Imaging
- Learnable Uncertainty under Laplace Approximations
- Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
- Contrastive Predictive Coding Done Right for Mutual Information Estimation
- Adversarially Adaptive Normalization for Single Domain Generalization
- More on verification of probability forecasts for football outcomes: score decompositions, reliability, and discrimination analyses
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- Administrative healthcare data applied to fracture risk assessment
- OpenMarket: A Synchronized Polymarket-Binance Dataset for High-Frequency Prediction-Market Research
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Mitigating Bias in Calibration Error Estimation
- Navigating prevalence shifts in image analysis algorithm deployment
- Advances in survival analyses: machine learning methods and model comparison
- Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
- Short-term spatio-temporal forecasting of human-caused wildland fire occurrence: an errors in variables approach
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- Exploring the Uncertainty Properties of Neural Networks' Implicit Priors\n in the Infinite-Width Limit
- Effect of the output activation function on the probabilities and errors in medical image segmentation
- Are Graph Neural Networks Miscalibrated?
- Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model
- The Pick-the-Winner-Picker Heuristic: Preference for Categorically Correct Forecasts
- Supraspecific Ecological Niche Models as a Tool for Predicting Burrowing Crayfish Habitat
- When Good Fit Goes Bad: Identifying and Minimising Overfitting in Ecological Niche Models
- Confidence Calibration and Predictive Uncertainty Estimation for Deep Medical Image Segmentation
- The Separation Plot: A New Visual Method for Evaluating the Fit of Binary Models
- Can tree cover extent be estimated directly from ICESat-2 spaceborne lidar data?
- Right Decisions from Wrong Predictions: A Mechanism Design Alternative to Individual Calibration
- Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
- Conditional Forecasts and Proper Scoring Rules for Reliable and Accurate Performative Predictions
- Instance-Adaptive Hypothesis Tests with Heterogeneous Agents
- CSU-PCAST: A Dual-Branch Transformer Framework for medium-range ensemble Precipitation Forecasting
- Capturing Intransitive Dominance in Tennis Forecasting: A Graph Neural Network Approach
- Calibration and Discrimination Optimization Using Clusters of Learned Representation
- No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- Uncertainty Estimation by Flexible Evidential Deep Learning
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- A modular framework for extreme weather generation
- Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?
- From Universal Approximation Theorem to Tropical Geometry of Multi-Layer Perceptrons
- Identity-Link IRT for Label-Free LLM Evaluation: Preserving Additivity in TVD-MI Scores
- Fast Visuomotor Policy for Robotic Manipulation
- Individual differences in reasoning: Implications for the rationality debate?
- ADVICE: Answer-Dependent Verbalized Confidence Estimation
- PENEX: AdaBoost-Inspired Neural Network Regularization
- The Fairness of Risk Scores Beyond Classification: Bipartite Ranking and the xAUC Metric
- A short history of philosophies of hydrological model evaluation and hypothesis testing
- Robust Estimation under Heavy Contamination using Enlarged Models
- Sample size for binary logistic prediction models: Beyond events per variable criteria
- Measuring Language Model Hallucinations Through Distributional Correctness
- Stronger Calibration Lower Bounds via Sidestepping
- Elicitability
- Neural Diffusion Processes for Physically Interpretable Survival Prediction
- Leveraging Clickstream Trajectories to Reveal Low-Quality Workers in Crowdsourced Forecasting Platforms
- Calibrating Verbalized Confidence with Self-Generated Distractors
- Assessing Large Language Models in Updating Their Forecasts with New Information
- Calibration Meets Reality: Making Machine Learning Predictions Trustworthy
- Bayesian and geometric analyses of power spectral densities of spin qubits in Si/SiGe quantum dot devices
- Demystifying the black box: A survey on explainable artificial intelligence (XAI) in bioinformatics
- C2GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
- Proper Proxy Scoring Rules
- Bézier Meets Diffusion: Robust Generation Across Domains for Medical Image Segmentation
- AEGIS: Authentic Edge Growth In Sparsity for Link Prediction in Edge-Sparse Bipartite Knowledge Graphs
- MCGrad: Multicalibration at Web Scale
- Low-Pathwidth GRAND: Exact Likelihood-Ordered Enumeration for BPSK Transmission over Correlated Gaussian Noise
- SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
- Memory compression and physical state augmentation favor different AMOC prediction tasks
- Uncertainty in Physics and AI: Taxonomy, Quantification, and Validation
- Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
- The State of Peptide Detectability in Computational Proteomics and Guidelines for AI Applications
- Measuring and comparing the accuracy of species distribution models with presence–absence data
- Probability Aggregation Methods in Geoscience
- Quantum Chemical Evaluation and QSAR Modeling of <i>N</i>-Nitrosamine Carcinogenicity
- The problem with the Brier score
- Temporal Drift in Privacy Recall: Users Misremember From Verbatim Loss to Gist-Based Overexposure
- Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
- Neural Earthquake Forecasting with Minimal Information: Limits, Interpretability, and the Role of Markov Structure
- On the Illusion of Success: An Empirical Study of Build Reruns and Silent Failures in Industrial CI
- Exploring Major Transitions in the Evolution of Biological Cognition With Artificial Neural Networks
- All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
- Similarity-Distance-Magnitude Activations
- Does Calibration Affect Human Actions?
- Epistemological Implementation of Social Choice Functions
- Deep Ensembles: A Loss Landscape Perspective
- Uncertainty-Aware Retinal Vessel Segmentation via Ensemble Distillation
- GrACE: A Generative Approach to Better Confidence Elicitation in Large Language Models
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Too Helpful, Too Harmless, Too Honest or Just Right?
- Stop using root-mean-square error as a precipitation target!
- Revisiting the Calibration of Modern Neural Networks
- ADHAM: Additive Deep Hazard Analysis Mixtures for Interpretable Survival Regression
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Predictive Inference Based on Markov Chain Monte Carlo Output
- The distribution of calibrated likelihood functions on the probability-likelihood Aitchison simplex
- Algorithm appreciation: People prefer algorithmic to human judgment
- Generalized Correlation Regression for Disentangling Dependence in Clustered Data
- Belief updating in AI‐risk debates: Exploring the limits of adversarial collaboration
- Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?
- ConfTuner: Training Large Language Models to Express Their Confidence Verbally
- Credence Calibration Game? Calibrating Large Language Models through Structured Play
- Regression Trees for Cumulative Incidence Functions
- SNAP-UQ: Self-supervised Next-Activation Prediction for Single-Pass Uncertainty in TinyML
- TCUQ: Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML
- Expert Incentives under Partially Contractible States
- How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
- Decomposing Global AUC into Cluster-Level Contributions for Localized Model Diagnostics
- Towards Unveiling Predictive Uncertainty Vulnerabilities in the Context of the Right to Be Forgotten
- Clinical Utility of the Automatic Phenotype Annotation in Unstructured Clinical Notes: ICU Use Cases
- Bayesian weighted discrete-time dynamic models for association football prediction
- Gender Differences in the Self-Assessment of Accuracy on Cognitive Tasks
- On Experiments
- Bayesian Conformal Prediction via the Bayesian Bootstrap
- Calibrated Language Models and How to Find Them with Label Smoothing
- EMORe: Motion-Robust 5D MRI Reconstruction via Expectation-Maximization-Guided Binning Correction and Outlier Rejection
- Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models
- A comparison of variable selection methods and predictive models for postoperative bowel surgery complications
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- Prediction Markets, Mechanism Design, and Cooperative Game Theory
- Generating Probabilities From Numerical Weather Forecasts by Logistic Regression
- Learning with Fenchel-Young Losses
- Transferable Calibration with Lower Bias and Variance in Domain Adaptation
- On Deep Neural Network Calibration by Regularization and its Impact on Refinement
- Measuring Forecasting Skill from Text
- Mind the Performance Gap: Examining Dataset Shift During Prospective Validation
- Screening of Informed and Uninformed Experts
- Calibrated Top-1 Uncertainty estimates for classification by score based models