VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY
1950/01/01 by GLENN W. BRIER, Glenn W. Brier · 5,337 citations
Environmental Science · Mathematics · #Climatology #Computer science #Environmental and Industrial Safety #Geography #Geology #Mathematics #Meteorology #Statistics
paper · doi:10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2
published in Monthly Weather Review 78(1), 1-3 (American Meteorological Society)
openalex publication_date 1950/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Cited by
- DNA methylation-based classification of central nervous system tumours
- In Favor of Logarithmic Scoring
- Multiomics, artificial intelligence, and precision medicine in perinatology
- Two-stage dynamic signal detection: A theory of choice, decision time, and confidence.
- Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for Dynamic Graph Systems
- Beyond scalar losses: calibrating segmentation models via gradient vector field surgery
- SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement
- Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
- The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
- FedDP-PALD: A Privacy-Preserving Federated Latent Diffusion Framework with Prototype Aggregation for Medical Data Synthesis
- ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
- Perturbation is All You Need for Extrapolating Language Models
- Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- A model-based restricted shapley value to measure the players’ contribution to shot actions in football
- Reliable Remediation Impact Prediction for Black-Box Security Ratings
- Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data
- Assing Preferential Sampling in Retail Survival Data: A Bayesian Joint LGCP and Spatial Probit Model for Mini-Supermarket Closure in Tokyo
- Model Uncertainty under Non-Gaussian Errors: Bayesian Model Averaging and Selection in Stochastic Frontier Models
- FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting
- Artificial intelligence for risk assessment and outcome prediction in malignant haematology
- Continuous Autoregressive Language Models
- Outcome-based Reinforcement Learning to Predict the Future
- High dimensional online calibration in polynomial time
- Rethinking Early Stopping: Refine, Then Calibrate
- Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
- proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference
- An Adaptive Glicko-2 Rating Framework for Probabilistic Football Forecasting and Season Simulation
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Graceful Degradation and Related Fields
- Consistent Estimation of the Expected Brier Score in General Survival Models with Right‐Censored Event Times
- Strategy-Proof Incentives for Predictions
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
- Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
- Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
- Spline-Based Probability Calibration
- Predicting Lockean from gradational accuracy
- Am I Confused or Is This Confusing?: Deep Ensembles for ENSO Uncertainty Quantification
- Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
- Adversarially Robust Detection of Harmful Online Content: A Computational Design Science Approach
- Do Large Language Models Know What They Don't Know? Kalshibench: A New Benchmark for Evaluating Epistemic Calibration via Prediction Markets
- EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
- The EEPAS Model Revisited: Statistical Formalism and a High-Performance, Reproducible Open-Source Framework
- WTNN: Weibull-Tailored Neural Networks for survival analysis
- Improving Multi-Class Calibration through Normalization-Aware Isotonic Techniques
- Multicalibration for LLM-based Code Generation
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- Knowing Your Uncertainty -- On the application of LLM in social sciences
- How (Mis)calibrated is Your Federated CLIP and What To Do About It?
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Comparing Variable Selection and Model Averaging Methods for Logistic Regression
- High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid
- Truthful Data Acquisition via Peer Prediction
- On Calibration and Out-of-domain Generalization
- MixMatch: A Holistic Approach to Semi-Supervised Learning
- Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
- Models, Markets, and the Forecasting of Elections
- Exploring the Watch-to-Warning Space: Experimental Outlook Performance during the 2019 Spring Forecasting Experiment in NOAA’s Hazardous Weather Testbed
- MEDIC: a network for monitoring data quality in collider experiments
- On classification, ranking, and probability estimation
- Why Calibration Error is Wrong Given Model Uncertainty: Using Posterior Predictive Checks with Deep Learning
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- Probability Calibration for Knowledge Graph Embedding Models
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- Boundary-Aware Adversarial Filtering for Reliable Diagnosis under Extreme Class Imbalance
- Transparent Early ICU Mortality Prediction with Clinical Transformer and Per-Case Modality Attribution
- Edge-aware baselines for ogbn-proteins in PyTorch Geometric: species-wise normalization, post-hoc calibration, and cost-accuracy trade-offs
- Measuring Calibration in Deep Learning
- Incentives for Federated Learning: a Hypothesis Elicitation Approach
- Multi-Loss Sub-Ensembles for Accurate Classification with Uncertainty Estimation
- Credit scoring using neural networks and SURE posterior probability calibration
- Stochastic Predictive Analytics for Stocks in the Newsvendor Problem
- Bayesian Few-Shot Classification with One-vs-Each Pólya-Gamma Augmented Gaussian Processes
- LT-Soups: Bridging Head and Tail Classes via Subsampled Model Soups
- Misaligned by Design: Incentive Failures in Machine Learning
- Machine Learning for Survival Analysis
- Calibrating Deep Neural Networks using Focal Loss
- AIA Forecaster: Technical Report
- Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
- Multivariate Variational Autoencoder
- Assessing win strength in MLB win prediction models
- Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
- Towards Continuous-variable Quantum Neural Networks for Biomedical Imaging
- Learnable Uncertainty under Laplace Approximations
- Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
- Contrastive Predictive Coding Done Right for Mutual Information Estimation
- Adversarially Adaptive Normalization for Single Domain Generalization
- More on verification of probability forecasts for football outcomes: score decompositions, reliability, and discrimination analyses
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- Administrative healthcare data applied to fracture risk assessment
- OpenMarket: A Synchronized Polymarket-Binance Dataset for High-Frequency Prediction-Market Research
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Mitigating Bias in Calibration Error Estimation
- Navigating prevalence shifts in image analysis algorithm deployment
- Advances in survival analyses: machine learning methods and model comparison
- Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
- Short-term spatio-temporal forecasting of human-caused wildland fire occurrence: an errors in variables approach
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width Limit
- Effect of the output activation function on the probabilities and errors in medical image segmentation
- Are Graph Neural Networks Miscalibrated?
- Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model
- The Pick-the-Winner-Picker Heuristic: Preference for Categorically Correct Forecasts
- Supraspecific Ecological Niche Models as a Tool for Predicting Burrowing Crayfish Habitat
- When Good Fit Goes Bad: Identifying and Minimising Overfitting in Ecological Niche Models
- Confidence Calibration and Predictive Uncertainty Estimation for Deep Medical Image Segmentation
- Online Learning with Continuous Ranked Probability Score
- Stable Discovery of Interpretable Subgroups via Calibration in Causal Studies
- The Separation Plot: A New Visual Method for Evaluating the Fit of Binary Models
- Can tree cover extent be estimated directly from ICESat-2 spaceborne lidar data?
- Right Decisions from Wrong Predictions: A Mechanism Design Alternative to Individual Calibration
- Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
- Conditional Forecasts and Proper Scoring Rules for Reliable and Accurate Performative Predictions
- Instance-Adaptive Hypothesis Tests with Heterogeneous Agents
- CSU-PCAST: A Dual-Branch Transformer Framework for medium-range ensemble Precipitation Forecasting
- Capturing Intransitive Dominance in Tennis Forecasting: A Graph Neural Network Approach
- Calibration and Discrimination Optimization Using Clusters of Learned Representation
- No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- Uncertainty Estimation by Flexible Evidential Deep Learning
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- A modular framework for extreme weather generation
- Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?
- From Universal Approximation Theorem to Tropical Geometry of Multi-Layer Perceptrons
- Identity-Link IRT for Label-Free LLM Evaluation: Preserving Additivity in TVD-MI Scores
- Fast Visuomotor Policy for Robotic Manipulation
- Individual differences in reasoning: Implications for the rationality debate?
- ADVICE: Answer-Dependent Verbalized Confidence Estimation
- PENEX: AdaBoost-Inspired Neural Network Regularization
- The Fairness of Risk Scores Beyond Classification: Bipartite Ranking and the xAUC Metric
- A short history of philosophies of hydrological model evaluation and hypothesis testing
- Robust Estimation under Heavy Contamination using Enlarged Models
- Sample size for binary logistic prediction models: Beyond events per variable criteria
- Measuring Language Model Hallucinations Through Distributional Correctness
- Stronger Calibration Lower Bounds via Sidestepping
- Elicitability
- Neural Diffusion Processes for Physically Interpretable Survival Prediction
- Leveraging Clickstream Trajectories to Reveal Low-Quality Workers in Crowdsourced Forecasting Platforms
- Calibrating Verbalized Confidence with Self-Generated Distractors
- Assessing Large Language Models in Updating Their Forecasts with New Information
- Calibration Meets Reality: Making Machine Learning Predictions Trustworthy
- Bayesian and geometric analyses of power spectral densities of spin qubits in Si/SiGe quantum dot devices
- Demystifying the black box: A survey on explainable artificial intelligence (XAI) in bioinformatics
- C2GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
- Proper Proxy Scoring Rules
- Bézier Meets Diffusion: Robust Generation Across Domains for Medical Image Segmentation
- AEGIS: Authentic Edge Growth In Sparsity for Link Prediction in Edge-Sparse Bipartite Knowledge Graphs
- MCGrad: Multicalibration at Web Scale
- Low-Pathwidth GRAND: Exact Likelihood-Ordered Enumeration for BPSK Transmission over Correlated Gaussian Noise
- SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
- Memory compression and physical state augmentation favor different AMOC prediction tasks
- Uncertainty in Physics and AI: Taxonomy, Quantification, and Validation
- Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
- The State of Peptide Detectability in Computational Proteomics and Guidelines for AI Applications
- Measuring and comparing the accuracy of species distribution models with presence–absence data
- Probability Aggregation Methods in Geoscience
- Quantum Chemical Evaluation and QSAR Modeling of N-Nitrosamine Carcinogenicity
- The problem with the Brier score
- Temporal Drift in Privacy Recall: Users Misremember From Verbatim Loss to Gist-Based Overexposure
- Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
- Neural Earthquake Forecasting with Minimal Information: Limits, Interpretability, and the Role of Markov Structure
- On the Illusion of Success: An Empirical Study of Build Reruns and Silent Failures in Industrial CI
- Exploring Major Transitions in the Evolution of Biological Cognition With Artificial Neural Networks
- All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
- Similarity-Distance-Magnitude Activations
- Does Calibration Affect Human Actions?
- Epistemological Implementation of Social Choice Functions
- Deep Ensembles: A Loss Landscape Perspective
- Uncertainty-Aware Retinal Vessel Segmentation via Ensemble Distillation
- GrACE: A Generative Approach to Better Confidence Elicitation in Large Language Models
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Too Helpful, Too Harmless, Too Honest or Just Right?
- Stop using root-mean-square error as a precipitation target!
- Revisiting the Calibration of Modern Neural Networks
- ADHAM: Additive Deep Hazard Analysis Mixtures for Interpretable Survival Regression
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Predictive Inference Based on Markov Chain Monte Carlo Output
- The distribution of calibrated likelihood functions on the probability-likelihood Aitchison simplex
- Algorithm appreciation: People prefer algorithmic to human judgment
- Generalized Correlation Regression for Disentangling Dependence in Clustered Data
- Belief updating in AI‐risk debates: Exploring the limits of adversarial collaboration
- Survival Analysis as Imprecise Classification with Trainable Kernels
- Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?
- ConfTuner: Training Large Language Models to Express Their Confidence Verbally
- Credence Calibration Game? Calibrating Large Language Models through Structured Play
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
- Regression Trees for Cumulative Incidence Functions
- SNAP-UQ: Self-supervised Next-Activation Prediction for Single-Pass Uncertainty in TinyML
- TCUQ: Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML
- Calibrated Reliable Regression using Maximum Mean Discrepancy
- Expert Incentives under Partially Contractible States
- How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- Strictly Proper Scoring Rules, Prediction, and Estimation
- Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
- Decomposing Global AUC into Cluster-Level Contributions for Localized Model Diagnostics
- Towards Unveiling Predictive Uncertainty Vulnerabilities in the Context of the Right to Be Forgotten
- Clinical Utility of the Automatic Phenotype Annotation in Unstructured Clinical Notes: ICU Use Cases
- Bayesian weighted discrete-time dynamic models for association football prediction
- Gender Differences in the Self-Assessment of Accuracy on Cognitive Tasks
- On Experiments
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Bayesian Conformal Prediction via the Bayesian Bootstrap
- Calibrated Language Models and How to Find Them with Label Smoothing
- EMORe: Motion-Robust 5D MRI Reconstruction via Expectation-Maximization-Guided Binning Correction and Outlier Rejection
- Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models
- A comparison of variable selection methods and predictive models for postoperative bowel surgery complications
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- Prediction Markets, Mechanism Design, and Cooperative Game Theory
- Generating Probabilities From Numerical Weather Forecasts by Logistic Regression
- Learning with Fenchel-Young Losses
- Transferable Calibration with Lower Bias and Variance in Domain Adaptation
- On Deep Neural Network Calibration by Regularization and its Impact on Refinement
- Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- Measuring Forecasting Skill from Text
- Threshold Choice Methods: the Missing Link
- Mind the Performance Gap: Examining Dataset Shift During Prospective Validation
- Screening of Informed and Uninformed Experts
- Calibrated Top-1 Uncertainty estimates for classification by score based models
- Temporal Probability Calibration
- Using Machine Learning Techniques to Identify Key Risk Factors for Diabetes and Undiagnosed Diabetes
- To Trust or Not to Trust: On Calibration in ML-based Resource Allocation for Wireless Networks
- Confidence Calibration in Vision-Language-Action Models
- Bayesian Deep Learning for Convective Initiation Nowcasting Uncertainty Estimation
- Know What You Don't Know: Uncertainty Calibration of Process Reward Models
- Machine learning-based clinical prediction modeling -- A practical guide for clinicians
- Deep Lifetime Clustering
- Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
- Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Use of the likelihood for measuring the skill of probabilistic forecasts
- Solving for multi-class: a survey and synthesis
- The Role of Individual Differences in the Accuracy of Confidence Judgments
- Protected probabilistic classification
- Early Warning with Calibrated and Sharper Probabilistic Forecasts
- Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
- Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
- An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
- Calibrating Deep Neural Network Classifiers on Out-of-Distribution Datasets
- Where are we with calibration under dataset shift in image classification?
- On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
- Stable discovery of interpretable subgroups via calibration in causal studies
- Property Elicitation on Imprecise Probabilities
- Personalized Federated Learning with Gaussian Processes
- An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks
- SSLayout360: Semi-Supervised Indoor Layout Estimation from 360-Degree Panorama
- Designing Informative Securities
- Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications
- Simulation and evaluation of local daily temperature and precipitation series derived by stochastic downscaling of ERA5 reanalysis
- Compressive Visual Representations
- Learning from Positive and Unlabeled Data by Identifying the Annotation Process
- On the Need of Preserving Order of Data When Validating Within-Project Defect Classifiers
- An exploration of potential risk factors for gastroschisis using decision tree learning
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- Sample Margin-Aware Recalibration of Temperature Scaling
- Collaborative Sampling in Generative Adversarial Networks
- The Finley Affair: A Signal Event in the History of Forecast Verification
- RobustiPy: An efficient next generation multiversal library with model selection, averaging, resampling, and explainable artificial intelligence
- On Equivariant Model Selection through the Lens of Uncertainty
- Causal Interventions in Bond Multi-Dealer-to-Client Platforms
- h-calibration: Rethinking Classifier Recalibration with Probabilistic Error-Bounded Objective
- Choice of Scoring Rules for Indirect Elicitation of Properties with Parametric Assumptions
- Conditional Probability Tree Estimation Analysis and Algorithms
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU Networks
- Robust Training with Data Augmentation for Medical Imaging Classification
- Attended Temperature Scaling: A Practical Approach for Calibrating Deep Neural Networks
- Reasoning under Uncertainty: Some Monte Carlo Results
- Technical Note: Towards ROC Curves in Cost Space
- Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
- Probabilistic patient risk profiling with pair-copula constructions
- Double or Nothing: Multiplicative Incentive Mechanisms for Crowdsourcing
- DeepGLEAM: A hybrid mechanistic and deep learning model for COVID-19 forecasting
- Adaptive Bayesian Very Short-Term Wind Power Forecasting Based on the Generalised Logit Transformation
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
- Equitable Discrimination in Survival Prediction: The Maximum Expected C-Index
- Evaluating Smartphone Pressure Observations for Mesoscale Analyses and Forecasts
- Impacts of Assimilating Smartphone Pressure Observations on Forecast Skill during Two Case Studies in the Pacific Northwest
- Performance of the HWRF Rapid Intensification Analog Ensemble (HWRF RI-AnEn) during the 2017 and 2018 HFIP Real-Time Demonstrations
- Probabilistic Prediction of Interactive Driving Behavior via Hierarchical Inverse Reinforcement Learning
- Meta-Embedding as Auxiliary Task Regularization
- Probabilistic Forecast of Real-Time LMP and Network Congestion
- Average Calibration Losses for Reliable Uncertainty in Medical Image Segmentation
- Balancing Accuracy, Calibration, and Efficiency in Active Learning with Vision Transformers Under Label Noise
- Online Learning: Beyond Regret
- Bayesian deep learning of affordances from RGB images
- Hybrid forecasting of geopolitical events†
- Improving model calibration with accuracy versus uncertainty optimization
- Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
- Privacy Preserving Recalibration under Domain Shift
- Position: Stop Chasing the C-index when Evaluating Survival Analysis Models
- Is Your Explanation Reliable: Confidence-Aware Explanation on Graph Neural Networks
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
- Multiple decision trees
- Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
- Towards a Fatality-Aware Benchmark of Probabilistic Reaction Prediction in Highly Interactive Driving Scenarios
- Knowing More About Questions Can Help: Improving Calibration in Question Answering
- Revisiting Reweighted Risk for Calibration: AURC, Focal, and Inverse Focal Loss
- Uncertainty Quantification with Proper Scoring Rules: Adjusting Measures to Prediction Tasks
- Post-processing of wind gusts from COSMO-REA6 with a spatial Bayesian hierarchical extreme value model
- Uncertainty Estimation for Heterophilic Graphs Through the Lens of Information Theory
- The Wisdom of the Crowd and Higher-Order Beliefs
- Identifying and Exploiting Structures for Reliable Deep Learning
- SquareχPO: Differentially Private and Robust χ2-Preference Optimization in Offline Direct Alignment
- Improving Classifier Confidence using Lossy Label-Invariant Transformations
- Uncertainty Quantification 360: A Holistic Toolkit for Quantifying and Communicating the Uncertainty of AI
- Maximum Likelihood Estimation of Flexible Survival Densities with Importance Sampling
- SLA2P: Self-supervised Anomaly Detection with Adversarial Perturbation
- Toward a Characterization of Loss Functions for Distribution Learning
- Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
- PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise
- Calibration of Neural Networks using Splines
- Urban transport systems shape experiences of social segregation
- Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
- Be Confident! Towards Trustworthy Graph Neural Networks via Confidence Calibration
- Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior
- Quantile Regularization: Towards Implicit Calibration of Regression Models
- Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
- Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift
- Reasoning Models Better Express Their Confidence
- Dating and localizing an invasion from post-introduction data and a coupled reaction–diffusion–absorption model
- LAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
- Multimodal domain adaptation under label shift and blockwise missing modalities
- Mitigating Sampling Bias and Improving Robustness in Active Learning
- Retail banking closures in the United Kingdom. Are neighbourhood characteristics associated with retail bank branch closures?
- Aggregating Concepts of Accuracy and Fairness in Prediction Algorithms
- Probabilistic approach to longitudinal response prediction: application to radiomics from brain cancer imaging
- Multimodal Cancer Modeling in the Age of Foundation Model Embeddings
- Continuous Visual Autoregressive Generation via Score Maximization
- Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
- Model Criticism of Bayesian Networks with Latent Variables
- Forecasting Future Language: Context Design for Mention Markets
- The Manokhin Probability Matrix: A Diagnostic Framework for Classifier Probability Quality
- Understanding Misunderstanding: Social Psychological Perspectives
- Asymmetric Penalties Underlie Proper Loss Functions in Probabilistic Forecasting
- Exploring the Impact of Explainable AI and Cognitive Capabilities on Users' Decisions
- Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
- Inside the Planning Fallacy: The Causes and Consequences of Optimistic Time Predictions
- Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI
- Predicting Recessions with Leading Indicators: Model Averaging and Selection over the Business Cycle
- CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction
- Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- Computationally Inferred Genealogical Networks Uncover Long-Term Trends in Assortative Mating
- Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory
- Leveraging Artificial Intelligence to Improve Chronic Disease Care: Methods and Application to Pharmacotherapy Decision Support for Type-2 Diabetes Mellitus
- Towards Fair Comparisons of AI- and Physics-Based Weather Models for Extreme Events via the Weighted Potential CRPS
- ForecastQA: A Question Answering Challenge for Event Forecasting with Temporal Text Data
- Martingale Doppelgänger-Eval: An Identification Framework for Auditing Candlestick Understanding in Vision-Language Models
- Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
- Multi-Scale Markov Switching GARCH
- Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets
- Comparison of non-homogeneous regression models for probabilistic wind speed forecasting
- Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment.
- Towards a balanced social psychology: Causes, consequences, and cures for the problem-seeking approach to social behavior and cognition
- Evaluating performance and potential clinical benefit of the Swedish On Scene Injury Severity Prediction (OSISP) model for prehospital field triage on Norwegian trauma data
- Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
- Spatial correlates of forest and land fires in Indonesia
- Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
- Diversity in Interpretations of Probability: Implications for Weather Forecasting
- Do those who know more also know more about how much they know?
- Support theory: A nonextensional representation of subjective probability.
- ForesightFlow: An Information Leakage Score Framework for Prediction Markets
- Metacognition in LLMs: Foundations, Progress, and Opportunities
- Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems
- Algorithmic Monitoring: Measuring Market Stress with Machine Learning
- Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
- Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review
- Like Goes with Like: The Role of Representativeness in Erroneous and Pseudo-Scientific Beliefs
- AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction
- High-order joint embedding for multi-level link prediction
- On the visualization, verification and recalibration of ternary probabilistic forecasts
- Extended-range statistical ENSO prediction through operator-theoretic techniques for nonlinear dynamics
- Beyond Cox Models: Assessing the Performance of Machine-Learning Methods in Non-Proportional Hazards and Non-Linear Survival Analysis
- A weakly informative default prior distribution for logistic and other regression models
- Optimising HEP parameter fits via Monte Carlo weight derivative regression
- Beyond the Correlation Coefficient in Studies of Self-Assessment Accuracy: Commentary on Zell & Krizan (2014)
- Clinical Versus Actuarial Judgment
- Heuristics and Biases
- Beyond Strictly Proper Scoring Rules: The Importance of Being Local
- Categorical Data Analysis Using a Skewed Weibull Regression Model
- Forecasting: theory and practice
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- Proper local scoring rules on discrete sample spaces
- Local proper scoring rules of order two
- Optimal distributions of growing‐type initial perturbations for ensemble forecasts: Theory and application in the Lorenz‐96 model
- A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools
- Why ex post peer review encourages high-risk research while ex ante review discourages it
- Forest-Based and Semiparametric Methods for the Postprocessing of Rainfall Ensemble Forecasting
- When Predictions Fail: The Dilemma of Unrealistic Optimism
- Skill of data based predictions versus dynamical models -- case study on\n extreme temperature anomalies
- The prediction of extreme uncertainty-production events in three-dimensional Navier-Stokes turbulence
- Adaptive Elicitation of Latent Information Using Natural Language
- Ambiguity and self-evaluation: The role of idiosyncratic trait definitions in self-serving assessments of ability.
- Sympathetic Magical Thinking: The Contagion and Similarity “Heuristics”
- Case mix, outcome and activity for patients admitted to intensive care units requiring chronic renal dialysis: a secondary analysis of the ICNARC Case Mix Programme Database. [europepmc]
- Outcomes following oesophagectomy in patients with oesophageal cancer: a secondary analysis of the ICNARC Case Mix Programme Database. [europepmc]
- Comparison of predictive modeling approaches for 30-day all-cause non-elective readmission risk. [europepmc]
- Determinants of the calibration of SAPS II and SAPS 3 mortality scores in intensive care: a European multicenter study. [europepmc]
- DNA methylation-based classification of central nervous system tumours. [europepmc]
- A machine learning approach to estimating preterm infants survival: development of the Preterm Infants Survival Assessment (PISA) predictor. [europepmc]
- Comparing Bayesian and non-Bayesian accounts of human confidence reports. [europepmc]
- Characterising risk of in-hospital mortality following cardiac arrest using machine learning: A retrospective international registry study. [europepmc]
- The prediction of suicide in severe mental illness: development and validation of a clinical prediction rule (OxMIS). [europepmc]
- The Brier score does not evaluate the clinical utility of diagnostic tests or prediction models. [europepmc]
- Estimating the success of re-identifications in incomplete datasets using generative models. [europepmc]
- Interrater Reliability of Experts in Identifying Interictal Epileptiform Discharges in Electroencephalograms. [europepmc]
- Development of Risk Prediction Equations for Incident Chronic Kidney Disease. [europepmc]
- Prognosticating for Adult Patients With Advanced Incurable Cancer: a Needed Oncologist Skill. [europepmc]
- Informative missingness in electronic health record systems: the curse of knowing. [europepmc]
- Accommodating individual travel history and unsampled diversity in Bayesian phylogeographic inference of SARS-CoV-2. [europepmc]
- Assessment of Machine Learning to Estimate the Individual Treatment Effect of Corticosteroids in Septic Shock. [europepmc]
- A meta-learning approach for genomic survival analysis. [europepmc]
- Use of Machine Learning Models to Predict Death After Acute Myocardial Infarction. [europepmc]
- Integrated multi-omics analysis of ovarian cancer using variational autoencoders. [europepmc]
- The relationship of smoking to cg05575921 methylation in blood and saliva DNA samples from several studies. [europepmc]
- Machine Learning-Based Models Incorporating Social Determinants of Health vs Traditional Models for Predicting In-Hospital Mortality in Patients With Heart Failure. [europepmc]
- Predictive Accuracy of Stroke Risk Prediction Models Across Black and White Race, Sex, and Age Groups. [europepmc]
- Development of machine learning-based models to predict 10-year risk of cardiovascular disease: a prospective cohort study. [europepmc]
- Predicting whether patients will achieve minimal clinically important differences following hip or knee arthroplasty. [europepmc]