vix.ing · top · new · best · stats · spec

Powerful A/B-Testing Metrics and Where to Find Them

2024/10/08 by Olivier Jeunen, Shubham Baweja, Neeti Pokharna +1 · 1 voice
Computer Science · Mathematics · #Gaussian Processes and Bayesian Inference #Statistical Methods and Inference #Statistical Methods in Clinical Trials

paper · pdf · doi:10.1145/3640457.3688036

openalex publication_date 2024/10/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

Online controlled experiments, colloquially known as A/B-tests, are the bread and butter of real-world recommender system evaluation. Typically, end-users are randomly assigned some system variant, and a plethora of metrics are then tracked, collected, and aggregated throughout the experiment. A North Star metric (e.g. long-term growth or revenue) is used to assess which system variant should be deemed superior. As a result, most collected metrics are supporting in nature, and serve to either (i) provide an understanding of how the experiment impacts user experience, or (ii) allow for confident decision-making when the North Star metric moves insignificantly (i.e. a false negative or type-II error). The latter is not straightforward: suppose a treatment variant leads to fewer but longer sessions, with more views but fewer engagements; should this be considered a positive or negative outcome?

Discussions

Related