2025/09/03 by Benjamin Heymann, Heymann, Benjamin, Otmane Sakhi +1
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Risk and Portfolio Optimization #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2509.03438
openalex publication_date 2025/09/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the problem of directly optimizing a non-linear function of an outcome, where this outcome itself is the sum of many small contributions. The non-linearity of the function means that the problem is not equivalent to the maximization of the expectation of the individual contribution. By leveraging the concentration properties of the sum of individual outcomes, we derive a scalable descent algorithm that directly optimizes for our stated objective. This allows for instance to maximize the probability of successful A/B test, for which it can be wiser to target a success criterion, such as exceeding a given uplift, rather than chasing the highest expected payoff.