vix.ing · top · new · best · stats · spec

What is the objective of reasoning with reinforcement learning?

2025/10/15 by Damek Davis, Benjamin Recht, Davis, Damek +1 · 1 citation
Computer Science · #Cognitive Science and Mapping #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Optimization and Control (math.OC)

paper · pdf · doi:10.48550/arxiv.2510.13651

openalex publication_date 2025/10/15 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28

Abstract

We show that several popular algorithms for reinforcement learning in large language models with binary rewards can be viewed as stochastic gradient ascent on a monotone transform of the probability of a correct answer given a prompt. In particular, the transformation associated with rejection sampling algorithms is the logarithm and that associated with the GRPO algorithm is the arcsine of the square root.

Citations

Cited by

Related