2024/08/17 by Min Dai, Dai, Min, Yu Sun +5 · 3 citations
Economics, Econometrics and Finance · Engineering · #FOS: Economics and business #FOS: Mathematics #Mathematical Finance (q-fin.MF) #Monetary Policy and Economic Impact #Optimization and Control (math.OC) #Pricing of Securities (q-fin.PR) #Reservoir Engineering and Simulation Methods #Stochastic processes and financial applications
paper · pdf · doi:10.48550/arxiv.2408.09242
openalex publication_date 2024/08/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study optimal stopping for diffusion processes with unknown model primitives within the continuous-time reinforcement learning (RL) framework developed by Wang et al. (2020), and present applications to option pricing and portfolio choice. By penalizing the corresponding variational inequality formulation, we transform the stopping problem into a stochastic optimal control problem with two actions. We then randomize controls into Bernoulli distributions and add an entropy regularizer to encourage exploration. We derive a semi-analytical optimal Bernoulli distribution, based on which we devise RL algorithms using the martingale approach established in Jia and Zhou (2022a). We establish a policy improvement theorem and prove the fast convergence of the resulting policy iterations. We demonstrate the effectiveness of the algorithms in pricing finite-horizon American put options, solving Merton's problem with transaction costs, and scaling to high-dimensional optimal stopping problems. In particular, we show that both the offline and online algorithms achieve high accuracy in learning the value functions and characterizing the associated free boundaries.