2025/07/16 by Ahmet Umur Özsoy, Özsoy, Ahmet Umur
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #FOS: Economics and business #Machine Learning (cs.LG) #Mathematical Finance (q-fin.MF) #Reinforcement Learning in Robotics #Smart Grid Energy Management
paper · pdf · doi:10.48550/arxiv.2507.12657
openalex publication_date 2025/07/16 · openalex created_date 2025/10/18 · openalex updated_date 2026/07/28
We reinterpret and propose a framework for pricing path-dependent financial derivatives by estimating the full distribution of payoffs using Distributional Reinforcement Learning (DistRL). Unlike traditional methods that focus on expected option value, our approach models the entire conditional distribution of payoffs, allowing for risk-aware pricing, tail-risk estimation, and enhanced uncertainty quantification. We demonstrate the efficacy of this method on Asian options, using quantile-based value function approximators.