2025/10/07 by Pei Xue, Xue, Pei, Yuanchun Ye +1
Decision Sciences · Economics, Econometrics and Finance · #91G10 (Primary) 68T05 #91G60 (Secondary) #Computational Engineering #FOS: Computer and information sciences #Finance #Financial Markets and Investment Strategies #Stock Market Forecasting Methods #and Science (cs.CE)
paper · pdf · doi:10.48550/arxiv.2510.06466
openalex publication_date 2025/10/07 · openalex created_date 2025/10/18 · openalex updated_date 2026/07/28
We develop a deep reinforcement learning framework for dynamic portfolio optimization that combines a Dirichlet policy with cross-sectional attention mechanisms. The Dirichlet formulation ensures that portfolio weights are always feasible, handles tradability constraints naturally, and provides a stable way to explore the allocation space. The model integrates per-asset temporal encoders with a global attention layer, allowing it to capture sector relationships, factor spillovers, and other cross asset dependencies. The reward function includes transaction costs and portfolio variance penalties, linking the learning objective to traditional mean variance trade offs. The results show that attention based Dirichlet policies outperform equal-weight and standard reinforcement learning benchmarks in terms of terminal wealth and Sharpe ratio, while maintaining realistic turnover and drawdown levels. Overall, the study shows that combining principled action design with attention-based representations improves both the stability and interpretability of reinforcement learning for portfolio management.