QLBS: Q-Learner in the Black-Scholes(-Merton) Worlds

This paper presents a discrete-time option pricing model that is rooted in Reinforcement Learning (RL), and more specifically in the famous Q-Learning method of RL. We construct a risk-adjusted Markov Decision Process for a discrete-time version of the classical Black-Scholes-Merton (BSM) model, where the option price is an optimal Q-function, while the optimal hedge is a second argument of thi…

Paper

Similar papers

© 2026 NYSGPT2525 LLC