Deep Deterministic Policy Gradient for Portfolio Management

Portfolio management is a financial problem that has been the subject of much research over the years. It is a planning task where an agent constantly redistributes resources across a set of assets in order to achieve investment objectives and thereby maximize return. However, it remains difficult to obtain an optimal strategy in an environment as complex and dynamic as the financial market. Our article focuses on solving this stochastic control problem in order to obtain an optimal strategy that would allow us to make profitable decisions by interacting directly with the environment. To do this, we explore the power of deep reinforcement learning which differs from traditional Machine Learning by combining the task of predicting stock behavior and analyzing the optimal course of action in a single unit, thus aligning the problem of Machine Learning with the investor's objectives. As a method, we propose to use the Deep Deterministic Policy Gradient which is an off-policy algorithm and is used for environments with continuous action spaces. The obtained results demonstrate that the model achieves a higher rate of return than the strategy of “Uniform Buy and Hold” stocks and the strategy of “Buy Best Stock in last month”.

Paper

Full text

PDF

Deep Deterministic Policy Gradient for Portfolio Management

Semantic Scholar · Computer Science · 2020

Abstract

Portfolio management is a financial problem that has been the subject of much research over the years. It is a planning task where an agent constantly redistributes resources across a set of assets in order to achieve investment objectives and thereby maximize return. However, it remains difficult to obtain an optimal strategy in an environment as complex and dynamic as the financial market. Our article focuses on solving this stochastic control problem in order to obtain an optimal strategy that would allow us to make profitable decisions by interacting directly with the environment. To do this, we explore the power of deep reinforcement learning which differs from traditional Machine Learning by combining the task of predicting stock behavior and analyzing the optimal course of action in a single unit, thus aligning the problem of Machine Learning with the investor's objectives. As a method, we propose to use the Deep Deterministic Policy Gradient which is an off-policy algorithm and is used for environments with continuous action spaces. The obtained results demonstrate that the model achieves a higher rate of return than the strategy of “Uniform Buy and Hold” stocks and the strategy of “Buy Best Stock in last month”.

Similar papers

© 2026 NYSGPT2525 LLC