This paper develops an end-to-end deep reinforcement learning (DRL) framework for long-short portfolio optimization in continuous-time trading environments. Two methodological advances underpin the approach. First, we formalise a realistic short-selling mechanism by constraining portfolio weights such that the sum of absolute weights equals one \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{(\sum |\omega _i|=1)}$$\end{document}, faithfully tracking capital usage and transaction dynamics over consecutive rebalancing periods. Second, we embed this mechanism in a fully data-driven DRL architecture that employs random episode sampling for robust training and couples specialised deep neural networks for high-dimensional time-series processing with a Sharpe-ratio-based reward, enabling the agent to learn optimal long-short allocations without prespecified trading rules. The framework is evaluated on two large equity universes — CSI 500 and S&P 500 constituents — using six independent random seeds to ensure robustness. Out-of-sample back-tests demonstrate that the learned strategy consistently outperforms a broad range of traditional and advanced portfolio optimization benchmarks, especially during high-volatility, high-risk market regimes. Block bootstrap significance tests confirm that the improvements in Sharpe ratio are statistically significant, with most results at the 1-5% levels and only isolated cases showing non-significance. Overall, the results show that integrating a principled short-selling constraint with DRL yields a resilient long-short portfolio strategy that adapts dynamically to changing market conditions, delivers superior risk-adjusted performance and materially advances the state of quantitative asset-allocation research.
Paper
References (79)
Scroll for more · 38 remaining