Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization
In modern financial markets, decision-makers increasingly rely on quantitative methods to navigate complex trade-offs among multiple, often conflicting objectives. This paper addresses constrained multi-objective optimization (MOO) with an application to portfolio optimization for minimizing risk and maximizing return. To this end, and to address existing gaps, we propose a novel reinforcement learning (RL)-guided non-dominated sorting genetic algorithm II (NSGA-II) enhanced with gray relational coefficients (GRC), termed RL-NSGA-II-GRC, which combines an RL agent controller and GRC-based selection to improve the convergence and diversity of the Pareto-optimal fronts. The agent adapts key evolutionary parameters online using population-level metrics of hypervolume, feasibility, and diversity, while the GRC-enhanced tournament operator ranks parents via a unified score simultaneously considering dominance rank, crowding distance, and geometric proximity to ideal reference. We evaluate the framework on the Kursawe and CONSTR benchmark problems and on a NASDAQ portfolio optimization application. On the benchmarks, RL-NSGA-II-GRC achieves convergence metric improvements of about 5.8% and 4.4% over the original NSGA-II, while preserving a well-distributed set of non-dominated solutions. In the portfolio application, the method produces a smooth and densely populated efficient frontier that supports the identification of the maximum Sharpe ratio portfolio (with annualized Sharpe ratio = 1.92), as well as utility-optimal portfolios for different risk-aversion levels. The main contributions of this work are three-fold: (1) we propose an RL-NSGA-II-GRC method that integrates an RL agent into the evolutionary framework to adaptively control key parameters using generational feedback; (2) we design a GRC-enhanced binary tournament selection operator that provides a comprehensive performance indicator to efficiently guide the search toward the Pareto-optimal front; (3) we demonstrate, on benchmark MOO problems and a NASDAQ portfolio case study, that the proposed method delivers improved convergence and well-populated efficient frontiers that support actionable investment insights.