LLM-based Personalized Portfolio Recommender: Integrating Large Language Models and Reinforcement Learning for Intelligent Investment Strategy Optimization
In modern financial markets, investors increasingly seek personalized and adaptive portfolio strategies that reflect their individual risk preferences and respond to dynamic market conditions. Traditional rule-based or static optimization approaches often fail to capture the nonlinear interactions among investor behavior, market volatility, and evolving financial objectives. To address these limitations, this paper introduces the LLM-based Personalized Portfolio Recommender (L-PPR), an integrated framework that combines Large Language Models (LLMs), reinforcement learning, and individualized risk preference modeling to support intelligent investment decision-making. The proposed system comprises three core components: (1) a Conversational Financial Agent (FinAgent) that engages with users through natural language, gathers behavioral feedback, and delivers interpretable advisory explanations; (2) a Personalization and Risk Modeling Module that infers investor-specific risk tolerance using Bayesian inference and behavioral imitation learning; and (3) a Strategy Recommendation Engine based on Proximal Policy Optimization (PPO) that generates personalized asset allocation strategies conditioned on user embeddings and real-time market states. Experimental results on a simulated multi-asset portfolio dataset show that L-PPR substantially outperforms established baselines, including Mean–Variance Optimization (MVO), Deep Reinforcement Learning Portfolio (DRL-PPO), and BERT-based Financial Advisor (BERT-FA). Specifically, L-PPR achieves a 73.8% improvement in annualized return and a 33.2% reduction in maximum drawdown relative to MVO. It also records the highest Sharpe Ratio (1.45), Information Ratio (0.78), and User Alignment Score (0.89). These findings demonstrate that L-PPR effectively enhances risk-adjusted performance while delivering superior personalization and user satisfaction.