Pareto-Constrained Deep Q-Learning for Multi-Objective News Recommendation with Ranking Optimization and Deep Engagement Modeling
Delivering personalized news content is essential for improving user interaction and satisfaction on digital media platforms. Balancing multiple, often conflicting, objectives such as click-through rate (CTR), dwell time, and content diversity presents a complex reward design challenge, further compounded by the sparsity and noise of implicit user feedback, which hinders precise modeling of user preferences in news recommendation tasks. To address the challenges of balancing multiple conflicting objectives and dealing with sparse user feedback, this study presents a Pareto-Constrained Deep Q-Learning framework for multi-objective news recommendation, which balances goals like CTR, diversity, and dwell time using ranking-based optimization and a deep learning dwell time predictor to improve personalization and user engagement. The proposed system achieves notable improvements over baselines like NAML, NRMS, and DAN, with a CTR of 0.099, Precision@K of 0.0984, nDCG@K of 0.3916, minimal training loss, and 99.5% constraint satisfaction. Thus, this study presents a Pareto-Constrained Deep Q-Learning framework for multi-objective news recommendation, which balances goals like CTR, diversity, and dwell time using ranking-based optimization and a deep learning dwell time predictor to improve personalization and user engagement.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex