Counterfactual Risk Minimization with IPS-Weighted BPR and Self-Normalized Evaluation in Recommender Systems

Learning and evaluating recommender systems from logged implicit feedback is challenging due to exposure bias. While inverse propensity scoring (IPS) corrects this bias, it often suffers from high variance and instability. In this paper, we present a simple and effective pipeline that integrates IPS-weighted training with an IPS-weighted Bayesian Personalized Ranking (BPR) objective augmented by a Propensity Regularizer (PR). We compare Direct Method (DM), IPS, and Self-Normalized IPS (SNIPS) for offline policy evaluation, and demonstrate how IPS-weighted training improves model robustness under biased exposure. The proposed PR further mitigates variance amplification from extreme propensity weights, leading to more stable estimates. Experiments on synthetic and MovieLens 100K data show that our approach generalizes better under unbiased exposure while reducing evaluation variance compared to naive and standard IPS methods, offering practical guidance for counterfactual learning and evaluation in real-world recommendation settings.

Paper

References (13)

07Adebiasedpairwiselearningframeworkforunbiasedrecommendersystem2021 · Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
08LightGCN:Simplifyingandpoweringgraphconvolutionnetworkfor recommendation2020 · Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
09Deep learning based recommender system: A survey and new perspectives2019 · Comput. Surveys
10Counterfactual risk minimization: Learning from logged bandit feedback2015 · Proceedings of the 32nd International Conference on Machine Learning . PMLR
12where log ( ) is the probability of showing item under the logging policy. This heavily favors high-index (popular) items. A.1.3Observed Interaction Generation

Scroll for more · 1 remaining

Similar papers

© 2026 NYSGPT2525 LLC