The feedback data of recommender systems are often subject to what was\nexposed to the users; however, most learning and evaluation methods do not\naccount for the underlying exposure mechanism. We first show in theory that\napplying supervised learning to detect user preferences may end up with\ninconsistent results in the absence of exposure information. The counterfactual\npropensity-weighting approach from causal inference can account for the\nexposure mechanism; nevertheless, the partial-observation nature of the\nfeedback data can cause identifiability issues. We propose a principled\nsolution by introducing a minimax empirical risk formulation. We show that the\nrelaxation of the dual problem can be converted to an adversarial game between\ntwo recommendation models, where the opponent of the candidate model\ncharacterizes the underlying exposure mechanism. We provide learning bounds and\nconduct extensive simulation studies to illustrate and justify the proposed\napproach over a broad range of recommendation settings, which shed insights on\nthe various benefits of the proposed approach.\n