Off-policy estimation for long-horizon problems is important in many\nreal-life applications such as healthcare and robotics, where high-fidelity\nsimulators may not be available and on-policy evaluation is expensive or\nimpossible. Recently, \\cite{liu18breaking} proposed an approach that avoids the\n\\emph{curse of horizon} suffered by typical importance-sampling-based methods.\nWhile showing promising results, this approach is limited in practice as it\nrequires data be drawn from the \\emph{stationary distribution} of a\n\\emph{known} behavior policy. In this work, we propose a novel approach that\neliminates such limitations. In particular, we formulate the problem as solving\nfor the fixed point of a certain operator. Using tools from Reproducing Kernel\nHilbert Spaces (RKHSs), we develop a new estimator that computes importance\nratios of stationary distributions, without knowledge of how the off-policy\ndata are collected. We analyze its asymptotic consistency and finite-sample\ngeneralization. Experiments on benchmarks verify the effectiveness of our\napproach.\n