We sincerely thank the referee for the positive feedback on the empirical nature and performance of our paper. Below, we address the remaining comments on our theoretical analysis.
- __Latent v.s. fixed variables.__
We will revise this as suggested. As outlined in the rebuttal, the error bound remains similar to Theorem 1, with an additional term of $c_2 R_\max (N^{-1/2}+T^{-1/2})$ up to some logarithmic factors, to account for the randomness of the latent factors. We hope this clarification resolves your concerns.
- __Clarification on the assumptions.__
This issue was not raised in the official review, but we believe Assumptions 1–4 are mild and achievable. We hope our response addresses your concerns.
* Assumptions 1 and 2 are standard in the reinforcement learning literature (e.g., [1], [2], [3], [4], [11], [12], [13], [14]). Though space limits citations, these assumptions are widely applied in offline policy optimization and off-policy evaluation.
* Assumption 3 is concerned with the estimation error of the transition function. This assumption is flexible, as $\varepsilon_{\mathcal{P},\delta}$ can be adjusted to a larger value to meet the condition. In our implementation, we use a conditional Gaussian model for the transition function, following [6]. The total variation bound, measured by $\varepsilon_{\mathcal{P},\delta}$, thus reduces to the estimation errors of the mean and covariance functions. Using neural networks for function approximation allows the estimator to achieve an optimal non-parametric convergence rate (e.g., [7], [8], [9]). This yields the specific form of $\varepsilon_{\mathcal{P},\delta}$. Even without the conditional Gaussian model, deep generative learning algorithms with theoretical guarantees can be employed, with error bounds provided in studies like [10].
* Assumption 4 is concerned with the estimation error of the latent confounders. Like Assumption 3, it is flexible, as $\varepsilon_{U,W,\delta}$ can be adjusted to a larger value. Additionally, under a two-way additive model assumption, the factors $\\{U\_i\\}\_i$ and $\\{W\_t\\}\_t$ can be estimated at orders of $\sqrt{T^{-1/2}\log (N/\delta)}$ and $\sqrt{N^{-1/2}\log (T/\delta)}$ with probabilities at least $1-O(\delta)$ (see, [15]). This provides the detailed forms of $\varepsilon_{U,W,\delta}$.
[1] Chen J, Jiang N. Information-theoretic considerations in batch reinforcement learning[C]//International Conference on Machine Learning. PMLR, 2019: 1042-1051. \
[2] Fan J, Wang Z, Xie Y, et al. A theoretical analysis of deep Q-learning[C]//Learning for dynamics and control. PMLR, 2020: 486-489.\
[3] Liu Y, Swaminathan A, Agarwal A, et al. Provably good batch off-policy reinforcement learning without great exploration[J]. Advances in neural information processing systems, 2020, 33: 1264-1274. \
[4] Uehara M, Sun W. Pessimistic model-based offline reinforcement learning under partial coverage[J]. arXiv preprint arXiv:2107.06226, 2021. \
[5] Hornik K, Stinchcombe M, White H. Multilayer feedforward networks are universal approximators[J]. Neural networks, 1989, 2(5): 359-366.\
[6] Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J.Y., Levine, S., Finn, C. and Ma, T., 2020. Mopo: Model-based offline policy optimization. Advances in Neural Information Processing Systems, 33, pp.14129-14142.\
[7] Schmidt-Hieber, J. (2020). Nonparametric regression using deep neural networks with ReLU activation function. The Annals of Statistics, 48(4).\
[8] Farrell, M. H., Liang, T., & Misra, S. (2021). Deep neural networks for estimation and inference. Econometrica, 89(1), 181-213.\
[9] Imaizumi, M., & Fukumizu, K. (2019, April). Deep neural networks learn non-smooth functions effectively. In The 22nd international conference on artificial intelligence and statistics (pp. 869-878). PMLR.\
[10] Zhou, Y., Shi, C., Li, L., & Yao, Q. (2023). Testing for the Markov property in time series via deep conditional generative learning. Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(4), 1204-1222.\
[11] Rashidinejad P, Zhu B, Ma C, et al. Bridging offline reinforcement learning and imitation learning: A tale of pessimism[J]. Advances in Neural Information Processing Systems, 2021, 34: 11702-11716.\
[12] Yin M, Wang Y X. Towards instance-optimal offline reinforcement learning with pessimism[J]. Advances in neural information processing systems, 2021, 34: 4065-4078.\
[13] Cui Q, Du S S. Provably efficient offline multi-agent reinforcement learning via strategy-wise bonus[J]. Advances in Neural Information Processing Systems, 2022, 35: 11739-11751.\
[14] Wang X, Cui Q, Du S S. On gap-dependent bounds for offline reinforcement learning[J]. Advances in Neural Information Processing Systems, 2022, 35: 14865-14877.\
[15] Bian, Z., Shi, C., Qi, Z., & Wang, L. (2023). Off-policy evaluation in doubly inhomogeneous environments. arXiv preprint arXiv:2306.08719.