We thank the reviewer for spending time reading our paper, and for all the comments.
We understand that this might be a last-minute/urgent review assignment for the reviewer, so we do appreciate all the comments and suggested references (we will include these in the future version of our paper). However, we believe the evaluation/criticism of a work on machine learning/deep learning should not be solely based on the experiments, but also on a large scope covering the modeling and theoretical contributions. To this end, we have created a TL;DR version of our positioning and contribution as in the general response. And we would really appreciate it if the reviewer could spend a bit more of time on reading these corresponding parts and evaluating these aspects of our paper, and in the following week, we are happy to address any follow-up comments/questions.
In the following, we address the raised questions by the reviewer:
"Does this method handle higher-dimensional environments or continuous action spaces?"
- First, we note that we are presenting a model, not only a method. The model aims to provide a theoretical foundation so that it will be compatible with different extensions. To incorporate higher-dimensional environments, our model is compatible with the existing theoretical analyses of learning methods involving linear function approximation of the state space which provide a solution for high dimension environments. For continuous action space, we have to develop a continuous model separately. Just like the case in generic RL, the related theoretical property for continuous action is similar to the discrete one but it is out of the current scope of this work.
"How does this method handle more complex adherence functions?"
- We first note that the proposed model can be naturally related to other theoretical extensions to solve the aforementioned problem. If the adherence is nonstationary, the problem can be solved by learning the adherence level at every time step, which is a standard extension for RL algorithms. If the adherence level changes from episode to episode, we can rely on a standard approach by introducing a variational budget (Besbes, et. al. 2014) for the adherence level over episodes and get the corresponding theoretical result. If the variation of human adherence level exceeds the variational budget, the corresponding decision-making problem remains an open problem for the general theoretical RL/online learning community.
"How does this method handle co-adapting advice takers?"
- As we understand, co-adapting means that through interaction, the human will adapt to the machine and change its behavior policy or adherence level. This corresponds to a learning problem where the underlying MDP is changing, and if we are not controlling the amount of change in the behavior policy/adherence level, there is no model/algorithm that can provide theoretical guarantees for the co-adapting setting. We understand that some data-driven methods without theoretical guarantees can be very effective and practical for this scenario. We agree with the reviewer that this is an important future direction, and we will discuss this point more in the future version of our paper.
Models and high-dimensional extensions: We appreciate the suggestions on extending the result to higher dimensions. We will mention this as a potential direction for future work.
Empirical Experiment vs Numerical Experiment: We agree that experiments in a more practical setting are important. However, these experiments sometimes cannot be explained by any model or theory. In this work, if we conduct empirical experiments as suggested, it will be detached from our model because these scenarios are not related to our model at all. Instead, simpler experiments can either corroborate the underlying theory or be explained by it, making our work more sound. Furthermore, to ensure a more concise presentation of our paper, it is sensible to conduct experiments that are more closely related to our theoretical framework.
We thank the reviewer again for the time spent reading our paper. We look forward to follow-up discussions.