Summary
This paper studies the infinite-horizon restless bandits problem and proposed a simulation-based framework, i.e., Follow-the-Virtual-Advice, to leverage single armed policy to solve a multi-armed problem, which gets rid of the difficult-to-verify condition, i.e., the uniform global attractor property. Removing this pre-condition is a significant improvement for RMABs.
Strengths
This paper studies the infinite-horizon restless bandits problem and proposed a simulation-based framework, i.e., Follow-the-Virtual-Advice, to leverage single armed policy to solve a multi-armed problem, which gets rid of the difficult-to-verify condition, i.e., the uniform global attractor property. Removing this pre-condition is a significant improvement for RMABs.
Weaknesses
1. The arguments on FTVA in Section 3.3 look to be a bit hands-waving since there is no clear evidence to support their claims. Based on the reviewer’s understandings, those arguments are possibly wrong. For example, in lines 181-184, the authors claim that even if the initial virtual state and real state of an arm are different, they will become identical in finite time by chance under mild assumptions in Section 4.1. However, Assumption 1 relies on the “observation” in lines 233-238, which is not true. All the transitions are stochastic, not deterministic. How could we guarantee that $S(t+1) =\hat{S}(t+1)$ if $S(t)$ is different with $\hat{S}(t)$? If Assumption 1 fails, the argument in lines 183-184 does not hold. The arguments in lines 185-189 are also not true. Even when the virtual process can always satisfy the budget constraint, how do we guarantee that the real process following the virtual actions does not violate the budget constraints.
2. The example in lines 194-215 is also not fully supported by evidence. The single-agent policy defined in eq. (8) is stationary stochastic policy, and agent selects actions according to certain probabilities. What do you mean by the preferred action at each state? Though this particular example may only have one action at each state, there is no preferred action in general. How to guarantee the descriptions in lines 207-209 to be true? This example shows that the real and virtual process has different state distributions, which contradicts Section 3.3. See comment 1.
3. A very important related work is missing. [Ghosh22] also gets rid of the global attractor assumption and considers a much more challenging setting with heterogeneous arms and multi-actions. An outstanding limitation of this work is that all arms must be the same.
Ghosh, A., Nagaraj, D., Jain, M., & Tambe, M. (2022). Indexability is Not Enough for Whittle: Improved, Near-Optimal Algorithms for Restless Bandits. arXiv preprint arXiv:2211.00112.
4. The references cited in this paper are not precise. For example, as far as the reviewer knows, [Ver16] does not consider the single-armed problem. Hence, the argument in lines 140-141 are not precise. Another example is in lines 123-124. The reviewer is not aware of the meaning of these sentences. How are they related to the RMAB problem in eqs. (1)-(2)?
Questions
See comments in Weaknesses.
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
no negative societal impact.