We consider two agents playing simultaneously the same stochastic three-armed bandit problem. The two agents are cooperating but they cannot communicate. We propose a strategy with no collisions at all between the players (with very high probability), and with near-optimal regret $O(\sqrt{T \log(T)})$. We also argue that the extra logarithmic term $\sqrt{\log(T)}$ should be necessary by proving a lower bound for a full information variant of the problem.
Paper
References (13)
11For i ∈ {1, 2, 3} and θ ∈ I i , we have i ∈ a(θ) ∪ b(θ)
12By Item 1 above, we know that for every θ ∈ I 1 , the arm 1 must be in exactly one of the two sets a(θ) and b(θ)
Scroll for more · 1 remaining