Decentralized Heterogeneous Multi-Player Multi-Armed Bandits with Non-Zero Rewards on Collisions
We consider a fully decentralized multi-player stochastic multi-armed bandit setting where the players cannot communicate with each other and can observe only their own actions and rewards. The environment may appear differently to different players, <inline-formula> <tex-math notation="LaTeX">$\textit {i.e.}$ </tex-math></inline-formula>, the reward distributions for a given arm are heterogeneous across players. In the case of a collision (when more than one player plays the same arm), we allow for the colliding players to receive non-zero rewards. The time-horizon <inline-formula> <tex-math notation="LaTeX">$T$ </tex-math></inline-formula> for which the arms are played is <italic>not</italic> known to the players. Within this setup, where the number of players is allowed to be greater than the number of arms, we present a policy that achieves near order-optimal expected regret of order <inline-formula> <tex-math notation="LaTeX">$O(\log ^{1 + \delta } T)$ </tex-math></inline-formula> for <inline-formula> <tex-math notation="LaTeX">$\delta >0$ </tex-math></inline-formula> (however small) over a time-horizon of duration <inline-formula> <tex-math notation="LaTeX">$T$ </tex-math></inline-formula>.