Adaptive Scheduling: A Reinforcement Learning Whittle Index Approach for Wireless Sensor Networks
We propose a Reinforcement Learning (RL)-based scheduling framework for Restless Multi-Armed Bandit (RMAB) problems, centred on a Whittle Index Q-Learning policy with Upper Confidence Bound (Whittle index Q-Learning (WIQL)-upper confidence bound (UCB)) exploration. Unlike existing approaches that rely on fixed or adaptive $\epsilon $ -greedy strategies and require careful hyperparameter tuning, the proposed method eliminates problem-specific tuning and is therefore more generalisable across diverse RMAB settings. We evaluate WIQL-UCB on standard RMAB benchmarks and on a practical sensor scheduling application based on the Age of Incorrect Information (AoII), using an edge-based estimation scheme that requires no prior knowledge of system dynamics. Experimental results show that WIQL-UCB achieves near-optimal reward performance and strong generalisation across different RMAB settings, consistently outperforming both non–Whittle-based reinforcement learning policies and existing Whittle-index learning baselines. In terms of efficiency, the proposed method exhibits memory and computational requirements comparable to other Whittle-index-based approaches, while remaining orders of magnitude more efficient than non–Whittle-based deep RL methods. At a representative problem size ( $N=15, M=3$ ), WIQL-UCB requires approximately 600 bytes of memory, compared to several kilobytes for tabular Q-learning and hundreds of kilobytes to megabytes for deep RL baselines, and achieves sub-millisecond per-decision runtimes. These results demonstrate that WIQL-UCB consistently outperforms both non–Whittle-based and Whittle-index learning baselines across diverse RMAB settings.