Unmanned Aerial Vehicles (UAVs) with Microwave Power Transfer (MPT)\ncapability provide a practical means to deploy a large number of wireless\npowered sensing devices into areas with no access to persistent power supplies.\nThe UAV can charge the sensing devices remotely and harvest their data. A key\nchallenge is online MPT and data collection in the presence of on-board control\nof a UAV (e.g., patrolling velocity) for preventing battery drainage and data\nqueue overflow of the sensing devices, while up-to-date knowledge on battery\nlevel and data queue of the devices is not available at the UAV. In this paper,\nan on-board deep Q-network is developed to minimize the overall data packet\nloss of the sensing devices, by optimally deciding the device to be charged and\ninterrogated for data collection, and the instantaneous patrolling velocity of\nthe UAV. Specifically, we formulate a Markov Decision Process (MDP) with the\nstates of battery level and data queue length of sensing devices, channel\nconditions, and waypoints given the trajectory of the UAV; and solve it\noptimally with Q-learning. Furthermore, we propose the on-board deep Q-network\nthat can enlarge the state space of the MDP, and a deep reinforcement learning\nbased scheduling algorithm that asymptotically derives the optimal solution\nonline, even when the UAV has only outdated knowledge on the MDP states.\nNumerical results demonstrate that the proposed deep reinforcement learning\nalgorithm reduces the packet loss by at least 69.2%, as compared to existing\nnon-learning greedy algorithms.\n