Deep Reinforcement Learning for Efficient Scheduling of Ground-based Astronomical Observations
Ground-based astronomical observations face inherent challenges from weather changes, target visibility window constraints, and observational requirements. Enhancing the efficiency and effectiveness of telescope operations has long been a key objective for many observatories because of the high cost of observational resources. In this study, we formalize observation scheduling as a time-dependent combinatorial optimization problem. To achieve this, we implement a pointer network with temporal attention that is capable of planning observations while accounting for time-varying factors such as moonlight interference, target altitude, and air mass, which impact the exposure time and image quality. To support the training of the deep neural network, we propose a scoring mechanism to evaluate the effectiveness of the observations, which is optimized through a refined REINFORCE algorithm with a baseline. Furthermore, an exposure time calculator and an equipment kinematic model are incorporated to dynamically estimate the time costs during the decision-making process. The simulation results demonstrated that the trained model significantly outperformed both manual scheduling and a greedy algorithm in terms of theoretical reward scores and the total number of scheduled targets. Observation experiments conducted using a dual-telescope system at Muztaga observatory further validated the superiority of our approach, demonstrating a 45.8% enhancement in the total signal-to-noise ratio across all observed targets and a 24.1% increase in the number of completed tasks under the same observing conditions.
Paper
Full text
Deep Reinforcement Learning for Efficient Scheduling of Ground-based Astronomical Observations
Semantic Scholar · Physics · 2025
Abstract
Ground-based astronomical observations face inherent challenges from weather changes, target visibility window constraints, and observational requirements. Enhancing the efficiency and effectiveness of telescope operations has long been a key objective for many observatories because of the high cost of observational resources. In this study, we formalize observation scheduling as a time-dependent combinatorial optimization problem. To achieve this, we implement a pointer network with temporal attention that is capable of planning observations while accounting for time-varying factors such as moonlight interference, target altitude, and air mass, which impact the exposure time and image quality. To support the training of the deep neural network, we propose a scoring mechanism to evaluate the effectiveness of the observations, which is optimized through a refined REINFORCE algorithm with a baseline. Furthermore, an exposure time calculator and an equipment kinematic model are incorporated to dynamically estimate the time costs during the decision-making process. The simulation results demonstrated that the trained model significantly outperformed both manual scheduling and a greedy algorithm in terms of theoretical reward scores and the total number of scheduled targets. Observation experiments conducted using a dual-telescope system at Muztaga observatory further validated the superiority of our approach, demonstrating a 45.8% enhancement in the total signal-to-noise ratio across all observed targets and a 24.1% increase in the number of completed tasks under the same observing conditions.