Continual Deep Reinforcement Learning for Decentralized Satellite Routing

This paper introduces a full solution for decentralized routing in Low Earth Orbit Satellite Constellations (LSatCs) based on continual Deep Reinforcement Learning (DRL), specifically designed for on-board implementation in satellites with limited computational and communication resources. This requires addressing multiple challenges, including the partial knowledge at the satellites and their continuous movement, and the time-varying sources of uncertainty in the system, such as traffic, communication links, or communication buffers. We follow a multi-agent approach, where each satellite acts as an independent decision-making agent, while acquiring a limited knowledge of the environment based on the feedback received from the nearby agents. The solution is divided into two phases. First, an offline learning phase relies on decentralized decisions and a global Deep Neural Network (DNN) trained with global experiences to learn the optimal paths. Then, the online phase with local, on-board, and pre-trained DNNs requires continual learning to evolve with the environment, which can be done in two different ways: 1) Model anticipation, where the predictable conditions of the constellation, resulting from its orbital dynamics, are exploited by each satellite sharing local model with the next satellite; and 2) Federated Learning (FL), where each agent’s model is merged first at the cluster level and then aggregated in a global Parameter Server (PS) at ground or at a geostationary orbit (GEO) satellite. Results from simulations with State-of-the-Art (SoA) constellations such show that the proposed approach converges in less than a second to similar end-to-end (E2E) latency than the shortest-path centralized approach with full knowledge of the network. Moreover, the Centered Kernel Alignment (CKA) metric quantifies the necessary alignment of the models when the dynamics of the environment change.

Paper

References (56)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC