A hardware-aware continual reinforcement learning framework designed for edge devices has been presented in this research. The framework couples an actor critic learner with a hybrid SRAM-RRAM memory system and a selective consolidation policy, preserving task rewards while making hardware costs explicit. The architecture couples a continual RL agent with an RRAM memory hierarchy to explicitly demonstrate the plasticity-stability-efficiency tradeoffs inherent in edge computing. The consolidation mechanism is formalized through a lightweight algorithm that buffers rapid updates in SRAM and selectively transfers important parameters to RRAM for long-term durable nonvolatile storage. Experiments on CartPole and Pendulum show baselinelevel performance, but RRAM energy is dominated by orders-of-magnitude, with respect to the energy expenditure and the endurance of the system. The proposed framework is evaluated jointly on learning performance and hardware-level costs, enabling a balanced view of algorithmic success versus system sustainability.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex