Non-Markovian quantum dynamics, characterized by information backflow from the environment to the system, has emerged as a potential resource for quantum technologies. A key challenge is therefore to control and enhance such memory effects. In this work, we investigate the use of reinforcement learning (RL) to maximize non-Markovianity in a driven two-level system coupled to a structured reservoir. We compare RL-based control strategies with standard optimal control theory (OCT). We show that OCT produces localized but relatively weak revivalsin the instantaneous non-Markovianity rate, whereas RL policies generate significantly stronger and better-timed information backflow by synchronizing the system dynamics with favorable memory intervals of the environment. This enhanced exploitation of memory effects leads to a higher total integrated non-Markovianity for RL than for OCT, with SAC achieving the largest overall enhancement and PPO delivering slightly lower but still strongly improved performance with smoother, experimentally attractive pulses. Our results contribute to the emerging view of non-Markovianity as an operational resource and illustrate how RL can serve as a flexible, model-free tool for non-Markovian quantum control.
Paper
References (56)
Scroll for more · 38 remaining