Reinforcement Learning (RL) provides a standard framework for sequential decision-making, but state-of-the-art Deep RL (DRL) methods are often sample-inefficient and struggle to generalize beyond small-scale training scenarios. We propose a neuro-symbolic DRL approach that integrates background symbolic knowledge to improve sample efficiency and generalization to more complex, unseen tasks. Partial policies learned in simple domains are transferred as logical rules and used for online reasoning to guide learning by biasing exploration and rescaling Q-values during exploitation. This integration enhances interpretability and accelerates convergence, particularly in sparse-reward and long-horizon settings. Experiments show superior performance over state-of-the-art reward machine methods.
Paper
References (35)
Scroll for more · 23 remaining