We propose a simple architecture for deep reinforcement learning by embedding\ninputs into a learned Fourier basis and show that it improves the sample\nefficiency of both state-based and image-based RL. We perform infinite-width\nanalysis of our architecture using the Neural Tangent Kernel and theoretically\nshow that tuning the initial variance of the Fourier basis is equivalent to\nfunctional regularization of the learned deep network. That is, these learned\nFourier features allow for adjusting the degree to which networks underfit or\noverfit different frequencies in the training data, and hence provide a\ncontrolled mechanism to improve the stability and performance of RL\noptimization. Empirically, this allows us to prioritize learning low-frequency\nfunctions and speed up learning by reducing networks' susceptibility to noise\nin the optimization process, such as during Bellman updates. Experiments on\nstandard state-based and image-based RL benchmarks show clear benefits of our\narchitecture over the baselines. Website at\nhttps://alexanderli.com/learned-fourier-features\n
Paper
References (52)
Scroll for more · 38 remaining