Given an environment with continuous state spaces and discrete actions, we investigate using a Double Deep Q-learning Reinforcement Agent to find optimal policies using the LunarLander-v2 OpenAI gym environment.
Paper
References (4)
04Hado van Hasselt -Double Q-learning2010 · Advances in Neural Information Processing Systems