Reinforcement learning agents that act under partial observability lack access to the environment state, and may need to account for the observable history to select actions optimally. However, offline training paradigms (e.g., training via simulator) are able to exploit state information during the training phase to improve learning performance. The literature contains a number of such methods that exploit state information during training, and empirically demonstrate superior performance during evaluation, e.g., asymmetric actor-critic methods that use state critics. However, such methods tend to be poorly motivated and lack a theoretically sound justification. In this work, we focus on the theoretical and practical consequences of using states to train partially observable agents, and propose interpretations to explain the role of state.
Paper
Full text
Role of State in Partially Observable Reinforcement Learning
Semantic Scholar · Computer Science · 2025
Abstract
Reinforcement learning agents that act under partial observability lack access to the environment state, and may need to account for the observable history to select actions optimally. However, offline training paradigms (e.g., training via simulator) are able to exploit state information during the training phase to improve learning performance. The literature contains a number of such methods that exploit state information during training, and empirically demonstrate superior performance during evaluation, e.g., asymmetric actor-critic methods that use state critics. However, such methods tend to be poorly motivated and lack a theoretically sound justification. In this work, we focus on the theoretical and practical consequences of using states to train partially observable agents, and propose interpretations to explain the role of state.