In practical implementations of the actor–critic framework, selecting optimal actions from a finite discrete action set has revealed inherent limitations, particularly in the challenge of predefining an appropriate action set tailored to specific tasks. To address this issue, this paper proposes a novel framework, namely Temporal Smoothing Exploration based PPO (TSE‐PPO) that enables agents to autonomously discover suitable actions within the actor–critic paradigm. Instead of relying on a fixed discrete action set, TSE‐PPO learns actions from a continuous action space, which inherently represents a diverse range of possibilities. Action selection is governed by probability densities derived from a Gaussian distribution conditioned on the agent's state, thereby promoting adaptive and context‐sensitive behavior. The effectiveness of the proposed algorithm is evaluated using the OpenAI Gym simulation environment, which provides a comprehensive continuous control setting. Experimental results indicate that the TSE‐PPO framework significantly improves the performance of the algorithm by facilitating more precise and diverse action selection, thereby offering a promising direction for enhancing agent learning in complex environments. © 2026 Institute of Electrical Engineers of Japan and Wiley Periodicals LLC.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex