The Performance of Actor–Critic Model Based on the Probability of Gaussian Distribution

In practical implementations of the actor–critic framework, selecting optimal actions from a finite discrete action set has revealed inherent limitations, particularly in the challenge of predefining an appropriate action set tailored to specific tasks. To address this issue, this paper proposes a novel framework, namely Temporal Smoothing Exploration based PPO (TSE‐PPO) that enables agents to autonomously discover suitable actions within the actor–critic paradigm. Instead of relying on a fixed discrete action set, TSE‐PPO learns actions from a continuous action space, which inherently represents a diverse range of possibilities. Action selection is governed by probability densities derived from a Gaussian distribution conditioned on the agent's state, thereby promoting adaptive and context‐sensitive behavior. The effectiveness of the proposed algorithm is evaluated using the OpenAI Gym simulation environment, which provides a comprehensive continuous control setting. Experimental results indicate that the TSE‐PPO framework significantly improves the performance of the algorithm by facilitating more precise and diverse action selection, thereby offering a promising direction for enhancing agent learning in complex environments. © 2026 Institute of Electrical Engineers of Japan and Wiley Periodicals LLC.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC