SECURE EXPLORATION FOR REINFORCEMENT LEARNING

Patent №

US 11,616,813

Granted

2023-03-28

Filed 2019

Owner

MICROSOFT TECHNOLOGY LICENSING, LLC

AI components

3

ml · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16554525

A secured exploration agent for reinforcement learning (RL) is provided. Securitizing an exploration agent includes training the exploration agent to avoid dead-end states and dead-end trajectories. During training, the exploration agent “learns” to identify and avoid dead-end states of a Markov Decision Process (MDP). The secured exploration agent is utilized to safely and efficiently explore the environment, while significantly reducing the training time, as well as the cost and safety concerns associated with conventional RL. The secured exploration agent is employed to guide the behavior of a corresponding exploitation agent. During training, a policy of the exploration agent is iteratively updated to reflect an estimated probability that a state is a dead-end state. The probability, via the exploration policy, that the exploration agent chooses an action that results in a transition to a dead-end state is reduced to reflect the estimated probability that the state is a dead-end state.

Machine learningPlanningAI hardwareG06N 3/08H04L 63/20G06N 3/006G06N 3/092G06N 5/043G06N 7/01G06N 20/00

AI classification

Machine learning1.00
Planning1.00
AI hardware0.96
Knowledge representation0.33
Vision0.18
Evolutionary computation0.11
Speech0.01
Natural language0.00

Ownership

MICROSOFT TECHNOLOGY LICENSING, LLC

assignment · 507310362

Assignors

FATEMI BOOSHEHRI, SEYED MEHDI, VAN SEIJEN, HARM HENDRICK

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC