REINFORCEMENT LEARNING USING ADVANTAGE ESTIMATES

Patent №

US 11,288,568

Granted

2022-03-29

Filed 2017

Owner

GOOGLE INC.

AI components

6

ml · vision · kr · planning · evo · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

15429088

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for computing Q values for actions to be performed by an agent interacting with an environment from a continuous action space of actions. In one aspect, a system includes a value subnetwork configured to receive an observation characterizing a current state of the environment and process the observation to generate a value estimate; a policy subnetwork configured to receive the observation and process the observation to generate an ideal point in the continuous action space; and a subsystem configured to receive a particular point in the continuous action space representing a particular action; generate an advantage estimate for the particular action; and generate a Q value for the particular action that is an estimate of an expected return resulting from the agent performing the particular action when the environment is in the current state.

AI classification

Planning1.00
Machine learning1.00
Knowledge representation1.00
AI hardware1.00
Evolutionary computation0.98
Vision0.98
Natural language0.05
Speech0.03

Ownership

GOOGLE INC.

assignment · 416970938

Assignors

GU, SHIXIANG, LILLICRAP, TIMOTHY PAUL, SUTSKEVER, ILYA, LEVINE, SERGEY VLADIMIR

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC