Patent №
US 10,860,920
Granted
2020-12-08
Filed 2019
Owner
DEEPMIND TECHNOLOGIES LIMITED
AI components
6
ml · vision · kr · planning · evo · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16508046
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting an action to be performed by a reinforcement learning agent interacting with an environment. A current observation characterizing a current state of the environment is received. For each action in a set of multiple actions that can be performed by the agent to interact with the environment, a probability distribution is determined over possible Q returns for the action-current observation pair. For each action, a measure of central tendency of the possible Q returns with respect to the probability distributions for the action-current observation pair is determined. An action to be performed by the agent in response to the current observation is selected using the measures of central tendency.
AI classification
Ownership
DEEPMIND TECHNOLOGIES LIMITED
assignment · 497940132
Assignors
GENDRON-BELLEMARE, MARC, DABNEY, WILLIAM CLINTON
On an employer assignment, the assignors are typically the inventors.