Patent №
US 11,627,165
Granted
2023-04-11
Filed 2020
Owner
DEEPMIND TECHNOLOGIES LIMITED
AI components
4
ml · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16752496
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a policy neural network having a plurality of policy parameters and used to select actions to be performed by an agent to control the agent to perform a particular task while interacting with one or more other agents in an environment. In one aspect, the method includes: maintaining data specifying a pool of candidate action selection policies; maintaining data specifying respective matchmaking policy; and training the policy neural network using a reinforcement learning technique to update the policy parameters. The policy parameters define policies to be used in controlling the agent to perform the particular task.
AI classification
Ownership
DEEPMIND TECHNOLOGIES LIMITED
assignment · 517700438
Assignors
SILVER, DAVID, VINYALS, ORIOL, JADERBERG, MAXWELL ELLIOT
On an employer assignment, the assignors are typically the inventors.