CONTROLLING AGENTS OVER LONG TIME SCALES USING TEMPORAL VALUE TRANSPORT

Patent №

US 11,769,049

Granted

2023-09-26

Filed 2020

Owner

DEEPMIND TECHNOLOGIES LIMITED

AI components

5

ml · kr · planning · evo · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17035546

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network system used to control an agent interacting with an environment to perform a specified task. One of the methods includes causing the agent to perform a task episode in which the agent attempts to perform the specified task; for each of one or more particular time steps in the sequence: generating a modified reward for the particular time step from (i) the actual reward at the time step and (ii) value predictions at one or more time steps that are more than a threshold number of time steps after the particular time step in the sequence; and training, through reinforcement learning, the neural network system using at least the modified rewards for the particular time steps.

Machine learningKnowledge representationPlanningEvolutionary computationAI hardwareG06N 3/08G06F 11/3037G06F 11/3072G06F 18/2193G06F 18/2413G06N 3/006G06N 3/044G06N 3/0442+9 more

AI classification

Machine learning1.00
AI hardware1.00
Knowledge representation1.00
Evolutionary computation1.00
Planning0.97
Vision0.43
Natural language0.02
Speech0.00

Ownership

DEEPMIND TECHNOLOGIES LIMITED

assignment · 549940357

Assignors

WAYNE, GREGORY DUNCAN, LILLICRAP, TIMOTHY PAUL, HUNG, CHIA-CHUN, ABRAMSON, JOSHUA SIMON

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC