SYSTEMS AND METHODS FOR LEARNING REUSABLE OPTIONS TO TRANSFER KNOWLEDGE BETWEEN TASKS
Patent №
US 11,511,413
Granted
2022-11-29
Filed 2020
Owner
HUAWEI TECHNOLOIES CO., LTD.
Lab
—
AI components
5
ml · vision · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16900291
A robot that includes an RL agent that is configured to learn a policy to maximize the cumulative reward of a task, to determine one or more features that are minimally correlated with each other. The features are then used as pseudo-rewards, called feature rewards, where each feature reward corresponds to an option policy, or skill, the RL agent learns to maximize. In an example, the RL agent is configured to select the most relevant features to learn respective option policies from. The RL agent is configured to, for each of the selected features, learn the respective option policy that maximizes the respective feature reward. Using the learned option policies, the RL agent is configured to learn a new (second) policy for a new (second) task that can choose from any of the learned option policies or actions available to the RL agent.
AI classification
Ownership
HUAWEI TECHNOLOIES CO., LTD.
assignment · 529320615
Assignors
MAVRIN, BORISLAV, GRAVES, DANIEL MARK
On an employer assignment, the assignors are typically the inventors.