LEXICOGRAPHIC DEEP REINFORCEMENT LEARNING USING STATE CONSTRAINTS AND CONDITIONAL POLICIES
Patent №
US 11,410,023
Granted
2022-08-09
Filed 2019
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
5
ml · vision · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16290413
A computer-implemented method is provided for modified Lexicographic Reinforcement Learning. The computer implemented method includes obtaining, by a hardware processor, a sequence of tasks. Each of the tasks corresponds to, and has a one-to-one correspondence with, a respective award from among set of rewards. The method further includes performing, by the hardware processor for each of the tasks, reinforcement learning and deep learning for both of (i) one or more policies and (ii) one or more value functions, with a plurality of sets of samples. A plurality of solutions in a form of the one or more policies and the one or more value functions are parametrized by a single neural network with a selector which selects an input of the single neural network from among the plurality of sets of samples.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 484830238
Assignors
AGRAVANTE, DON JOVEN R., MUNAWAR, ASIM, TACHIBANA, RYUKI
On an employer assignment, the assignors are typically the inventors.