CLOSED LOOP MODEL-BASED ACTION LEARNING WITH MODEL-FREE INVERSE REINFORCEMENT LEARNING
Patent №
US 11,468,334
Granted
2022-10-11
Filed 2018
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
5
ml · kr · planning · evo · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16012229
A computer-implemented method is provided for learning an action policy. The method includes obtaining, by a processor, environment dynamics including triplets of a state, an action, and a next state. The state in each of the triplets is an expert state. The method further includes training, by the processor using the environment dynamics as training data, a dynamics model which obtains a pair of the state and the action as an input and outputs, for each next state, state-transition probabilities. The method also includes learning, by the processor, the action policy using trajectories of expert states according to a supervised learning technique by back-propagating error gradients through the trained dynamics model.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 461310277
Assignors
CHAUDHURY, SUBHAJIT, KIMURA, DAIKI, INOUE, TADANOBU, TACHIBANA, RYUKI
On an employer assignment, the assignors are typically the inventors.