METHOD AND APPARATUS FOR IMPROVED REWARD-BASED LEARNING USING ADAPTIVE DISTANCE METRICS
Patent №
US 9,298,172
Granted
2016-03-29
Filed 2007
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
4
ml · planning · evo · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
11870661
The present invention is a method and an apparatus for reward-based learning of policies for managing or controlling a system or plant. In one embodiment, a method for reward-based learning includes receiving a set of one or more exemplars, where at least two of the exemplars comprise a (state, action) pair for a system, and at least one of the exemplars includes an immediate reward responsive to a (state, action) pair. A distance metric and a distance-based function approximator estimating long-range expected value are then initialized, where the distance metric computes a distance between two (state, action) pairs, and the distance metric and function approximator are adjusted such that a Bellman error measure of the function approximator on the set of exemplars is minimized. A management policy is then derived based on the trained distance metric and function approximator.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 200700892
Assignors
TESAURO, GERALD J., WEINBERGER, KILLIAN Q.
On an employer assignment, the assignors are typically the inventors.