METHOD AND APPARATUS FOR IMPROVED REWARD-BASED LEARNING USING ADAPTIVE DISTANCE METRICS

Patent №

US 9,298,172

Granted

2016-03-29

Filed 2007

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

4

ml · planning · evo · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11870661

The present invention is a method and an apparatus for reward-based learning of policies for managing or controlling a system or plant. In one embodiment, a method for reward-based learning includes receiving a set of one or more exemplars, where at least two of the exemplars comprise a (state, action) pair for a system, and at least one of the exemplars includes an immediate reward responsive to a (state, action) pair. A distance metric and a distance-based function approximator estimating long-range expected value are then initialized, where the distance metric computes a distance between two (state, action) pairs, and the distance metric and function approximator are adjusted such that a Bellman error measure of the function approximator on the set of exemplars is minimized. A management policy is then derived based on the trained distance metric and function approximator.

Machine learningPlanningEvolutionary computationAI hardwareG05B 13/0265G06F 18/24G06F 18/24147G06N 5/02G06N 20/00

AI classification

Machine learning1.00
Planning1.00
AI hardware1.00
Evolutionary computation0.54
Vision0.31
Knowledge representation0.04
Speech0.01
Natural language0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 200700892

Assignors

TESAURO, GERALD J., WEINBERGER, KILLIAN Q.

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC