SMOOTHED SARSA: REINFORCEMENT LEARNING FOR ROBOT DELIVERY TASKS

Patent №

US 8,326,780

Granted

2012-12-04

Filed 2009

Owner

HONDA MOTOR CO., LTD.

Lab

AI components

4

ml · planning · evo · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

12578574

The present invention provides a method for learning a policy used by a computing system to perform a task, such delivery of one or more objects by the computing system. During a first time interval, the computing system determines a first state, a first action and a first reward value. As the computing system determines different states, actions and reward values during subsequent time intervals, a state description identifying the current sate, the current action, the current reward and a predicted action is stored. Responsive to a variance of a stored state description falling below a threshold value, the stored state description is used to modify one or more weights in the policy associated with the first state.

AI classification

Machine learning1.00
AI hardware1.00
Evolutionary computation0.99
Planning0.96
Knowledge representation0.07
Vision0.01
Natural language0.00
Speech0.00

Ownership

HONDA MOTOR CO., LTD.

assignment · 239130024

Assignors

GUPTA, RAKESH, RAMACHANDRAN, DEEPAK

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC