Patent №
US 8,326,780
Granted
2012-12-04
Filed 2009
Owner
HONDA MOTOR CO., LTD.
Lab
—
AI components
4
ml · planning · evo · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
12578574
The present invention provides a method for learning a policy used by a computing system to perform a task, such delivery of one or more objects by the computing system. During a first time interval, the computing system determines a first state, a first action and a first reward value. As the computing system determines different states, actions and reward values during subsequent time intervals, a state description identifying the current sate, the current action, the current reward and a predicted action is stored. Responsive to a variance of a stored state description falling below a threshold value, the stored state description is used to modify one or more weights in the policy associated with the first state.
AI classification
Ownership
HONDA MOTOR CO., LTD.
assignment · 239130024
Assignors
GUPTA, RAKESH, RAMACHANDRAN, DEEPAK
On an employer assignment, the assignors are typically the inventors.