Patent №
US 9,536,191
Granted
2017-01-03
Filed 2015
Owner
OSARO, INC.
Lab
—
AI components
5
ml · vision · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
14952540
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for reinforcement learning using confidence scores. One of the methods includes receiving a current observation; for each of multiple actions: determining a respective value function estimate that is an estimate of a return resulting from the agent performing the action in response to the current observation, determining a respective confidence score that is a measure of confidence that the respective value function estimate for the action is an accurate estimate of the return that will result from the agent performing the action in response to the current observation, adjusting the respective value function estimate for the action using the respective confidence score for the action to determine a respective adjusted value function estimate; and selecting an action to be performed by the agent in response to the current observation using the respective adjusted value function estimates.
AI classification
Ownership
OSARO, INC.
assignment · 373220228
Assignors
AREL, ITAMAR, KAHANE, MICHAEL, ROHANIMANESH, KHASHAYAR
On an employer assignment, the assignors are typically the inventors.