OPTIMAL POLICY DETERMINATION USING REPEATED STACKELBERG GAMES WITH UNKNOWN PLAYER PREFERENCES
Patent №
US 8,545,332
Granted
2013-10-01
Filed 2012
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
5
ml · kr · planning · evo · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
13364843
A system, method and computer program product for planning actions in a repeated Stackelberg Game, played for a fixed number of rounds, where the payoffs or preferences of the follower are initially unknown to the leader, and a prior probability distribution over follower types is available. In repeated Bayesian Stackelberg games, the objective is to maximize the leader's cumulative expected payoff over the rounds of the game. The optimal plans in such games make intelligent tradeoffs between actions that reveal information regarding the unknown follower preferences, and actions that aim for high immediate payoff. The method solves for such optimal plans according to a Monte Carlo Tree Search method wherein simulation trials draw instances of followers from said prior probability distribution. Some embodiments additionally implement a method for pruning dominated leader strategies.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 276430674
Assignors
MARECKI, JANUSZ, TESAURO, GERALD J., SEGAL, RICHARD B.
On an employer assignment, the assignors are typically the inventors.