Contextual Bandit Algorithms with Supervised Learning Guarantees

We address the problem of competing with any large set of N policies in the nonstochastic bandit setting, where the learner must repeatedly select among K actions but observes only the reward of the chosen action.

Paper

Similar papers

© 2026 NYSGPT2525 LLC