On-Policy and Off-Policy Learning for Large Action Spaces

This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected…

Paper

Similar papers

© 2026 NYSGPT2525 LLC