Learning Structured Predictors from Bandit Feedback for Interactive NLP

Structured prediction from bandit feed-back describes a learning scenario where instead of having access to a gold standard structure, a learner only receives partial feedback in form of the loss value of a predicted structure. We present new learning objectives and algorithms for this interactive scenario, focusing on convergence speed and ease of elicitability of feed-back. We present supervised-to-bandit simulation experiments for several NLP tasks (machine translation, sequence labeling, text classification), showing that bandit learning from relative preferences eases feedback strength and yields improved empirical convergence.

Paper

Full text

PDF

Learning Structured Predictors from Bandit Feedback for Interactive NLP

Semantic Scholar · Computer Science · 2016

Abstract

Structured prediction from bandit feed-back describes a learning scenario where instead of having access to a gold standard structure, a learner only receives partial feedback in form of the loss value of a predicted structure. We present new learning objectives and algorithms for this interactive scenario, focusing on convergence speed and ease of elicitability of feed-back. We present supervised-to-bandit simulation experiments for several NLP tasks (machine translation, sequence labeling, text classification), showing that bandit learning from relative preferences eases feedback strength and yields improved empirical convergence.

References (63)

07Online adaptation to post-edits for phrase-based statistical machine translation2014 · Machine Translation

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC