Recently, several studies have investigated active learning (AL) for natural\nlanguage processing tasks to alleviate data dependency. However, for query\nselection, most of these studies mainly rely on uncertainty-based sampling,\nwhich generally does not exploit the structural information of the unlabeled\ndata. This leads to a sampling bias in the batch active learning setting, which\nselects several samples at once. In this work, we demonstrate that the amount\nof labeled training data can be reduced using active learning when it\nincorporates both uncertainty and diversity in the sequence labeling task. We\nexamined the effects of our sequence-based approach by selecting weighted\ndiverse in the gradient embedding approach across multiple tasks, datasets,\nmodels, and consistently outperform classic uncertainty-based sampling and\ndiversity-based sampling.\n