To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging
Leveraging large amounts of unlabeled data using Transformer-like\narchitectures, like BERT, has gained popularity in recent times owing to their\neffectiveness in learning general representations that can then be further\nfine-tuned for downstream tasks to much success. However, training these models\ncan be costly both from an economic and environmental standpoint. In this work,\nwe investigate how to effectively use unlabeled data: by exploring the\ntask-specific semi-supervised approach, Cross-View Training (CVT) and comparing\nit with task-agnostic BERT in multiple settings that include domain and task\nrelevant English data. CVT uses a much lighter model architecture and we show\nthat it achieves similar performance to BERT on a set of sequence tagging\ntasks, with lesser financial and environmental impact.\n
Paper
References (38)
Scroll for more · 26 remaining