To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging

Leveraging large amounts of unlabeled data using Transformer-like\narchitectures, like BERT, has gained popularity in recent times owing to their\neffectiveness in learning general representations that can then be further\nfine-tuned for downstream tasks to much success. However, training these models\ncan be costly both from an economic and environmental standpoint. In this work,\nwe investigate how to effectively use unlabeled data: by exploring the\ntask-specific semi-supervised approach, Cross-View Training (CVT) and comparing\nit with task-agnostic BERT in multiple settings that include domain and task\nrelevant English data. CVT uses a much lighter model architecture and we show\nthat it achieves similar performance to BERT on a set of sequence tagging\ntasks, with lesser financial and environmental impact.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC