Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning

State-of-the-art natural language understanding classification models follow\ntwo-stages: pre-training a large language model on an auxiliary task, and then\nfine-tuning the model on a task-specific labeled dataset using cross-entropy\nloss. However, the cross-entropy loss has several shortcomings that can lead to\nsub-optimal generalization and instability. Driven by the intuition that good\ngeneralization requires capturing the similarity between examples in one class\nand contrasting them with examples in other classes, we propose a supervised\ncontrastive learning (SCL) objective for the fine-tuning stage. Combined with\ncross-entropy, our proposed SCL loss obtains significant improvements over a\nstrong RoBERTa-Large baseline on multiple datasets of the GLUE benchmark in\nfew-shot learning settings, without requiring specialized architecture, data\naugmentations, memory banks, or additional unsupervised data. Our proposed\nfine-tuning objective leads to models that are more robust to different levels\nof noise in the fine-tuning training data, and can generalize better to related\ntasks with limited labeled data.\n

Paper

References (66)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC