On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and Saliency
A well-calibrated neural model produces confidence (probability outputs)\nclosely approximated by the expected accuracy. While prior studies have shown\nthat mixup training as a data augmentation technique can improve model\ncalibration on image classification tasks, little is known about using mixup\nfor model calibration on natural language understanding (NLU) tasks. In this\npaper, we explore mixup for model calibration on several NLU tasks and propose\na novel mixup strategy for pre-trained language models that improves model\ncalibration further. Our proposed mixup is guided by both the Area Under the\nMargin (AUM) statistic (Pleiss et al., 2020) and the saliency map of each\nsample (Simonyan et al.,2013). Moreover, we combine our mixup strategy with\nmodel miscalibration correction techniques (i.e., label smoothing and\ntemperature scaling) and provide detailed analyses of their impact on our\nproposed mixup. We focus on systematically designing experiments on three NLU\ntasks: natural language inference, paraphrase detection, and commonsense\nreasoning. Our method achieves the lowest expected calibration error compared\nto strong baselines on both in-domain and out-of-domain test samples while\nmaintaining competitive accuracy.\n
Paper
References (36)
Scroll for more · 24 remaining