Large-scale pretrained language models are the major driving force behind\nrecent improvements in performance on the Winograd Schema Challenge, a widely\nemployed test of common sense reasoning ability. We show, however, with a new\ndiagnostic dataset, that these models are sensitive to linguistic perturbations\nof the Winograd examples that minimally affect human understanding. Our results\nhighlight interesting differences between humans and language models: language\nmodels are more sensitive to number or gender alternations and synonym\nreplacements than humans, and humans are more stable and consistent in their\npredictions, maintain a much higher absolute performance, and perform better on\nnon-associative instances than associative ones. Overall, humans are correct\nmore often than out-of-the-box models, and the models are sometimes right for\nthe wrong reasons. Finally, we show that fine-tuning on a large, task-specific\ndataset can offer a solution to these issues.\n
Paper
References (51)
Scroll for more · 38 remaining