A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English
Transformer-based language models achieve high performance on various tasks,\nbut we still lack understanding of the kind of linguistic knowledge they learn\nand rely on. We evaluate three models (BERT, RoBERTa, and ALBERT), testing\ntheir grammatical and semantic knowledge by sentence-level probing, diagnostic\ncases, and masked prediction tasks. We focus on relative clauses (in American\nEnglish) as a complex phenomenon needing contextual information and antecedent\nidentification to be resolved. Based on a naturalistic dataset, probing shows\nthat all three models indeed capture linguistic knowledge about grammaticality,\nachieving high performance. Evaluation on diagnostic cases and masked\nprediction tasks considering fine-grained linguistic knowledge, however, shows\npronounced model-specific weaknesses especially on semantic knowledge, strongly\nimpacting models' performance. Our results highlight the importance of (a)model\ncomparison in evaluation task and (b) building up claims of model performance\nand the linguistic knowledge they capture beyond purely probing-based\nevaluations.\n