Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction
When language models process syntactically complex sentences, do they use\ntheir representations of syntax in a manner that is consistent with the grammar\nof the language? We propose AlterRep, an intervention-based method to address\nthis question. For any linguistic feature of a given sentence, AlterRep\ngenerates counterfactual representations by altering how the feature is\nencoded, while leaving intact all other aspects of the original representation.\nBy measuring the change in a model's word prediction behavior when these\ncounterfactual representations are substituted for the original ones, we can\ndraw conclusions about the causal effect of the linguistic feature in question\non the model's behavior. We apply this method to study how BERT models of\ndifferent sizes process relative clauses (RCs). We find that BERT variants use\nRC boundary information during word prediction in a manner that is consistent\nwith the rules of English grammar; this RC boundary information generalizes to\na considerable extent across different RC types, suggesting that BERT\nrepresents RCs as an abstract linguistic category.\n