Syntactic Perturbations Reveal Representational Correlates of Hierarchical Phrase Structure in Pretrained Language Models

While vector-based language representations from pretrained language models\nhave set a new standard for many NLP tasks, there is not yet a complete\naccounting of their inner workings. In particular, it is not entirely clear\nwhat aspects of sentence-level syntax are captured by these representations,\nnor how (if at all) they are built along the stacked layers of the network. In\nthis paper, we aim to address such questions with a general class of\ninterventional, input perturbation-based analyses of representations from\npretrained language models. Importing from computational and cognitive\nneuroscience the notion of representational invariance, we perform a series of\nprobes designed to test the sensitivity of these representations to several\nkinds of structure in sentences. Each probe involves swapping words in a\nsentence and comparing the representations from perturbed sentences against the\noriginal. We experiment with three different perturbations: (1) random\npermutations of n-grams of varying width, to test the scale at which a\nrepresentation is sensitive to word position; (2) swapping of two spans which\ndo or do not form a syntactic phrase, to test sensitivity to global phrase\nstructure; and (3) swapping of two adjacent words which do or do not break\napart a syntactic phrase, to test sensitivity to local phrase structure.\n Results from these probes collectively suggest that Transformers build\nsensitivity to larger parts of the sentence along their layers, and that\nhierarchical phrase structure plays a role in this process. More broadly, our\nresults also indicate that structured input perturbations widens the scope of\nanalyses that can be performed on often-opaque deep learning systems, and can\nserve as a complement to existing tools (such as supervised linear probes) for\ninterpreting complex black-box models.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC