Spying on your neighbors: Fine-grained probing of contextual embeddings for information about surrounding words

Although models using contextual word embeddings have achieved\nstate-of-the-art results on a host of NLP tasks, little is known about exactly\nwhat information these embeddings encode about the context words that they are\nunderstood to reflect. To address this question, we introduce a suite of\nprobing tasks that enable fine-grained testing of contextual embeddings for\nencoding of information about surrounding words. We apply these tasks to\nexamine the popular BERT, ELMo and GPT contextual encoders, and find that each\nof our tested information types is indeed encoded as contextual information\nacross tokens, often with near-perfect recoverability-but the encoders vary in\nwhich features they distribute to which tokens, how nuanced their distributions\nare, and how robust the encoding of each feature is to distance. We discuss\nimplications of these results for how different types of models breakdown and\nprioritize word-level context information when constructing token embeddings.\n

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC