RAW-C: Relatedness of Ambiguous Words--in Context (A New Lexical Resource for English)

Most words are ambiguous--i.e., they convey distinct meanings in different\ncontexts--and even the meanings of unambiguous words are context-dependent.\nBoth phenomena present a challenge for NLP. Recently, the advent of\ncontextualized word embeddings has led to success on tasks involving lexical\nambiguity, such as Word Sense Disambiguation. However, there are few tasks that\ndirectly evaluate how well these contextualized embeddings accommodate the more\ncontinuous, dynamic nature of word meaning--particularly in a way that matches\nhuman intuitions. We introduce RAW-C, a dataset of graded, human relatedness\njudgments for 112 ambiguous words in context (with 672 sentence pairs total),\nas well as human estimates of sense dominance. The average inter-annotator\nagreement (assessed using a leave-one-annotator-out method) was 0.79. We then\nshow that a measure of cosine distance, computed using contextualized\nembeddings from BERT and ELMo, correlates with human judgments, but that cosine\ndistance also systematically underestimates how similar humans find uses of the\nsame sense of a word to be, and systematically overestimates how similar humans\nfind uses of different-sense homonyms. Finally, we propose a synthesis between\npsycholinguistic theories of the mental lexicon and computational models of\nlexical semantics.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC