Low Anisotropy Sense Retrofitting (LASeR) : Towards Isotropic and Sense Enriched Representations
Contextual word representation models have shown massive improvements on a\nmultitude of NLP tasks, yet their word sense disambiguation capabilities remain\npoorly explained. To address this gap, we assess whether contextual word\nrepresentations extracted from deep pretrained language models create\ndistinguishable representations for different senses of a given word. We\nanalyze the representation geometry and find that most layers of deep\npretrained language models create highly anisotropic representations, pointing\ntowards the existence of representation degeneration problem in contextual word\nrepresentations. After accounting for anisotropy, our study further reveals\nthat there is variability in sense learning capabilities across different\nlanguage models. Finally, we propose LASeR, a 'Low Anisotropy Sense\nRetrofitting' approach that renders off-the-shelf representations isotropic and\nsemantically more meaningful, resolving the representation degeneration problem\nas a post-processing step, and conducting sense-enrichment of contextualized\nrepresentations extracted from deep neural language models.\n