A Cluster-based Approach for Improving Isotropy in Contextual Embedding Space

The representation degeneration problem in Contextual Word Representations\n(CWRs) hurts the expressiveness of the embedding space by forming an\nanisotropic cone where even unrelated words have excessively positive\ncorrelations. Existing techniques for tackling this issue require a learning\nprocess to re-train models with additional objectives and mostly employ a\nglobal assessment to study isotropy. Our quantitative analysis over isotropy\nshows that a local assessment could be more accurate due to the clustered\nstructure of CWRs. Based on this observation, we propose a local cluster-based\nmethod to address the degeneration issue in contextual embedding spaces. We\nshow that in clusters including punctuations and stop words, local dominant\ndirections encode structural information, removing which can improve CWRs\nperformance on semantic tasks. Moreover, we find that tense information in verb\nrepresentations dominates sense semantics. We show that removing dominant\ndirections of verb representations can transform the space to better suit\nsemantic applications. Our experiments demonstrate that the proposed\ncluster-based method can mitigate the degeneration problem on multiple tasks.\n

Paper

References (33)

Scroll for more · 21 remaining

Similar papers

© 2026 NYSGPT2525 LLC