SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation

We present SimLex-999, a gold standard resource for evaluating distributional\nsemantic models that improves on existing resources in several important ways.\nFirst, in contrast to gold standards such as WordSim-353 and MEN, it explicitly\nquantifies similarity rather than association or relatedness, so that pairs of\nentities that are associated but not actually similar [Freud, psychology] have\na low rating. We show that, via this focus on similarity, SimLex-999\nincentivizes the development of models with a different, and arguably wider\nrange of applications than those which reflect conceptual association. Second,\nSimLex-999 contains a range of concrete and abstract adjective, noun and verb\npairs, together with an independent rating of concreteness and (free)\nassociation strength for each pair. This diversity enables fine-grained\nanalyses of the performance of models on concepts of different types, and\nconsequently greater insight into how architectures can be improved. Further,\nunlike existing gold standard evaluations, for which automatic approaches have\nreached or surpassed the inter-annotator agreement ceiling, state-of-the-art\nmodels perform well below this ceiling on SimLex-999. There is therefore plenty\nof scope for SimLex-999 to quantify future improvements to distributional\nsemantic models, guiding the development of the next generation of\nrepresentation-learning architectures.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC