A Study of Metrics of Distance and Correlation Between Ranked Lists for Compositionality Detection

Compositionality in language refers to how much the meaning of some phrase\ncan be decomposed into the meaning of its constituents and the way these\nconstituents are combined. Based on the premise that substitution by synonyms\nis meaning-preserving, compositionality can be approximated as the semantic\nsimilarity between a phrase and a version of that phrase where words have been\nreplaced by their synonyms. Different ways of representing such phrases exist\n(e.g., vectors [1] or language models [2]), and the choice of representation\naffects the measurement of semantic similarity.\n We propose a new compositionality detection method that represents phrases as\nranked lists of term weights. Our method approximates the semantic similarity\nbetween two ranked list representations using a range of well-known distance\nand correlation metrics. In contrast to most state-of-the-art approaches in\ncompositionality detection, our method is completely unsupervised. Experiments\nwith a publicly available dataset of 1048 human-annotated phrases shows that,\ncompared to strong supervised baselines, our approach provides superior\nmeasurement of compositionality using any of the distance and correlation\nmetrics considered.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC