WiC-TSV: An Evaluation Benchmark for Target Sense Verification of Words in Context

We present WiC-TSV, a new multi-domain evaluation benchmark for Word Sense\nDisambiguation. More specifically, we introduce a framework for Target Sense\nVerification of Words in Context which grounds its uniqueness in the\nformulation as a binary classification task thus being independent of external\nsense inventories, and the coverage of various domains. This makes the dataset\nhighly flexible for the evaluation of a diverse set of models and systems in\nand across domains. WiC-TSV provides three different evaluation settings,\ndepending on the input signals provided to the model. We set baseline\nperformance on the dataset using state-of-the-art language models. Experimental\nresults show that even though these models can perform decently on the task,\nthere remains a gap between machine and human performance, especially in\nout-of-domain settings. WiC-TSV data is available at\nhttps://competitions.codalab.org/competitions/23683\n

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC