MCL@IITK at SemEval-2021 Task 2: Multilingual and Cross-lingual Word-in-Context Disambiguation using Augmented Data, Signals, and Transformers

In this work, we present our approach for solving the SemEval 2021 Task 2:\nMultilingual and Cross-lingual Word-in-Context Disambiguation (MCL-WiC). The\ntask is a sentence pair classification problem where the goal is to detect\nwhether a given word common to both the sentences evokes the same meaning. We\nsubmit systems for both the settings - Multilingual (the pair's sentences\nbelong to the same language) and Cross-Lingual (the pair's sentences belong to\ndifferent languages). The training data is provided only in English.\nConsequently, we employ cross-lingual transfer techniques. Our approach employs\nfine-tuning pre-trained transformer-based language models, like ELECTRA and\nALBERT, for the English task and XLM-R for all other tasks. To improve these\nsystems' performance, we propose adding a signal to the word to be\ndisambiguated and augmenting our data by sentence pair reversal. We further\naugment the dataset provided to us with WiC, XL-WiC and SemCor 3.0. Using\nensembles, we achieve strong performance in the Multilingual task, placing\nfirst in the EN-EN and FR-FR sub-tasks. For the Cross-Lingual setting, we\nemployed translate-test methods and a zero-shot method, using our multilingual\nmodels, with the latter performing slightly better.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC