The ability to correctly model distinct meanings of a word is crucial for the\neffectiveness of semantic representation techniques. However, most existing\nevaluation benchmarks for assessing this criterion are tied to sense\ninventories (usually WordNet), restricting their usage to a small subset of\nknowledge-based representation techniques. The Word-in-Context dataset (WiC)\naddresses the dependence on sense inventories by reformulating the standard\ndisambiguation task as a binary classification problem; but, it is limited to\nthe English language. We put forward a large multilingual benchmark, XL-WiC,\nfeaturing gold standards in 12 new languages from varied language families and\nwith different degrees of resource availability, opening room for evaluation\nscenarios such as zero-shot cross-lingual transfer. We perform a series of\nexperiments to determine the reliability of the datasets and to set performance\nbaselines for several recent contextualized multilingual models. Experimental\nresults show that even when no tagged instances are available for a target\nlanguage, models trained solely on the English data can attain competitive\nperformance in the task of distinguishing different meanings of a word, even\nfor distant languages. XL-WiC is available at\nhttps://pilehvar.github.io/xlwic/.\n