Cognates are present in multiple variants of the same text across different\nlanguages (e.g., "hund" in German and "hound" in English language mean "dog").\nThey pose a challenge to various Natural Language Processing (NLP) applications\nsuch as Machine Translation, Cross-lingual Sense Disambiguation, Computational\nPhylogenetics, and Information Retrieval. A possible solution to address this\nchallenge is to identify cognates across language pairs. In this paper, we\ndescribe the creation of two cognate datasets for twelve Indian languages,\nnamely Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu,\nPunjabi, Bengali, Marathi, and Malayalam. We digitize the cognate data from an\nIndian language cognate dictionary and utilize linked Indian language Wordnets\nto generate cognate sets. Additionally, we use the Wordnet data to create a\nFalse Friends' dataset for eleven language pairs. We also evaluate the efficacy\nof our dataset using previously available baseline cognate detection\napproaches. We also perform a manual evaluation with the help of lexicographers\nand release the curated gold-standard dataset with this paper.\n
Paper
References (35)
Scroll for more · 23 remaining