Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments
There is growing interest in ASR systems that can recognize phones in a\nlanguage-independent fashion. There is additionally interest in building\nlanguage technologies for low-resource and endangered languages. However, there\nis a paucity of realistic data that can be used to test such systems and\ntechnologies. This paper presents a publicly available, phonetically\ntranscribed corpus of 2255 utterances (words and short phrases) in the\nendangered Tangkhulic language East Tusom (no ISO 639-3 code), a Tibeto-Burman\nlanguage variety spoken mostly in India. Because the dataset is transcribed in\nterms of phones, rather than phonemes, it is a better match for universal phone\nrecognition systems than many larger (phonemically transcribed) datasets. This\npaper describes the dataset and the methodology used to produce it. It further\npresents basic benchmarks of state-of-the-art universal phone recognition\nsystems on the dataset as baselines for future experiments.\n