Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview
This paper presents an overview of a program designed to address the growing\nneed for developing freely available speech resources for under-represented\nlanguages. At present we have released 38 datasets for building text-to-speech\nand automatic speech recognition applications for languages and dialects of\nSouth and Southeast Asia, Africa, Europe and South America. The paper describes\nthe methodology used for developing such corpora and presents some of our\nfindings that could benefit under-represented language communities.\n