We demonstrate a program that learns to pronounce Chinese text in Mandarin, without a pronunciation dictionary. From non-parallel streams of Chinese characters and Chinese pinyin syllables, it establishes a many-to-many mapping between characters and pronunciations. Using unsupervised methods, the program effectively deciphers writing into speech. Its token-level character-to-syllable accuracy is 89%, which significantly exceeds the 22% accuracy of prior work.
Paper
References (15)
10From image to translation: Processing the endangered Nyushu script2016 · ACM Trans. Asian Low-Resource Lang. Inf. Process., 15.
Scroll for more · 3 remaining