CUCHILD: A Large-Scale Cantonese Corpus of Child Speech for Phonology\n and Articulation Assessment
This paper describes the design and development of CUCHILD, a large-scale\nCantonese corpus of child speech. The corpus contains spoken words collected\nfrom 1,986 child speakers aged from 3 to 6 years old. The speech materials\ninclude 130 words of 1 to 4 syllables in length. The speakers cover both\ntypically developing (TD) children and children with speech disorder. The\nintended use of the corpus is to support scientific and clinical research, as\nwell as technology development related to child speech assessment. The design\nof the corpus, including selection of words, participants recruitment, data\nacquisition process, and data pre-processing are described in detail. The\nresults of acoustical analysis are presented to illustrate the properties of\nchild speech. Potential applications of the corpus in automatic speech\nrecognition, phonological error detection and speaker diarization are also\ndiscussed.\n