Speech Representation Learning Combining Conformer CPC with Deep Cluster\n for the ZeroSpeech Challenge 2021

We present a system for the Zero Resource Speech Challenge 2021, which\ncombines a Contrastive Predictive Coding (CPC) with deep cluster. In deep\ncluster, we first prepare pseudo-labels obtained by clustering the outputs of a\nCPC network with k-means. Then, we train an additional autoregressive model to\nclassify the previously obtained pseudo-labels in a supervised manner. Phoneme\ndiscriminative representation is achieved by executing the second-round\nclustering with the outputs of the final layer of the autoregressive model. We\nshow that replacing a Transformer layer with a Conformer layer leads to a\nfurther gain in a lexical metric. Experimental results show that a relative\nimprovement of 35% in a phonetic metric, 1.5% in the lexical metric, and 2.3%\nin a syntactic metric are achieved compared to a baseline method of CPC-small\nwhich is trained on LibriSpeech 460h data. We achieve top results in this\nchallenge with the syntactic metric.\n

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC