Unsupervised Discovery of Recurring Speech Patterns Using Probabilistic Adaptive Metrics

Unsupervised spoken term discovery (UTD) aims at finding recurring segments\nof speech from a corpus of acoustic speech data. One potential approach to this\nproblem is to use dynamic time warping (DTW) to find well-aligning patterns\nfrom the speech data. However, automatic selection of initial candidate\nsegments for the DTW-alignment and detection of "sufficiently good" alignments\namong those require some type of pre-defined criteria, often operationalized as\nthreshold parameters for pair-wise distance metrics between signal\nrepresentations. In the existing UTD systems, the optimal hyperparameters may\ndiffer across datasets, limiting their applicability to new corpora and truly\nlow-resource scenarios. In this paper, we propose a novel probabilistic\napproach to DTW-based UTD named as PDTW. In PDTW, distributional\ncharacteristics of the processed corpus are utilized for adaptive evaluation of\nalignment quality, thereby enabling systematic discovery of pattern pairs that\nhave similarity what would be expected by coincidence. We test PDTW on Zero\nResource Speech Challenge 2017 datasets as a part of 2020 implementation of the\nchallenge. The results show that the system performs consistently on all five\ntested languages using fixed hyperparameters, clearly outperforming the earlier\nDTW-based system in terms of coverage of the detected patterns.\n

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC