Speaker-Independent Dysarthria Severity Classification using Self-Supervised Transformers and Multi-Task Learning
Dysarthria, characterised by slurred speech, is a hallmark of many neurological disorders and brain trauma. Clinical assessment requires an audio-visual investigation by a trained healthcare expert, who evaluates criteria such as respiration, phonation, articulation, resonance, and prosody during speech. Quantitative assessment of dysarthria is challenging due to its complexity, variability, and the subjective nature of human-observation-based scoring methods. We present a novel machine-learning framework using transformers for stratifying and monitoring patient speech. Our framework integrates a wav2vec 2.0 model, pre-trained on raw speech data from healthy individuals. To reduce reliance on speaker-specific characteristics and effectively manage the intrinsic intra-class variability of dysarthric speech, we employ a contrastive learning strategy with a multi-task objective: cross-entropy loss for classifying dysarthria severity, and triplet margin loss to ensure latent embeddings are grouped by severity rather than by speaker. This Speaker-Agnostic Latent Regularisation (SALR) framework provides an objective, accessible, and cost-effective alternative to traditional assessments. On the UA-Speech dataset, SALR achieved 70.5% accuracy and 59.2% F1 using leave-one-subject-out cross-validation—a 16.5% absolute (30% relative) improvement over prior benchmarks. Explainability analysis indicates that our multi-task objective enhances the ordinal structure of the latent space, reducing dependence on speaker-specific cues and demonstrating robustness and generalisability. In conclusion, this proof-of-concept study demonstrates the potential of the SALR framework for speaker-independent dysarthria severity classification, with potential implications for broader clinical applications in automated dysarthria assessments.
Paper
References (45)
Scroll for more · 33 remaining