Comparing Natural Language Processing Techniques for Alzheimer's Dementia Prediction in Spontaneous Speech

Alzheimer's Dementia (AD) is an incurable, debilitating, and progressive\nneurodegenerative condition that affects cognitive function. Early diagnosis is\nimportant as therapeutics can delay progression and give those diagnosed vital\ntime. Developing models that analyse spontaneous speech could eventually\nprovide an efficient diagnostic modality for earlier diagnosis of AD. The\nAlzheimer's Dementia Recognition through Spontaneous Speech task offers\nacoustically pre-processed and balanced datasets for the classification and\nprediction of AD and associated phenotypes through the modelling of spontaneous\nspeech. We exclusively analyse the supplied textual transcripts of the\nspontaneous speech dataset, building and comparing performance across numerous\nmodels for the classification of AD vs controls and the prediction of Mental\nMini State Exam scores. We rigorously train and evaluate Support Vector\nMachines (SVMs), Gradient Boosting Decision Trees (GBDT), and Conditional\nRandom Fields (CRFs) alongside deep learning Transformer based models. We find\nour top performing models to be a simple Term Frequency-Inverse Document\nFrequency (TF-IDF) vectoriser as input into a SVM model and a pre-trained\nTransformer based model `DistilBERT' when used as an embedding layer into\nsimple linear models. We demonstrate test set scores of 0.81-0.82 across\nclassification metrics and a RMSE of 4.58.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC