[TD‐P‐016]: STUDYING NEURODEGENERATION WITH AUTOMATED LINGUISTIC ANALYSIS OF SPEECH DATA
Ellen A. Korcovelos, Kathleen Fraser, Jed A. Meltzer, Graeme Hirst, Frank Rudzicz · Alzheimer s & Dementia · 2017
Recorded changes in the language and speech of aging individuals offer a new means of quantifying neurodegeneration. By analyzing linguistic features such as parts-of-speech, word length, word frequency, and acoustic variables, automated techniques in computational linguistics make it possible to classify groups of differing linguistic ability. We extract features from audio recordings, and their respective transcripts, of participants recalling the narrative of Cinderella. These features identify significant characteristics for each of four populations: aphasic stroke (ST; N=19), primary progressive aphasia (PPA; N=11), mild cognitive impairment and Alzheimer's disease (MCI/AD; N=9 and N=2, respectively), and healthy elderly controls (CT; N=26). We then use these features to train a machine-learning classifier to correctly distinguish healthy individuals from patients (CT vs. ST+PPA+MCI), MCI/AD patients from ST and PA patients, and controls from each individual patient group (e.g., CT vs. ST). Our decision tree model is able to classify CT versus ST+PPA+MCI with 76.1% accuracy. We classify controls from MCI/AD patients with 89.2% accuracy, controls from PPA with 91.9% accuracy, and controls from stroke patients with 71.1% accuracy. Finally, the MCI/AD patients versus the combined stroke and PPA groups were classified with 80.5% accuracy. Word length and filled pauses were found to serve as prominent features in identifying pathology; however, when comparing controls and the MCI/AD group, acoustic features were selected more often than for the other populations’ feature sets. Binary classification between groups was between 13% and 21% more accurate than baseline values, and 4-way classification was 14.9% better. It appeared that linguistic features yielded better predictions than did the addition of acoustic features. Ongoing work aims to explain these phenomena and further evaluate the possible use of speech to serve as diagnostic criteria.