Reliability of wrist‐worn accelerometry devices and algorithms for sleep detection

Luis R. Peraza, Richard Joules, Yves A Dauvilliers, Robin Wolz · Alzheimer s & Dementia · 2020

Abstract Background Wearable devices have shown great potential for obtaining measures of real world evidence in clinical trials, but standardization and variability between different devices remain one of the barriers for systematic deployment. In this investigation, we present a neural network algorithm for sleep detection, Deep Learning Sleep (DLS), and compare its reliability across different datasets and two widely used sleep algorithms. Method A publically available dataset TUD (Borazio; ICHI 2014, N=45) and the CONTEXT dataset collected in Montpellier, IXI (Wolz; Alzheimer’s & Dementia 2017, N=45) were used to compare algorithm performance. Each dataset provided one‐night accelerometry data alongside polysomnography (PSG) per participant. DLS performance was compared against two other algorithms for sleep segment detection; Cole‐Kripke, CK (Cole; Sleep 1992), and estimation of stationary segments or ESS (Borazio; ICHI 2014). Then, cross‐validations were implemented with the datasets; e.g. training with dataset TUD and testing on dataset IXI, and vice versa. Result The cross‐validations experiments showed no significant differences in performance (sleep detection: 85‐89% sensitivity, 43‐45% specificity and 72‐74% accuracy) for the DLS algorithm. On the contrary, CK and ESS showed significant differences in their performances across experiments, and CK showed the highest difference for specificity (mean specificity = 27‐52%), demonstrating that CK and ESS algorithms are susceptible to different devices and/or data sources (Figure 1). Conclusion The CK and ESS algorithms showed susceptibility to different devices or data sources (e.g. clinical groups) and this raises questions about their reliability. In conclusion, we presented a deep learning based algorithm (DLS) that showed high robustness against different data sources as well as high accuracy when compared against the gold standard PSG.

Read the paper · More papers on PaperTik