Scaling Personalized Machine Learning through DTW Clustering: Predicting Glycemia Levels as an Example

Abdelmounaam Rezgui · 2023

Personalized Machine Learning (PML) is a branch of ML where models are built to make predictions for an individual by analyzing the behavior of people similar to that individual. In healthcare applications, the data used for training PML models are often given as a large number (e.g., corresponding to thousands of patients) of very long, unsynchronized time series, e.g., of vital signs of patients. Capturing the biological similarities between humans (in terms of body vital signs) is much more intricate than capturing similarities between their preferences towards things such as books, airlines, grocery items, cars, sports teams, etc. Therefore, a good accuracy of PML models may sometimes be achieved only if models are trained with the data of each individual separately. This, however, translates into a prohibitive computational cost in the case of applications dealing with large populations. In this paper, we evaluate the tradeoffs between accuracy and computational cost of PML models when used in the context of healthcare. Specifically, we consider the case of predicting glycemia levels using a data set of time series of glucose values of 16 people. We use dynamic time warping (DTW) to generate three clusters of the given 16 time series. For each individual, we train two LSTM models, one using the original time series and one using the centroid of the cluster containing the original time series of that individual. Our study reveals that DTW clustering is a good alternative to generate clusters of time series that can be used to train PML models with an accuracy comparable to individually trained PML models.

Read the paper · More papers on PaperTik