Deep Imputation of Temporal Data

Chao Yan, Cheng Gao, Xinmeng Zhang, You Chen, Bradley Malin · 2019

Predictive modeling in healthcare has shown promise in various settings, such as early diagnosis, discovery of genotypephenotype associations, and the optimization of medical resource allocations [1]. Due to their data-driven nature, the effectiveness of these studies heavily relies on the quality of the collected data. However, clinical domain is often driven by business processing needs, such that ancillary information, which would be useful for building predictive models, is often collected in a haphazard fashion. As a consequence, the data are often incomplete. This is a challenge for learning effective predictive models, which only compounds in complexity in multivariate time series. This is because 1) filtering out instances that contain missing values (e.g., the feature vector associated with a particular time step) can induce bias in the model learning process [2], 2) in practice, it is difficult for learned models to make ad hoc predictions on test cases with multiple missing features, and 3) applying ineffective imputation strategies can result in the loss of noteworthy patterns [3] and learning models that are simply incorrect.

Read the paper · More papers on PaperTik