A machine learning method to quantify completeness of curated data sets
Douglas G. Howe · 2017
Many biological data repositories gather data through expert curation of published literature. New data types are added to these systems as research methods and priorities change, and resource limitations can result in incomplete curation of published papers or data types. Either way, expertly curated data sets can be incomplete when compared to what has been published, and knowing which data sets are incomplete or how incomplete they are remains a challenge. Knowing that a data set may be incomplete can justify further exploration of published literature for additional data of interest. In this work, the tested hypothesis was that machine learning methods could be used to identify genes in the Zebrafish Model Organism Database (ZFIN) that had incomplete curated gene expression data sets. A strong linear correlation was observed between the number of gene expression experiment records and the total number of journal publications associated with the gene. Starting with 36655 gene records from ZFIN, a data aggregation, cleansing, and filtering process reduced the set to 9870 gene records suitable for building and testing a predictive model for the number of expression experiments per gene. Feature selection and engineering reduced relevant features to the total number of journal publications, the number of journal publications already attributed for gene expression annotation, the percent of journal publications already attributed for expression data, and the number of transgenic constructs associated with each gene. These features were used to train a linear regression model to predict how many gene expression experiments each gene should have. Twenty five percent of the available gene records (2483 genes) were used to train the model. The remaining 7387 genes were used to test the model. One hundred and twenty two of the 7387 tested genes had a residual expression experiment count outside the model 95% confidence interval, suggesting they were missing expression annotations. Journal publications not already annotated for expression data from a random sample of 100 genes each inside or outside the 95% confidence interval and having a negative residual were examined in chronological order starting with the oldest papers for missing expression annotations. Genes were scored as missing expression data as soon as one paper was found that contained uncurated expression data for that gene. The model identified genes with published, but unannotated, expression data with a precision of 0.97 and recall of 0.71. This method can be used to reliably identify specific genes that are likely to be missing curation of published expression data and to help gauge whether to look further for published data to augment the existing expertly curated information.