Correcting Sampling Bias in Structural Genomics through Iterative Selection of Underrepresented Targets

Kang Peng, Slobodan Vučetić, Zoran Obradović · 2005

In this study we proposed an iterative procedure for correcting sampling bias in labeled datasets for supervised learning applications. Given a much larger and unbiased unlabeled dataset, our approach relies on training contrast classifiers to iteratively select unlabeled examples most highly underrepresented in the labeled dataset. Once labeled, these examples could greatly reduce the sampling bias present in the labeled dataset. Unlike active learning methods, the actual labeling is not necessary in order to determine the most appropriate sampling schedule. The proposed procedure was applied on an important bioinformatics problem of prioritizing protein targets for structural genomics projects. We show that the procedure is capable of identifying protein targets that are underrepresented in current protein structure database, the Protein Data Bank (PDB). We argue that these proteins should be given higher priorities for experimental structural characterization to achieve faster sampling bias reduction in current PDB and make it more representative of the protein space.

Read the paper · More papers on PaperTik