Gene function analysis of semi-supervised multi-label learning

Suqun Cao · Caai Transactions on Intelligent Systems · 2008

Conventional machine learning is used only for single label learning, implying that every sample has only one label. However, in bioinformatics, a gene has more than one function, so it needs more than one label. Therefore, multi-label learning is more effective for identifying gene groups than conventional learning approach. Current research mainly focuses on supervised multi-label learning. The problem of effective semi-supervised multi-label learning strategies for labeled examples and unlabeled examples of gene expression datasets still remains unsolved. In this paper, a semi-supervised multi-label learning algorithm, named SML_SVM, is presented as an effective multi-label learner for analysis of gene expressions with at least one function. First, the proposed SML_SVM algorithm transforms the semi-supervised multi-label learning into corresponding semi-supervised single-label learning by the PT4 method, then it labels unlabeled examples using the maximum a posteriori (MAP) principle in combination with the K-nearest neighbor method, and finally, it solves the corresponding single-label learning problem using SVM. The distinctive characteristic of the proposed algorithm is its efficient integration of SVM-based single-label learning with MAP and K-nearest neighbor methods. Experimental results with a real Yeast gene expression dataset and a Genbase protein dataset show that the proposed SML_SVM algorithm outperforms the PT4-based MLSVM method and self-training MLSVM.

Read the paper · More papers on PaperTik