Drug related webpages classification using images and text information based on multi-kernel learning

Ruiguang Hu, Liping Xiao, Wenjuan Zheng · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2015

In this paper, multi-kernel learning(MKL) is used for drug-related webpages classification. First, body text and image-label text are extracted through HTML parsing, and valid images are chosen by the FOCARSS algorithm. Second, text based BOW model is used to generate text representation, and image-based BOW model is used to generate images representation. Last, text and images representation are fused with a few methods. Experimental results demonstrate that the classification accuracy of MKL is higher than those of all other fusion methods in decision level and feature level, and much higher than the accuracy of single-modal classification.

Read the paper · More papers on PaperTik