Applying One-Versus-One SVMs to classify multi-label data with large labels using spark
Suthipong Daengduang, Peerapon Vateekul · 2017
Multi-Label classification aims to classify an example that can belong to many classes. Although One-versus-All (OVA) is the most common approach, our prior work has shown that the proposed One-versus-One (OVO) always gives higher prediction accuracy than OVA. However, OVO requires an extremely high computational cost when there are a large number of labels. In this paper, we apply our OVO SVMs on the proposed Spark framework along with a mechanism to split a job into a set of small jobs and then process them in parallel. The framework can induce OVO SVMs very fast, while maintaining the prediction accuracy even though there are a large number of classes. The experiment was conducted on six standard benchmarks. The result shows that our framework can really reduce computing time on Spark environment, while significantly outperforms OVA in terms of F1 on all data.