Compound Classification Using the scikit‐learn Library

Jenny Balfer, Jürgen Bajorath, Martin Vogt · 2017

This chapter includes a tutorial demonstrating the use of different machine learning models for binary classification problems including naïve Bayes, decision trees, and support vector machines, and discussing their applicability to nonlinear data, interpretability, use for multi-class problems, and computational complexity. The software used by the tutorial is Python with packages NumPy, scikit-learn, and optionally pydot. The tutorial shows how to derive different models from labeled training data, and to apply these models on the test data. While naïve Bayes is a generative model, meaning that it derives full probabilities for all variables, decision trees and support vector machines are discriminative models. The naïve Bayes classifier can easily be applied without many parameter choices, whereas decision trees and support vector machines require careful parameter selection. In terms of computational complexity, support vector machines usually require much longer training time than naïve Bayes and decision tree classifiers, which can be trained in linear time.

Read the paper · More papers on PaperTik