Getting Started with Scikit‐learn for Machine Learning

Wei-Meng Lee · 2019

This chapter introduces Scikit-learn, a Python library that implements the various types of machine learning algorithms, such as classification, regression, clustering, decision tree, and more. Using Scikit-learn, implementing machine learning is now simply a matter of calling a function with the appropriate data so that the readers can fit and train the model. Scikit-learn comes with a few standard sample datasets, which makes learning machine learning easy. Kaggle is the world's largest community of data scientists and machine learners. For learners of machine learning, readers can make use of the sample datasets provided by Kaggle. The University of California, Irvine Machine Learning Repository is a collection of databases, domain theories, and data generators that are used by the machine learning community for the empirical analysis of machine learning algorithms. The chapter shows some of the common tasks the coders need to perform when cleaning their data.

Read the paper · More papers on PaperTik