Data Collection and Preprocessing

Abhishek Mishra · 2019

This chapter explains some of the techniques that can be used to explore data, impute missing values, and engineer features. A number of sources of datasets can be used to create machine learning models. Some of the popular sources are the UCI machine learning repository, Kaggle.com, and AWS public datasets. The UCI machine learning repository is a public collection of over 450 datasets that is maintained by the Center for Machine Learning and Intelligent Systems at UC Irvine. The datasets are contributed by the general public and vary in the level of preprocessing developers will need to perform in order to use them for model building. The datasets can be downloaded onto developers' local computer and then processed using tools like Pandas and Scikit-learn. The chapter uses the Titanic dataset in a Jupyter notebook with NumPy and Pandas. It also explains how to convert categorical features into numeric features using one-hot encoding.

Read the paper · More papers on PaperTik