Machine Learning for Data Science Applications
Ravindra B. Keskar, Mansi Anup Radke · 2021
Machine Learning (ML) involves use of various mathematical and statistical methods to draw useful conclusions and insights from the given data. With data growing exponentially, its analysis becomes extremely crucial. In this chapter, the theory and utility of various machine learning algorithms for data science applications is discussed. The data can be distinguished based on whether it is labeled or not. The major supervised and unsupervised algorithms considered in this chapter are linear regression, decision trees, naïve Bayes, support vector machines (SVMs), and clustering techniques like K-means. For the sake of completeness, this chapter also presents techniques for making better predictions on the new ( unseen ) data with performance metrics used for evaluating machine learning algorithms. These include concepts like splitting the given data into train/test/validation datasets, cross-validation, precision, recall, F/F1 measures, purity of clustering, and so on. Feature engineering and dimensionality reduction is a very important step for successfully modeling a data science problem. The chapter concludes with some discussion on practical issues and how to resolve them using ensemble learning and meta-learning techniques like bagging, boosting, and stacking.