Data Preparation
Vineet Raina, Srinath Krishnamurthy · Apress eBooks · 2021
This chapter is dedicated to the data preparation step of the data science process. The captured data is typically explored to understand it better. Such exploration may reveal that the captured data is in a form which cannot be directly used to build models – we saw one such case in Chapter 6 where the data captured for predicting the category of an email consisted of just emails and their folders. This data had to go through a lot of preparation before it could be used for building models. In some other cases, it may seem that the data could be given to ML algorithms directly, but preparing the data in various ways might result in more effective models. We saw such a case in one of the examples we discussed in Chapter 1, where the captured data contained the timestamps and sale amounts for transactions at the checkout counter of a store. In this case, we felt that sale amount might have some trends based on what day of the week it is, what month it is, etc. Hence, we transformed the data so that it contains the hour, day, month, etc., along with the corresponding aggregated sale amount hoping that it will enable ML algorithms to find such trends. Hence, preparing the data in such ways might result in better models.