Study of Distributed Framework Hadoop and Overview of Machine Learning using Apache Mahout

Raxitkumar Solanki, Sree Harsha Ravilla, Doina Bein · 2019

The amount of data generated every day in digital format is overwhelming, so we need storage mechanisms to store it and manage it. The technical solutions to store and manage the data should be scalable to allow extraction of relevant information and analysis. We describe the initial steps on using Apache Mahout to find out the total number of books written by authors of different age groups, analyzing the patterns in the authors' age, and predicting which age group have authored the highest number of books in different calendar years. Publishing houses and literary agents can use our proposed software.

Read the paper · More papers on PaperTik