Apache Spark

Jason Bell · Machine Learning · 2020

The Apache Spark project was created by the AMPLab at UC Berkeley as a data analytics cluster computing framework. This chapter is a quick overview of the Scala language and its use within the Spark framework. With the first Spark application created, it expands on this knowledge and looks at some of the other APIs that Spark provides: SparkSQL, Spark Streaming, and MLLib. The introduction of SparkSQL gave us a system to run high-performance queries against large datasets. Since the first edition of this book, the Spark SQL system went through an overhaul that introduced data frames and a more robust way of handling queries. The chapter explains how to build up a Java Spark application to load CSV data and run queries against it. It also shows three machine learning types: decision trees, K-means clustering, and association rules with FP-Growth.

Read the paper · More papers on PaperTik