Data Pipelines and Structured Spark Applications

Scott G. Haines · Apress eBooks · 2022

There is a central processing paradigm that exists behind the scenes and can help connect just about everything you build as a data engineer. The processing paradigm is a physical as well as a mental model for effectively moving and processing data, known as the data pipeline. We first touched on the data pipeline in Chapter 1 , while introducing the history and common components driving the modern data stack. This chapter will teach you how to write, test, and compile reliable Spark applications that can be weaved directly into the data pipeline.

Read the paper · More papers on PaperTik