Advanced Analytics with Spark Stateful Structured Streaming

Scott G. Haines · Apress eBooks · 2022

In the last chapter, you learned to use Apache Spark’s powerful aggregation and analytics functions, from the agg operator that enabled powerful columnar aggregation capabilities directly off a grouped dataset, to the analytical window functions that allowed you to partition and analyze datasets using these unique windowing capabilities. This gave you the ability to look back (lag) or forward (lead) across many rows from your current position in an iteration. You learned to use lag over to create row-by-row average deltas and similar techniques to create running cumulative totals.

Read the paper · More papers on PaperTik