Vortex: A Stream-oriented Storage Engine For Big Data Analytics

Pavan Edara, Jonathan Forbesj, Bigang Li · 2024

Organizations are increasingly looking for ways to simplify collection and transformation of vast amounts of data collected from a highly connected internet. Data analytics over continuous streams of data enables interactive applications and reduces time to insights. Traditionally, streaming data collection and analysis has been either achieved by building systems, or using data warehouses built for batch processing. In this paper, we present Vortex, a storage engine that we built inside Google BigQuery to support real-time analytics. Vortex is a streaming-first storage system that supports both streaming and batch data analytics. Today, BigQuery uses Vortex to support petabyte scale data ingestion with sub-second data freshness and query latency.

Read the paper · More papers on PaperTik