Deploy Distributed Real-Time Data Pipelines Across Cloud and Edge to Process Rig Data for Easy Data Fusion and Processing

A. Wang, P. Acosta, Robert McL. Whitney · SPE Annual Technical Conference and Exhibition · 2024

Abstract The need for the distributed real-time data pipeline arises from the requirement to process disparate, asynchronous data streams that emanate from edge devices deployed to remote field locations, like drilling rigs, and in the cloud. These data streams need to be combined in real-time to serve advanced analytics applications like Artificial Intelligence (AI) and digital twins. Because of its distributed nature, this data pipeline not only needs to process data but also manage deployment, containers, and physical and virtual computing appliances. Due to this complexity, today's distributed real-time data pipelines are nearly all built as custom software solutions, incurring expensive development costs, long development time, and introducing an extensive maintenance burden. This paper presents a new framework that builds the data pipelines as interconnected, graphical functional blocks in a single software user interface. Each functional block can be assigned to execute in multiple remote locations. Furthermore, each functional block can be parametrized to support unique properties of each rig. This architecture makes it significantly simpler and faster to build, deploy and scale distributed real-time data pipelines while accelerating data engineering and data science projects that require high-speed, high-volume, distributed real-time data.

Read the paper · More papers on PaperTik