Data Engineering At Scale: Streaming Analytics with Cloud and Apache Spark
Santhosh Kumar Pendyala · Journal of Artificial intelligence and Machine Learning · 2025
This Study in modern healthcare systems, efficient data engineering is critical for processing vast amounts of real-time data generated by hospitals and medical devices. This article explores the transformative potential of integrating cloud-based technologies, specifically AWS and Databricks, with Apache Spark for real-time streaming analytics. Leveraging Databricks’ Lakehouse architecture and Unity Catalog enhances data governance and security through Identity and Access Management (IAM) and encryption mechanisms. This framework addresses challenges such as fragmented data pipelines, compliance concerns, and the latency of traditional data processing systems. Apache Spark's distributed computing and AWS's robust infrastructure provide scalable, highperformance analytics pipelines. Unity Catalog ensures secure, unified data access, meeting stringent healthcare compliance requirements like HIPAA. For example, patient admission and vital data streaming through Spark’s structured streaming enabled a 40% reduction in hospital response times. With increasing adoption of AI in healthcare, the proposed architecture bridges the gap between raw data ingestion and real-time actionable insights, enhancing patient outcomes. The methodology and results underscore the framework's scalability and its potential to revolutionize healthcare data engineering.