Enhancing Data Throughput and Latency in Distributed In-Memory Systems for AI-Driven Applications across Public Cloud Infrastructure
Thulasiram Yachamaneni, Uttam Kotadiya, Amandeep Singh Arora · International Journal of AI BigData Computational and Management Studies · 2021
Data processing systems are exposed to inordinate pressure to provide real-time computation in the era of artificial intelligence (AI), especially in distributed cloud computing. Distributed In-Memory Systems (DIMS) have also become a crucial infrastructure for supporting AI-based applications, which require both low latency and high throughput. The paper presents an improvement of data throughput and final latency in DIMS to serve AI workload in the famous cloud sites, such as AWS, Azure, and Google Cloud. We explore how existing systems cannot architecturally support performance bottlenecks, and we present a model of hybrid in-memory data distribution that utilises adaptive caching, smart sharding of data, and intelligent data placement based on the proximity principle. On simulations and deployment to benchmark AI applications, the proposed methodology shows considerable performance improvements. Our solution is a layered architecture with modular components to address the issues of scalability, consistency, and fault tolerance, which is backed by efficient methods of memory management. The paper is accompanied by a comparative study with baseline models, such as Apache Ignite, Redis Cluster, and Memcached, which implement these models on the public cloud fringe. We present test results indicating that the enhancements lower average latency by 35 percent and raise data throughput by 47 percent on a variety of AI workloads such as image classification, natural language processing, and predictive analytics. The paper will conclude with a discussion on the implications of this research for large, scalable, AI-enabled cloud computing infrastructures, as well as the extensive work that can be done in the future