Communication Latency Optimization for Mesos-based Cloud Computing Systems
Shixin Huang, Chao Chen, Sifan Zhang, Jinhan Xin, Zheng Wang, Zhibin Yu · 2021
Apache Mesos is a resource manager for clouding computing systems and it provides effective resource isolation and sharing across distributed computing frameworks such as Apache Hadoop, Spark, and Flink. Continuously increasing computing demands require a larger number of nodes in cloud computing nowadays, and it has even become normal for a master node to process tens of thousands of messages from slave nodes each second. This creates considerable communication pressure for the master node, which may dramatically affect the communication latency between master and slave nodes. To reduce the latency, we propose a novel method which consists of three steps: 1) we build an environment to simulate the scenario of 10,000 slave nodes that simultaneously send messages to the master node. 2) We systematically evaluate how the heartbeat interval and message size affect the communication latency through a series of experiments. 3) We adopt Zlib compression algorithm to effectively reduce the communication cost. Evaluation results indicate that compared to default Mesos-based cloud computing systems, our method can reduce the communication latency by up to 56.34%.