A Survey of Coflow Scheduling Schemes for Data Center Networks
Shuo Wang, Jiao Zhang, Tao Huang, Jiang Liu, Tian Gong Pan, Yunjie Liu · IEEE Communications Magazine · 2018
Cluster computing applications, such as MapReduce and Spark, have been widely deployed in data centers to support commercial applications and scientific research. These applications often involve a collection of parallel flows generated by two groups of machines, and the slowest flow will determine the completion of applications. However, existing network-level optimizations are agnostic to the special communication pattern of the applications. Most of them only focus on improving the completion time of an individual flow instead of a collection of parallel flows. The recently proposed coflow abstraction exactly expresses the requirements of cluster computing applications and creates new opportunities to reduce the completion time of jobs in cluster computing. Due to the wide variations in coflow characteristics, coflow scheduling faces great challenges in data center networks. Much work has been proposed to solve one or some of the various challenges. Therefore, in this article, we survey the latest development in coflow scheduling for data center networks. We hope that this article will help readers quickly understand the causes of each problem and learn about current research progress, so as to serve as a guide and motivation toward further research in this area.