Precise and Adaptable: Leveraging Deep Reinforcement Learning for GAP-based Multipath Scheduler

Binbin Liao, Guangxing Zhang, Zulong Diao, Gaogang Xie · 2020 IFIP Networking Conference (Networking) · 2020

Using multiple network interfaces to accelerate data transfer is an attractive feature on multi-homed endhost. The key component in a multipath system such as MPTCP is the scheduler, which determines how to distribute the packets over multiple paths. In this paper, we propose GAPS, a new multipath scheduler aiming at decreasing the out-of-order queue size (OQS) under heterogeneous paths. In GAPS, a deep reinforcement learning agent monitors the network states, adjusts the GAP value of each path, and maximizes MPTCP’s reward utility. GAPS is precise in that it searches the optimal action policy, and only has 1.2% to 3.3% deviation from the true GAP. It is adaptable in that it performs better in varying network conditions and congestion control algorithms. In the controllable and realistic experiments, GAPS decreases the subflows’ 99th percentile OQS by up to 68.3%. It allows an increase of 12.7% in application goodput with bulk traffic while reducing application delay by 9.4% as compared to the state-of-the-art schedulers. For the short and long MPTCP flows, GAPS has the flow completion time (FCT) reduction by 13% and 21%, respectively.

Read the paper · More papers on PaperTik