Learning to Coordinate Video Codec with Transport Protocol for Mobile Video Telephony
Anfu Zhou, Huanhuan Zhang, Guangyuan Su, Leilei Wu, Ruoxuan Ma, Zhen Meng, Xinyu Zhang, Xiufeng Xie, Huadóng Ma, Xiaojiang Chen · 2019
Despite the pervasive use of real-time video telephony services, the users' quality of experience (QoE) remains unsatisfactory, especially over the mobile Internet. Previous work studied the problem via controlled experiments, while a systematic and in-depth investigation in the wild is still missing. To bridge the gap, we conduct a large-scale measurement campaign on \appname, an operational mobile video telephony service. Our measurement logs fine-grained performance metrics over 1 million video call sessions. Our analysis shows that the application-layer video codec and transport-layer protocols remain highly uncoordinated, which represents one major reason for the low QoE. We thus propose ame, a machine learning based framework to resolve the issue. Instead of blindly following the transport layer's estimation of network capacity, ame reviews historical logs of both layers, and extracts high-level features of codec/network dynamics, based on which it determines the highest bitrates for forthcoming video frames without incurring congestion. To attain the ability, we train ame with the aforementioned massive data traces using a custom-designed imitation learning algorithm, which enables ame to learn from past experience. We have implemented and incorporated ame into \appname. Our experiments show that ame outperforms state-of-the-art solutions, improving video quality while reducing stalling time by multi-folds under various practical scenarios.