DeepCon: Improving Distributed Deep Learning Model Consistency in Edge-Cloud Environments via Distillation

Bin Qian, Jiaxu Qian, Zhenyu Wen, Di Wu, Shibo He, Jiming Chen, Rajiv Kumar Ranjan · IEEE Transactions on Cognitive Communications and Networking · 2025

In a typical distributed Deep Learning (DL) based application, models are configured differently to meet the requirements of resource constraints. For instance, a large ResNet56 model is deployed on the cloud server while a small lightweight MobileNet model is more suitable for the end-user device with fewer computation resources. However, the heterogeneity of the model architectures and configurations may bring a systemic problem - models may produce different outputs when given the same input. This inconsistency problem may cause severe system failure of prediction agreement inside the application. Current research has not studied the systemic design for efficiently detecting and reducing the inconsistency among models in distributed DL applications. With the increasing scale of distributed DL applications, the challenges of inconsistency mitigation should consider both algorithm and system design. To this end, we design and implementDeepCon, an adaptive deployment system across the edge-cloud layer with over-the-air model updates. We implement ASRS sampling for efficiently sampling data to reveal the real data distribution as well as model prediction inconsistency. Then, we implement DMML-Par, an asynchronous parallel training algorithm for quickly updating the models and reducing inconsistency.DeepConimplements over-the-air updates with a set of APIS to enable seamless inconsistency detection and reduction in such deep learning applications. Our experiment results on both vision and language tasks demonstrate that DMML could improve the model consistency up to 4%, 7%, and 13% at CIFAR10/100 and IMDB datasets without sacrificing the accuracy of individual models. We also show that the ASRS sampling can save 90% network bandwidth of data transmission and that DMML-Par is up to 60% faster compared to simple synchronous parallel training.

Read the paper · More papers on PaperTik