Network optimization for distributed machine learning over networks

Yuezhou Liu · 2023

Significant advances in edge and mobile computing capabilities enable machine learning (ML) and artificial intelligence (AI) to occur at geographically diverse locations in networks, e.g., cloud, edge, and mobile devices. The training data needed in ML may not be fully generated locally. Moreover, some promising distributed learning paradigms enable devices to collaboratively train a model, requiring communication among the devices to exchange necessary information. Thus, optimizing network strategies for the transmission/exchange of ML/AI ingredients (e.g., input data, model parameters, gradients) is important for facilitating efficient distributed ML. While many works use ML to optimize network operation strategies, few works study optimized networks that boost ML performance. This dissertation tries to fill the gap by studying several network optimization problems for distributed ML. Different from classic network optimization problems for data delivery or edge computing that optimize energy consumption, delay, throughput, etc., we also care about ML-driven metrics such as model accuracy and training convergence time. We first study network optimization for federated learning (FL), a distributed paradigm for clients to collaboratively learn ML models without having clients disclose their private data. We propose to use caching to improve FL efficiency with respect to the model training time for convergence. In each iteration, instead of having all clients download the latest global model from a parameter server, we select a subset of clients to access, with smaller delays, a somewhat stale global model stored in caches. We propose CacheFL - a cache-enabled variant of FedAvg, and provide theoretical convergence guarantees of it in the general setting where the local data is imbalanced and heterogeneous. With this result, we determine the caching strategies that minimize total wall-clock training time at a given convergence threshold for both stochastic and deterministic communication/computation delays. We then propose an experimental design network paradigm, wherein learner nodes train ML models via consuming data streams generated by data source nodes over a network. The goal is to efficiently use the network resources to transmit the most valuable data to the learners, where the value of data messages to model training is quantified by the experimental design objectives which capture the uncertainties of model parameters after training with the data messages. We formulate a social welfare optimization problem that maximizes the sum of (expected) experimental design objectives of individual learners by optimizing data transmission strategies subject to network constraints. We show that, by assuming linear regression models and Poisson data streams, the global objective is continuous DR-submodular, which enables the design of efficient approximate algorithms with approximation guarantees. We further extend our framework to incorporate more practical applications, i.e., ML with arbitrary nonlinear models. This requires the use of different experimental design objectives which usually have no closed form for general nonlinear models. We discuss methods for estimating the experimental design objectives, and show that by using expected information gain as the design objective and estimating it by nested Monte Carlo sampling, efficient algorithms can be constructed to solve the optimization problem with approximation guarantees.--Author's abstract

Read the paper · More papers on PaperTik