DNN Partitioning and Inference Task Offloading in 6G Resource-Constrained Networks

Dimitrios Kafetzis, Iordanis Koutsopoulos · 2024

In the emerging 6G networks landscape, edge computing applications are tasked with performing Machine-Learning (ML) inference using Deep Neural Networks (DNNs) on resource-constrained devices in terms of computational power, memory, and energy. This paradigm shift, brought forth by 6G’s promise of ultra-reliable low-latency communication necessitates novel approaches for managing DNN tasks. This paper studies the problem of DNN partitioning and selective offloading of ML inference tasks to a Base Station (BS) or an edge gateway node equipped with substantial computational resources. We consider a set of resource-constrained devices, each of which has an inference task (abstracted as a DNN) to execute. DNN partitioning determines the part of the neural network to run locally, and the part to offload to the BS. We are interested to determine a suitable task partition for each device, namely determine the neural network layer up to which computation will take place locally on the device. A certain task partition on a device affects the inference delays of tasks of other devices due to task coupling, when task scheduling for processing is decided at the BS. We formulate the optimization problems of DNN partitioning for minimizing total execution delays of tasks and for providing fair treatment to DNN inference tasks, by ensuring balanced task execution delays. We propose a greedy heuristic algorithm to solve the problems, and we evaluate it through numerical simulations, using input from experiment measurements on Raspberry Pi devices. The proposed approach is shown to perform much better than the baseline approach where each task is entirely executed locally on the device.

Read the paper · More papers on PaperTik