MUTA: Enabling Multi-Task Neural Network Inference in Programmable Data-Planes
Kaiyi Zhang, Changgang Zheng, Nancy Samaan, Ahmed Karmouch, Noa Zilberman · 2025
The need for real-time inference of large volumes of data led to the development of in-network machine learning. Programmable network switches can now execute various machine learning models in the data-plane at line rate. While a stream of data may require several prediction tasks, such as predicting bit rate, flow size, or traffic class, current solutions only support separate models for each task. This places a significant burden on the data-plane and leads to substantial resource consumption when deploying multiple tasks. To solve this problem, we introduce MUTA; a novel in-network multi-task learning solution. MUTA enables executing multiple inference tasks concurrently in the data-plane, without exhausting available resources. It introduces a data-plane mapping methodology to fit non-binarized multi-task neural networks within network switches. MUTA is deployed on P4-based hardware switches, and is shown to reduce memory requirements by × 10.5 and improve accuracy by up to 9.14% using limited training data, compared with state-of-the-art single-task learning solutions.