One-shot Bottleneck Size Optimization in CNN for Split Computing
Yutaro Horikawa, Takayuki Nishio · 2025
Split computing (SC) is a distributed inference method designed to balance computational load and reduce latency by partitioning a neural network into head and tail networks, deployed on a mobile device and a server, respectively. Due to the limited bandwidth, the data size at the split point must be reduced. For this reason, layers with small channel sizes, called bottlenecks, are often introduced in the model. However, the degree to which the channel size should be reduced depends on the overall structure of the model and the difficulty of the task, so a search is necessary. In this paper, we propose a training method that simultaneously reduces the number of bottleneck channels and optimizes model parameters for SC. This method introduces weights for the bottleneck layer into the loss function of training using cross-entropy. This makes the output value of each channel in the bottleneck layer closer to zero, and reduces the size of the transmitted data by filling in the values close to zero at the server side instead of sending them by the client during inference.