Inference Efficient Source Separation Using Input-dependent Convolutions
Shogo Seki, Li Li · 2024
This paper proposes a large-capacity but inference-efficient source separation model that considers input mixtures. Convolutional neural networks (CNNs) are fundamental elements for building robust separation models, and CNN-based separation models are yet attractive in their lower computational complexity than recent competitive separation models. One premise of these models is that a single kernel is applied to the input at each convolution. This does not take into account different input mixtures, such as those with speakers of the same or different genders, resulting in suboptimal separation performance. To overcome the issue, the proposed method employs an input-dependent convolution called conditionally parameterized convolution (CondConv) for CNN-based separation models. CondConv contains multiple kernels and generates an aggregated single kernel depending on the input. This can increase the network capacity while maintaining the computation complexity during inference as a standard convolution, enabling separation models to account for the input mixtures. Through the experimental evaluations under a speech separation task, the proposed input-dependent convolution approach consistently improves several CNN-based separation models while maintaining a negligible increase in inference.