Convolutional GRU Networks based Singing Voice Separation

Harshit Harsh, Akhil Indraganti, Sunny Dayal Vanambathina, Bharat Siva Yaswanth Ramanam, Vagicherla Sai Chandu, Hari Kishan Kondaveeti · 2022 2nd International Conference on Artificial Intelligence and Signal Processing (AISP) · 2022

Toned voice study is gaining importance due to advancement in the music industry. The breaking down of toned voice and its backtracking is similar to carrying images from the source domain to the target domain while preserving its content representation. For our case, the mixed voice prints were transformed into their constituent component. The drawback of U-Net convolutional architecture is that the learning rate may come down in the middle layers for deeper models, so there is some risk if the network learning is ignored in some cases where the abstract features are represented in those layers. In this work, we proclaim the methodology CGRUN for the task of singing voice division. It leads to a causal system that is naturally suitable for real-time processing applications. The speech processing application is the segregation of toned voices for voice mixing. Through software evaluation, this experiment confirms the use of CGRUN for toned voice separation. The technical term used for toned voice segregation and its backtracking is Music Information Retrieval (MIR).

Read the paper · More papers on PaperTik