Singing Detection System Based on RNN and CNN Depth Features
Sihan Wu · 2023
With the constant development of the internet music service platform and social media, it becomes common to obtain and hear the music at any time and place, and the vocal composition plays an important role in the popular music. The music is released and spread in the form of the digital audio and it gradually changes the interaction between the users and music, including creation, enjoyment and exploration. As the main force at the background of the school's drama club, the musical instrument will be selected to imitate the relevant sounds on the platform and the system will find the relevant sound effects of the computer immediately. However, I doubt how the computer identifies different sounds, and whether the computer could distinguish the sounds made by human and musical instruments? I have checked some materials and understood that the music recognition is associated with the technologies about the music information retrieval. To cope with the growing digital audio, so that the users could better find, organize and analyze the required music files from the music library, it requires the intelligent calculation methods and tools based on the music contents, which is the core contents to be studied in the music information retrieval. The music information retrieval is a major subject, including the singing information processing, music search, audio information security, such as singing detection, humming/singing retrieval, automatic music and music emotion computing, etc. It is an emerging interdisciplinary subject that combines music with the computer field, including the singing information processing, music search and audio information security, which plays an important role in the music performance, music creation, auxiliary medical treatment and psychotherapy. The singing sound detection is a classification task and it determines whether there is the singing voice in the given audio clips. This process is a crucial pre-processing step and it could be selected to improve the performance of other tasks. Although the algorithm proposed has shown the high performance, but it has a large room for improvement to construct a more comprehensive singing sound detection system. The main problem to be solved is to distinguish the sound of songs and instruments. In the process of making some musical instruments, the sound principle is simulated based on human sound process and it is more difficult to recognize them and the human sound. The singing and non-singing are not judged comprehensively in a single feature, and it lacks of the robustness. Hence, the multi-feature fusion could solve the robustness of the singing detection system efficiently. In this paper, the author proposed the multi-feature singing detection system based on the deep CNN and RNN. The CNN could be applied to extract the spatial relationship features in the two-dimensional spectrum, and RNN could be selected to extract the temporal relationship features in the audio spectrum. In addition, the customized feature MFCC is combined to carry out the variance statistics and extraction, and the random forest is combined to acquire the dynamic features of singing and non-singing. The open DS Jamendo experiment is carried out and the experimental results indicate that the proposed method has a better detection performance than the baseline method. Wherein, F1 value could reach 0.93.