Optimized Voice Activity Detection for Audio Signal Processing
Haritha Palanichamy, Shanmugavadivu Pichai · 2025
The advancement in virtual assistant tools has augmented the scope and development of the speech processing application. In any such application, speech signals undergo pre-processing stages such as noise cancellation and silence removal. This study aims to develop a method to reconstruct a given audio recordings by removing silences, in order to improve the quality of speech signal. In this article, a novel silence removal technique termed as Optimized Voice Activity for Audio Signal Processing (OVAD-ASP) is presented. This technique is observed to effectively remove the silences embedded in audio signals, with assured consistency, scalability and reliability. The is designed to use Standard Deviation for the silence removal. The proposed OVAD-ASP was developed using MATLAB. The proposed OVAD-ASP method has a salient feature of determining the threshold dynamically for silent removal. Therefore, the thresholds were selected and tested using random thresholds and statistical thresholds, specifically mean and standard deviation are considered to determine the dynamic threshold from the input audio files, of which standard deviation yielded more encouraging results than the conventional approaches, namely Zero Crossing Rate (ZCR), Short Time Energy (STE), and Root Mean Square Energy (RMSE). It is observed that the results of OVAD-ASP are encouraging, as endorsed by Average Mean Squared Error of 0.0046 and Average Similarity Ratio of 89.3%. Therefore, the OVAD-ASP plays a vital and inevitable role in speech preprocessing.