Single-Channel Speech Enhancement Based on Sparse Regressive Deep Neural Network
海霞 孙 · Software Engineering and Applications · 2017
语音增强可以改进语音质量,抑制、降低噪声干扰,提高信噪比,在手机等语音通信设备中广泛应用。近年来,由于深度神经网络学习的语音增强技术,可有效克服传统神经网络语音消噪算法易陷于局部最优的不足,取得更好的语音消噪效果,成为语音增强技术领域的研究热点。本文针对已有深度神经网络模型泛化能力较弱、存储开销较大等问题,研究提出一种基于稀疏回归深度神经网络的语音增强算法。该算法通过在预训练阶段引入丢弃法(Dropout)和稀疏约束正则化技术改进训练模型保持预训练和调优阶段模型结构一致性,提升模型泛化能力。通过权值共享和权值量化进行网络压缩,降低存储开销。用谱减法进行后处理,有效去除稳态噪声,提高语音质量。仿真实验结果表明,改进算法可达到较高的语音性能评价指标,取得较好的语音增强效果,可满足语音增强处理要求。 Speech enhancement is a mean to improve the quality and intelligibility by noise suppression and enhancing the SNR at the same time, which has been widely applied in voice communication equipments. In recent years, Deep Neural Network (DNN) has become a research hot point due to its powerful ability to avoid local optimum, which is superior to the traditional neural network. However, the existed DNN costs storage and has a bad generalization. Now, this document puts forward a sparse regression DNN model to solve the above problems. First, we will take two regularization skills called Dropout and sparsity constraint to strengthen the generalization ability of the model. Obviously, in this way, the model can reach the consistency between the pre-training model and the training model. Then network compression by weights sharing and quantization is taken to reduce storage cost. Next, spectral subtraction is used in post-processing to overcome stationary noise. The result proofs that the improved framework gets a good effect and meets the requirement of the speech processing.