A Comprehensive Investigation on Image Caption Generation using Deep Neural Networks
Andreza Patricia Batista, Lucas Alexandre de Souza COSTA, Demóstenes Zegarra Rodríguez · LA Referencia (Red Federada de Repositorios Institucionales de Publicaciones Científicas) · 2022
Currently, Voice over IP (VoIP) is one of the most used communication services, however, itsquality is related to several external factors that cause various types of degradation of the voice signal,directly affecting the quality of experience (QoE) of users. In order to classify the quality of the voicesignal transmitted in a VoIP communication affected by packet loss, two deep learning network models(DL - Deep Learning) were implemented. The models were developed using a deep neural networkmodel (DNN), through which the analysis of the voice signal affected by the packet loss rate (PLR) ofthe degraded signals, so it was possible to classify them into four different classs according to the user’sexperience. Thus, two databases were prepared, each containing four distinct classs. One of these wasprepared with the ITU-T P.862 recommendation database files with different packet loss rates, and theother database was prepared with the ITU-T P.501 recommendation files according to the index MOS ofMean Opinion Score (MOS) of each degraded file. The results obtained from the model for the databaseprepared by the packet loss rate was 94% accuracy in model validation, while the model results for thedatabase prepared by MOS the result obtained was 91% of accuracy. In a comparison with the resultsobtained by the P.563 algorithm and the results obtained by the P.862 algorithm, it was possible to obtainan average of 53.21% accuracy for the P.563 algorithm in comparison with the classification results ofthe algorithm P.862. Through the results obtained, it can be concluded that the generated models wereable to classify the packet loss rate and the MOS index in a non-intrusive way and with a great accuracyrate. Concluding that the generated models are able to determine the MOS of the degraded voice filesmore efficiently than the P.563 algorithm.Keywords: VoIP, Voice Quality, ITU-T P.862, ITU-T P.563, ITU-T P.501, Deep Learning, MachineLearning