Assessment of Speech Quality in VoIP

Zdeněk Bečvář, Michal Vondra, Lukas Novak · InTech eBooks · 2011

In VoIP (Voice over Internet Protocol), the voice is transmitted over the IP networks in the form of packets. This way of voice transmission is highly cost effective since the communication circuit need not to be permanently dedicated for one connection; however, the communication band is shared by several connections. On the other hand, the utilization of IP networks causes some drawbacks that can result to the drop of the Quality of Service (QoS). The QoS is defined by ITU-T E.800 recommendation (ITU-T E.800, 1994) as a group of characteristics of a telecommunication service which are related to the ability to satisfy assumed requirements of end users. The overall QoS of the telecommunication chain (denoted as end-to-end QoS) depends on contributions of all individual parts of the telecommunication chain including users, end devices, access networks, and core network. Each part of the chain can introduce some effects which lead to the degradation of overall speech quality. Lower speech quality causes user’s dissatisfaction and consequently shorter duration of calls (Holub et al., 2004) which reduces profit of telecommunication operators. Therefore, both sides (users as well as operators or providers) are discontented. The end device decreases speech quality by coding and/or compression of the speech signal. The speech quality can be also influenced by a distortion of the speech by its processing in the end device e.g. in the manner of filtering. It can lead to the saturation of the speech, insertion of a noise, etc. The processed speech is carried in packets via routers in the networks. Individual packets are routed to the destination as conventional data packets. Therefore, the packets can be delayed or lost. According to ITU-T G.114 recommendation (ITU-T G.114, 2003), the delay of speech should be lower than 150 ms to ensure high quality of the speech. Each packet is routed independently; therefore the delay of packets can vary in time. The variation of packet delay is usually denoted jitter. The impact of all above mentioned effects on the speech quality can be evaluated either by subjective or objective tests. The first group, subjective tests, uses real assessments of the speeches by users. Therefore it cannot be performed in real-time. The second set of tests, objective tests, tries to estimate the speech quality by speech processing and evaluation. The rest of chapter is organized as follows. The next section gives an overview on the related work in the field of VoIP speech quality. The third one describes basic principles of the speech quality assessment. The speech processing for all performed tests are described in section four. Section five presents the results of realized assessments of the speech quality. Last section sums up the chapter and provides major conclusions.

Read the paper · More papers on PaperTik