A Modern Version of the Old Theory of Telephone Quality

Harvey Fletcher · The Journal of the Acoustical Society of America · 1947

In 1921 a method was formulated for calculating the articulation—the percent of speech sounds interpreted correctly during a telephone conversation—from the physical characteristics of the telephone system and the noise environment of the listener. This paper describes this empirical theory as recently revised to correspond to modern techniques. A quantity A is calculated which is called the articulation index and which is related to the various kinds of measured values of articulation. In particular, if s is the percent of the fundamental sounds correctly perceived by the listener A = −p log (l-s), where p is a factor depending upon the skill of the crew making the articulation tests. For a well-practiced crew the value of p is between 0.5 and 0.6. The value of A is calculated by the equation A = F⋅V⋅E, where F is the maximum value of A, and V a factor which is zero at threshold intensity of the speech received by the listener and grows as the received speech intensity increases reaching unity for intensities between 60 and 70 db above threshold. The factor E is unity below this intensity but decreases from unity to 0.85 as the intensity levels increase to 60 db above these levels. The factor F is calculated by the equation F = ∫0∞D⋅Wdf, where D is a function of the frequency f, and W is a function of the relative response of the system. The functions D and W have been determined directly from articulation data. The factors V(x) and E(x) are determined directly from articulation data. The value of x for systems in a quiet place is equal to a weighted average of the response of the system. When a noise is present which produces a masking M at each frequency at the listener's ear, the same functions are used for determining F and V, but the response R at each frequency is replaced by R − M. However, for this case the x in E(x) is increased by an amount Δ, a quantity calculated directly from the masking M of the noise and the intensity of the received speech.

Read the paper · More papers on PaperTik