Why current speech technology is false phonetics

Sven E. G. Öhman · 2009

Some of the best known physical-acoustic models of speech production that have been proposed by Speech Technologists such as G. Fant (Mouton 1960) at KTH Stockholm, Sweden and by K. N. Stevens (MIT Press 2000) at MIT the US are discussed this paper light of the principle which has been chosen as its motto, a principle I attribute to Fant, since his very concrete and practical guidance as my thesis advisor once made me aware of its fundamental importance modern and Technology. I argue that the models question are not adequate as intellectual tools for classical phonetics. As such, they are fact inferior to the tools already available to phoneticians, but which they have been taken by many to replace. I try to show that this last interpretation of the new models is due to misunderstandings. 1 Background During the Second Worid War the American linguist Mai t Joos made his military service on a submarine the U S Navy where he was trained to use special teclmical equipment that had been developed for the purpose of detecting enemy submarines on the basis of the under-water sound they emitted. M u c h owmg to Joos' efforts this equipment was later made commercially available to linguistics and phonetics researchers by the K a y Electric Corporation as the so-called SonaGraph. In 1948 Joos published the results of his phonetic war time experiments as a supplement to the international Journal Language under the title Acoustic Phonetics. The most impressive of these results was Joos' demonstration, inspired by the writings o f the German physicist and physiologist H . v. Helmholtz' (1821 1894), that the Vowel Triangle o f Classical Phonetics could be approximately reproduced by a two-dimensional plot of the formants F l against F2 . (This triangle is of course best known today as the Quadrangle of the International Phonetics Association, I P A , for identification research and elsewhere.) Many linguists includmg Joos liimself saw, this correlation, a promise for phonetics to actually become the exact science that it had striven to be the final two decades of the 19* centuiy. A few years later Potter, Kopp and Green at Be l l Labs published a book containing a collection o f Sonagraphized speech samples with the title Visible Speech (1947 N . Y . U S A ) which also happened to be the name o f a speech transcription system that had been worked out the 1880's by the famous Scottish-American phonetician and inventor (of the telephone) Alexander Graham Bell whose son had also given his name to the Bell Telephone Company (later part of A . T & T.). The impetus that this gave to the then brand new engineering research o f Acoustic Speech Studies thus had an orientation toward Phonetics fiom the very beginning. For instance, the French phonetician Pierre W H Y C U R R E N T S P E E C H T E C H N O L O G Y IS F A L S E PHONETICS 181 Delattre was able to synthesize vowels and even the stop consonants p, t and k by means o f a by modem standards rather primhive speech synthesis de\'ice at the Haskins Laboratories, then N e w York. This orientation toward phonetics, especially Joos' acoustic variety o f it, was challenged m the Monumental Acoustic Theory of Speech Production published by G . Fant 1960. Here Fant, usmg mathematical-acoustic theories that had been developed at M I T when Fant was a student there, argued that the formantcavity affiliations assumed by Joos, which had of course been the key to his and several other people's claim that the lowest-frequency fomiant ( F l ) reflected the big production, and the higher-frequency one the smaU cavity, were not justified when a model derived by a more detailed physical analysis was used. In fact by this model the so-called vowel spectrum consisted of an infinity of formant frequencies none of which was more affiliated whh any one cavity than with any of the other cavities! In the eyes of many linguists, this meant a death blow to Acoustic Phonetics as a method of interest to the of Linguistics (and Phonetics). To them it became a pure engmeering concern, and it occurred to few critics that Fant's results need not be interpreted this way. A more reasonable conclusion is that Fanfs results left Phonetics unaffected by physical acoustics, and that Joos' somewhat amateurish choice o f model was Fant's primary target. Yet many people among them Roman Jakobson and Morr is Halle, and later, also K . N . Stevens concluded that further progress phonetic research would be nearly impossible without building on the most recent engineering approach. The submissive attitude on the part of linguists and phoneticians to the new approaches should be seen the light of the ideological climate at that time. The anticommunist currents m the Western including Western Europe (with Sweden) reached its climax the Cold War. The U S , as the leading miUtary and teclmological power of the world, started to take over the role as World Police which had been the prerogative of Great Britain before the war. The U S defined hself as the power whose duty it was to make the world safe for democracy. Antidemocratic movements anywhere around the world were considered to threaten American national security. Therefore the U S started to invest huge sums of money into annaraents, as well as science and technology, not only at home, but also abroad, e. g. Sweden. It is my impression that the linguists of the world were so overwhelmed by this quite new shuation o f Wealth for Physics and Technology and relative poverty for the Humanities and Social Sciences that they were struck by what might be called teclinoplegia, a condition which a prostrated attitude becomes natural! This condition actually still prevails many places. One of hs symptoms was the attempt by Chomsky and his followers to rewrite linguistics as an earlier unheard-of Natural Science working with quasi-mathematical methods which included a generative phonology, and which was ctlso to replace classical phonetics. Classical Phonetics was beguining to be overtaken by two scientific competitors Technological Speech Research and Generative Phonology! The most recent sign of this movement is K . N . Stevens' Acoustic Phonetics, at M I T Press 2000 which the influence from Generative Phonology is evident on almost every page. 1.1 The intended of classical phonetics models Referring back to the motto concerning model evaluations -The test of a model is its performance its intended applications we must consider the intended for the classical quadrangle and other ideas of articulatory and auditory, i . e.. Classical Phonetics. They were meant as tools that individual language students should be able to use in the field i.e. study situations which one encountered events of sound

Read the paper · More papers on PaperTik