A stochastic parts program and noun phrase parser for unrestricted text
Kenneth Church · 1988
alice!k-wcIt is well-known that part of speech depends on context.The word "table," for example, can be a verb in some contexts (e.g., "He will table the motion") and a noun in others (e.g., "The table is ready").A program has been written which tags each word in an input sentence with the most likely part of speech.The program produces the following output for the two "table" sentences just mentioned:• He/PPS will/lVlD table/VB the/AT motion/NN ./.• The/AT table]NN is/BEZ ready/J/./.(PPS = subject pronoun; MD = modal; V'B = verb (no inflection); AT = article; NN = noun; BEZ ffi present 3rd sg form of "to be"; Jl = adjective; notation is borrowed from [Francis and Kucera, pp.6-8])Part of speech tagging is an important practical problem with potential applications in many areas including speech synthesis, speech recognition, spelling correction, proof-reading, query answering, machine translation and searching large text data bases (e.g., patents, newspapers).The author is particularly interested in speech synthesis applications, where it is clear that pronunciation sometimes depends on part of speech.Consider the foUowing three examples where pronunciation depends on part of speech.First, there are words like "wind" where the noun has a different vowel than the verb.That is, the noun "wind" has a short vowel as in "the wind is strong," whereas the verb "wind" has a long vowel as in "Don't forget to wind your watch."Secondly, the pronoun "that" is stressed as in "Did you see THAT?" unlike the complementizer "that," as in "It is a shame that he's leaving."Thirdly, note the difference between "oily FLUID" and "TRANSMISSION fluid"; as a general rule, an adjective-noun sequence such as "oily FLUID" is typically stressed on the fight whereas a noun-noun sequence such as "TRANSMISSION fluid" is typically stressed on the left.These are but three of the many constructions which would sound more natural if the synthesizer had access to accurate part of speech information.