Practical Speech Translation Systems will Integrate Human Expertise, Multimodal Communication, and Interactive Disambiguation
Christian Boitet · 1993
Summary It has always been remarkably difficult to build really practical Machine Translation (MT) and Speech Processing (SP) systems. As Speech Translation (ST) combines the difficulties of both endeavours, it should come as no surprise that the first prototypes, although well-researched and brilliantly demonstrated, cannot be extended towards practical systems. Dramatic progress in both MT and SP technology is not being likely to be witnessed in the near future. Besides necessary but inherently limited improvements in the component technologies, the construction of practical ST systems will require better user-friendliness, achievable through the introduction of a human expert (interpreter), multimodal communication facilities between the expert and the speakers, and various control and feed-back facilities. Because of the quality and coverage required of the Speech Recognition (SR) and Natural Language Analysis (NLA) components in realistic applications and their inherent difficulty, however, it will also be necessary to involve the end users (the speakers) in these processes, by encouraging them to control their own voice, and asking them to help through multimodal active and passive disambiguation.