Integrating Knowledge Bases and Statistics in MT

Kevin K. Knight, Ishwar Chander, Matthew Haines, Vasileios Hatzivassiloglou, Eduard H. Hovy, Masayo Iida, Steve K. Luk, Akitoshi Okumura, Richard Whitney, Kenji Yamada · 1994

We summarize recent machine translation (MT) research at the Information Sciences Institute of USC, and we describe its application to the development of a Japanese-English newspaper MT system. Our work aims at scaling up grammar-based, knowledge-based MT techniques. This scale-up involves the use of statistical methods, both in acquiring e ective knowledge resources and in making reasonable linguistic choices in the face of knowledge gaps. Knowledge-based machine translation (KBMT) techniques have yielded high quality MT systems in narrow problem domains (Nirenburg et al. 1992; Nyberg & Mitamura 1992). This high quality is delivered by algorithms and resources that permit some access to the meaning of texts. But can KBMT be scaled up to unrestricted newspaper articles? We believe it can, provided we address two additional questions: 1. In constructing a KBMT system, how can we acquire knowledge resources (lexical, grammatical, conceptual) on a large scale? 2. In applying a KBMT system, what do we do when de nitive knowledge is missing? There are many approaches to these questions. Our working hypotheses are (1) a great deal of useful knowledge can be extracted from online dictionaries and text; and (2) statistical methods, properly integrated, can e ectively ll knowledge gaps until better knowledge bases or linguistic theories arrive. This paper describes completed and ongoing research on these hypotheses. This research is tightly coupled with our development e ort on a large-scale Japanese-English MT system, as part of the ARPAsponsored

Read the paper · More papers on PaperTik