Corpus-based speech and language research in the Institute of Systems Science

Horng Jyh Paul Wu, Jin Guo, Ho Chung Lui, Hwee Boon Low · 2002

This paper describes the ongoing and planned research projects on speech and language modeling in the Institute of Systems Science. Four main areas of work have been concentrated and targeted: (1) intonation unit modeling using prosodic features; (2) identification and acquisition of lexical compounds; (3) stochastic dependency grammar parsing; and (4) factual information extraction. These research topics cover full-range of issues from the speech prosody level to the language discourse level. None the less, one consistent theme hinges together requirements from these different levels of processing-that is the so called corpus-based statistical approach. As revealed to us by applying this approach to various application systems, two related characteristics of a practical natural language processing (NLP) system emerge as rather crucial: (1) to prepare a high quality and large amount of tagged corpora as training examples; (2) to identify of a set of tag features most relevant to an application domain.>

Read the paper · More papers on PaperTik