From parsing to database generation: applying natural language systems

Paul S. Jacobs · 2002

A new architecture for database generation that combines several levels of language processing is described and the author examines the relevant technique in each layer. The new methods for database generation break the task roughly into three stages, emphasizing different types of processing in each phase. The preprocessing, or corpus-driven tasks, use empirical results from large text samples to guide processing. The analysis tasks perform traditional parsing and semantic interpretation. Postprocessing actually produces the database templates. Syntactic, semantic, and domain knowledge can affect processing in each state. The author explains why database generation needs these separate stages of analysis. Some of the critical methods in each category are also described.>

Read the paper · More papers on PaperTik