Knowledge acquisition from real-world texts: some lessons learned

Fernando Gómez · 2002

In recent years, natural language processing (NLP) has experienced a dramatic research shift by focusing on the processing of real-world texts, rather than on restricted domains. This shift of focus has been an acid test for the core components of NLP, namely parsing and semantic interpretation. For the last few years, we have been working on knowledge acquisition from texts (F. Gomez, 1985; F. Gomez and C. Segani, 1989). The research started as a set of theoretical ideas and, then, gradually we built a system that embodies the theory. The system could be at first called a "toy system", that works in restricted domains. Recently, we have extended every component of the system to handle "real-world" texts. The model has been implemented in a program that reads unedited texts from The World Book Encyclopedia (1994), and acquires new concepts and conceptual relations about topics dealing with the dietary habits of animals, their classifications and habitats. The program is also able to answer an ample set of questions about the database that it has automatically acquired. Some of the major lessons that derive from the research are presented.>

Read the paper · More papers on PaperTik