Kernels for Structured Data for Natural Language Processing
Jun Suzuki · Institutional Repositories DataBase (IRDB) · 2005
The recent success of statistical natural language processing (NLP) has made feasible the development of challenging applications.For example, text classification traditionally classified texts into 'topics' but now researchers attempt to classify texts according to 'sentiments', such as 'intention' and 'polarity'.Text summarization traditionally only extracted important sentences from single documents but now strives to automatically generate an abstract from multiple-source documents.These trends indicate that recent tasks have increasingly demanded a text to be interpreted semantically or contextually.In other words, solving recent tasks with higher performance has required methods that can handle richer types of linguistic information.Conventionally, in the field of NLP, a set of words, called bag-of-words, is the most widely used model for representing features of texts.However, it is known that bag-of-words models lack many of the linguistic features found in texts.It is widely accepted that the lack of structures in bag-of-words models has led to inadequate performance in recent tasks.Therefore, methods are needed that are capable of handling richer linguistic information within texts.For these reasons, this dissertation proposes a methodology that is capable of handling richer structural information derived from syntactic and semantic analysis.I formalize all of my proposed methods within the framework of kernel