Kernel methods, syntax and semantics for relational text categorization

Alessandro Moschitti · 2008

Previous work on Natural Language Processing for Information Retrieval has shown the inadequateness of semantic and syntac-tic structures for both document retrieval and categorization. The main reason is the high reliability and effectiveness of language models, which are sufficient to accurately solve such retrieval tasks. However, when the latter involve the computation of relational se-mantics between text fragments simple statistical models may re-sult ineffective. In this paper, we show that syntactic and semantic structures can be used to greatly improve complex categorization tasks such as determining if an answer correctly responds to a ques-tion. Given the high complexity of representing semantic/syntactic structures in learning algorithms, we applied kernel methods along with Support Vector Machines to better exploit the needed rela-tional information. Our experiments on answer classification on Web and TREC data show that our models greatly improve on bag-of-words.

Read the paper · More papers on PaperTik