Partial language analysis using support vector learning

Hiroyasu Yamada · Institutional Repositories DataBase (IRDB) · 2002

The difficulty of analyzing full language often hampers the problem to develop practical natural language processing (NLP) systems.In this thesis, we focus on two fundamental partial language analysis, 1) Japanese named entity extraction and 2) Partial parsing in English.Named entity extraction is the task for extracting information such as proper nouns and numerical expressions from a document and classifying these expressions into some categories such as person, location, organization, and date.Partial parsing takes a sentence as an input and interprets it as a parsed sub-tree which does not include ambiguities.Both techniques are useful for not only a wide range of applications such as machine translation and information retrieval field but also preprocessing of full language analysis.We present corpus-based method for named entity extraction and partial parsing using a machine learning technique, Support Vector Machines(SVMs) and especially show how SVMs can work well in partial language analysis.

Read the paper · More papers on PaperTik