Robust N-gram Based Syntactic Analysis Using Segmentation Words

Nobuo Inui, Yoshiyuki Kotani · Institutional Repositories DataBase (IRDB) · 2001

We describe an N-gram based syntactic analysis using a dependency grammar. Instead of generalizing syntactic rules, N-gram information of parts of speech is used to segment a sequence of words into two clauses. A special part of speech, called segmentation word, which corresponds to the beginning or end symbol of clauses is introduced to express a sentence structure. Segmentation words for each clause were learned using the hill climbing method and a small bracketed corpus. Experimental results for Japanese sentences showed that N-gram based syntactic parser achieved 72.2 % recall, which is about the same level of performance as a probabilistic context-free grammar based parser with human-made language-dependent information.

Read the paper · More papers on PaperTik