Machine Learning Based Approach to S-clause Segmentation

Mi-Young Kim, Jong-Hyeok Lee · International Journal of Computer Processing Of Languages · 2004

When a dependency parser analyzes long sentences with fewer subjects than predicates, it is difficult to recognize which predicate governs which subject. To handle such syntactic ambiguity between subjects and predicates, we define a "S(ubject)-clause" as a group of words containing several predicates and their common subject. This paper proposes a method to segment S-clauses and perform syntactic analysis of long sentences using S-clause segmentation. To segment S-clauses, various machine learning methods were applied. We found that the Logitboost method produced the best performance.In our experimental evaluation, S-clause information turned out to be effective in determining the governor of a subject and that of a predicate in dependency parsing. Further syntactic analysis using S-clauses achieved an improvement in precision by 5.25 percent.

Read the paper · More papers on PaperTik