Shallow parsing with pos taggers and linguistic features
Beáta Megyesi · 2002
Three data-driven publicly available part-of-speech taggers are applied to shallow parsing of Swedish texts. The phrase structure is represented by n00 types of phrases in a hierarchical structure containtu labels for every con81E1wN t type the token belongs to in the parse tree. The enw din is based on the con0E91wN086 of the phrase tags on the path from lowest to highern des. Various linsw2---E2 features are used in learning -- the taggers are trained on the basis of lexical incalwB62R on , part-of-speech on0 ,and a combination -- of both, to predict the phrase structure of the token with or without part-of-speech. Special attention is directed to the taggers' senrsitivity to diffrent types of lin0E0wN0 in0E0wN02 in0E0 in learning as well as the taggers' senrsitivity to the size and the various types of training data sets. The method can be easily transferred to otherlanguages. Keywords: Chunking, Shallow parsing, Part-of-speech taggers,Hidden Markov models, Maximum entropy learning TranERwN2E2E8wnwn learnRw 1. Introduction Machine learning-- techniques in the last decade have permeated several areas of natura1 language processing (NLP). The reason is that a vast number of machine learning algorithms have proved to be able to learn from nomw29 language data given a relatively small correctly annotated corpus. Therefore, machine learning algorithms make it possible to within a short period of time develop language resources -- data ansourc on various linsw1B8--- levels---that are nwB---19Ew for numerous application in natural language processing. One of the most popular NLP areas that machine learning algorithms have been successfully applied to is part-of-speech (PoS) tagging, i.e., the annotation of words with their textually appropriate PoS tags, often including morphological features. The data driven algorithms that have been su...