A ProbabilisticWord Class Tagging Module Based On Surface Pattern Matching
Robert Eklund · DSpace repository (University of Tartu) · 2015
This paper 1 treats automatic, probabilistic tagging. First, residual, untagged, output from the lexical analyser SWETWOL 2 is described and discussed. A method of tagging residual output is proposed and implemented: the left-stripping method. This algorithm, employed by the module ENDTAG, recursively strips a word of its leftmost letter, and looks up the remaining ‘ending ’ in a dictionary. If the ending is found, ENDTAG tags it according to the information found in the dictionary. If the ending is not found in the dictionary, a match is searched in ending lexica containing statistical information about word classes associated with the ending and the relative frequency of each word class. If a match is found in the ending lexica, the word is given graded tagging according to the statistical information in the ending lexica. If no match is found, the ending is stripped of what is now its left-most letter and is recursively searched in dictionary and ending lexica (in that order). The ending lexica – containing the statistical information – employed in this paper are obtained from a reversed version of Nusvensk Frekvensordbok (Allén 1970), and contain endings of one to seven letters. Success rates for ENDTAG as a standalone module are presented. 1