Semantic Typing for Corpus Pattern Analysis

Rogelio Nazar · International Journal of Lexicography · 2026

Abstract For the ongoing creation of a database of Spanish verbs, we developed a methodological proposal for the automatic tagging of semantic types in running text according to Hanks’ Corpus Pattern Analysis (CPA) guidelines. In this task, a text document is the input and the output is the tagging of each noun, noun phrase or proper noun with one of the semantic types in the CPA Ontology. The present proposal is based on a combination of algorithms for automatic ontology population, named entity recognition and, most importantly, word sense disambiguation, to assign the appropriate type to a noun according to the context. The paper includes an evaluation of the method tagging a random sample of 200 Wikipedia pages in Spanish and English. Evaluation figures by a panel of three experts show 84% precision and 88% recall in Spanish and 83% precision and 93% recall in English. These are competitive results considering the simplicity and computational efficiency of the algorithm.

Read the paper · More papers on PaperTik