Combining statistics on n-grams for automatic term recognition

Almudena Ballester, Ángel Martín Municio, Fernando Pardos, Jordi Porta, Rafael Jesús Ruiz Ureña, Fernando Sánchez León · 2002

This paper presents the work-in-progress in the development of an automatic term recognition (ATR) system built around the Corpus Científico-Técnico (CCT).Terms are modeled using three non-correlated dimensions: unithood, domainhood and usage, applied to a set of -grams automatically extracted from the corpus.These dimensions are combined with a supervised machine learning algorithm in order to classify -grams as terms or non-terms.Results of both noise and silence are promising given the paucity of data employed for training.Moreover, error analysis on noise reveals that other information dimensions can be used for significantly reducing noise.

Read the paper · More papers on PaperTik