Design of Domain-specific Term Extractor Based on Multi-strategy

Ruzhan Lu · Jisuanji gongcheng · 2005

This paper designs a multi-strategy based term extracting algorithm combining both statistics-based and rule-based methods. Withmultiple statistics measuring relationship between words in a string, it firstly uses a threshold classifier to extract two-word candidates from rawcorpus, extends these candidates left and right to obtain multi-word candidate terms and at last filters these terms to get domain-specific terms, thefinal result. It implements an extractor with an unprocessed corpus as input and domain-specific terms as output according to this algorithm. Aftersome experiments on corpora from multiple domains, the paper analyzes the results, figures out problems in it and finally does some expectations.

Read the paper · More papers on PaperTik