Design of Domain-specific Term Extractor Based on Multi-strategy
Ruzhan Lu · Jisuanji gongcheng · 2005
This paper designs a multi-strategy based term extracting algorithm combining both statistics-based and rule-based methods. Withmultiple statistics measuring relationship between words in a string, it firstly uses a threshold classifier to extract two-word candidates from rawcorpus, extends these candidates left and right to obtain multi-word candidate terms and at last filters these terms to get domain-specific terms, thefinal result. It implements an extractor with an unprocessed corpus as input and domain-specific terms as output according to this algorithm. Aftersome experiments on corpora from multiple domains, the paper analyzes the results, figures out problems in it and finally does some expectations.