Updating Thesaurus via Extracting Keywords from Metadata
Jun Wang · Zhongwen xinxi xuebao · 2005
The application of thesauri in digital libraries is seriously constrained by the manual nature of current thesaurus maintenance mechanism which cannot keep up with the rapid evolvement of knowledge.This paper proposes a statistical method of extracting new terms from titles of metadata and settling them into the thesaurus.The settlement is based on the subject indexing coded in the metadata records.An experiment was conducted on the Chinese Classification and Thesaurus and a corpus of 5 thousands bibliographic data of computing domain.The successful result demonstrates that the techniques proposed are effective and can be applied to the corpus of large size and foreign language.