Towards using Wikipedia as a substitute corpus for topic detection and metadata generation in e-learning

Marek Meyer, Christoph Rensing, Ralf Steinmetz · 2006

Metadata is crucial for reuse of Learning Resources. Only with good metadata, there is a chance that a Learning Resource can be successfully found in a repository. However, many Learning Resources are still delivered with no or little attached metadata. Automatic metadata generation is used to put things right- either as assistance for the author, or as part of a repository’s retrieval functionality. Among the various metadata fields, those that cover the topic of a Learning Resource are the most important ones-especially keywords and categorization information. This paper presents a novel approach for domainindependent classification and keyword extraction by utilizing the immense knowledge that is gathered in the free Wikipedia encyclopedia. Wikipedia is proposed as a substitute corpus for classification methods in E-Learning. To support this proposal, the co-occurrence of matching topics and statistical similarity between Learning Resources and Wikipedia articles is analyzed. An algorithm for keyword generation based on the Wikipedia encyclopedia has been implemented and is described in detail in this paper. First results of the algorithm are presented and discussed. 1.

Read the paper · More papers on PaperTik