Automatic Classification and Relationship Extraction for Multi-Lingual and Multi-Granular Events from Wikipedia
Daniel Hienert, Dennis Wegener, Heiko Paulheim · TUbilio (Technical University of Darmstadt) · 2012
Wikipedia is a rich data source for knowledge from all domains.As part of this knowledge, historical and daily events (news) are collected for different languages on special pages and in event portals.As only a small amount of events is available in structured form in DBpedia, we extract these events with a rule-based approach from Wikipedia pages.In this paper we focus on three aspects: (1) extending our prior method for extracting events for a daily granularity, (2) the automatic classification of events and (3) finding relationships between events.As a result, we have extracted a data set of about 170,000 events covering different languages and granularities.On the basis of one language set, we have automatically built categories for about 70% of the events of another language set.For nearly every event, we have been able to find related events.