Cross Language Information Extraction for Digitized Textbooks of Specific Domains

Wenhao Zhu, Laihu Luo, Chaoyou Ju, Bofeng Zhang · 2012

While the influence of the digitization movement is getting wider and wider, more and more countries have initiated their own digital library projects to preserve the culture by digitize millions of books. Together with all kinds of digital resources, such as videos, audios, images etc., the digital library can provide advanced services far more than reading and browsing. Information extraction is one of the fundamental methods to get structured information out of the digital books. Therefore, due to its importance for content integration and knowledge discovery, information extraction for different languages is becoming a key problem for the development of digital library. In this paper, we present a domain-related information extraction framework that suits for digitized textbooks of different languages. To achieve cross language adaptation, we introduce language independent features and simple language dependent features that bind with domain characters to generate extractors. Finally, we present two preliminary experiments to show the feasibility of this framework.

Read the paper · More papers on PaperTik