RDLL at CrossLink Anchor Extraction Considering Ambiguity in CLLD.

Fuminori Kimura, Kensuke Horita, Yuuki Konishi, Hisato Harada, Akira Maeda · 2013

In this paper, we describe our work in NTCIR-10 on the task of cross-lingual link discovery (CLLD). Our proposed method is focused mainly on two aspects in order to accomplish this task: how to find important anchors from an original article in order to crosslink and how to find the correct links to articles in the target language for the original articles. The system first uses online data collected from Japanese Wikipedia articles in order to build a basic crosslink database. These data will be applied in order to identify the anchors and find out the relevant corresponding English articles. We carried out this task in three steps. First, we parsed the Japanese articles and extracted the candidate anchors. Second, we ranked anchors on the basis of the weights of their importance. Third, we determined the correct English articles for each anchor. We marked LMAP 0.151 with manual assessment.

Read the paper · More papers on PaperTik