Learning Indexing Patterns from One Language for the Benefit of Others

Udo Hahn, Kornél Markó, Stefan Schulz · 2004

Using language technology for text analysis and light-weight ontologies as a content-mediating level, we acquire index-ing patterns from vast amounts of indexing data for English-language medical documents. This is achieved by statisti-cally relating interlingual representations of these documents (based on text token bigrams) to their associated index terms. From these ‘English ’ indexing patterns, we then induce the associated index terms for German and Portuguese docu-ments when their interlingual representations match those of English documents. Thus, we learn from past English in-dexing experience and transfer it in an unsupervised way to non-English texts, without ever having seen concrete index-ing data for languages other than English.

Read the paper · More papers on PaperTik