Interlingual Indexing across Different Languages
Kornél Markó, Udo Hahn, Stefan Schulz, Philipp Daumke, Percy Nohama · FreiDok plus (Universitätsbibliothek Freiburg) · 2004
We present two methods for automatic indexing, which are based on an interlingual layer of content description. In the first approach, we acquire indexing patterns from English documents by statistically relating interlingual representations of English documents (based on text token bigrams) to their associated index terms. Given such indexing patterns, we then induce the associated index terms when the same interlingual representations turn up for documents of other natural languages (viz. German and Portuguese). Hence, we 'learn' from the past English indexing experience and transfer it in an unsupervised way to non-English languages, without ever having seen any concrete indexing data for languages other than English. In the second approach, documents from the three different languages are heuristically matched with a sophisticated medical thesaurus (the English MeSH) after both, documents and the thesaurus, have been transformed into the interlingua. The combination of the statistical and heuristical method in a fully automated indexing system achieves 56% to 68% of the human indexing performance for each of the three languages.