Approximative Indexierungstechnik für historische deutsche Textvarianten

Heller, Markus · Social Science Open Access Repository (GESIS – Leibniz Institute for the Social Sciences) · 2006

Historical documents have specific propertieswhich make life hard for traditional information retrievaltechniques. The missing notion of orthography and a generalhigh degree of variation in the phonetic-graphemic representation,as well as in derivational morphology obstruct thepossibility to find documents upon the entry of a modernword as the search term. The following paper gives anoverview of existing string approximation technologies asused in bioinformatics, but also of phonetic approximationalgorithms. It proposes an architecture of combining bothnotions, while using Jörg Michael’ phonet program to deductfrom graphemes to a phonetic representation and alevenshtein automaton to allow for fast approximativematching. The final part of the paper evaluates the suitabilityof the approach, while using the levenshtein algorithm inits non-automaton-based implementation.

Read the paper · More papers on PaperTik