IMBIT: Resources and Search Strategy.

Claudia Wortmann · 2002

The working group "Search Technologies " at the University of Hildesheim has its main research focus in the field of pattern recognition (pattern matching, pattern completion, pattern extraction). The present participation of IMBIT in the CLEF project was based on a particular specification of our engine, namely towards "text-pattern recognition". Here, we mainly differentiate two kinds of patterns aiming at lexicographic and semantic similarities. The first kind is expressed in form of word lists where items are similar in writing (with respect to the input string). This type often is called error- (or fault-) tolerant retrieval. The second kind automatically groups word clusters-without grammar- which define concepts, ideas, events or processes. Such items normally form the content of articles, however following grammatical rules there. The technical basis is a particular artificial neural network, the SpaCAM (Sparsely Coded Associative Memory). SpaCAM has proven to work fine in many applications with text-data (like fulltext retrieval, translation memory tasks, terminology extraction etc.), but also in other contexts (like DNA retrieval, signature recognition, machine control tasks etc). Another module is called DCC-Mindmap and allows the graphic display of such word groups which mirror ideas or events, described in the normal text in underlying articles or documents. The distance within the (multidimensional) word clusters or between them-expressed via their position in the mindmap- relates to their appearance and

Read the paper · More papers on PaperTik