Automatic index creation to support navigation in lexical graphs encoding part_of relations
Michael Zock, Debela Tesfaye · 2012
We describe here the principles underlying the automatic creation of a semantic map to support navigation in a lexicon, our target group being authors (speakers, writers) rather than readers. While machines can generally access information that it has stored, this does not always hold for people. A speaker may very well know a word, yet still be (occasionally) unable to access it. To help authors to overcome word-finding problems one could add to an existing electronic resource an index based on the (age-old) notion of association. Since ideas or their expressive forms (words) are related, they may evoke each other (lemon-yellow), but the likelihood for doing so varies over time and with the context. For example, the word 'piano' may prime 'instrument' or 'weight', but which of the two gets evoked depends on the context: 'concert' vs. 'house moving'. Given this dynamic aspect of the human brain, we should build the index automatically, computing the relation of terms and their weights on the fly. This dynamic creation of the index could be done via a corpus. This latter representing ideally the dictionary users' world knowledge, and the way how the prominence of words and ideas varies over time. Another important point are link-names, i.e. the type of relationship holding between two associates: [(rose) <--color (red)]. Given the fact that any query (e.g. 'India') may yield many hits, hits whose weights may be misleading, it makes sense to group the output according to some (other) category, for example, link names (color, city_of, instrument, ...). Yet, important as they may be, links or relations are hard to extract and to name. This is why we have decided to start with a very small sub-set, meronymic-, i.e. part-of relations (x is part of y, x has y, etc.).