Establishing and Retrieving Domain Knowledge from Semi-Structural Corpora

Hsien-Chang Wang, Pei-Chin Yang, Chen-Chieh Li · Machine Learning · 2010

In this study, we proposed an approach to establish and retrieve domain knowledge automatically. The domain knowledge is established by combining the method of linguistic processing and frame-based representation. The features of descriptions consist of two major types: literal vectors and fuzzy vectors. The cosine and overlap measure is chosen to compute the similarity between literal vectors and fuzzy vectors respectively. According to our study, several results were observed: 1. The proposed approach for domain knowledge processing is useful for establishing and retrieving eco-knowledge. 2. For some birds, its features maybe marked directly on the figures in the book, a few descriptions may be missed in the text data. This will cause some mismatch in the experiment. 3. If an experienced bird watcher wants to use the inquiry system, the literal weighting should be increased. Experiment results showed that the weighting factor could be set as 0.9. 4. For a naive use to user the inquiry system, the literal weighting should be decreased. The weighting factor could be set as 0.2. For queries made by expert, it seems that only lexical matching is enough. However, for naive people who have no expertise on how to use specialized wording for the description of birds, combining lexical vector score with the fuzzy ones is a good choice. Since color attributes are essential for discrimination of birds, it plays an important role in the visual cognition of birds. Currently, our study adopted only the eleven basic colors, more sophisticate color membership determination should be considered to obtain better results. The further interesting research topic will be discovering the commonality and difference between book-style knowledge and knowledge collected from large amount of spontaneous description about objects.

Read the paper · More papers on PaperTik