Explaining Data Incompleteness in Knowledge Aggregation
Honglei Zeng, Richard E. Fikes · 2005
Knowledge aggregation is the problem of taking information from multiple heterogeneous sources and aggregating it into a unified knowledge base. One of the main challenges in that work has been dealing with data incompleteness because data sources seldom contain complete answers to a user’s query. Current approaches leverage users’ preferences over data sources when trying to aggregate incomplete data. Nevertheless, these approaches are not adequate to satisfy users’ needs to trust aggregated data before they can use them with confidence in the presence of incomplete information. We believe such trust may be earned by providing users with the explanations for incomplete data. In this paper, we build a decision tree-based classification system to acquire context knowledge about the sources and present techniques for applying the knowledge to explain incomplete data. Our experiments suggest the decision trees we built being 87% accurate in predicting unseen data. Further, context knowledge provides good characterizations of sources that we show to be valuable and often critical to users.