Avoiding bias when aggregating relational data with degree disparity
David D. Jensen, Jennifer Neville, Michael Hay · 2003
A common characteristic of relational data sets leads many relational learning algorithms to discover misleading correlations. This characteristic—degree disparity—occurs when entities of one class participate in systematically higher numbers of relations than entities of another class. In such cases, the aggregation functions that are used in many relational learning algorithms (e.g., AVG, MODE, SUM, EXISTS, COUNT, MAX, MIN) will result in misleading correlations and added complexity in models. We examine this problem through a combination of simulations and experiments. We show how two novel significance testing procedures can adjust for the effects of using aggregation functions in the presence of degree disparity. 1.