A Resampling Technique for Relational Data Graphs

Hoda Eldardiry, Jennifer Neville · 2008

Resampling (a.k.a. bootstrapping) is a computationallyintensive statistical technique for estimating the sampling distribution of an estimator. Resampling is used in many machine learning algorithms, including ensemble methods, active learning, and feature selection. Resampling techniques generate pseudosamples from an underlying population by sampling with replacement from a single sample dataset. It is straightforward to sample with replacement from propositional data that are independent and identically distributed (i.i.d.). However, it is not clear how to sample with replacement from an interconnected relational data graph with dependencies among related instances. In this paper, we develop a novel method for resampling from relational data that uses a subgraph sampling approach to preserve the local relational dependencies while generating a pseudosample with sufficient global variance. We evaluate our approach on synthetic data, showing that compared to an i.i.d. resampling approach it results in significantly lower error when used to estimate the variance of feature scores. We also evaluate our approach on a real-world relational classification task, showing that it improves the accuracy of bagging when compared with i.i.d. resampling. 1.

Read the paper · More papers on PaperTik