Subspace Search for Community Detection and Community Outlier Mining in Attributed Graphs.

Emmanuel Müller · LWA · 2014

Attributed graphs are widely used for the representation of social networks, gene and protein interactions, communication networks, or product co-purchase in web stores. Each object is represented by its relationships to other objects (edge structure) and its individual properties (node attributes). For instance, social networks store friendship relations as edges and age, income, and other properties as attributes. These relationships and properties seem to be dependent on each other and exploiting these dependencies is beneficial, e.g. for community detection and community outlier mining. However, state-of-the-art techniques highly rely on this dependency assumption. In particular, community outlier mining [2] is able to detect an outlier node if and only if connected nodes have similar values in all attributes. Such assumptions are generally known as homophily [4] and are widely used. However, looking at multivariate spaces, one can observe that not all given attributes have high dependencies with the graph structure. For example, social properties such as income or age have strong dependencies with the graph structure of social networks [4]. In contrast, properties such as gender are rather independent from it. Consequently, recent graph mining algorithms degenerate for multivariate attribute spaces that lack dependency with the graph structure in some of the attributes. This calls for a general pre-processing step that selects subspaces, i.e. subsets of the attributes, showing dependencies with the graph. This talk covers several methods for the selection of such relevant subspaces in attributed graphs: As first method, ConSub [3] proposes the statistical selection of congruent subspaces, i.e. subsets of attributes showing a dependency with the graph structure. A core challenge in selecting these subspaces lies in the modeling of dependence between graph structure and attribute values. Further, one has to ensure that congruent subspaces are selected only if there is sufficient evidence on this dependence. ConSub addresses all those problems by: (1) a novel measure for the degree of congruence between a set of node attributes and a graph by means

Read the paper · More papers on PaperTik