Relationship discovery and tracking in dynamic noisy environments

Kanad Ghose, James Rosswog · 2013

This dissertation is focused on the problem of detecting and tracking groups of related entities in large, noisy, time-varying data sets. We considered two specific instances of this problem: a) Detecting and tracking small groups of individuals moving in a coordinated manner through a large crowd b) Identifying computers infected with botnet malware form network traffic, where the sparse traffic related to botnets are inundated by normal network traffic. In this dissertation, we propose, develop, and evaluate techniques that are able to accurately and efficiently detect and track persistent relationships in very noisy environments. These relationships are used to identify groups of coordinated entities in environments that contain much larger sets of noise entities that essentially disguise the coordination we are attempting to detect. The main contributions of this dissertation are as follows: (1) The introduction of the notion of persistent relationships to discriminate casual relationships from stable or persistent relationships that characterize the behaviors of the entities to be detected. Specifically, we target noisy environments where the number of casual relationships far outnumber the actual persistent relationships of interest. Persistent relationships are represented using a relationship graph. (2) The development of techniques and relevant algorithms for detecting and tracking clusters of entities that have a persistent relationship using a dynamically updated relationship graph. The techniques developed are highly accurate and are capable of real-time performance. (3) The experimental evaluation of the proposed algorithms and techniques on data representing individuals moving in a coordinated fashion within a crowd and real botnet traces superposed on campus network data collected over a 2-week period. Our experimental evaluations demonstrate that the proposed techniques, especially techniques based on the use of support vector machines, are capable of meeting our original objectives with both a high-accuracy and low false positive rates.

Read the paper · More papers on PaperTik