Crowdsourcing of network data
Ding Wang, Prakash Mandayam Comar, Pang‐Ning Tan · 2016
A key requirement for supervised learning is the availability of sufficient amount of labeled data to build an accurate prediction model. However, obtaining labeled data can be manually tedious and expensive. This paper examines the use of crowdsourcing technology to acquire labeled examples for classifying network data. Unfortunately, creating human intelligence tasks (HITs) to enable crowdsourcing is cumbersome for network data and may even be prohibitive for privacy reasons. To overcome this limitation, we present a novel framework called surrogate learning to transform the network data into a new representation (i.e., images) so that the labeling task can be completed even by non-domain experts. We analyze the reconstruction error of the transformation and use the theoretical insights to provide guidance on how to develop an effective surrogate learning approach for any given network and source image corpus. We also performed extensive experiments using Amazon Mechanical Turk to demonstrate the efficacy of our approach on node classification problems.