Amazon Mechanical Turk for Subjectivity Word Sense Disambiguation

Cem Akkaya, Alexander Conrad, Janyce Wiebe, Rada F. Mihalcea · University of North Texas Digital Library (University of North Texas) · 2010

Amazon Mechanical Turk (MTurk) is a marketplace for so-called “human intelligence tasks ” (HITs), or tasks that are easy for humans but currently difficult for automated processes. Providers upload tasks to MTurk which workers then complete. Natural language annotation is one such human intelligence task. In this paper, we investigate using MTurk to collect annotations for Subjectivity Word Sense Disambiguation (SWSD), a coarse-grained word sense disambiguation task. We investigate whether we can use MTurk to acquire good annotations with respect to gold-standard data, whether we can filter out low-quality workers (spammers), and whether there is a learning effect associated with repeatedly completing the same kind of task. While our results with respect to spammers are inconclusive, we are able to obtain high-quality annotations for the SWSD task. These results suggest a greater role for MTurk with respect to constructing a large scale SWSD system in the future, promising substantial improvement in subjectivity and sentiment analysis. 1

Read the paper · More papers on PaperTik