Collecting and archiving tweets: a DataPool case study
Steve Hitchcock · 2013
Information presented to a user via Twitter is variously called a ‘stream’, that is, a constant flow of data passing the viewer or reader. Where the totality of information passing through Twitter at any moment is considered, the flow is often referred to as a ‘firehose’, in other words, a gushing torrent of information. Blink and you’ve missed it. But does this information have only momentary value or relevance? Is there additional value in collecting, storing and preserving these data? This short report describes a small case study in archiving collected tweets by, and about, a research data project, DataPool at the University of Southampton. It explains the constraints imposed by Twitter on the use of such collections, describes how a service for collections evolved within these constraints, and illustrates the practical issues and choices that resulted in an archived collection. The second version of the report adds a short postscript on rights, ethics and privacy of archiving Twitter data, prompted by a Twitter dialogue on this report. Two additional references of related work at Southampton are provided towards the end of sections 1 and 2.