hashtagged English tweet collection filtered from Web Archive's downloads 26.03.2013-24.05.2017
Gevorg S. Poghosyan · Figshare · 2019
The 2.1GB .tar.gz file contains a 6.6GB (uncompressed) .txt file with 346'613'181 rows, each row of which is a tweet ID.This dataset is distilled from the The Internet Archive’s downloaded Twitter “Spitzer” stream (https://archive.org/details/twitterstream) by filtering the English tweets containing at least a single hashtag. This collection has periods when there were no tweets downloaded due to technical issues on the Web Archive's servers. These periods are 2014.01.01 - 2014.02.05, 2015.01.01 - 2015.03.22, 2015.11.16 - 2015.11.17, 2015.11.22 - 2015.11.30, 2015.12.06 - 2015.12.15, 2016.02.10 - 2016.02.14, 2016.02.27 - 2016.03.16. Hydrate the tweet ids with Twarc (https://github.com/edsu/twarc) and write to a file. You will need to provide Twarc with a set of Twitter API keys. twarc.py --hydrate tweet_ids.txt > tweets.json It is probably not a good idea to hydrate all the tweets in one go, and may be better to split the file into chunks and hydrate the tweets chunk-by-chunk.