Harvard CGA Geotweet Archive v2.0

Benjamin Lewis, Jain, Devika · Harvard Dataverse · 2012

Geotweet Archive v2.0 The Harvard Center for Geographic Analysis (CGA) maintains the Geotweet Archive, a global record of tweets spanning time, geography, and language. The primary purpose of the Archive is to make a comprehensive collection of geo-located tweets available to the academic research community. The Archive extends from 2010 to July 12, 2023 when Twitter stopped allowing free access to its API, transitioning API access to a paid model. The number of tweets in the collection totals approximately 10 billion, and it is stored on Harvard University’s High Performance Computing (HPC) cluster. The Harvard HPC supports many applications for working with big spatio-temporal datasets, including two geospatial tools recently deployed by the CGA: Heavy.ai, and PostGIS. The Geotweet Archive consists of tweets which carry two types of geospatial signature: 1) GPS-based longitude/latitude generated by the originating device 2) Place-name-centroid-based longitude/latitude from the bounding box provided by Twitter, based on the user-define place designation (typically a town name). Any tweet which carries one or both of these signatures is included in the Archive. Approximately 1-2% of all tweets contain such geographic coordinates, (this percentage needs verification and may vary over time). The current version of the Archive is Version 2.0. The original Version 1.0 archive began in 2012 as part of a project started by Ben Lewis of CGA and then Harvard graduate student Todd Mostak, to develop a GPU-powered spatial database. The database needed an interesting, large, spatio-temporal dataset to show off its capabilities. So Todd and Ben built a harvester with the goal of harvesting all tweets containing GPS coordinates coming out of the Twitter firehose. The first version of the GPU database was called GEOPS, and it powered TweetMap which ran within WorldMap and represented the first vector-based geospatial big data query and display platform. Eventually when GEOPS was not made open source, Ben Lewis and Apache developer David Smiley developed an open source version of GEOPS built on 2D faceting within Solr which was called "The BOP" (https://gis.harvard.edu/billion-object-platform-bop). Later GEOPS formed the basis for technology startup MapD Technologies, which then became OmniSci, and then Heavy.ai. Heavy.ai software now runs on Harvard’s High Performance Computing (HPC) environment to support interactive exploration and analytics with the Geotweet Archive and any other large datasets. Version 2.0 of the geotweet archive represents the results of a merge between the CGA archive, and an archive developed by the Department of Geoinformatics at the University of Salzburg in Austria, as well as several other archives lead by Ben Lewis of Harvard CGA. Clemens Havas and Bernd Resch at University of Salzburg, worked with Devika Jain of Harvard CGA, to deploy Version 2.0. ======================================================== Schema of Geotweet Archive v2.0 Field name____TYPE____Description message_id----BIGINT----Tweet ID tweet_date----TIMESTAMP----Date and time of tweet from Twitter (utc) tweet_text----TEXT ENCODING----Text content of tweet tags----TEXT ENCODING DICT----Tweet hashtags tweet_lang----TEXT ENCODING DICT----Language that the tweet is in source ----TEXT ENCODING DICT----Operating system or application type used to create the tweet place*----TEXT ENCODING NONE----The geographic place as defined by the user, usually a town name. A bounding box determined by Twitter based on this field, from which centroids (see longitude and latitude fields) and the spatial_error field are derived, and used when not overridden by a GPS coordinate. See Twitter tweet object for place. retweets ----SMALLINT----Number of retweets as of last time it was checked tweet_favorites----SMALLINT----Now known as ‘likes’ photo_url----TEXT ENCODING DICT----URL of any image referenced quoted_status_id ----BIGINT----ID number for quote status user_id ----BIGINT----User ID number user_name----TEXT ENCODING NONE----User name user_location*----TEXT ENCODING NONE----User defined location, usually a city or town. See Twitter user object. followers ----SMALLINT----Followers as of the last time checked friends ----SMALLINT----Number of users followed by this user user_favorites----INT----Number of topics the user is interested in status----INT----Code for what user is doing as of last time it was checked user_lang----TEXT ENCODING DICT----User defined language latitude----FLOAT----Latitude from GPS or bounding box based on Place field longitude----FLOAT----Longitude from GPS or bounding box based on Place field data_source*----TEXT ENCODING DICT----The source crawler or dataset for the tweet gps----TEXT ENCODING DICT----Flag for whether lon/lat is from GPS or town name bounding box (SRID – 4326). When both are present, the GPS coordinate takes priority. spatialerror----FLOAT----Estimate in meters horizontal error for lon/lat coordinate. 10m for GPS coordinates, error for bounding boxes calculated as radius of circle with area of bounding box. ===================================================== *data_source____Code U. Salzburg REST API crawler----1 Harvard CGA streaming crawler----2 U. Salzburg streaming API crawler----3 Ryan Qi Wang and Harvard Medical School datasets----4 U. Heidelberg dataset----5 Archive.org dataset----6 ---------------------------------------------------------------------------------------------- Note: Before April of 2015 the default for GPS coordinate capture was turned on for Twitter users. After this date users have had to opt-in to share their precise location. This is one reason for the large decrease in volume of geotweets after this date. A number of automated tweet-bots have been discovered which generate tweets with (apparently) randomly spoofed coordinates. These bots appear to make up no more than a few percent of harvested geotweets. A list of bot sender names so far discovered is here. Tweets from these sender names were not generated by a human with a mobile device since they are randomly scattered across the globe:see our current list of tweet bots. If you are interested in accessing the archive please, please fill out CGA contact form. Before requesting or receiving Tweet IDs, requestors must agree to Twitter's Terms of Service, Twitter's Privacy Policy, Developer Agreement" and Developer Policy. Tweet data provided by CGA may only be used for not-for-profit research and for academic purposes. Recipients may not share CGA provided tweet IDs or tweets or content derived from them without written permission from the CGA. CITATIONS: If you use the Geotweet Archive in your research please reference it: "Harvard Center for Geographic Analysis Geotweet Archive, (https://doi.org/10.7910/DVN/3NCMB6). For examples of geospatial systems capable of querying, visualizing, analyzing millions or even billions of objects, please see the Heavy.ai Enterprise and PostGIS database platforms which are running on the Harvard FAS Research Cluster . Note for Non-Harvard Researchers : We are unable to share the full raw tweets with anyone outside of Harvard as per Twitter's content redistribution policy. We are not longer harvesting since July 12, 2023 due to change in Twitter's policy.

Read the paper · More papers on PaperTik