Googleology is Bad Science
Adam Kilgarriff · Computational Linguistics · 2007
The World Wide Web is enormous, free, immediately available, and largely linguistic.As we discover, on ever more fronts, that language analysis and generation benefit from big data, so it becomes appealing to use the Web as a data source.The question, then, is how.The low-entry-cost way to use the Web is via a commercial search engine.If the goal is to find frequencies or probabilities for some phenomenon of interest, we can use the hit count given in the search engine's hits page to make an estimate.People have been doing this for some time now.Early work using hit counts include Grefenstette (1999), who identified likely translations for compositional phrases, and Turney (2001), who found synonyms; perhaps the most cited study is Keller and Lapata (2003), who established the validity of frequencies gathered in this way using experiments with human subjects.Leading recent work includes Nakov and Hearst (2005), who build models of noun compound bracketing.The initial-entry cost for this kind of research is zero.Given a computer and an Internet connection, you input the query and get a hit count.But if the work is to proceed beyond the anecdotal, a range of issues must be addressed.First, the commercial search engines do not lemmatize or part-of-speech tag.To take a simple case: To estimate frequencies for the verb-object pair fulfil obligation, Keller and Lapata make 36 queries (to cover the whole inflectional paradigm of both verb and noun and to allow for definite and indefinite articles to come between them) to each of Google and Altavista.It would be desirable to be able to search for fulfil obligation with a single search.If the research question concerns a language with more inflection, or a construction allowing more variability, the issues compound.Secondly, the search syntax is limited.There are animated and intense discussions on the CORPORA mailing list, the chief forum for such matters, on the availability or otherwise of wild cards and 'near' operators with each of the search engines, and cries of horror when one of the companies makes changes.(From my reading of the CORPORA list, these changes seem mainly in the direction of offering less metalanguage.)Thirdly, there are constraints on numbers of queries and numbers of hits per query.Google only allows automated querying via its API, limited to 1,000 queries per user per day.If there are 36 Google queries per single 'linguistic' query, we can make just 27 linguistic queries per day.Other search engines are currently less restrictive but that may arbitrarily change (particularly as corporate mergers are played out), and also Google has (probably) the largest index-and size is what we are going to the Internet for.Fourthly, search hits are for pages, not for instances.Working with commercial search engines makes us develop workarounds.We become experts in the syntax and constraints of Google, Yahoo, Altavista, and so on.We become 'googleologists'.The argument that the commercial search engines provide