Looking for Shakespeare: The Bard in Big Data
Aaron Rodriguez · Shakespeare · 2026
This essay examines the intersection of Shakespearean studies and generative AI. It contends that LLMs do not ‘know’ Shakespeare; rather, they computationally amplify the statistical biases of a literary ecosystem in which he is already overrepresented. By examining the threat of model collapse, the essay cautions against a probabilistic flattening of textual culture. To understand this training environment, the research presents an empirical study of Shakespeare’s statistical resonance within the 23.6-billion-word NOW Corpus (2010–2025). Using Python-driven n-gram analysis, this study identifies the frequency of Shakespearean phrasing in contemporary digital discourse. Findings challenge long-standing qualitative judgments: the line ‘to be or not to be’ is eclipsed by the idiom ‘too much of a good thing’. This demonstrates that Shakespeare’s enduring presence in big data is defined less by conscious literary reference and more by unconscious linguistic utility. The essay advocates for humanists to reclaim their role as ‘word scientists’, actively shaping the development of linguistic tools through interdisciplinary expertise in nuance, reference, and historical usage.