813Tongue, Language or Noise? Word Sense Disambiguation in Ancient Greek with Corpus-Based Methods
Wouter Mercelis, Toon Van Hal, Alek Keersmaekers · 2025
Corpus-based methods are underutilized in both intellectual history and the history of linguistics. This paper endeavors to demonstrate the potential for an automatically annotated corpus of Ancient Greek to enrich our understanding of intellectual history. It focuses on disambiguating the meaning of Ancient Greek words related to the concept of language by using corpus and natural language processing (NLP) methods. We adopt both a semasiological (meaning-focused) and onomasiological (word-focused) approach, with a primary focus on the terms γλῶττα and φωνή. To differentiate between their primary meanings, we employ both supervised and unsupervised techniques, relying on an ELECTRA model tailored to the Ancient Greek language. The results of our supervised approach indicate that a sample size of 150 sentences is sufficient to achieve stable precision and recall (around 0.90) in distinguishing the two main meanings of both γλῶττα and φωνή. Our initial attempt at using unsupervised techniques failed to clearly distinguish the two meanings of γλῶττα: the clusters formed were based on formal and morphological criteria, rather than semantic meaning. However, by applying a transformation to the original sentences, we were eventually able to plot fairly clear clusters based on meaning. This study is only a small step forward in the application of corpus-based methods to intellectual history. Further progress in unsupervised methods is necessary to further explore onomasiological approaches, offering promising perspectives for corpus-based investigations into intellectual history.