Pop Lyrics through Time: Challenges in Corpus-Based Modeling of Linguistic and Emotional Dynamics in German Pop Lyrics

Roman Schneider · 2026

This paper presents a large-scale diachronic analysis of German pop lyrics based on a linguistically rich, TEIencoded monitoring corpus.We describe multi-layer annotation and reproducible workflows for deriving higherlevel features at scale, including lexical diversity indices, a pronoun-based subjectivity measure, modal particle density, and a length-normalized sentiment intensity score.Particular attention is paid to the development and evaluation of pipelines for two notoriously challenging phenomena: modal particles and sentiment.For modal particles, we build a manually curated gold standard and train sequence models whose performance we relate to inter-annotator agreement.For sentiment, we integrate a lexicon-based resource with a dedicated human annotation experiment to assess reliability and alignment with expert judgments.On this basis, we investigate how structural and affective features co-vary in the corpus and how they change over time, showing, among other trends, declining lexical diversity and sentiment intensity alongside a slight increase in first-and second-person pronouns.Beyond the empirical findings, the paper highlights practical challenges in managing culturally specific corpora, and makes evaluation materials available to support transparent, reusable corpus-based research on popular music and related domains.

Read the paper · More papers on PaperTik