Computational narratology with transformer embeddings
William Rathje · Computational Humanities Research · 2026
Abstract This article presents a new method for analyzing narrative structure using sentence embeddings. We embed chunked text sequences to create narrative plot trajectory curves. This enables us to apply commonly used unsupervised word embedding techniques, such as semantic axes projection, to narrative sequence analysis at the scale of entire books. Our aim is to prototype several simple examples of the kinds of narratology techniques this method might enable. We apply a pretrained sentence transformer model to books, represented as embeddings of chunked sequences, from nearly 5 percent of the Project Gutenberg English Books Corpus filtered to works in the fiction category (totaling 2,965 texts). Texts date from the 1500s through the early 1900s, with the majority of texts concentrated in the nineteenth century. This period and dataset offer an opportunity to investigate the rise and development of the early novel form. Our analysis measures narrative trajectories along seven binary oppositions – class, gender, morality, collectivity/individual, nature/artifice, order/disorder, emotional/rational, and segments texts by decade, genre, and dialogue gender. We find trends broadly relevant to hypotheses from postclassical and feminist narratological theory and also present novel aggregate findings on narrative presentation and story shape.