Methods For Text Style Processing
Tikhonov, Aleksei · Qucosa (Saxon State and University Library Dresden) · 2025
This dissertation investigates how to define, represent, and manipulate text style so that it can be used reliably across modern NLP applications. It treats 'style' as the distinctive manner of writing that spans multiple linguistic levels -- phonology, morphology, syntax, semantics, and discourse -- and argues that meaningful progress requires separating stylistic signal from semantic content. The work surveys and systematizes existing definitions, then proposes computational models that operationalize style for downstream tasks such as generation, summarization, and translation, with an emphasis on personalization in human–computer interaction. Methodologically, the thesis contrasts two dominant paradigms. The first is classifier-based modeling, which encodes style through predictive features (e.g., sentence length, lexical richness, punctuation patterns) and evaluates success by a classifier`s ability to distinguish styles or detect successful style transfer. The second is neural representation learning, which leverages RNNs and transformer architectures to learn vector spaces in which style can be isolated from meaning, enabling controlled editing of stylistic attributes without distorting content. The dissertation analyzes the strengths and limitations of both approaches, showing where hand-engineered features provide interpretability and where neural models offer flexibility and scalability. A central technical challenge addressed is disentanglement: ensuring that modifications to style leave semantic content intact. To that end, the work introduces novel style transfer methods that enforce content preservation through architectural constraints and training objectives, and it proposes new metrics that jointly assess (i) fidelity of style change, (ii) content consistency, and (iii) overall text quality and fluency. These metrics aim to reduce overreliance on proxy scores and to align evaluation with human judgments. Empirically, the dissertation studies a broad spectrum of style facets -- sentiment, formality, politeness, temporal and geospatial markers, authorship, domain, and tone -- acknowledging their frequent overlap in real texts. It documents the practical role of curated corpora for specific styles and the risks of confounding topic with style. Extensive experiments compare neural architectures, including transformers, for their capacity to encode and transfer stylistic attributes, and they highlight trade-offs between control, interpretability, and robustness. Beyond technical contributions, the thesis argues that style-aware NLP can make machine-generated text more authentic, persuasive, and culturally resonant, with tangible benefits for personalization in dialogue systems and assistive technologies. By clarifying definitions, advancing modeling techniques, and standardizing evaluation, the dissertation lays groundwork for principled, controllable style in NLP systems.