Content style transfer using RAG
FreiDok plus (Universitätsbibliothek Freiburg) · 2026
Retrieval-Augmented Generation (RAG) pipelines are traditionally used to augment Large Language Models (LLMs) with external, factual knowledge. However, their potential applications in Content Style Transfer , the process of preserving a message’s meaning while adapting it to different style attributes remain largely unexplored. This thesis explores In-Context Learning by dynamically retrieving stylistic context, allowing the model to adapt to an author’s unique text style without the need for computationally expensive fine-tuning. Using a subset of the Enron email corpus, we establish an experimental setup that compares two reconstruction pipelines against original ground truth emails. The first, a Vanilla Pipeline, reconstructs emails using only the de-stylised semantic content. The second, a RAG Pipeline that induces style with retrieved relevant historical emails and stylistic attributes such as lexical preferences and syntactic structures extracted from the author’s historical writing. We evaluate these pipelines using an LLM-as-a-Judge framework, measuring Style transfer accuracy, content preservation and naturalness of the generated text. The research further validates this approach through a user study, in which participants generate emails using the pipeline and rate them based on three core metrics (Style Transfer, Meaning Preservation and Naturalness) for a specific recipient. The results demonstrate that RAG-driven stylistic augmentation markedly improves the Style Transfer that standard models often fail to capture. This work demonstrates that retrieval mechanisms can effectively condition generation on user specific stylistic preferences, extending their applicability beyond factual grounding. The code for experimentation with the Enron dataset and the web application for the user study is publicly available at https://github.com/BebopCode/content-style-transfer-using-rag.