Learning Interpretable Style Embeddings via Prompting LLMs
Ajay Patel, Delip Rao, Ansh Kothary, Kathleen R. McKeown, Chris Callison-Burch · 2023
Style representation learning builds contentindependent representations of author style in text.To date, no large dataset of texts with stylometric annotations on a wide range of style dimensions has been compiled, perhaps because the linguistic expertise to perform such annotation would be prohibitively expensive.Therefore, current style representation approaches make use of unsupervised neural methods to disentangle style from content to create style vectors.These approaches, however, result in uninterpretable representations, complicating their usage in downstream applications like authorship attribution where auditing and explainability is critical.In this work, we use prompting to perform stylometry on a large number of texts to generate a synthetic stylometry dataset.We use this synthetic data to then train humaninterpretable style representations we call LISA embeddings.We release our synthetic dataset (STYLEGENOME) and our interpretable style embedding model (LISA) as resources.Generation: The author is using a conversational style of grammar.The author is using short, simple sentences.The author is using language that is informal and direct.The author is expressing enthusiasm for the topic in a straightforward manner.The author is using contractions, such as "I'll".The author is using a casual tone.The author is emphasizing their interest in the topic with the phrase "really cool".The author is using the present tense to express anticipation for the future.