Computational style processing
Robert Levinson, Foaad Khosmood · 2011
Our main thesis is that computational processing of natural language styles can be accomplished using corpus analysis methods and language transformation rules. We demonstrate this first by statistically modeling natural language styles, and second by developing tools that carry out style processing, and finally by running experiments using the tools and evaluating the results. Specifically we present a model for style in natural languages, and demonstrate style processing in three ways: Our system analyzes styles in quantifiable terms according to our model (analysis), associates documents based on stylistic similarity to known corpora (classification); and manipulates texts to match a desired target style (transformation). In our model, we view style markers as the building blocks of styles. Style markers are features that can be extracted from the text which directly or indirectly reflect choices made by the authors. In order to perform stylistic analysis, we have developed a full marker taxonomy, a large library of marker instances and associated marker extraction routines. To perform style-based classification, we extract a vector of markers from documents and corpora. Using nearest neighbor methods, we compare vectors to each other and calculate distances, and use the distances to perform supervised classifications. We derive a weighted distribution of markers for like-labeled document collections which we consider the mathematical representation of their collective style. We further experiment with varying the number of target classes and find a general decline in training convergence ability with increased number of classes. In addition we find that some markers are more persistently useful than others in our experiments. We also demonstrate machine transformation of texts with detectable styles. Much like Statistical Machine Translation (SMT) from one language to another, we demonstrate that machine transformation of one style to another is achievable. A number of monolingual text-to-text transformation routines (transforms) must be available and applied intelligently and systematically to a document. We have developed a small number of such transforms to demonstrate the possibilities for style transformation. We show that a finite number of transformations on a document will lead to a change in the document's overall style. Our experimental results support our thesis that automatic style processing is empirically possible. As a final step we demonstrate significant correlations between our model of style and the common human understanding of style with a 19 subject user study.