The Wheres and Whyfores for Studying Textual Genre Computationally

Jussi Karlgren · 2004

Stylistic variation in text is not incidental but an integral part of the intended and understood communication: the content and form of the message cannot be divorced. Discussing style, however is difficult, without descending into specifics or idiosyncracies. Mostly, readers report stylistic differences in terms of genres. Genres, while vague and undefined, are wellestablished and talked about: very early on, readers learn to distinguish genres. Differences readers are aware of are mostly based on utility- not on textual characteristics per se. By textual statistics, it is easy enough to establish that there are observable differences between genres we find in collections of textual material. However, using the textual, lexical, and other linguistic features we find to cluster the collection without anchoring information in usage will risk finding statistically stable categories of data without explanatory power or utility. This chapter argues for more informed target metrics for the statistical processing of text collections. Much as operationalized relevance proved a useful goal to strive for in information retrieval, research in textual stylistics, whether application oriented or philologically inclined, needs goals formulated in terms of pertinence, relevance, and utility — notions that agree with reader experience of text. This brief paper gives an example of statistical stylistic experimentation and argues for more informed measures of variation and choice and more informed measures of readership analysis to be able to posit dimensions of textual variation usefully. Variation in text Texts are much more than what they are about. Authors make choices when they write a text: they decide how to organize the material they have planned to introduce; they make choices between synonyms and syntactic constructions; they choose an intended audience for the text. Authors will make these choices in various ways and for various reasons: based on personal preferences, on their view of the reader, and on what they know and like about other similar texts. These choices are governed by a range of constraints,

Read the paper · More papers on PaperTik