Are Three Words All We Need? Recognizing Genre at the Sub-Sentential Level - eScholarship
Philip M. McCarthy, Stephen W. Briner, John C. Myers, Arthur C. Graesser, Danielle S. McNamara · Proceedings of the Annual Meeting of the Cognitive Science Society · 2008
Are Three Words All We Need? Recognizing Genre at the Sub-Sentential Level Philip M. McCarthy ([email protected]) Department of English, The University of Memphis, Memphis. TN 38152 Stephen W. Briner ([email protected]) Department of Psychology, DePaul University, Chicago, IL 60614 John C. Myers ([email protected]) Arthur C. Graesser ([email protected]) Danielle S. McNamara ([email protected]) Department of Psychology, The University of Memphis, Memphis. TN 38152 comprehension. If readers are indeed using different strategies to process different genres of text, then it is important to understand this process and potential information constraints during the course of genre identification. We ask five questions in the current study. First, how quickly (in terms of number of words) do readers identify the genre of a text? Second, what types of errors (i.e., genre misclassifications) do readers make when identifying genres? Third, does the process of genre identification depend on reading skill? Fourth, what textual features (e.g., syntax, lexical choice) influence genre identification? And fifth, can a computational model categorize genre as humans do, using information available in only the initial words of sentences? In McCarthy and McNamara (2007), we conducted a pilot study to provide a preliminary answer to our first two questions. Three experts (i.e. published authors) in the psychology of discourse processing were asked to identify the genre of isolated sentences culled from a corpus of narrative, history, and science texts. The experts had high inter-rater agreement (min = 90%) and required less than half the words in the sentence to accurately identify genres (accuracy as measured by F1, a standard index that considers both recall and precision: Narrative = .82; History = .84; Science = .82). The results further showed that these experts often classified many history and science sentences as narrative, suggesting that expository texts tend to be composed of a notable number of narrative-like sentences. On the other hand, science- like sentences were the least likely to be misclassified into other genres, suggesting the science-like sentences seldom occur in the non-science genres. The current study builds on the study conducted by McCarthy and McNamara (2007) by including a larger sample of participants, an independent assessment of reading ability, a measure of time on task, and recording accuracy in terms of number of words used. We also construct a computational model based on our results. We use the model to investigate what information could be present in the initial words of sentences such that it can provide participants with sufficient information to make a genre evaluation. The question of whether or not we could Abstract Genre identification is a critical facet of text comprehension, but very little is known about the process and information constraints of classifying texts by genres. In this study, higher- skill and lower-skill participants read 210 sentences from three genres. The words in the sentences were presented sequentially, one at a time. With each new word, participants decided whether the sentences came from a narrative, science, or history text. Both groups were able to correctly identify the genre by the third word of the sentence. Higher-skilled readers made their genre decisions more quickly and more accurately, and were also more precise in their selection of narrative texts. The study includes a computational model that uses text features from only the first three words of the sentences. The model reflects key features of the participants’ genre classifications. Keywords: genre recognition; reading skill; categorization Introduction Reading comprehension is greatly influenced by the genre of the text. Whether a text is a narrative, history, or science text influences the characteristics of the text, how the text is read, and can have a substantial influence on how well it will be understood (Bhatia, 1997; Graesser, Olde, & Klettke, 2002; Zwaan, 1993). More skilled readers utilize different strategies depending on the genre of the text (van Dijk & Kintsch, 1983; Zwaan, 1993) and training readers to recognize text structure helps to improve their comprehension (Meyer & Wijekumar, 2007; Oakhill & Cain, 2007; Williams, 2007). Once text genre is identified, it guides the reader’s memory activations, expectations, inferences, depth of comprehension, evaluation of truth and relevance, pragmatic ground-rules, and other psychological mechanisms. For example, readers are more likely to scrutinize whether an event actually occurred in a history text, whereas that is not a particularly relevant consideration in most narrative fiction (Coleridge, 1985; Gerrig, 1993). In contrast, stylistic attributes are more important in literary narratives than expository texts (Zwaan, 1993). Better understanding the nature of text genre and its effects on comprehension is important for theories of text comprehension as well as interventions to improve