Using LSA to Automatically Identify Givenness and Newness of Noun Phrases in Written Discourse

Zhiqiang Cai, David F. Dufty, Arthur C. Graesser, Christian F. Hempelmann, Philip M. McCarthy, Danielle S. McNamara · eScholarship (California Digital Library) · 2005

Identifying given and new information within a text has long been addressed as a research issue.However, there has previously been no accurate computational method for assessing the degree to which constituents in a text contain given versus new information.This study develops a method for automatically categorizing noun phrases into one of three categories of givenness/newness, using the taxonomy of Prince (1981) as the gold standard.The central computational technique used is span (Hu et al., 2003), a derivative of latent semantic analysis (LSA).We analyzed noun phrases from two expository and two narrative texts.Predictors of newness included span as well as pronoun status, determiners, and word overlap with previous noun phrases.Logistic regression showed that span was superior to LSA in categorizing noun-phrases, producing an increase in accuracy from 74% to 80%.

Read the paper · More papers on PaperTik