Hidden Markov Models for Coding Story Recall Data

Michael A. Durbin, Jason Earwood, Richard M. Golden · eScholarship (California Digital Library) · 2000

Current methods of coding recall, summarization, talk-aloud, and question-answering data are inherently unreliable and not effectively documented.If the process of coding protocol data could even be partially automated, this would be an important scientific advance in the field of text comprehension.Twenty-four human subjects read and recalled each of four short texts.Half of the human recall data (the ''training data'') was coded by a human coder and then used to estimate the parameters of a set of Hidden Markov Models (HMMs) where each HMM was associated with a particular complex proposition in the text.The Viterbi algorithm was then used to assign the ''most probable'' complex proposition to humancoder specified text segments in the remaining half of the human recall data (the ''test data'').The HMM algorithm made coding decisions which agreed well with a human coder's decision on the test data indicating that the HMM is indeed capable of formally representing a human coder's "theory'' of how text segments should be mapped into complex propositions for simple texts.

Read the paper · More papers on PaperTik