Preserving subsegmental variation in modeling word segmentation (or, the raising of baby Mondegreen)
C. Anton Rytting · OhioLink ETD Center (Ohio Library and Information Network) · 2007
Many computational models have been developed to show how infants breakapart utterances into words prior to building a vocabulary-the "word segmentation task."Most models assume that infants, upon hearing an utterance, represent this input as a string of segments.One type of model uses statistical cues calculated from the distribution of segments within the child-directed speech to locate those points most likely to contain word boundaries.However, these models have been tested in relatively few languages, with little attention paid to how different phonological structures may affect the relative effectiveness of particular statistical heuristics.