Using imitation to learn infant-adult acoustic mappings
G. Ananthakrishnan, Giampiero Salvi · 2011
Abstract This paper discusses a model which conceptually demonstrateshow infants could learn the normalization between infant-adultacoustics. Themodelproposesthatthemappingcanbeinferredfrom the topological correspondences between the adult and in-fant acoustic spaces, that are clustered separately in an unsu-pervised manner. The model requires feedback from the adultin order to select the right topology for clustering, which is acrucial aspect of the model. The feedback is in terms of anoverall rating of the imitation effort by the infant, rather thana frame-by-frame correspondence. Using synthetic, but contin-uous speech data, we demonstrate that clusters, which have agood topological correspondence, are perceived to be similarby a phonetically trained listener.IndexTerms: infantspeechacquisition, unsupervisedlearning,self organizing maps 1. Introduction An infant learning to communicate with speech has been asource of intrigue and interest, both from the point of view ofpsychology and medicine, and from the point of view of speechresearch. Understanding this phenomenon is especially diffi-cult, given that infants do not remember the process once theygrowupandtheremaybelimitedmeansofcommunicatingwithinfants while the process of learning takes place.There are several challenges an infant faces when tryingto acquire the ability to speak. The first and the most impor-tant challenge is learning the sensorimotor mapping betweenthe acoustics and the articulatory configurations [1, 2, 3]. Mostofthetheoriesintheabovestudiesmakeuseofthephenomenonofbabblingtoexplainthisprocess. Secondly, aninfantlearningto speak needs to learn how to categorize the different soundsthat he/she produces or hears from the adults [4, 5].An infant also needs to understand how to correlate the dif-ferentsoundsproducedduringbabblingtothesoundspresentinadult speech which is the main focus of our study. There havebeen several proposals about which acoustic correlates are in-variant between infant and adult speech (e.g., [6]). These mea-sureswereevaluatedby[7]and[8]andwereshowntobeusefulinacquiringtheabilitytoimitatethesoundsproducedbyadults.However, are these invariant acoustic features that normalizeadult speech, intrinsically and instinctively known to children,or do they learnt it? While most studies have assumed this to beinnate to children, Plummer