Development challenges in automatic speech recognition for computer assisted pronunciation teaching and language learning
James Salsman · 2014
Automation has been improving the productivity of almost all human endeavors, recently including language instruction, but also has been displacing human labor, and threatens to exacerbate limited but pedagogically essential student-teacher interaction opportunities and teacher demand. Several methodological challenges in the development of automatic speech recognition for computer assisted pronunciation teaching (CAPT) range from the need to support instead of displace human teachers, software engineering issues, the availability of appropriate technology to support clinical speech language pathology applications, addressing vastly larger numbers of less literate students, and the use of open source self-study applications to attract and retain students for human tutoring. Each of these challenges can be addressed, but potential solutions range widely in difficulty, time and resource requirements, costs and benefits. Software engineering issues such as testing and validation of enhancements and new features have occupied more time and required more effort than most of our other development issues combined. The use of physiologically neighboring phonemes, diphones, and segment durations are all likely to help improve learner outcomes. The language instructor’s experience of computer-assisted pronunciation assessment can be enhanced by offering comparisons of students’ utterances to exemplar pronunciations for each of those attributes.