Novelty Detection as a Tool for Automatic Detection of Orthographic Transcription Errors

Michele Gubian, Barbara Schuppler, Joost van Doremalen, Eric P. Sanders, Lou Boves · 2009

Making accurate orthographic transcriptions is very time-consuming and in the case of extemporaneous speech of native and non-native speakers the task is extremely difficult. While previous research focused on evaluating phonemic transcrip-tions, the goal of our research is the automatic detection of tran-scription errors on the orthographic level, which degrade the quality of every following annotation level. Since it is hard to statistically characterize a bad transcription, we use a Novelty Detection approach to model accurate transcriptions only and use models of good transcriptions to reject all inputs that do not fit. A hand-segmented corpus of spontaneous speech is used to build models of correct transcriptions. The speech material is first subjected to a forced alignment; then two features, viz. du-ration and acoustic score from the ASR aligner, are extracted from each aligned phone and used for training and detection. A simple likelihood threshold method is employed on the align-ment data in order to flag an utterance as incorrectly transcribed. We compare two different lexicons and discuss different issues with our approach to error detection. 1.

Read the paper · More papers on PaperTik