Automatic language identification using sub-word models
Roger Tucker, Michael J. Carey, E.S. Parris · 2002
The paper describes initial experiments on automatic language identification with the particular aim of discriminating languages in the same language group. Subword models were built from the English, Dutch and Norwegian sections of the EUROM1 database using fully automatic segmentation based on TIMIT-derived models. Three techniques were then examined. In the first technique only acoustic differences between the phonemes of each language were used. The second technique relied on the relative frequencies of the phonemes of each language, while the third technique combined the two sources of information. The latter technique proved the best giving 97% accuracy for English vs. Dutch, and 90% across the three languages.>