A multilingual corpus for language identification

Lori F Lamel, Gilles Adda, Martine Adda‐Decker, Cristobal Corredor-Ardoy, Jean-Jacques Gangolf, Jean‐Luc Gauvain · 1998

In this paper we describe the design, recording, and transcrip-tion of a large, multilingual (French, English, German and Span-ish) corpus of telephone speech for research in automatic language identification. The corpus contains over 250 calls from native speakers of each language from their home country, and an ad-ditional 50 calls per language from another country. Although the same recording protocol was used for all languages, slight modi-fications were necessary to account for language or country speci-ficities. Issues in designing comparable corpora in different lan-guages are addressed, including how to interact with callers so as to obtain the desired responses.

Read the paper · More papers on PaperTik