Collecting and Sharing Bilingual Spontaneous Speech Corpora: the ChinFaDial Experiment
Georges Fafiotte, Christian Boitet, Mark Seligman, Chengqing Zong · 2004
We describe here the three main platforms in the ERIM family of Web-based environments for human interpreting, two of them in more details, ERIM-Interp and ERIM-Collect, then ERIM-Aid.Each platform supports an aspect of the collecting or study of spontaneous bilingual dialogues, translated by an interpreter.ERIM-Interp is the core environment, providing mediated communication between speakers and human interpreters over the network.Using ERIM-Collect, French-Chinese interpreting data have been collected within the 3-year "ChinFaDial" project supported by LIAMA, a French-Chinese laboratory in Beijing.These "raw" speech data will be made available in the spring of 2004 on an open-access basis, using the DistribDial server, on a CLIPS-GETA website.Our goal is to extend such corpora, on a collaborative scheme, to allow other research groups to contribute to the site whatever annotations they may have created, and to share them under the same conditions (GPL).An ERIM-Aid variant is intended to provide focused machine aids to Web-based human interpreters, or to monolingual distant speakers conversing in different languages.