A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous Interpretation

Seunghee Han, Gary Geunbae Lee, Hung Soon Kim, Sun‐Hee Kim, Minhwa Chung · 2025

This paper presents a novel multilingual speech translation corpus for complex, domain-specific content in Korean, English, Spanish, and Japanese. The corpus contains 4,000 hours of parallel speech, including 1,000 hours of Korean audio with simultaneous sight interpretations in the other three languages by 294 professionals (242 interpreters and 52 Korean voice actors). It also includes transcriptions, translations, and annotations for all languages. The Dewey Decimal Classification was adapted to balance knowledge representation, and speech tasks were conducted in a controlled studio environment to ensure data consistency. Translation, transcription, and annotation workflows were managed through a custom-built platform. The corpus captures nuanced contexts, cultural sensitivities, and domain-specific terminology, addressing linguistic challenges like structural differences between SOV (Korean, Japanese) and SVO languages (English, Spanish). Preliminary evaluations indicate its potential to enhance end-to-end speech translation models, support cross-lingual transfer learning, and tackle real-time translation issues.

Read the paper · More papers on PaperTik