Comparing automatic rich transcription for Portuguese, Spanish and English Broadcast News
Fernando Batista, Isabel M. Trancoso, Nuno Mamede · 2009
This paper describes and evaluates a language independent approach for automatically enriching the speech recognition output with punctuation marks and capitalization information. The two tasks are treated as two classification problems, using a maximum entropy modeling approach, which achieves results within state-of-the-art. The language independence of the approach is attested with experiments conducted on Portuguese, Spanish and English broadcast news corpora. This paper provides the first comparative study between the three languages, concerning these tasks.