End-to-End Speech Processing: Denoising, Transcription, and Sentiment-Guided Translation with Domain-Specific Named Entity Recognition
D BHOOSHAN, Salini P · 2025
This paper presents an integrated pipeline for processing noisy speech, combining audio denoising, auto matic speech recognition (ASR) using Vosk, Wav2Vec 2.0, and Whisper, sentiment-guided translation, and domain-specific named entity recognition (NER) restoration. The application of DNS64 denoising significantly enhances transcription accuracy, leading to a 79.6% reduction in character error rate (CER) for the Wav2Vec model, lowering it from 0.4064 to 0.0829. For Whisper, it reduces the Levenshtein distance by 32.8%, decreasing from 506 to 340 edit operations. The refined output from Wav2Vec, once denoised, is translated into German with a formal tone, enriched with domainspecific terminology (e.g., "Jahrhunderte") and corrected for punctuation. Evaluation metrics include CER, WER, BLEU, perplexity, and Levenshtein distance for transcription, and METEOR (0.5488) and ROUGE-L (0.5477) for translation, both surpassing the baseline performance of mBART (ME-TEOR: 0.5200, ROUGE-L: 0.5100). Key contributions of the system involve adaptive voice activity detection (VAD), detailed Levenshtein-based error analysis, and the integration of domain-specific named entity recognition (NER), making the approach highly effective for multilingual speech-to-text tasks in historical domains.