On-Device Evaluation of Joint Speech-Text Embeddings for ASR, TTS, and Speaker Recognition Tasks
Michael Gian Gonzales, Peter Corcoran, Naomi Harte, Michael Schukat · IEEE Access · 2025
Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) are key elements of speech based features and user interfaces (UI) for many modern consumer technologies, ranging from embedded AI assistants to AI chatbots and most recently the latest generation of wearables such as smart-glasses. Growing demand for privacy-preserving, speech-enabled consumer devices is pushing ASR, TTS and speaker verification from the cloud to the edge, where memory and compute budgets are tight. This paper introduces a single, joint speech-text architecture that unifies acoustic and textual representations in a shared embedding space and supports three downstream tasks: ASR, TTS and speaker recognition, without task-specific encoders. The model replaces the Connectionist Temporal Classification (CTC) decoder used in earlier work with an RNN-Transducer, boosting ASR accuracy while reducing overall parameter count to 86.08 M. Trained on 460 h of LibriTTS, it attains 10.8% word error rate (WER) / 3.96% character error rate (CER), 78.1% speaker-ID accuracy, and a 7.14 dB MCD with predicted mean opinion score (MOS) on naturalness of 2.9. On a Raspberry Pi 5 (CPU-only) the system runs close to real time (RTF 0.83 for ASR; 1.34 for TTS), demonstrating practical on-device deployment. Beyond the primary tasks, the learned embeddings achieve 90% audio-text retrieval accuracy on clean LibriTTS/Librispeech splits and 98.3% digit classification on AudioMNIST, highlighting their versatility. These results show that a single, compact model can deliver competitive multimodal speech performance while staying within the strict resource envelope of embedded hardware.