IoT system for real-time audio information processing
Oleh Osadchuk, Igor B. Olenych · Advances in Cyber-Physical Systems · 2025
This paper presents the development and inves- tigation of a speech-to-text conversion and speaker identi- fication system based on a Raspberry Pi microcomputer, designed for local audio data processing in environments with limited network connectivity. The system integrates Silero and WebRTC models for voice activity detection, SpeechBrain for speaker identification, and the Whisper family of models for speech recognition. In particular, a comparative analysis has been conducted on the efficiency of local speech processing using Whisper Tiny and Whisper Large 2 models versus cloud-based processing through the Whisper-1 and Whisper-1-en APIs (the latter applied exclu- sively to English-language speech). The study evaluates the impact of sentence length, processing time, memory consum- ption, and recognition accuracy on system performance. The advantages and resource-related limitations of the models in local and cloud-based IoT environments has been analyzed, and the feasibility of their application in real-time and data privacy contexts has been determined. Performance metrics of the models under various conditions has been used for the analysis.