Real-Time Multilingual Speech Translation for Peer Communication

B K Meenakshi, Mohammed Wajahat Hussain, Mittapalli Arvind Sai · International Research Journal on Advanced Engineering Hub (IRJAEH) · 2025

Language continues to be a major obstacle to effective communication in a world that is becoming more interconnected by the day. This paper presented a real-time audio translation system that facilitates multilingual communication during peer-to-peer video calls. The application enables natural communication in the user’s preferred language by utilizing WebRTC for low-latency media transmission and incorporating sophisticated AI models such as Whisper for speech-to-text, GPT for language translation, and gTTS for text-to-speech synthesis. In addition to allowing real-time subtitle overlays and translated audio playback during conversations, the system supports five other languages: English, Hindi, Tamil, Telugu, and German. Low latency, scalability, and user-centric design are prioritized in the architecture, which is constructed with a Fast API backend and a React-based front-end. We address issues such as translation delays, synchronization, and audio buffering, and assess the system using user experience, latency benchmarks, and qualitative performance.

Read the paper · More papers on PaperTik