Integrated Parallel System for Audio Conferencing Voice Transcription and Speaker Identification

Ke Miao, Oloff C. Biermann, Zhen Miao, Simon Leung, Jianhong Wang, Keke Gai · 2020

In response to the request from a well-known international financial corporation, an integrated system prototype was architected and implemented to automatically record corporate audio conferencing, transcribe the recordings to text while identifying speakers, and compile the transcription and identification results into text-based meeting minutes, which then gets sent as meeting summary email attachments as well as saved into a meeting management database. Three technology focuses of this integrated system are discussed in this paper 1) Selection of a 3rd-party audio transcription and identification API (Audio API) through prototyping, factor comparison, and considering the existing technology environment at the corporation. 2) Optimize the adoption of the selected Audio API based on knowledge of Natural Language Process (NLP) methods. 3) Support asynchronous scheduling and processing of concurrent meetings using parallel computing architecture methods. The completed system was evaluated and shown to have met all the requirements from the corporation, perform well in audio language intelligent processing and multi-threaded parallel execution. With further enhancements, we foresee this system solution has good commercial values and has potential to be adopted widely among other businesses.

Read the paper · More papers on PaperTik