Comparative Analysis of Speech Recognition Open API Error Rate

Ju‐Young Kim, Dai Yeol Yun, Oh Seok Kwon, Seok-Jae Moon, Chi-Gon Hwang · International journal of advanced smart convergence · 2021

Speech recognition technology refers to a technology in which a computer interprets the speech language spoken by a person and converts the contents into text data. This technology has recently been combined with artificial intelligence and has been used in various fields such as smartphones, set-top boxes, and smart TVs. Examples include Google Assistant, Google Home, Samsung's Bixby, Apple's Siri and SK's NUGU. Google and Daum Kakao offer free open APIs for speech recognition technologies. This paper selects three APIs that are free to use by ordinary users, and compares each recognition rate according to the three types. First, the recognition rate of numbers and secondly, the recognition rate of Ga Na Da Hangul are conducted, and finally, the experiment is conducted with the complete sentence that the author uses the most. All experiments use real voice as input through a computer microphone. Through the three experiments and results, we hope that the general public will be able to identify differences in recognition rates according to the applications currently available, helping to select APIs suitable for specific application purposes.

Read the paper · More papers on PaperTik