Telugu Vakyalu: Spoken Telugu Sentences for IoT Applications
Parabattina Bhagath, V. Sree Lasya, P Dhyeya, Pradip K. Das · 2023
Spoken data sets are very critical components for speech research in various languages. The new developments in this area heavily relies on the availability of data sets. The unavailability of data is not prominent in well-developed spoken languages, but it is seen in the case of low and under resourced languages. Indian languages are considered low-resource languages with respect to the accessibility of speech. Internet of Things (IoT) is a field of research where speech processing can contribute a large portion in developing the interfaces. This requires speech recognition frameworks that deal with the IoT related problems. This paper releases a speech corpus that can help researchers to develop spoken dialog or interfaces for smart devices. The paper discusses the methodology used in developing spoken Telugu sentence dataset.