A Speech2Text based Recognizer and Probabilistic Parser Application for TeluguSentences
Surisetty Hima Varshini, Gottimukkala Sarayu Varma, Meena Belwal · 2024
Telugu, the largest Dravidian language, is one of the Indian languages that is predominantly spoken in Andhra Pradesh, Telangana and a few parts of southern India. The proposed work presents a speech-to-text based recognizer cum parser for some simple sentences in Telugu which are commonly used in day-to-day conversations. The Probabilistic Cocke-Younger-Kasami (P-CYK) algorithm has been employed in order to recognize these simple sentences and generate the corresponding probabilistic parse for these sentences. These parses will be useful in gaining a better understanding about the grammatical structure of a given sentence. At present, there are not many linguistic processing methodologies which are integrated with speech-based inputs for regional languages like Telugu. Hence, an effort has been made in order to efficiently use speech as the form of input and obtain the parse for some basic sentences with the corresponding probability of that parse. The probabilistic parser that has been developed in this work takes speech in Telugu as the input and identifies both grammatically correct and incorrect sentences and produce the probabilistic parse of the grammatically correct sentences. This system was then assessed by making use of the metrics like accuracy score, precision score and recall score. Also, a user interface was developed using Streamlit for the same.