Preliminary Tasks of Word Embeddings Comparison of Unaligned Audio and Text Data for the Kazakh Language
Zhanibek Kozhirbayev, Talgat Islamgozhayev, Altynbek Sharipbay, Arailym Serkazyyeva, Zhandos Yessenbayev · 2023
In this paper, we present our preliminary work on the topological analysis of audio and text data for unsupervised speech processing for the Kazakh language. Our research is motivated by the advancements in generative models and their potential impact in this domain. We make the assumption that phoneme frequencies and contextual relationships exhibit similarity in both the acoustic and text domains for a given language. In order to delve into this hypothesis, we carried out experiments utilizing the variational autoencoder (VAE) framework at the word level. This was undertaken as an initial step for comparing word embeddings in unaligned audio and text data specific to the Kazakh language.. We trained VAE models on both acoustic and text data, with the aim of extracting the encoding part from the acoustic VAE and the decoding part from the text VAE. In the next stage of our planned research, we intend to apply persistent homology methods to analyze the topological structure of these two latent vector spaces. This analysis will provide insights into the underlying patterns and relationships between the audio and text data.