Harnessing GANs for Innovative Voice Generation in AI Applications
R. Manimegalai, S. Sugumaran, S. Lakshmi, T. Lakshmibai, T. Dinesh Kumar, Manne Archana · 2025
Generative Adversarial Networks (GANs) have revolutionized artificial intelligence by enabling high-fidelity data generation across multiple domains. In speech processing, GANbased voice generation techniques have demonstrated remarkable improvements in realism, expressiveness, and adaptability. This paper explores the advancements in GAN-based voice synthesis, focusing on architectures such as WaveGAN, VoiceGAN, and MelGAN. We investigate their applications in text-to-speech (TTS) systems, voice cloning, and speech enhancement while addressing challenges related to training stability, mode collapse, and dataset biases. Additionally, we propose an optimized GAN framework for high-quality voice generation, validated through objective and subjective evaluations. Experimental results show that our method outperforms traditional models in naturalness, intelligibility, and diversity. This work contributes to the growing field of voice-based generative AI by providing insights into future trends and applications in human-computer interaction, assistive technologies, and digital content creation.