P-GELU: A Novel Activation Function to Optimize Whisper for Darija Speech Translation
Maria Labied, Abdessamad Belangour, Mouad Banane · IEEE Access · 2025
Activation functions play a critical role in optimizing deep learning models, directly influencing gradient flow, convergence stability, and overall translation accuracy. In this work, we investigate their impact within the Whisper-Turbo model, a speech-to-text Transformer trained from scratch on the Darija-C dataset for Moroccan Darija speech translation. Our study begins by evaluating baseline activation functions—GELU, Swish, and Mish—demonstrating that while GELU is widely used in Transformer-based architectures, it may not be optimal for dialectal speech translation, particularly in low-resource settings. To address this limitation, we introduce Parameterized GELU (P-GELU), a novel activation function that extends GELU by incorporating trainable parameters (±, ²), allowing the model to dynamically adjust its non-linearity across layers and training phases. Through extensive experiments, we demonstrate that P-GELU outperforms standard GELU, with superior convergence speed and generalization. Furthermore, P-GELU reduces training loss, improves feature retention, and enhances linguistic adaptability, making it a more effective alternative for speech translation tasks involving phonetic variability, code-switching, and limited training data. The proposed P-GELU offers a promising balance between computational efficiency and performance gains, presenting a viable solution for enhancing Transformer-based speech models in low-resource language scenarios.