Harmonizing AI: A GAN–Transformer fusion for expressive multimodal music synthesis in IoT systems

Yang Zhang, Shu Yu · Alexandria Engineering Journal · 2025

Synthesizing music that sounds convincingly realistic continues to pose significant challenges, largely due to the complex timbral nuances and expressive variations inherent in live performance. While traditional rule-based methods and many existing deep learning approaches have made progress, they often struggle to capture the subtle spectral textures and dynamic shifts that give real music its natural feel. In this work, we present an integrated framework for music sound synthesis that combines Generative Adversarial Networks (GANs), Transformer architecture, and multimodal data fusion to overcome the limitations of unimodal generative approaches. Unlike conventional methods that focus exclusively on either symbolic representations or audio signals, our model concurrently processes MIDI and audio data, extracting complementary features such as Mel-Frequency Cepstral Coefficients (MFCCs), pitch contours, and spectral envelopes. The alignment between MIDI and audio is carried out using beat-synchronous timestamp mapping. Specifically, note onset times in the MIDI data are aligned with corresponding audio frames based on temporal annotations provided by the dataset. This approach ensures accurate cross-modal synchronization, which is essential for effective feature fusion. These features are independently encoded through modality-specific encoders and subsequently fused using a cross-attention mechanism by allowing the system to capture fine-grained dependencies across modalities. The fused representation is then passed to a Transformer-based generator, which models long-range temporal dynamics essential for coherent musical structure. A GAN-based discriminator concurrently evaluates the authenticity of the synthesized output by guiding the generator towards more expressive and perceptually realistic audio. This combination of cross-modal alignment, temporal modeling, and adversarial training forms the core of a synthesis framework well suited to the demands of Internet of Things (IoT) applications, including smart speakers, context-aware sound controllers, and interactive ambient systems. Empirical evaluation across multiple benchmark datasets demonstrates the effectiveness of the proposed model. It achieved a Mean Opinion Score (MOS) of 8.1, indicating high perceptual quality, and a Fréchet Audio Distance (FAD) of 1.4, signifying close alignment with real audio distributions. The model further attained a pitch accuracy of 91.5% and an F1-score of 92.2%, confirming its ability to faithfully translate symbolic input into musically accurate output. Complementary performance indicators, including low spectral distance, minimal cycle-consistency loss, and high perceptual and diversity scores, such as reinforce the robustness of the approach. Our system combines symbolic and audio modalities to generate expressive music that closely aligns with human perception and performance.

Read the paper · More papers on PaperTik