Cross-Modal Generation of Visual and Auditory Content

Feng Gao, Mengting Liu, Ying Zhou · 2024

In the early stages of Artificial Intelligence (AI) development, research mainly focused on areas such as image classification, object detection, machine translation, and sentiment analysis. With the emergence of AI technologies such as Generative Adversarial Networks (GAN), Transformer, Contrastive Language-Image Pre-training (CLIP), Diffusion, pre-trained models, and multimodal learning, Artificial Intelligence Generated Content (AIGC) has a significant impact on basic productivity tools and has gradually attracted public attention. AIGC aims to achieve novel content generation through deep learning and computer technologies that mainly involve visual and auditory modalities, including image, video, speech, music, etc. Visual and auditory modalities play important roles in human perception and processing of real-world information. Endowing intelligent agents with more understanding and perception ability of different modalities has become a new challenge in AIGC. With the vigorous development of cross-modal technologies, visual and auditory content generation has made remarkable progress in text-to-vision, text-to-audio, and vision-to-audio tasks. From the perspective of classic challenges of cross-modal generation, this chapter summarizes the unimodal representation, multimodal representation, and alignment of different modalities. This chapter also systematically reviews the research progress and main methods of cross-modal generation of visual and auditory content in recent years according to various tasks. Finally, to deepen the understanding of cross-modal generation from the data perspective, this chapter introduces high-quality vision-related and audio-related datasets.

Read the paper · More papers on PaperTik