Multimodal Content Generation: A Project for AI Creativity
Geetha M P, Hariharan S, T Jeevanprasanth, Lavanya Sanjay · 2024
21st-century interactions involve the consumption of content, whether in the form of text, audio, images, or any other medium. This paper presents a multimodal content generation (MCG) system that seeks to ease the content production process by providing integrated solutions for video, image, audio and text content production among others. Large language models (LLMs) have shown a great promise in enhancing the creativity process in various fields such as media and education where mediums of communication interweave. The system takes into account the current data, situation, user needs, and composes an appropriate and suitable content for them. At the same time, conventional AI models, such as transformers, are combined with multimodal models enhanced with optimization strategies such as transfer learning and fine-tuning. This strategy greatly augments the content generation potential and scalability of the system and thus, is of great importance in the design of systems for the content production process. The quantum- inspired elements of the hybrid model are the ones targeted towards making changes in decision and feature selection within the creative processes. The effectiveness of MCG in this research was compared to other content creation systems, and the results of the research revealed that MCG is more creative, efficient, and satisfying for the users than the existing systems; therefore, it has a great potential for changing the practices of producing content in a multimedia environment.