A Review of the Development of Multimodal Large Models

<p>Zheng Miaomiao<sup>1</sup>, Gao Yi<sup>1,2,3</sup></p> · The Frontiers of Society Science and Technology · 2024

With the continuous advancement of deep learning technology, multimodal large language models built on large-scale language models and large-scale vision models have been making breakthroughs and achieving significant accomplishments in the field of natural language processing. The concept of general artificial intelligence and the explosive popularity of ChatGPT have brought large language models into people's daily lives. These models are typically based on the Transformer architecture, enabling them to handle and generate large amounts of text data while demonstrating strong language understanding and generation capabilities. As multimodal large models progressively enhance their language understanding and reasoning abilities, the application of instruction fine-tuning, context learning, and chain-of-thought tools has become increasingly widespread. This paper mainly analyzes the key technologies and development trends of multimodal large models, as well as the numerous challenges they face.

Read the paper · More papers on PaperTik