Large Language Models and Multimodal Systems

J. M. Bishop, Gabriel Seiberth · 2025

A large language model (LLM) is a machine learning model, typically based on a transformer architecture, that is constructed to generate natural language text from a prompt using large amounts of training data. 1 By analogy, imagine that an LLM is like a highly skilled librarian who has read millions of books in many languages. The librarian does not need to understand each word in every book because they have seen enough examples of how words and sentences usually fit together. When you ask a question or provide a prompt, the librarian does not simply look for an exact answer in one book. Instead, they use patterns and knowledge they have gained from all the books they have read to piece together the most relevant and coherent response. Similarly, LLMs use patterns from enormous amounts of data they&s;ve been pretrained on to generate cohesive and appropriate responses based on the prompts they receive. Although initially developed for natural language processing applications, the process of encoding words and tokens as embeddings can also be applied to other modalities (images, video, audio, etc.). This adaptability explains why LLMs are—as multimodal systems—particularly significant in the context of AVs.

Read the paper · More papers on PaperTik