LLMs and GPTs
Robert H. Chen, Chelsea Chen · 2024
A LargeLlanguage Model (LLM) is a giant artificial neural network (ANN) that has trillions of parameters and performs accelerated self-supervised learning from data searched from the Internet using a Common Crawl search engine. Generative Pre-trained Transformer-3 (GPT-3), an LLM that unsupervised scours the Internet in response to a user&s;s textual request (prompts). Transformer architecture encoder–decoder maps an input sequence to a series of representations, assemblages of related information that can broaden the scope, reduce the search, and better fit the context of the word. Transformer architecture is based on cognitive attention by distinguishing “soft” weights for each word in context or sequentially. “Soft weights” can change after each epoch, but “hard weights” are pretrained and frozen. The generative output of the model assumes its linear dependence on its own previous values and a stochastic term to form a recurrence relation autoregressive model with discriminative and NLP fine-tuning to compose prose and poetry, translate, and generate images, music, computer programs, and almost any literal thing that a user requests. ChatGPT has 1.6 × 10 12 parameters and 800GB of memory to perform those tasks.