Transformers and Generative AI
Gabriela Arriagada-Bruneau, Claudia López, Marcelo Mendoza · 2025
In this chapter we introduce the Transformer architecture, which underpins modern advancements in Generative AI. Originating with applications like machine translation, text summarization, and classification, Transformers have propelled innovations in Generative AI, enabling technologies such as chatbots and the creation of hyper-realistic images and videos. At the core of this success is the self-attention mechanism, which allows Transformers to encode long-range dependencies and process inputs efficiently through parallelisation and multi-head attention. The architecture’s versatility extends to encoding (BERT) and decoding (GPT) applications. Accordingly, here we examine key developments like BERT’s masked language modelling for context-dependent embeddings, and GPT’s causal language modelling for generative tasks, and we further highlight the integration of human feedback through Reinforcement Learning from Human Feedback (RLHF). After highlighting its technical advancements and transformative potential, we also discuss the risks of transformers and generative AI, including issues about bias, misinformation, privacy violations, and overreliance on AI.