On opportunities and challenges of large multimodal foundation models in education
Stefan Küchemann, Karina E. Avila, Yavuz Dinc, Chiara Hortmann, Natalia Revenga, Verena Ruf, Niklas Stausberg, Steffen Steinert, Frank Fischer, Martin Rudolf Fischer, Enkelejda Kasneci, Gjergji Kasneci, T. Kuhr, Gitta Kutyniok, Sarah Malone, Michael F. Sailer, Albrecht Schmidt, Matthias J. Stadler, J. Weller, Jochen Kühn · npj Science of Learning · 2025
Recently, the option to use large language models as a middleware connecting various AI tools and other large language models led to the development of so-called large multimodal foundation models, which have the power to process spoken text, music, images and videos. In this overview, we explain a new set of opportunities and challenges that arise from the integration of large multimodal foundation models in education.