Development of an Aerospace Engineering Evaluation Set for Large Language Model Benchmarking

Brian J. Connolly · 2025

Large language models (LLMs) have demonstrated human-level intelligence and problem solving that can be applied to many engineering tasks. These tasks have, until recently, been the sole domain of humans. Before widespread adoption of artificial intelligence for general engineering work, accurate assessments must be made of LLMs’ ability to robustly solve classes of problems. This work proposes the creation of a multimodal benchmark designed to evaluate Large Language Models’ (LLMs’) capabilities on a sampling of aerospace engineering test questions and tasks. The benchmark consists of True/False and short response questions. Various LLMs and prompting strategies, including different retrieval augmented generation (RAG) approaches, will be evaluated using this benchmark. The intent is to provide an approach to undertake a holistic assessment of LLM performance on a wide range of tasks that a human aerospace engineer might undertake.

Read the paper · More papers on PaperTik