Board 268: Enhancing Zero-Shot Learning of Large Language Models for Early Forecasting of STEM Performance
Ahatsham Hayat, Sharif Akil, Helen Martinez, Bilal Khan, Mohammad Rashedul Hasan · 2024
This paper introduces an innovative application of conversational Large Language Models (LLMs), such as OpenAI's ChatGPT and Google's Gemini, for the early prediction of student performance in STEM education, circumventing the need for extensive data collection or specialized model training.Utilizing the intrinsic capabilities of these pre-trained LLMs, we develop a cost-efficient, training-free strategy for forecasting end-of-semester outcomes based on initial academic indicators.Our research investigates the efficacy of these LLMs in zero-shot learning scenarios, focusing on their ability to forecast academic outcomes from minimal input.By incorporating diverse data elements, including students' background, cognitive, and non-cognitive factors, we aim to enhance the models' zero-shot forecasting accuracy.Our empirical studies on data from first-year college students in an introductory programming course reveal the potential of conversational LLMs to offer early warnings about students at risk, thereby facilitating timely interventions.The findings suggest that while fine-tuning could further improve performance, our training-free approach presents a valuable tool for educators and institutions facing resource constraints.The inclusion of broader feature dimensions and the strategic design of cognitive assessments emerge as key factors in maximizing the zero-shot efficacy of LLMs for educational forecasting.Our work underscores the significant opportunities for leveraging conversational LLMs in educational settings and sets the stage for future advancements in personalized, data-driven student support.However, the prerequisites for such fine-tuning-extensive datasets and computational resources, along with domain-specific machine learning expertise-often pose significant barriers.This paper explores an alternative pathway by investigating the use of pre-trained LLMs to infer STEM students' end-of-semester performance without the need for extensive data collection or model fine-tuning.Our focus is on a training-free approach that leverages the inherent zero-shot learning capabilities of LLMs, which, despite their proven effectiveness across various tasks [17,18,19], have yet to be thoroughly examined within the context of STEM education forecasting.Employing conversational LLMs, specifically OpenAI's ChatGPT 3.5 and 4.0 [20] and Google's Gemini [21], we aim to demonstrate how these tools can forecast STEM students' performance early in the semester with minimal cost.Our research is driven by two primary research questions (RQs):• RQ1: To what extent are conversational LLMs effective as zero-shot learners in STEM education, i.e., in the early forecasting of STEM performance?• RQ2: How can the zero-shot learning performance of LLMs be enhanced in the STEM education domain?We collected data from 48 first-year college students enrolled in an introductory programming course.This dataset includes a range of features, from students' background information and socio-economic status to their cognitive and non-cognitive attributes, all of which we translate into natural language text suitable for LLM processing.Our analysis assesses the capability of LLMs to make early semester performance predictions based on data sequences of varying lengths and at different granularity levels.Addressing RQ1, we undertake a comprehensive analysis of the zero-shot predictive power of LLMs within the academic sphere, focusing on two key dimensions.Primarily, we explore the temporal aspect of prediction-determining how early in a semester LLMs can provide accurate forecasts of student performance.To this end, we analyze data spanning three distinct timeframes: 2-week, 4-week, and 8-week intervals.The selection of these intervals allows us to assess the efficacy of LLMs at different stages of the semester.Secondly, we aim to identify the optimal granularity for performance prediction by LLMs at these specified intervals.We categorize student performance into following three progressively detailed levels, facilitating a nuanced analysis of LLMs' ability to differentiate between students who are at risk and those who are prone to risk, as well as their capacity to discern various degrees of risk among students.This approach enables us to address critical questions regarding the timing and precision of LLM-based forecasts, such as the earliest point at which LLMs can effectively predict student risk levels, and how accurately they can distinguish between students at different risk levels.• Two types: at-risk or prone-to-risk (grade below B-), and average or outstanding (grade Bor above)• Three types: at-risk or prone-to-risk (grade below B-), average (grade B-or above but below A-), and outstanding (grade A-or above)