Exploring Large Language Models for Automated Essay Grading in Finance Domain

Garima Malik, Mücahit Çevik, Sojin Lee · 2024

This study explores the application of large language models (LLMs) for the automated grading of essays in the finance domain. The focus is on generating grades for six Assessment Indicators (AIs) related to finance and accounting for each essay. Our research highlights the potential of LLMs and showcases custom prompt engineering's effectiveness in a domain-specific Automated Essay Scoring (AES) task. We propose two distinct prompting techniques: unified and discrete. The unified technique generates grades for all AIs using a single comprehensive prompt, while the discrete technique employs separate prompts for each AI. To enhance the effectiveness of these models, we apply In-Context learning through One-shot and Few-shot methods. Through extensive experimentation, we show that LLMs outperform fine-tuned BERT-like baselines, demonstrating consistency and generalizability in their results. However, challenges remain with output post-processing and the cost of processing input tokens.

Read the paper · More papers on PaperTik