Evaluating the Performance of Large Language Models in Classifying Numerical Data

Divya Mary Biji, Yong‐Woon Kim · 2024

In recent years, large language models (LLMs) such as BERT, DistilBERT, RoBERTa, and XLNet have revolutionized natural language processing (NLP) tasks due to their powerful representation capabilities. This research investigates the efficacy of these LLMs in the novel domain of numerical data classification by transforming numerical data into textual formats. The study involves converting numerical data points into descriptive strings and leveraging the advanced text processing capabilities of LLMs to classify them into predefined categories. We conduct a comparative analysis of BERT, DistilBERT, RoBERTa, and XLNet, evaluating their performance using standard metrics such as accuracy, precision, recall, and F1-score. Our experimental results demonstrate the potential of these models in accurately classifying numerical data, highlighting the strengths and limitations of each model. This research opens new avenues for applying LLMs beyond traditional NLP tasks, providing insights into their applicability for structured data classification.

Read the paper · More papers on PaperTik