Efficiency Comparison of Dataset Generated by LLMs using Machine Learning Algorithms

Premraj Pawade, Mohit Rameshchandra Kulkarni, Shreya Naik, Aditya Raut, Kishor S. Wagh · 2024

The constantly expanding field of Large Language Models (LLMs) offers exciting opportunities for various domains. These powerful models, such as GPT-3.5, Bard, and Bing, can produce massive amounts of text-based data, creating new avenues for generating synthetic datasets. The primary focus of this research is to explore the effectiveness of LLMs in creating high-quality, structured datasets for different ML applications. Specifically, this study concentrates on password strength prediction. It compares the performance of three prominent LLMs - Bard, ChatGPT, and BingAI - in generating datasets of text-based passwords with their corresponding strength levels. This research uses a diverse set of ML models, including traditional algorithms like XGBoost, Random Forest, etc., to evaluate the generated datasets. The evaluation process assesses their performance, generalization, and adaptability. This research contributes to the growing field of LLM-based data generation by demonstrating their effectiveness in creating valuable datasets for specific machine learning applications. The findings of this study pave the way for further exploration of LLMs' capabilities for diverse data types and tasks, potentially unlocking new avenues for advancements in various machine learning domains.

Read the paper · More papers on PaperTik