Generating Evaluation Criteria of Domain-Specific Large Language Model Using Word Vector Clustering

Jingyu Yang, Xu Hao, Rongxiao Wang, Xuran Ming, Shoubin Li · 2024

In recent years, the use of large language models for domain-specific tasks has gained increasing attention. However, the performance of current domain-specific large language models tends to be assessed through manual evaluation, which limits the broader application of this technology. Addressing the limitations of existing methods for evaluating domain-specific large language models, we introduce a frame-work capable of generating evaluation criteria from multiple perspectives using large language models. This framework elicits a plethora of candidate outcomes from large language models through multi-turn dialogues. It employs word vector clustering methods to eliminate outliers caused by hallucination issues, thereby identifying the most valuable generated results. We illustrate our approach by applying it to financial and legal consulting tasks. Through experimentation, we identify the dimensionality reduction and clustering methods that yield optimal performance in these tasks. An experiment demonstrated the process of this framework, proving its feasibility. The results indicate that our framework is capable of effectively generating comprehensive evaluation criteria, providing guidance for further testing of Domain-Specific Large Language Models.

Read the paper · More papers on PaperTik