Enhanced Large Language Models
J. R. V. Jeny, G. Bharath Babu, S. Jaganmohan Reddy, N. Muthukumaran, K. Abhiram · 2024
This work studies ways for improving the effectiveness of a language model in creating and understanding dialogue summaries. Two primary approaches are investigated, in the Complete Fine-Tuning step, a pre-trained FLAN-T5 model is fine-tuned using a dataset of dialog-summary pairs. The dataset is pre-processed to provide explicit instructions to the model, and the fine-tuning process involves training the entire model on the dataset. Qualitative and quantitative evaluations are then conducted to assess the effectiveness of the fine-tuned model in generating summaries. In the PEFT phase, a more computationally efficient method is employed, utilizing Low-Rank Adaptation (LoRA) to train a new layer/parameter adapter while keeping the underlying FLAN-T5 model frozen. This approach significantly reduces computational resources compared to Complete Fine-Tuning while maintaining comparable performance. The resulting model, equipped with the PEFT adapter, is evaluated qualitatively and quantitatively to measure its summarization capabilities. Both approaches sheds light on the compromises between computing efficiency and model performance.