SimpilAI: Fine-Tuning Large Language Models to Generate Humanized Educational Content for Computer Science Topics

Abhay Pratap, Arun Chauhan · 2024

Observations indicate that language models (LLMs) demonstrate strong performance in generating text that is both grammatically correct and contextually appropriate. However, a significant limitation of LLMs, particularly regarding educational applications in computer science, is the ineffective teaching disposition inherent in their design. While LLMs can produce technically correct information, they often fail to present content in a manner reflective of an educator's approach, particularly one with expertise in the field. Educational websites typically convey information in a structured and decisive manner, employing various proportional methods. Articles authored by content experts are crafted to transform complex subjects into accessible steps, anticipating the challenges that students may encounter and emphasizing concepts over mere terminology. This deficiency in LLM-generated content lies in the infrequent addressing of learners' preconceptions, often leading to the explanation of narrow, technical aspects through fragmented and uneducational methods. This paper introduces SimpilAI, a fine-tuned large language model based on TinyLlama-1.1B parameters, specifically developed for this educational purpose. SimpilAI has been trained on a curated dataset of over 3,000 computer science articles and blogs, designed to simulate the central objectives of the discipline. With effective exposure to such content, the model is capable of organizing topics logically, engagingly, and didactically for the audience, specifically the learners.

Read the paper · More papers on PaperTik