F-LoRA-QA: Finetuning LLaMA Models with Low-Rank Adaptation for French Botanical Question Generation and Answering
Institut de Systématique, Evolution, Biodiversité (ISYEB), Sorbonne Université, Paris, France, Ayoub Nainia, Muséum National d’Histoire Naturelle, Paris, France, CNRS, EPHE-PSL, Université des Antilles, Paris, France, Régine Vignes‐Lebbe, Institut de Systématique, Evolution, Biodiversité (ISYEB), Sorbonne Université, Paris, France, Muséum National d’Histoire Naturelle, Paris, France, CNRS, EPHE-PSL, Université des Antilles, Paris, France, HAJAR MOHAMED MOUSANNIF, Jihad Zahir, UMMISCO, IRD, France · International conference Recent advances in natural language processing · 2025
Despite recent advances in large language models (LLMs), most question-answering (QA) systems remain English-centric and poorly suited to domain-specific scientific texts.This linguistic and domain bias poses a major challenge in botany, where a substantial portion of knowledge is documented in French.We introduce F-LoRA-QA, a fine-tuned LLaMA-based pipeline for French botanical QA, leveraging Low-Rank Adaptation (LoRA) for efficient domain adaptation.We construct a specialized dataset of 16,962 question-answer pairs extracted from scientific flora descriptions and fine-tune LLaMA models to retrieve structured knowledge from unstructured botanical texts.Expert-based evaluation confirms the linguistic quality and domain relevance of the generated responses.Compared to baseline LLaMA models, F-LoRA-QA achieves a four-fold improvement in BLEU, a 70% ROUGE-1 F1 gain, a 16.8% increase in BERTScore F1, and an Exact Match improvement from 2.01% to 23.57%.These results demonstrate the effectiveness of adapting LLMs to low-resource scientific domains and highlight the potential of our approach for automated trait extraction and biodiversity data structuring.