A RAG-Augmented LLM for Yunnan Arabica Coffee Cultivation

Zheng Chen, Zihao Jiang, Jianping Yang · Agriculture · 2025

Foundation models for agriculture often suffer from fragmented and stale knowledge, making it difficult to deliver stable, traceable answers. We present an evidence-grounded retrieval-augmented generation (RAG) system for Yunnan Arabica coffee cultivation. First, we curate a lightweight knowledge base (approximately 250k Chinese characters) from cultivation textbooks, technical guidelines, and reports. Second, we adopt a retrieve–rerank–generate workflow: semantic-aware chunking with stable identifiers [docid#cid]; hybrid retrieval fused by reciprocal rank fusion (RRF); cross-encoder reranking on top; and final answer generation by DeepSeek v3.1 with mandatory inline evidence tags. In addition, we use GPT-5 Thinking to synthesize 346 gold QA items on the corpus with document-/chunk-level citations, and we evaluate with citation-level per-sample macro precision/recall/F1. On this gold set, our optimized system attains a citation-level per-sample macro F1 of 0.813 (81.3%), significantly outperforming a Simple RAG baseline that reads only a vector store (0.573; 57.3%). Error analysis shows that residual errors are dominated by fragment mismatch and missing evidence; latency analysis indicates that end-to-end delay is primarily driven by generation, whereas retrieval, fusion, and reranking incur sub-0.1 s overhead. The workflow preserves traceability and verifiability, supports hot updates via index rebuilding rather than model fine-tuning, and we release scripts for corpus construction, ablation, and citation-based evaluation to facilitate reproducibility.

Read the paper · More papers on PaperTik