Remind: Recall Enhanced Memory Integration for Natural Language Dialogue Systems
Bo-Sheng Huang, Ted T. Kuo, Li-Jen Wang, Chia‐Yu Lin · 2025
Large Language Models (LLMs) lack robust memory management for multi-turn dialogues, limiting their effectiveness in personalized applications. We introduce REMIND, a lightweight, modular framework integrating Short-Term Memory (STM) for session-specific recall and Long-Term Memory (LTM) for persistent knowledge retention. REMIND employs a hybrid retrieval mechanism that combines vector search, hierarchical grouping, iterative query refinement, and keyword filtering to enhance recall precision without modifying LLM architectures. To benchmark REMIND, we introduce MemoRA, an open dataset for evaluating dynamic memory retrieval in multi-turn dialogues. Our evaluation combines static metrics with Kernel Density Estimation (KDE) and GPT-40-mini as a relevance judge, showing that REMIND outperforms MemGPT in recall scenarios. Results indicate that LLaMA3.3:70B demonstrates weaker grouping capabilities, while DeepSeek-V3 is a cost-effective reasoning engine. By open-sourcing the MemorA dataset [1] and REMIND's implementation [2], we provide resources to advance research in memory-augmented dialogue systems, enabling personalized, scalable applications in healthcare, tutoring, and customer support. While current evaluations focus on the MemoRA and DMR datasets, future work will extend assessments to broader benchmarks.