LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

Di Wu, Wang, Hongwei, Wenhao Yu, Zhang, Yuwei, Kai‐Wei Chang, Yu, Dong · arXiv (Cornell University) · 2024

Audit kit for the memory-retrieval benchmark of Mnemosyne OS, a local-firstmemory system for AI assistants. This archive contains everything needed to audit the published scoreindependently: the per-question ledger for every run, the scoring andverification scripts, the methodology, and the raw run logs. Headline result: 72.9% on LongMemEval-M, full-haystack variant. It isreported as a LOWER BOUND, not a leaderboard score — full-haystack is theharder setting and is rarely the one reported. The kit exists so the numbercan be audited rather than trusted: verify.js re-runs the scoring over theraw ledger. Contents:- verification-kit/results/*.jsonl — per-question ledgers (48q baseline, 8q engine multi-session, 12q local sovereign)- verification-kit/verify.js and scoring.js — scoring and verification- verification-kit/METHODOLOGY.md — protocol, and what the result does not claim- fullhaystack-2026-07/ — raw run logs and campaign summary Licensing: documentation and data under CC BY 4.0; code under MIT.

Read the paper · More papers on PaperTik