Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Yadav, Deshraj · arXiv (Cornell University) · 2025

Agent memory systems that scale overwrite protection by a domain volatility prior V_d treat two different quantities as one number: a slow belief about how often a kind of fact changes, and a fast residual about whether this particular observation was unexpected. This paper separates them. A homeostatic law charges V_d in both the evidence score and the threshold. An allostatic law drops V_d from the score and instead scales the threshold by residual surprise — distance from predicted mismatch rather than from stored text. A composite gate uses the allostatic law only for an explicit high-mismatch correction or an unexpected residual against a learned expectation, and otherwise keeps the homeostatic insurance. On scripted probes with oracle domain and mismatch, dropping V_d from the score recovers explicit recency-shift (an entrenched career change) as a cliff at exponent p=0, not a blend. The same drop produces 20% more false updates under the classifier’s real error structure, because the double V_d charge was insurance against a mislabeled stable trait paired with weak evidence. The composite gate matches the homeostatic false-update rate (94.9%, 0.93 false updates / 18) while keeping the recency-shift win. Defining surprise as leftover mismatch after anticipation makes a predicted weak stream go quiet; catching that stream is a sleeptime job on time-decayed belief mass, not a live EMA of raw mismatch — sixteen daily weak mentions supersede overnight, while the same sixteen spread monthly do not. The overwrite law is only reached after a match. On end-to-end remember(), decision error once routed is about six points; similarity and linking dominate the error budget. Topic similarity is non-separable for must-link versus must-not-link pairs. A two-stage recall-then-verify step with a conservative local model takes irreversible errors to zero on a combined update-plus-coexist harness. We do not claim a public-benchmark win. We claim a measured decomposition: prior and residual are different jobs, a switch beats a blend, and linking sits in front of both laws. This is an empirical companion to “Volatility-Adjusted Memory Protection” (doi:10.5281/zenodo.21962419). Results use VoltMem 0.4.0. This work does not claim to implement consciousness or allostasis in Sterling’s physiological sense, and it is not a new continual-learning algorithm.

Read the paper · More papers on PaperTik