Can it remember across many sessions?
LongMemEval measures whether information remains useful across extended, interactive histories.
Research
HOM is measured on remembering across long conversations, answering from that history, and resisting information designed to mislead retrieval. The results are published.
What we are testing
Three complementary evaluations ask whether the memory works in ordinary use and when an attacker tries to corrupt what gets retrieved.
LongMemEval measures whether information remains useful across extended, interactive histories.
LoCoMo tests whether the system can recover facts, preferences, and events from lengthy conversations.
PoisonedRAG examines whether targeted malicious content can displace the evidence a user should receive.
MutMem
HOM’s first paper studies a practical problem: useful memory must adapt when evidence changes, but that adaptation should not become an invisible rewrite. MutMem focuses on making those changes traceable.
Publication standard
Results are presented with complete run sizes, model roles, dataset identity, timing, and cryptographic verification artifacts. Partial runs remain labeled as partial. The repository and paper are public: results regenerate from a self-hashed aggregate bound by SHA-256.
Independent scrutiny is welcome
The repository and paper are public. Clone the system, inspect the evidence chain, and run the benchmarks yourself.