Can it remember across many sessions?
LongMemEval measures whether information remains useful across extended, interactive histories.
Research
HOM is being tested on whether it can remember across long conversations, answer from that history, and resist information designed to mislead retrieval.
What we are testing
Three complementary evaluations ask whether the memory works in ordinary use and when an attacker tries to corrupt what gets retrieved.
LongMemEval measures whether information remains useful across extended, interactive histories.
LoCoMo tests whether the system can recover facts, preferences, and events from lengthy conversations.
PoisonedRAG examines whether targeted malicious content can displace the evidence a user should receive.
MutMem
HOM’s first paper studies a practical problem: useful memory must adapt when evidence changes, but that adaptation should not become an invisible rewrite. MutMem focuses on making those changes traceable.
Publication standard
Results will be presented with complete run sizes, model roles, dataset identity, failures, timing, and cryptographic verification artifacts. Partial tests will remain labeled as partial.
Independent scrutiny is welcome
Researchers can follow the technical work now and review the complete artifacts when the repository and paper are released.