Research

Memory should be measured. Not merely promised.

HOM is measured on remembering across long conversations, answering from that history, and resisting information designed to mislead retrieval. The results are published.

What we are testing

Useful over time. Reliable under pressure.

Three complementary evaluations ask whether the memory works in ordinary use and when an attacker tries to corrupt what gets retrieved.

LONG-TERM MEMORY

Can it remember across many sessions?

LongMemEval measures whether information remains useful across extended, interactive histories.

Published — 459/500, 91.8% LLM-judged
CONVERSATION MEMORY

Can it answer from a long relationship?

LoCoMo tests whether the system can recover facts, preferences, and events from lengthy conversations.

Published — 74.12% judged; 58.20 token F1
MEMORY SECURITY

Can it resist poisoned information?

PoisonedRAG examines whether targeted malicious content can displace the evidence a user should receive.

Published — 0/100 poisoned disclosures

MutMem

When memory changes, the history should remain accountable.

HOM’s first paper studies a practical problem: useful memory must adapt when evidence changes, but that adaptation should not become an invisible rewrite. MutMem focuses on making those changes traceable.

Earlier evidenceOriginal assessment
New evidenceRevised assessment
Both states remain inspectable

Publication standard

A number is not enough. The evidence must travel with it.

Results are presented with complete run sizes, model roles, dataset identity, timing, and cryptographic verification artifacts. Partial runs remain labeled as partial. The repository and paper are public: results regenerate from a self-hashed aggregate bound by SHA-256.

Independent scrutiny is welcome

Read the method. Reproduce the run. Challenge the result.

The repository and paper are public. Clone the system, inspect the evidence chain, and run the benchmarks yourself.