Research

Memory should be measured. Not merely promised.

HOM is being tested on whether it can remember across long conversations, answer from that history, and resist information designed to mislead retrieval.

What we are testing

Useful over time. Reliable under pressure.

Three complementary evaluations ask whether the memory works in ordinary use and when an attacker tries to corrupt what gets retrieved.

LONG-TERM MEMORY

Can it remember across many sessions?

LongMemEval measures whether information remains useful across extended, interactive histories.

Evaluation in progress
CONVERSATION MEMORY

Can it answer from a long relationship?

LoCoMo tests whether the system can recover facts, preferences, and events from lengthy conversations.

Evaluation scheduled
MEMORY SECURITY

Can it resist poisoned information?

PoisonedRAG examines whether targeted malicious content can displace the evidence a user should receive.

Protocol prepared

MutMem

When memory changes, the history should remain accountable.

HOM’s first paper studies a practical problem: useful memory must adapt when evidence changes, but that adaptation should not become an invisible rewrite. MutMem focuses on making those changes traceable.

Earlier evidenceOriginal assessment
New evidenceRevised assessment
Both states remain inspectable

Publication standard

A number is not enough. The evidence must travel with it.

Results will be presented with complete run sizes, model roles, dataset identity, failures, timing, and cryptographic verification artifacts. Partial tests will remain labeled as partial.

Independent scrutiny is welcome

Read the method. Reproduce the run. Challenge the result.

Researchers can follow the technical work now and review the complete artifacts when the repository and paper are released.