MemStrata vs Mem0, Graphiti, and Letta: Same-Stack Comparison on Temporal Axes
Apples-to-apples local 7B evaluation: MemStrata’s deterministic supersession hits 1.000 on MemArch supersession/poisoning/TEMPO axes where flat RAG and several agent-memory systems still serve stale values.
Why same-stack measurement matters
Vendor memory scores are often incomparable: different models, judges, and harnesses. Our comparison grid runs competitors on the same local model (Qwen2.5-Coder-7B via Ollama, nomic-embed, temp 0) with fresh collections / wiped graphs per run.
We care about axes where architecture shows: evolving code mutation, world-knowledge updates, MemArch supersession, poisoning resistance, TEMPO as-of-T queries, and deterministic abstention on unknown facts.
What the grid shows
On MemStrata’s moat columns (code mutation, MemArch supersede, as-of point-in-time, poisoning, TEMPO), the default config is config-invariant at 1.000 accuracy with 0.000 stale where measured — because flags never touch deterministic active/retired routing.
Measured competitor snapshots (same stack, see FINAL_COMPARISON_MATRIX):
- Mem0 — strong published long-context recall; supersession on our MemArch export ~0.05 accuracy (serves stale ~95%) without a true retire model.
- Graphiti / Zep-class — bi-temporal graph story, but local 7B runs lag on code mutation and world-knowledge evolving facts.
- Letta / MemGPT — agentic archival memory; setup-heavy; no structural supersession → 0 by construction on some axes; strong on cue-rich synthetic reads where raw passages help.
- Naive / advanced RAG — competitive static recall; structural failure on evolving knowledge and stale-fact-error.
Honest caveats (we publish them)
Cue-rich synthetic temporal exports can be “read” by models that keep raw passages — they do not alone prove a memory time model. Marker-free and code-mutation suites are the harder tests. LongMemEval on a local 7B is a weak-baseline floor (naive RAG ~0.25), not comparable to frontier published numbers.
MemStrata’s differentiation is not “best at every recall leaderboard.” It is determinism on temporal validity: never confidently wrong on superseded facts, with a structural UNKNOWN tier for facts never asserted.
Product takeaway
If you need long-dialogue reminiscence, several open systems compete. If you need an agent that will not resurrect last week’s API path after a migration, you need a ledger with supersession — not a bigger embedding index.