LoCoMo Benchmark Results 2026: OctaMem reaches 93.51% accuracy.
LoCoMo tests whether a memory system can answer questions about facts buried deep in a long, multi-session conversation, without stuffing the whole history into a prompt. Here is where OctaMem landed — including the part that did not work.

93.51%
1,440 of 1,540 questions answered correctly, computed directly from per-question judgements, with nothing rescaled or smoothed.
The trajectory
This is our third evaluation, and it is not a smooth creep.
It is one fix landing. Run over run, accuracy moved 5.21 percentage points, and almost all of that arrived at once.
| Run | Accuracy | Correct |
|---|---|---|
| Initial evaluation | 88.30% | 1,360 / 1,540 |
| Second evaluation | 90.62% | ≈1,396 / 1,540 |
| Final run (this report) | 93.51% | 1,440 / 1,540 |
Where the fix landed
Temporal reasoning was our worst category. It is now our best.
Both earlier evaluations flagged the same weakness: ordering, duration, and “as of when” questions. Temporal accuracy moved from 83.70% to 97.20%, a 13.5-point jump that accounts for almost the entire run-over-run gain. Every other category held within a point of where it was. The release did one thing, and did it without breaking anything else.
Error concentration
One conversation is carrying almost all our error.
We are not going to bury this, because it is the most useful number in the report. The 100 remaining errors are not spread across the benchmark: 95 of them sit in a single conversation out of ten, which scores 63/158 — 39.9%. Eight of the others score a perfect 100.0% and the ninth scores 96.8%. Outside that one conversation, the run answers 1,377 of 1,382 questions correctly, or 99.64%.
Conversation
A near-perfect run sitting next to one conversation that fails badly is not a diffuse capability gap. It is a signature that points somewhere specific. We are auditing it end to end, and applying the same scrutiny to the eight conversations that scored 100.0%, because a perfect score deserves exactly as much suspicion as a bad one.
What it costs
An accuracy number without a cost figure is half a claim.
Both of our earlier reports flagged missing cost telemetry as a gap. This run closes it.
3,519
Mean tokens / question
2,306
Answer-stage tokens
$0.00076
Mean cost / question
$1.17
Full 1,540-question run
Where OctaMem sits in the field
There is no comparison chart on this page. That is deliberate.
LoCoMo is the benchmark most memory vendors report against, including Zep, Mem0, Supermemory, and Letta, and it has become a crowded, noisy leaderboard. Before considering a comparison chart, we checked how those numbers are actually sourced. Published LoCoMo scores for the same system can vary by fifteen points or more depending on who measured it and under what harness, and new entrants publish their own “beats everyone” claim every few weeks. Supermemory’s public figure is a different metric altogether: recall rather than overall accuracy, so it cannot sit on the same axis as the others regardless of harness.
Until we have reproduced results against these systems in a shared, neutral setup, we are not putting a ranking chart next to our name. When that work is done, we will publish the comparison, numbers and methodology together.
We’d rather be the company whose numbers are never wrong than the company that was first.
Questions
What people ask about this score.
What is the LoCoMo benchmark?
LoCoMo (Long Conversational Memory) tests whether a memory system can answer questions about facts buried deep in long, multi-session conversations without loading the entire history into the prompt. It covers 1,540 questions across ten conversations, in four categories: single-hop, multi-hop, open-domain, and temporal.
What did OctaMem score on LoCoMo?
93.51% overall — 1,440 of 1,540 questions answered correctly. By category: temporal 97.20%, single-hop 94.29%, open-domain 93.75%, and multi-hop 86.88%. The figure is computed directly from per-question judgements, with nothing rescaled or smoothed.
How much does a LoCoMo run cost on OctaMem?
$0.00076 per question, or $1.17 for the full 1,540-question run, at a mean of 3,519 tokens per question with 2,306 of those in the answer stage.
Where do OctaMem's remaining LoCoMo errors come from?
They are highly concentrated. Of 100 remaining errors, 95 sit in a single conversation which scores 63/158, or 39.9%. Excluding that conversation, the run answers 1,377 of 1,382 questions correctly — 99.64%. We are auditing that conversation end to end.
Full methodology and per-question data are available on request. If you are evaluating memory layers, we will walk you through the harness.
Request the methodologyNew to these benchmarks? What LoCoMo, LongMemEval and BEAM actually measure explains what each one tests and how to read a published score. For the buyer’s view, see how to choose an AI memory layer.
