Notes
LongMemEval, with the conditions on the figures
LongMemEval asks whether an agent can find the right session, and then whether a reader can answer from it. The two figures this site prints are below. They are the same rows as the benchmarks page, from the same source. They are not a rank.
What the two rows are
One row is session recall: did the right session come back at a small k. The other is end-to-end question answering on a locked run, with a reader and a judge named in the detail line. Both are maintainer-reported. The harnesses are in the repository. The runs were frozen in July 2026. Nobody outside the project has reproduced them. This note does not compare either figure with a score from another product.
What this note is not
It is not a new measurement. Nothing was re-run for this page. The conditions sit next to the numbers on purpose: maintainer-reported, frozen July 2026, harnesses open, no independent reproduction yet. Evidence recall on LoCoMo is a different benchmark and a different note.
The figures this site still prints
- Maintainer-reported
- Frozen July 2026
- Harnesses open
- No independent reproduction yet
- 97.6% LongMemEval-S session recall. at k=8 · 488 of 500 Source: benchmarks/results/longmemeval-colab-v2-full-2026-07-04.json. Opens the repository in a new tab.
- 97.4% LongMemEval end-to-end QA. 487 of 500 · locked run, gpt-4o reader and judge Source: benchmarks/results/e2e-cert-paper-v2-2026-07-07.json. Opens the repository in a new tab.
The same rows, with the other measurements, are on the benchmarks page. The harness file is BENCHMARKS.md.