# LongMemEval, with the conditions on the figures

Maintainer-reported LongMemEval session recall and end-to-end QA for FluctlightDB, frozen July 2026. Open harnesses, no independent reproduction.

LongMemEval asks whether an agent can find the right session, and then whether a reader can answer from it. The two figures this site prints are below. They are the same rows as the benchmarks page, from the same source. They are not a rank.

## What the two rows are

One row is session recall: did the right session come back at a small k. The other is end-to-end question answering on a locked run, with a reader and a judge named in the detail line. Both are maintainer-reported. The harnesses are in the repository. The runs were frozen in July 2026. Nobody outside the project has reproduced them. This note does not compare either figure with a score from another product.

## What this note is not

It is not a new measurement. Nothing was re-run for this page. The conditions sit next to the numbers on purpose: maintainer-reported, frozen July 2026, harnesses open, no independent reproduction yet. Evidence recall on LoCoMo is a different benchmark and a different note.

## Figures

- Maintainer-reported
- Frozen July 2026
- Harnesses open
- No independent reproduction yet

- 97.6% — LongMemEval-S session recall. at k=8 · 488 of 500
- 97.4% — LongMemEval end-to-end QA. 487 of 500 · locked run, gpt-4o reader and judge

## Related

- [The same figures, on the benchmarks page](/benchmarks)
- [The LoCoMo note, including the withdrawn headline](/notes/locomo)
- [The word, defined](/glossary#longmemeval)
- [BENCHMARKS.md](https://github.com/voxmastery/FluctlightDB/blob/main/docs/BENCHMARKS.md)
