FluctlightDB

Notes

The headline that counted turns the engine never retrieved

LoCoMo evidence recall asks whether the right conversation turn came back. An earlier headline for that measurement was withdrawn. The figures this site still prints are below, from the same source as the benchmarks page, with the same conditions.

What the withdrawn protocol did

The retracted figure expanded every retrieved turn by three neighbours on each side, then counted neighbours the engine never retrieved. A plain BM25 baseline also reaches about that same figure under the protocol, so it distinguished nothing. It is not a result. The wording the site uses for this is the retraction note under the table, and it is the only place that number should appear.

What the two remaining figures mean

The high-k figure is a lenient ceiling: a large candidate set, on ten conversations, with MiniLM-384, frozen in July 2026. The tight-k figure is the operational one, because a real prompt does not get the whole ceiling. Both are maintainer-reported. The harnesses are in the repository. Nobody outside the project has reproduced them. Evidence recall is not an LLM judge scoring an answer, and this note does not compare the two.

The figures this site still prints

  • Maintainer-reported
  • Frozen July 2026
  • Harnesses open
  • No independent reproduction yet

99.0% LoCoMo evidence recall at k=150. The retracted figure: The old figure expanded every retrieved turn by three neighbours on each side, then counted neighbours the engine never retrieved. A plain BM25 baseline also reaches about 99% under that protocol, so it distinguished nothing. It is not the headline any more, and it is not a number we will defend. Source: README.md:70 · CHANGELOG.md:49. Opens the repository in a new tab.

The same rows, with the other measurements, are on the benchmarks page. The harness file is BENCHMARKS.md.