FluctlightDB

The full catalogue entry

Everything the landing page states briefly

The comparison, the three verbs with a runnable sample, the read path of one activate() call, every measurement with the conditions it was taken under, and the complete condition report. Every figure keeps the repository path it was read from.

Plate I

A brain for agents,
not a database for documents.

An agent restarts and forgets. So you bolt a session table to a vector index, write the glue that keeps the two agreed, and hope the embedding matches when the user phrases the question a different way. What you have built is retrieval. What the agent needed was memory.

The gap is not storage, it is the question being asked. A relational engine answers which rows match. A vector index answers what is nearest. Neither can answer the question an agent actually has, which is what it learned, and what this particular cue should bring back.

Agents today remember by scraping context windows and vector stores. That is lookup, not life.

Manifesto
An engraving of a library card-catalogue drawer pulled open, packed with index cards, a few of them fallen in front of it.
What four kinds of data store are natively asked, and the API each exposes
Store Unit The question it answers
Postgres rows, fixed schema Which rows match? SELECT
Chroma, Qdrant vectors What is nearest? vector_search()
Mem0-style layer facts scraped from chat What should we extract? app pipeline + index
FluctlightDB engrams with provenance What did the agent learn, and what should this cue recall? experience() / activate()
Plate II

Three verbs, and the third one is why it survives.

The whole contract is a write, a read and a commit. Everything else the engine does — consolidation, forgetting, provenance ranking — happens underneath those three calls rather than in a pipeline you assemble.

  1. experience()

    Write what happened

    A tool result, a user fact, an observation. It carries context, salience and provenance, not just text. Nothing is extracted from a transcript later; the agent records the moment it happens.

  2. activate(cue)

    Recall from a cue

    Full-text search, vector neighbours and graph spreading activation are fused inside one call. The question is not "what is nearest" but "what should this cue bring back".

  3. checkpoint()

    Survive the restart

    A write-ahead log and atomic checkpoints put the brain directory on disk. Crash recovery is exercised in CI with SIGKILL mid-write and torn WAL tails.

An engraving of three objects in a row: a quill pen, an upright brass seal stamp, and an ornate iron key.
from fluctlightdb import connect_embedded

brain = connect_embedded("./my-agent-brain")   # one directory per agent, in-process
brain.experience("User prefers dark mode", context="settings", salience=0.7)
brain.experience("Wallet balance is $42.10", context="ledger:wallet",
                 verified=True, source_uri="file://wallet.json")

hits = brain.activate("dark mode")             # cue -> ranked recalls
brain.checkpoint()                             # durable; survives restart
Run against the published package before this page shipped. Recall offline is lexical: activate("dark mode") returns the stored memory, activate("theme preference") does not, unless you pass a vector from your own embedder.
Plate III

What activate() actually does.

One call, three retrieval mechanisms, one ranking. Full-text search finds the literal tokens. An approximate vector index finds the neighbours. Spreading activation walks the graph outward from both, so a memory can be recalled by something it was merely stored beside. The results are fused once, not concatenated. Source: README.md:113. Opens the repository in a new tab.

An engraving of fine branches radiating outward from a single central point and thinning as they go, like nerve dendrites.

This is the read path from the paper, drawn to the same shape. Seeds are cheap and shallow; the spread is what makes recall feel associative rather than literal; the fuse is where provenance is applied, so a verified file outranks a remembered sentence.

The read path of a single activate call A cue enters at the left. It seeds two retrievers in parallel: a full-text FTS5 seed and an HNSW vector seed. Both light up engrams in the middle. Activation then spreads through the memory graph to neighbouring engrams. Finally all candidates are fused into one ranked list, with the top-ranked memory marked in red. cue FTS5 HNSW 1 2 3 4 the question seeds engrams graph spread fused rank
Figure — the read path of one activate(cue). Redrawn from the paper. Source: papers/figures/01-brain-hero.png. Opens the repository in a new tab. Run a cue against the engine
Plate V

Measurements, with the conditions attached.

These are the numbers the maintainer measured, on open harnesses, frozen in July 2026. Nobody outside the project has reproduced them yet. That is stated here because it is the first thing you would want to know. Source: docs/REPRODUCIBILITY.md:107. Opens the repository in a new tab.

Read the tight-k row before the headline row. Retrieving 150 candidates covers roughly eighteen per cent of a conversation and is a lenient ceiling; a real retrieval turn uses the top five to twenty. 72.6% at k=5 is the operational number, and it is the one worth arguing with.

Evidence recall is not question-answering accuracy. They are different measurements, and comparing this table against an LLM-judged QA score from another memory system would be meaningless. Source: docs/BENCHMARKS.md:151. Opens the repository in a new tab.

An engraving of a brass balance scale with two empty pans, a pair of calipers lying beside it.
Maintainer-reportedFrozen July 2026Harnesses openNo independent reproduction yet
Maintainer-reported benchmark results, frozen July 2026
96.8% LoCoMo evidence recall Source: benchmarks/results/locomo-invented-stack-engine-2026-07-13.json. Opens the repository in a new tab. at k=150 · 10 conversations, 1,982 questions · MiniLM-384
72.6% LoCoMo evidence recall Source: benchmarks/results/locomo-invented-stack-engine-2026-07-13.json. Opens the repository in a new tab. at k=5 — the operational number
97.6% LongMemEval-S session recall Source: benchmarks/results/longmemeval-colab-v2-full-2026-07-04.json. Opens the repository in a new tab. at k=8 · 488 of 500
97.4% LongMemEval end-to-end QA Source: benchmarks/results/e2e-cert-paper-v2-2026-07-07.json. Opens the repository in a new tab. 487 of 500 · locked run, gpt-4o reader and judge
0.646 BEIR SciFact nDCG@10 Source: benchmarks/results/paper-2026-07-09.json. Opens the repository in a new tab. against Chroma 0.645 in a shared harness
99.0% LoCoMo evidence recall at k=150 Source: README.md:70 · CHANGELOG.md:49. Opens the repository in a new tab. Retracted

The old figure expanded every retrieved turn by three neighbours on each side, then counted neighbours the engine never retrieved. A plain BM25 baseline also reaches about 99% under that protocol, so it distinguished nothing. It is not the headline any more, and it is not a number we will defend.

Plate VI

Condition report.

This is beta software with one maintainer. Here is the part of the page that most products leave out, written the way a collection catalogue records the state of an object.

An engraving of a single figure working at a bench beneath one hanging lamp in a vast unfinished vaulted hall, with scaffolding rising into the dark.

Stable

Semantic versioning applies. Pin to these.

  • connect_embedded(), connect_agent() and the other connect_* entry points
  • experience(), activate(), checkpoint(), recall()
  • The v4 brain directory on disk
  • POST /api/v1/experience and /api/v1/activate

Experimental

May change without a major version.

  • The Recall Fabric behind FLUCTLIGHT_FABRIC=1, off by default
  • Every neuroscience-named internal module
  • Governance, snapshots and retention policy
  • Auth and multi-tenancy — adversarially tested in CI, never third-party audited

Full policy in the stability document.

Plate VII

Where it plugs in.

It is a library first. One pip install, one directory per agent, no server to run and no account to make. The framework adapters are thin wrappers over the same three verbs.

pip install "fluctlightdb[native]" Linux, macOS and Windows · Python 3.9 to 3.13
  • LangChain FluctlightMemory, chat history
  • LlamaIndex FluctlightLlamaMemory
  • OpenAI Agents SDK remember and search handlers
  • MCP Cursor, Claude and Codex
  • Docker ghcr.io/voxmastery/fluctlightdb
  • Node over HTTP from fluctlight-serve

Adapter details in the integrations document.

An engraving of an assortment of brass pipe couplings and threaded connector fittings.
Two workers, one shared memory — 54 seconds, no sound. A recorded run of the Codex swarm demo. Two parallel workers receive the same verified facts and the same warnings, but FluctlightDB gives each a different episodic strategy. A worker citing its peer’s memory is rejected. A worker verifying its own success is rejected. Then the coordinator is stopped and restarted, and the finished run comes back from the write-ahead log.
Read the narration instead

This scripted visualization shows Fluctlight Swarm Memory as a Codex workflow. The root starts one durable coordinator run, and two illustrated workers act against the same project state.

Both workers receive the same verified truth and failure warning, but FluctlightDB assigns each a different episodic strategy. The API worker gets the transaction-boundary memory. The test worker gets the crash-recovery memory. Their allocation overlap is zero.

When one worker tries to cite its peer’s memory, the coordinator rejects it. A worker also cannot verify its own success. Trusted test evidence applies credit only to the memory actually cited.

Finally, the public one-command demo proves all four behaviors, stops the coordinator, restarts it, and recovers the completed run from the WAL and version-four checkpoint.

Parallel Codex agents can now remember together without thinking the same way.