Notes
Kill the writer. Open the brain. See if it verifies.
Crash recovery in this repository is a Rust test, not a slogan. It was read for this note and not re-run. What was re-run is the Python contract: checkpoint(), then a new process.
What the test does
crates/fluctlightdb/tests/chaos_jepsen.rs, test chaos_subprocess_sigkill_mid_write, is Unix-only. It opens a brain, checkpoints, and drops it. It then spawns fluctlight-chaos-worker, which writes experiences in a loop and only checkpoints after the loop. The parent sleeps briefly and sends SIGKILL. It asserts the child died from the signal or otherwise failed. It reopens the brain, asserts the engram set is non-empty, and calls verify_path, which must report ok.
The worker source says it writes without checkpointing until it is killed or it finishes the loop. The test’s comment says the WAL must recover a committed prefix or reject a torn tail without corrupting the store. The assertion in the test is the non-empty reopen plus verify_path. It does not publish a count of how many of the killed writes came back. This note will not invent one.
A sibling test on Windows skips the signal and points at the crash-recovery unit tests. The file header calls the suite Jepsen-style and says the engine is not a distributed consensus database.
What the Python API did when asked
experience() followed by checkpoint(), then a new interpreter on the same directory, printed the stored decision. That run is the crash-safe use case . turn_end(flush=True) without checkpoint() reported a committed turn in-process and was invisible to the next process. The documented line stands: nothing is durable until checkpoint() returns. The Rust test is about killing a writer mid-write and still opening a store that verifies. Those are related and they are not the same sentence.