A cascade has no end worth recording; a run does. Parameters go in, the graph executes until it drains, and the result is kept — which is what an ML experiment is and what a CI-style job is, so both are one entity. Each run gets a state backend namespaced to itself, so two runs of one flow cannot overwrite each other's messages; that is a constructor argument rather than a change to the pipeline, because every key the engine keeps already goes through the state backend. Its record is written by the driver thread rather than folded off the event bus, which drops what it cannot keep up with. Its own Redis stream wakes an engine up, and from the claim onwards the database row is the truth: redelivering hours of training because an acknowledgement was late is not recovery, so a stale lease is what marks a run whose engine died. Flows gain mode: batch, which are built and validated but never activated, and nodes gain a device label for the worker that must run them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AD8SfVhzXBG2nAfFcVh3iD
Generic single-database configuration.