Measured with `make bench-engine` against a real Redis: 103.6 -> 164.4
messages a second on a five-node chain (p50 latency 2125 -> 1171 ms) and
34.8 -> 63.2 on a fan-out of twenty. Against the memory backend, which is
what a pip install runs on, 262 -> 626.
The two that bought most of it:
- `StateBackend.record` puts a published value, its timestamp, its series
and its version counter in one round trip. They were four calls building
four pipelines, and a value crossing an edge pays them twice. A released
rate-limit hold rides along instead of a DEL per port.
- the readiness check reads a node's inputs and hands them to the node,
rather than reading the triggering ones to count them and having the node
read the same keys again a moment later.
`apply_outputs` was a second copy of `_record_outputs` and is now the same
code plus the event that distinguishes it.
The rest, each small:
- `_derive` builds a node-by-id map and a `consumes` index, so dispatching
an item and publishing a value stop scanning every node in the
installation.
- `read_all` is memoised against the store revision — it sits on the
publish path, so a dashboard slider was reading and validating every
flow file per value. Same mechanism `_wiring` already uses.
- the `message_value` source block is built once per node instead of per
emission.
- both timer threads ask the queue to promote only when something is
actually due, which takes an idle engine from ~4 Redis round trips a
second to one.
- the shared httpx client is bounded (32 connections, one retry); its
default pool is 100 with no per-host cap, so one slow endpoint could
take it and every other sender node with it.
- the MQTT and delay nodes no longer log a line per message at INFO.
Robustness, in the same pass:
- `MemoryWorkQueue._done` was a set nothing ever removed from — one entry
per non-idempotent node per item, for the life of the process, in the
default configuration. Capped, the way the Redis side expires its
markers.
- a saturated engine can claim from the due lane past the cascade limit.
The capacity gate sits in front of the claim, so the due lane's priority
— decided inside it — did not apply while every slot was held: a motor's
stop was not behind the long nodes, it was unread. Only after a slot has
genuinely failed to free for half a second, and briefly, so the backlog
is not starved in turn.
- `reclaim_stale` dispatches through that same gate. It could return sixty
entries and push in-flight far past the limit the gate exists to hold.
- a flow's nodes are stopped together rather than one after another. Each
gets `NODE_STOP_TIMEOUT`, so a flow whose broker was unreachable took
five seconds per node — long enough to outlast `REBUILD_WAIT` and 503
the deploy.
- the worker pool and the HTTP client are closed on a thread, not on the
event loop, and a run closes the state backend it built (on Redis, a
client and a connection pool per run).
- the five background tasks say something when they die. Each catches
exceptions inside its loop, so one raised anywhere else left the engine
serving with no metrics, no alerts or no artifact sweep, silently.
`tests/flow/test_round_trips.py` counts the state operations one message
costs — four, where it was about eleven — because none of the above would
fail a behavioural test if it were undone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M6hPWS6YEbT1P8LxhhFb2T
Three faults with one root: the stored body of a code-defined node is an
import shim, and nothing that mattered was ever read from the code itself.
- The run stamp could not identify what ran. The shim imports whatever is on
disk when the worker starts, and an uncommitted tree stamps <commit>-dirty
for every run it ever produces. Run.code_digest hashes the repository's .py
files, memoized on their stat state, and it is read again when the run is
actually claimed -- so a sweep queued for hours records the code each of its
runs executed, not the code that was there when it was submitted.
- The stage cache adopted code that was too new. The fingerprint hashed the
shim, which is invariant under any edit to the imported function or anything
it calls into, so a re-run was served from cache and answered without the
outputs the edit added. It now carries the repo digest and the node's
declared ports. Every fingerprint changes once, which invalidates the
existing cache; a canvas flow has no repository and keys as before.
- An interrupted sync looked like a hand-edited canvas. The engine answers a
new-node template for a node with no stored body, and the template carries
no marker, so the drift check read "somebody edited this" and demanded
--force -- for the one state that re-running the sync is the fix for.
NodeSource.missing states the fact, and sync skips those and reuses the
bodies it read instead of asking for each one twice.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`FlowStore._git` now passes `-c gc.auto=0`, so `git commit` no longer forks a
background `gc --auto` that reparents onto PID 1 and stays there unreaped. With
auto-gc off nothing packs on its own, so `_commit` runs a foreground `git gc`
every 500 commits — the trigger a long-lived seeding session actually reaches.
`docker/compose.yml` takes the frontend's `VITE_API_URL` build arg from the
environment, keeping `https://api.${DOMAIN}` only as the fallback. `setup.sh`
already derives the scheme from `ENVIRONMENT`, so an `up --build` that does not
layer `compose.local.yml` stops shipping a bundle that calls `https://api.localhost`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StpRc2C6au1WJ1EUU7fsfu
- A node's height follows the ports on its busiest side. It is a function of
the document, so `layoutGraph` reserves exactly what is drawn and nothing
measured is fed back into the layout.
- The three status controls now sit in slots that are there whether the
control is or not. A node running many times a second mounted and unmounted
the stop button on every execution, resizing the card each time.
- A port bound to another flow's message is drawn as a label, naming the node
at the far end and its type. Only the opposite direction was answered
before. The scan behind both is now cached on the store's commit counter
rather than reading every flow per request.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ZeGnqVsf5VHQqvz4HdUhN
A `pip install` on a locked-down host — the case the CLI exists for —
may have no git, and the store shelled out to it while building the flow
repository, so the engine refused to start at all. The store is files;
git is their history. Missing it is now one warning and no commits
rather than a stack trace, which is the difference between a machine
that runs your experiments and one that does not.
Found by installing the wheels into a bare python:3.12-slim and pairing
it with the portal: `fluksio enroll` took the code, `fluksio serve`
dialled out, and the hub was proxying requests through the tunnel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A wheel whose top-level module is `app` collides with anything else in a
user's venv, so the package that is about to be published takes the name
it is published under. Only the Python package moves; the repo, the
Docker WORKDIR and the compose project keep theirs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>