Measured with `make bench-engine` against a real Redis: 103.6 -> 164.4
messages a second on a five-node chain (p50 latency 2125 -> 1171 ms) and
34.8 -> 63.2 on a fan-out of twenty. Against the memory backend, which is
what a pip install runs on, 262 -> 626.
The two that bought most of it:
- `StateBackend.record` puts a published value, its timestamp, its series
and its version counter in one round trip. They were four calls building
four pipelines, and a value crossing an edge pays them twice. A released
rate-limit hold rides along instead of a DEL per port.
- the readiness check reads a node's inputs and hands them to the node,
rather than reading the triggering ones to count them and having the node
read the same keys again a moment later.
`apply_outputs` was a second copy of `_record_outputs` and is now the same
code plus the event that distinguishes it.
The rest, each small:
- `_derive` builds a node-by-id map and a `consumes` index, so dispatching
an item and publishing a value stop scanning every node in the
installation.
- `read_all` is memoised against the store revision — it sits on the
publish path, so a dashboard slider was reading and validating every
flow file per value. Same mechanism `_wiring` already uses.
- the `message_value` source block is built once per node instead of per
emission.
- both timer threads ask the queue to promote only when something is
actually due, which takes an idle engine from ~4 Redis round trips a
second to one.
- the shared httpx client is bounded (32 connections, one retry); its
default pool is 100 with no per-host cap, so one slow endpoint could
take it and every other sender node with it.
- the MQTT and delay nodes no longer log a line per message at INFO.
Robustness, in the same pass:
- `MemoryWorkQueue._done` was a set nothing ever removed from — one entry
per non-idempotent node per item, for the life of the process, in the
default configuration. Capped, the way the Redis side expires its
markers.
- a saturated engine can claim from the due lane past the cascade limit.
The capacity gate sits in front of the claim, so the due lane's priority
— decided inside it — did not apply while every slot was held: a motor's
stop was not behind the long nodes, it was unread. Only after a slot has
genuinely failed to free for half a second, and briefly, so the backlog
is not starved in turn.
- `reclaim_stale` dispatches through that same gate. It could return sixty
entries and push in-flight far past the limit the gate exists to hold.
- a flow's nodes are stopped together rather than one after another. Each
gets `NODE_STOP_TIMEOUT`, so a flow whose broker was unreachable took
five seconds per node — long enough to outlast `REBUILD_WAIT` and 503
the deploy.
- the worker pool and the HTTP client are closed on a thread, not on the
event loop, and a run closes the state backend it built (on Redis, a
client and a connection pool per run).
- the five background tasks say something when they die. Each catches
exceptions inside its loop, so one raised anywhere else left the engine
serving with no metrics, no alerts or no artifact sweep, silently.
`tests/flow/test_round_trips.py` counts the state operations one message
costs — four, where it was about eleven — because none of the above would
fail a behavioural test if it were undone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M6hPWS6YEbT1P8LxhhFb2T
Node stop paths cancelled their background task and then caught
CancelledError around the await, which swallows a cancellation aimed at
the caller — the trap Supervisor._cancel already documents. One shared
Node._cancel_task now waits the way the supervisor does; mqtt's publisher
and subscription and delay's cron call it.
The api container also collected zombie python workers: orphaned when
--reload replaces the process holding their handle, they reparent onto a
PID 1 that reaps nothing but its own. `init: true` on the backend service.
The gates have never gone green on the new runners. Three separate reasons:
- backend/Dockerfile shipped Python 3.10 while the code imports typing.Self
and datetime.UTC, so the container exited on import and the suite could not
even load its conftest. The image moves to 3.13 and the packages declare
>=3.12, which is the floor the tests actually pass on; ruff's target follows
and rewrites timezone.utc and asyncio.TimeoutError accordingly. Relocking
drops the 3.10 branch, which bumps FastAPI and so regenerates the SDK.
- frontend/README.md had no trailing newline and two dashboard widgets used
arbitrary text-[…] sizes. Both are em-relative on purpose, so they move to
the inline style the neighbouring ramp already uses.
- Every commit left its own run queued: without a concurrency group a runner
that was offline for a while works through a backlog nobody reads. A stack
that fails to come up now prints its logs before the teardown removes it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four small things, each with a device behind it.
An MQTT filter now routes what it subscribed to. `+` and `#` reached the
broker and were then looked up in an exact-match dict, so every message a
wildcard subscription received was dropped in silence.
`json_key` lifts a value out of the object a device wraps it in — Victron
publishes `{"value": 47}` on every path, which was otherwise a Python node
per port.
The trigger node learned `passthrough` and `wait_port`, because how long to
wait can be a value rather than a constant: a rollershutter takes 26 seconds
up and 28 down. A wait of zero sends nothing afterwards and still cancels
what the last message scheduled, which is how a stop is commanded once
instead of forever.
The HTTP sender takes fixed `query` parameters, so an API key is a secret
reference rather than a message on the canvas, and `send_inputs` off for a
request whose inputs are only a trigger.
Also: `delay` accepts fractional seconds, and `TZ` reaches the container, so
a cron expression means local time. Left unset it is UTC, as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A wheel whose top-level module is `app` collides with anything else in a
user's venv, so the package that is about to be published takes the name
it is published under. Only the Python package moves; the repo, the
Docker WORKDIR and the compose project keep theirs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>