Four things the python SDK turned up, each fixed where every client sees it. A key no port declares is now an error rather than a silent drop, on the return, the yield and the emit alike — the contract the docs already stated. The SDK reads literal yields at sync time, so a typo fails before anything runs, and an emission of one fails the call rather than being logged where nobody looks. NaN and infinity are refused at the port. JSON cannot spell either, so one that travelled came back as a 500, a socket frame that stopped the canvas, or a metric batch the database dropped whole. An artifact input takes `@run:<id>.<output>` or a bare digest, resolved on the engine — so the CLI, the run dialog and a python caller mean the same thing, and a sweep can pass one at all. Node timeouts are off by default. The clock measured silence, which a training node is full of, and remote workers had already stopped enforcing it — their heartbeat reset it. Now a heartbeat proves the agent rather than the node, ninety seconds of nothing fails the call either way, and the engine touches work it is still running so a long node is not redelivered at sixty seconds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019V5bsYGNxcgPs4xXmTPx69
5.5 KiB
Remote workers
The engine runs where the automations are. The GPU is somewhere else, the Raspberry Pi with the relays is in a shed, and neither of them is on the same network as the other.
A worker is a process that runs the code of nodes marked for it. It dials out to the engine over one authenticated websocket, so nothing on that machine has to be reachable — and nothing has to expose the engine's state backend across hosts, which it never should.
Install and attach
pip install fluksio-worker
fluksio-worker \
--url wss://api.example.com/api/v1/workers/attach \
--token "$FLUKSIO_WORKER_TOKEN" \
--labels gpu,cuda12 \
--python /opt/torch-venv/bin/python
fluksio-worker is its own distribution — the agent, the node runner, and
websockets. Nothing of the engine, so a GPU box does not install a database
driver in order to run a training step. An engine host already has it, and
fluksio worker … is the same program.
| Option | Default | What it does |
|---|---|---|
--url |
required | wss://…/api/v1/workers/attach |
--token |
$FLUKSIO_WORKER_TOKEN |
the credential, minted on the engine |
--name |
this host's name | how it shows up in the worker list |
--labels |
none | comma-separated; what a node's device matches |
--python |
this interpreter | the interpreter node code runs on |
--parallel |
1 |
how many node calls it will take at once |
--artifact-url |
derived from --url |
where the artifact store is, if not beside the socket |
--python is the important one. It is how this machine keeps its own wheels —
the CUDA build, the vendor SDK, the thing that will not install anywhere else —
without the engine ever installing them or knowing about them.
Mint the token
On the engine, as a superuser:
curl -X POST $FLUKSIO/workers/tokens -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"name": "gpu-dev"}'
Shown once, valid for a year — a worker is a machine somebody sets up and leaves running. It is signed with the same keypair agent tokens use, so rotating that key revokes every worker along with them.
??? note "A host where pip is not an option"
The two files work copied into one directory and run with `python agent.py
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime` —
it is the same module its own local workers run, deliberately standard
library only.
Send a node to it
A node declares the label of the machine it needs:
{
"id": "train",
"device": "gpu",
"device_policy": "require",
"timeout": 7200
}
device_policy |
Behaviour when nothing carrying the label is attached |
|---|---|
require (default) |
the run stays queued and says what it is waiting for |
prefer |
it runs on the engine instead |
prefer is what makes a flow work before the GPU box exists. require is what
you want once it does.
!!! note "Set from the API"
`device` and `device_policy` are not yet fields in the node panel. Set them
with `PUT /flows/{name}`.
What follows from this
- The node's source travels with every call. Nothing has to be deployed to the worker, and changing a node's code takes effect on the next execution.
- A node bound to a device is compiled on that machine. A node importing
torchis correct on the GPU box and a missing module on the engine, so checking it here would fail something that is fine. import fluksioinside a node is the worker's own reporter.emit,save_artifact,load_artifact— installed before your code runs, so an installedfluksiopackage on that box never shadows it.- Cancelling a run kills what it is executing, there or here, and leaves other runs of the same node alone.
- If the worker disappears mid-call, the run fails in seconds with
worker went away mid-callrather than waiting out its timeout. - A worker sends a heartbeat every ten seconds while it executes, so a long node is distinguishable from a dead socket. Ninety seconds of nothing at all — not even a heartbeat — fails the call as gone. A heartbeat says the agent is alive and nothing about the node, so it never satisfies a node's own timeout: one set to thirty seconds fires after thirty seconds of the node reporting nothing, wherever it runs.
Artifacts across machines
An artifact reference names content by its hash, not a location, so it stays valid wherever the store is reachable from. A worker that shares the engine's filesystem writes to it directly; one that does not fetches and uploads over HTTP, using the artifact endpoint beside the socket it already has. Either way your node code is the same two calls.
Seeing what is attached
curl -s $FLUKSIO/workers -H "Authorization: Bearer $TOKEN" | jq
Name, labels, how many calls it will take at once, how many are in flight, when it attached, when it was last seen, its Python version, and a digest of its environment.
What a worker is not
It is not a second engine. Subscriptions, schedules, webhooks, the dashboards and the run queue all stay in one process — that is what keeps a message having one definition and a cron tick happening once. A worker executes node bodies.
Running two engines against one data directory is not supported. Distribute work with workers.