Make the docs state things rather than argue them
Docs / docs (push) Successful in 37s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m35s
Playwright Tests / test-playwright (2, 2) (push) Failing after 17s
pre-commit / pre-commit (push) Failing after 2m8s
Test Backend / test-backend (push) Failing after 2m48s
Compose Smoke Test / test-compose (push) Failing after 13s
Playwright Tests / merge-reports (push) Failing after 2m25s

The site read as a design journal: rationale paragraphs, hedges
("deliberately", "on purpose", "genuinely"), meta-commentary about the docs
themselves, and one em-dash every ten lines carrying an aside.

Roughly twenty rationale blocks are gone or reduced to what a reader needs
in order to use the thing. Em-dashes go from 507 to 135, and what is left is
structural rather than prose: list and definition separators, table cells,
and four inside code blocks that quote what the CLI actually prints.

Also: api.example.com becomes api.fluksio.com (the emails stay, since
bootstrap.py really defaults to admin@example.com and RFC 2606 reserves it);
the mqtt table gains the two settings it had drifted behind on and inject's
wording matches the engine; llms.txt lists the two connector pages that were
in the nav but not in it; and the two device/device_policy notes now agree.

Builds clean under `zensical build --strict`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YrQnKV3bnQd4K342y8tKj
This commit is contained in:
2026-08-31 10:49:58 +02:00
co-authored by Claude Opus 5
parent 2422a9b22b
commit bdad6d7fc2
25 changed files with 450 additions and 479 deletions
+28 -30
View File
@@ -6,7 +6,7 @@ network as the other.
A **worker** is a process that runs the code of nodes marked for it. It dials
*out* to the engine over one authenticated websocket, so nothing on that
machine has to be reachable and nothing has to expose the engine's state
machine has to be reachable, and nothing has to expose the engine's state
backend across hosts, which it never should.
## Install and attach
@@ -15,13 +15,13 @@ backend across hosts, which it never should.
pip install fluksio-worker
fluksio-worker \
--url wss://api.example.com/api/v1/workers/attach \
--url wss://api.fluksio.com/api/v1/workers/attach \
--token "$FLUKSIO_WORKER_TOKEN" \
--labels gpu,cuda12 \
--python /opt/torch-venv/bin/python
```
`fluksio-worker` is its own distribution the agent, the node runner, and
`fluksio-worker` is its own distribution: the agent, the node runner, and
`websockets`. Nothing of the engine, so a GPU box does not install a database
driver in order to run a training step. An engine host already has it, and
`fluksio worker …` is the same program.
@@ -40,13 +40,13 @@ driver in order to run a training step. An engine host already has it, and
| `--ram-mb` | what the job or the machine has | memory to advertise, in MB |
| `--max-idle` | never | stop after this many seconds with nothing running |
`--python` is the important one. It is how this machine keeps its own wheels
the CUDA build, the vendor SDK, the thing that will not install anywhere else
`--python` is the important one. It is how this machine keeps its own wheels
(the CUDA build, the vendor SDK, the thing that will not install anywhere else)
without the engine ever installing them or knowing about them.
### What it says it has
A worker reports its inventory when it attaches cores, GPUs and memory and
A worker reports its inventory when it attaches (cores, GPUs and memory) and
the engine schedules against it: a node asking for two cores and a GPU goes to
a machine that has them free, not merely to one carrying the right label.
@@ -58,12 +58,12 @@ or `FLUKSIO_WORKER_GPUS`) or something you say with `--gpus`. A worker that
reports nothing still attaches and is scheduled by its label alone, as every
worker was before any of them reported anything.
The engine tells each call what it may use thread caps, and the devices it
The engine tells each call what it may use: thread caps, and the devices it
may see. The worker starts a process per call, so it applies them at the only
moment a numerical library still reads them: before the import.
`--max-idle` is for a worker something else started for one job a batch
scheduler, say. It exits when nothing has run for that long, so the allocation
`--max-idle` is for a worker something else started for one job, a batch
scheduler say. It exits when nothing has run for that long, so the allocation
goes back rather than idling until its walltime.
## Mint the token
@@ -75,15 +75,15 @@ curl -X POST $FLUKSIO/workers/tokens -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"name": "gpu-dev"}'
```
Shown once, valid for a year a worker is a machine somebody sets up and
Shown once, valid for a year, since a worker is a machine somebody sets up and
leaves running. It is signed with the same keypair agent tokens use, so
rotating that key revokes every worker along with them.
??? note "A host where pip is not an option"
The two files work copied into one directory and run with `python agent.py
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime`
it is the same module its own local workers run, deliberately standard
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime`.
It is the same module its own local workers run, deliberately standard
library only.
## Send a node to it
@@ -107,10 +107,10 @@ A node declares the label of the machine it needs:
`prefer` is what makes a flow work before the GPU box exists. `require` is what
you want once it does.
!!! note "Set from the API"
!!! note "Not in the panel yet"
`device` and `device_policy` are not yet fields in the node panel. Set them
with `PUT /flows/{name}`.
`device` and `device_policy`, which machine a node's code runs on, are set
through the API rather than the panel, with `PUT /flows/{name}`.
## What follows from this
@@ -120,15 +120,15 @@ you want once it does.
`torch` is correct on the GPU box and a missing module on the engine, so
checking it here would fail something that is fine.
- **`import fluksio` inside a node is the worker's own reporter.** `emit`,
`save_artifact`, `load_artifact` installed before your code runs, so an
`save_artifact`, `load_artifact`, installed before your code runs, so an
installed `fluksio` package on that box never shadows it.
- **Cancelling a run kills what it is executing**, there or here, and leaves
other runs of the same node alone.
- **If the worker disappears mid-call**, the run fails in seconds with
`worker went away mid-call` rather than waiting out its timeout.
- **A worker sends a heartbeat every ten seconds while it executes**, so a long
node is distinguishable from a dead socket. Ninety seconds of nothing at all
not even a heartbeat fails the call as gone. A heartbeat says the *agent* is
node is distinguishable from a dead socket. Ninety seconds of nothing at all,
not even a heartbeat, fails the call as gone. A heartbeat says the *agent* is
alive and nothing about the node, so it never satisfies a node's own timeout:
one set to thirty seconds fires after thirty seconds of the node reporting
nothing, wherever it runs.
@@ -154,8 +154,8 @@ environment, and what it says it has: cores, GPUs and memory.
## Upgrading
The engine and the worker speak a version-matched protocol, and a worker
announcing anything else is refused rather than half-understood. Protocol 2
the one that carries inventory is `fluksio-worker` 0.2.0. An older agent is
announcing anything else is refused rather than half-understood. Protocol 2,
the one that carries inventory, is `fluksio-worker` 0.2.0. An older agent is
told so on the socket and stops, rather than retrying against an engine that
will never accept it; `pip install -U fluksio-worker` on that host is the whole
upgrade. Nothing changed in the runner served at `GET /api/v1/workers/runtime`,
@@ -163,7 +163,7 @@ so a host that copies its two files copies the same one as before.
## Machines from a batch scheduler
A cluster is not a machine that attaches and stays it is a queue somebody else
A cluster is not a machine that attaches and stays; it is a queue somebody else
owns. So Fluksio does not submit *nodes* to Slurm. It submits a job whose payload
is an ordinary worker dialling back in, and from there everything works the way
it already does: the same protocol, the same artifacts, the same cancellation.
@@ -176,7 +176,7 @@ Write the clusters into `provisioners.json` beside the flows:
"name": "hpc",
"login": "me@login.cluster",
"ssh_key": "/secrets/hpc_ed25519",
"engine_url": "wss://api.example.com/api/v1/workers/attach",
"engine_url": "wss://api.fluksio.com/api/v1/workers/attach",
"max_idle_s": 300,
"provision_timeout_s": 900,
"profiles": [{
@@ -192,12 +192,10 @@ Write the clusters into `provisioners.json` beside the flows:
A **profile** is what the scheduler is asked for, where a flavor is what a node
asks for. They are separate on purpose, and agree when you set them up to.
When a node needs a machine nothing attached can give, and a profile would fit,
the engine `sbatch`es one over ssh — the system `ssh`, so nothing new is
installed — and the run waits meanwhile, saying so. `prerun` owns the
environment: a `module load`, a venv with `fluksio-worker` already in it. There
is deliberately no `pip install` in the generated script, because what is
installed on a cluster is somebody's decision and not this program's.
When a node needs a machine nothing attached can give, and a profile fits, the
engine `sbatch`es one over the system `ssh`, and the run waits meanwhile, saying
so. `prerun` owns the environment: a `module load`, or a venv with
`fluksio-worker` already in it. The generated script runs no `pip install`.
One outstanding request per profile, however often it is asked for. A job that
never attaches within `provision_timeout_s` is `scancel`led, as is anything
@@ -219,8 +217,8 @@ queue does, which is what a cluster that queues overnight needs, and
## What a worker is not
It is not a second engine. Subscriptions, schedules, webhooks, the dashboards
and the run queue all stay in one process — that is what keeps a message having
one definition and a cron tick happening once. A worker executes node bodies.
and the run queue all stay in one process, which keeps a message having one
definition and a cron tick happening once. A worker executes node bodies.
Running two engines against one data directory is not supported. Distribute
work with workers.