Make the docs state things rather than argue them
Docs / docs (push) Successful in 37s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m35s
Playwright Tests / test-playwright (2, 2) (push) Failing after 17s
pre-commit / pre-commit (push) Failing after 2m8s
Test Backend / test-backend (push) Failing after 2m48s
Compose Smoke Test / test-compose (push) Failing after 13s
Playwright Tests / merge-reports (push) Failing after 2m25s
Docs / docs (push) Successful in 37s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m35s
Playwright Tests / test-playwright (2, 2) (push) Failing after 17s
pre-commit / pre-commit (push) Failing after 2m8s
Test Backend / test-backend (push) Failing after 2m48s
Compose Smoke Test / test-compose (push) Failing after 13s
Playwright Tests / merge-reports (push) Failing after 2m25s
The site read as a design journal: rationale paragraphs, hedges
("deliberately", "on purpose", "genuinely"), meta-commentary about the docs
themselves, and one em-dash every ten lines carrying an aside.
Roughly twenty rationale blocks are gone or reduced to what a reader needs
in order to use the thing. Em-dashes go from 507 to 135, and what is left is
structural rather than prose: list and definition separators, table cells,
and four inside code blocks that quote what the CLI actually prints.
Also: api.example.com becomes api.fluksio.com (the emails stay, since
bootstrap.py really defaults to admin@example.com and RFC 2606 reserves it);
the mqtt table gains the two settings it had drifted behind on and inject's
wording matches the engine; llms.txt lists the two connector pages that were
in the nav but not in it; and the two device/device_policy notes now agree.
Builds clean under `zensical build --strict`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YrQnKV3bnQd4K342y8tKj
This commit is contained in:
+28
-30
@@ -6,7 +6,7 @@ network as the other.
|
||||
|
||||
A **worker** is a process that runs the code of nodes marked for it. It dials
|
||||
*out* to the engine over one authenticated websocket, so nothing on that
|
||||
machine has to be reachable — and nothing has to expose the engine's state
|
||||
machine has to be reachable, and nothing has to expose the engine's state
|
||||
backend across hosts, which it never should.
|
||||
|
||||
## Install and attach
|
||||
@@ -15,13 +15,13 @@ backend across hosts, which it never should.
|
||||
pip install fluksio-worker
|
||||
|
||||
fluksio-worker \
|
||||
--url wss://api.example.com/api/v1/workers/attach \
|
||||
--url wss://api.fluksio.com/api/v1/workers/attach \
|
||||
--token "$FLUKSIO_WORKER_TOKEN" \
|
||||
--labels gpu,cuda12 \
|
||||
--python /opt/torch-venv/bin/python
|
||||
```
|
||||
|
||||
`fluksio-worker` is its own distribution — the agent, the node runner, and
|
||||
`fluksio-worker` is its own distribution: the agent, the node runner, and
|
||||
`websockets`. Nothing of the engine, so a GPU box does not install a database
|
||||
driver in order to run a training step. An engine host already has it, and
|
||||
`fluksio worker …` is the same program.
|
||||
@@ -40,13 +40,13 @@ driver in order to run a training step. An engine host already has it, and
|
||||
| `--ram-mb` | what the job or the machine has | memory to advertise, in MB |
|
||||
| `--max-idle` | never | stop after this many seconds with nothing running |
|
||||
|
||||
`--python` is the important one. It is how this machine keeps its own wheels —
|
||||
the CUDA build, the vendor SDK, the thing that will not install anywhere else —
|
||||
`--python` is the important one. It is how this machine keeps its own wheels
|
||||
(the CUDA build, the vendor SDK, the thing that will not install anywhere else)
|
||||
without the engine ever installing them or knowing about them.
|
||||
|
||||
### What it says it has
|
||||
|
||||
A worker reports its inventory when it attaches — cores, GPUs and memory — and
|
||||
A worker reports its inventory when it attaches (cores, GPUs and memory) and
|
||||
the engine schedules against it: a node asking for two cores and a GPU goes to
|
||||
a machine that has them free, not merely to one carrying the right label.
|
||||
|
||||
@@ -58,12 +58,12 @@ or `FLUKSIO_WORKER_GPUS`) or something you say with `--gpus`. A worker that
|
||||
reports nothing still attaches and is scheduled by its label alone, as every
|
||||
worker was before any of them reported anything.
|
||||
|
||||
The engine tells each call what it may use — thread caps, and the devices it
|
||||
The engine tells each call what it may use: thread caps, and the devices it
|
||||
may see. The worker starts a process per call, so it applies them at the only
|
||||
moment a numerical library still reads them: before the import.
|
||||
|
||||
`--max-idle` is for a worker something else started for one job — a batch
|
||||
scheduler, say. It exits when nothing has run for that long, so the allocation
|
||||
`--max-idle` is for a worker something else started for one job, a batch
|
||||
scheduler say. It exits when nothing has run for that long, so the allocation
|
||||
goes back rather than idling until its walltime.
|
||||
|
||||
## Mint the token
|
||||
@@ -75,15 +75,15 @@ curl -X POST $FLUKSIO/workers/tokens -H "Authorization: Bearer $TOKEN" \
|
||||
-H 'Content-Type: application/json' -d '{"name": "gpu-dev"}'
|
||||
```
|
||||
|
||||
Shown once, valid for a year — a worker is a machine somebody sets up and
|
||||
Shown once, valid for a year, since a worker is a machine somebody sets up and
|
||||
leaves running. It is signed with the same keypair agent tokens use, so
|
||||
rotating that key revokes every worker along with them.
|
||||
|
||||
??? note "A host where pip is not an option"
|
||||
|
||||
The two files work copied into one directory and run with `python agent.py
|
||||
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime` —
|
||||
it is the same module its own local workers run, deliberately standard
|
||||
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime`.
|
||||
It is the same module its own local workers run, deliberately standard
|
||||
library only.
|
||||
|
||||
## Send a node to it
|
||||
@@ -107,10 +107,10 @@ A node declares the label of the machine it needs:
|
||||
`prefer` is what makes a flow work before the GPU box exists. `require` is what
|
||||
you want once it does.
|
||||
|
||||
!!! note "Set from the API"
|
||||
!!! note "Not in the panel yet"
|
||||
|
||||
`device` and `device_policy` are not yet fields in the node panel. Set them
|
||||
with `PUT /flows/{name}`.
|
||||
`device` and `device_policy`, which machine a node's code runs on, are set
|
||||
through the API rather than the panel, with `PUT /flows/{name}`.
|
||||
|
||||
## What follows from this
|
||||
|
||||
@@ -120,15 +120,15 @@ you want once it does.
|
||||
`torch` is correct on the GPU box and a missing module on the engine, so
|
||||
checking it here would fail something that is fine.
|
||||
- **`import fluksio` inside a node is the worker's own reporter.** `emit`,
|
||||
`save_artifact`, `load_artifact` — installed before your code runs, so an
|
||||
`save_artifact`, `load_artifact`, installed before your code runs, so an
|
||||
installed `fluksio` package on that box never shadows it.
|
||||
- **Cancelling a run kills what it is executing**, there or here, and leaves
|
||||
other runs of the same node alone.
|
||||
- **If the worker disappears mid-call**, the run fails in seconds with
|
||||
`worker went away mid-call` rather than waiting out its timeout.
|
||||
- **A worker sends a heartbeat every ten seconds while it executes**, so a long
|
||||
node is distinguishable from a dead socket. Ninety seconds of nothing at all —
|
||||
not even a heartbeat — fails the call as gone. A heartbeat says the *agent* is
|
||||
node is distinguishable from a dead socket. Ninety seconds of nothing at all,
|
||||
not even a heartbeat, fails the call as gone. A heartbeat says the *agent* is
|
||||
alive and nothing about the node, so it never satisfies a node's own timeout:
|
||||
one set to thirty seconds fires after thirty seconds of the node reporting
|
||||
nothing, wherever it runs.
|
||||
@@ -154,8 +154,8 @@ environment, and what it says it has: cores, GPUs and memory.
|
||||
## Upgrading
|
||||
|
||||
The engine and the worker speak a version-matched protocol, and a worker
|
||||
announcing anything else is refused rather than half-understood. Protocol 2 —
|
||||
the one that carries inventory — is `fluksio-worker` 0.2.0. An older agent is
|
||||
announcing anything else is refused rather than half-understood. Protocol 2,
|
||||
the one that carries inventory, is `fluksio-worker` 0.2.0. An older agent is
|
||||
told so on the socket and stops, rather than retrying against an engine that
|
||||
will never accept it; `pip install -U fluksio-worker` on that host is the whole
|
||||
upgrade. Nothing changed in the runner served at `GET /api/v1/workers/runtime`,
|
||||
@@ -163,7 +163,7 @@ so a host that copies its two files copies the same one as before.
|
||||
|
||||
## Machines from a batch scheduler
|
||||
|
||||
A cluster is not a machine that attaches and stays — it is a queue somebody else
|
||||
A cluster is not a machine that attaches and stays; it is a queue somebody else
|
||||
owns. So Fluksio does not submit *nodes* to Slurm. It submits a job whose payload
|
||||
is an ordinary worker dialling back in, and from there everything works the way
|
||||
it already does: the same protocol, the same artifacts, the same cancellation.
|
||||
@@ -176,7 +176,7 @@ Write the clusters into `provisioners.json` beside the flows:
|
||||
"name": "hpc",
|
||||
"login": "me@login.cluster",
|
||||
"ssh_key": "/secrets/hpc_ed25519",
|
||||
"engine_url": "wss://api.example.com/api/v1/workers/attach",
|
||||
"engine_url": "wss://api.fluksio.com/api/v1/workers/attach",
|
||||
"max_idle_s": 300,
|
||||
"provision_timeout_s": 900,
|
||||
"profiles": [{
|
||||
@@ -192,12 +192,10 @@ Write the clusters into `provisioners.json` beside the flows:
|
||||
A **profile** is what the scheduler is asked for, where a flavor is what a node
|
||||
asks for. They are separate on purpose, and agree when you set them up to.
|
||||
|
||||
When a node needs a machine nothing attached can give, and a profile would fit,
|
||||
the engine `sbatch`es one over ssh — the system `ssh`, so nothing new is
|
||||
installed — and the run waits meanwhile, saying so. `prerun` owns the
|
||||
environment: a `module load`, a venv with `fluksio-worker` already in it. There
|
||||
is deliberately no `pip install` in the generated script, because what is
|
||||
installed on a cluster is somebody's decision and not this program's.
|
||||
When a node needs a machine nothing attached can give, and a profile fits, the
|
||||
engine `sbatch`es one over the system `ssh`, and the run waits meanwhile, saying
|
||||
so. `prerun` owns the environment: a `module load`, or a venv with
|
||||
`fluksio-worker` already in it. The generated script runs no `pip install`.
|
||||
|
||||
One outstanding request per profile, however often it is asked for. A job that
|
||||
never attaches within `provision_timeout_s` is `scancel`led, as is anything
|
||||
@@ -219,8 +217,8 @@ queue does, which is what a cluster that queues overnight needs, and
|
||||
## What a worker is not
|
||||
|
||||
It is not a second engine. Subscriptions, schedules, webhooks, the dashboards
|
||||
and the run queue all stay in one process — that is what keeps a message having
|
||||
one definition and a cron tick happening once. A worker executes node bodies.
|
||||
and the run queue all stay in one process, which keeps a message having one
|
||||
definition and a cron tick happening once. A worker executes node bodies.
|
||||
|
||||
Running two engines against one data directory is not supported. Distribute
|
||||
work with workers.
|
||||
|
||||
Reference in New Issue
Block a user