Make the docs state things rather than argue them
Docs / docs (push) Successful in 37s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m35s
Playwright Tests / test-playwright (2, 2) (push) Failing after 17s
pre-commit / pre-commit (push) Failing after 2m8s
Test Backend / test-backend (push) Failing after 2m48s
Compose Smoke Test / test-compose (push) Failing after 13s
Playwright Tests / merge-reports (push) Failing after 2m25s

The site read as a design journal: rationale paragraphs, hedges
("deliberately", "on purpose", "genuinely"), meta-commentary about the docs
themselves, and one em-dash every ten lines carrying an aside.

Roughly twenty rationale blocks are gone or reduced to what a reader needs
in order to use the thing. Em-dashes go from 507 to 135, and what is left is
structural rather than prose: list and definition separators, table cells,
and four inside code blocks that quote what the CLI actually prints.

Also: api.example.com becomes api.fluksio.com (the emails stay, since
bootstrap.py really defaults to admin@example.com and RFC 2606 reserves it);
the mqtt table gains the two settings it had drifted behind on and inject's
wording matches the engine; llms.txt lists the two connector pages that were
in the nav but not in it; and the two device/device_policy notes now agree.

Builds clean under `zensical build --strict`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015YrQnKV3bnQd4K342y8tKj
This commit is contained in:
2026-08-31 10:49:58 +02:00
co-authored by Claude Opus 5
parent 2422a9b22b
commit bdad6d7fc2
25 changed files with 450 additions and 479 deletions
+1 -1
View File
@@ -43,7 +43,7 @@ A machine that should only *run nodes* for an engine elsewhere installs less:
```sh
pip install fluksio-worker
fluksio-worker --url wss://api.example.com/api/v1/workers/attach --token "$TOKEN" --labels gpu
fluksio-worker --url wss://api.fluksio.com/api/v1/workers/attach --token "$TOKEN" --labels gpu
```
## Getting started
+12 -12
View File
@@ -11,16 +11,16 @@ build, publish and run flows.
MCP_ENABLED=true
```
It is off by default, because switching it on opens client registration to
whoever can reach the host. Turn it on deliberately.
It is off by default: switching it on opens client registration to whoever can
reach the host.
The endpoint is `https://api.${DOMAIN}/mcp` or `http://api.localhost/mcp` on
a local stack.
The endpoint is `https://api.${DOMAIN}/mcp`, or `http://api.localhost/mcp` on a
local stack.
## Connecting a client
```sh
claude mcp add --transport http fluksio https://api.example.com/mcp
claude mcp add --transport http fluksio https://api.fluksio.com/mcp
```
The OAuth flow does the rest: the client registers itself, you approve it in
@@ -30,7 +30,7 @@ the browser, and it exchanges the grant for a token. Nothing to paste.
An agent acts as **the person who approved it**, and only through the same REST
API the dashboard uses. It cannot do anything you could not do in the browser,
and it cannot do anything *you* cannot do — an agent approved by a non-superuser
and it cannot do anything *you* cannot do. An agent approved by a non-superuser
is not one either.
Its tokens are signed with a keypair of their own, kept beside the flow store.
@@ -53,7 +53,7 @@ Read-only:
| `get_message_history` | one message's recent numeric values |
| `get_graph` | every flow as one graph |
| `list_shared_nodes` | the shared library and its usages |
| `list_secrets` | secret **names** never values |
| `list_secrets` | secret **names**, never values |
| `get_modules` | the packages node code can import |
| `get_health` | flows, nodes, queue and loop lag |
| `get_metrics` | executions, errors and timings per minute |
@@ -87,10 +87,10 @@ Operating:
The tools do not reach into the engine. Each one calls the same REST endpoint
the dashboard calls, over an in-process transport.
That keeps exactly one description of what a flow is and how it may be changed
the validation, the draft/publish split, the version check that stops two
clients overwriting each other. An agent that saved a flow through a private
back door would be an agent that could write one the canvas cannot open.
That keeps exactly one description of what a flow is and how it may be changed:
the validation, the draft/publish split, and the version check that stops two
clients overwriting each other. An agent saving a flow through a private back
door could write one the canvas cannot open.
The caller's token rides along on every hop, so the API sees the agent's own
identity rather than a service account. The audit trail on Home names it.
@@ -105,7 +105,7 @@ you what the canvas would have told you.
## Limits
- MCP is not currently reachable through a [portal](../interface/portal.md)
tunnel the proxy forwards `/api/v1/` only. Connect an agent on the same
tunnel: the proxy forwards `/api/v1/` only. Connect an agent on the same
network as the instance.
- Secrets are never readable, by an agent or by anyone else. `list_secrets`
returns names.
+13 -14
View File
@@ -1,7 +1,7 @@
# The HTTP API
Everything the browser does, the API does first. The dashboard is a generated
client of this schema, not a privileged path into the engine so anything you
client of this schema, not a privileged path into the engine, so anything you
can click, you can script.
Base URL: `https://api.${DOMAIN}/api/v1`, or `http://127.0.0.1:8000/api/v1` for
@@ -24,7 +24,7 @@ A bearer token, valid for eight days. `POST /login/test-token` checks one.
Password recovery and reset are `POST /password-recovery/{email}` and
`POST /reset-password/`.
Agents authenticate differently see [Agents over MCP](agents.md).
Agents authenticate differently; see [Agents over MCP](agents.md).
!!! tip "The interactive schema"
@@ -33,9 +33,8 @@ Agents authenticate differently — see [Agents over MCP](agents.md).
`/api/v1/openapi.json`. That is the authoritative reference; this page is
the tour.
It is closed in production on purpose: the schema enumerates every endpoint
the instance serves, including the paths webhook nodes mounted at
runtime.
It is closed in production: the schema enumerates every endpoint the
instance serves, including paths webhook nodes mounted at runtime.
## Flows
@@ -81,8 +80,8 @@ shared sources and which flows use each.
| `POST` | `/flows/{name}/nodes/{id}/acknowledge` | clear the failure the canvas is marking |
`POST /flows/{name}/run` on a **batch** flow submits a run instead, because
that is what running one means the parameters, the series and the result are
the point, and a call that quietly did something else would be a trap.
that is what running one means: the parameters, the series and the result are
the point.
## Messages
@@ -95,7 +94,7 @@ curl -s $FLUKSIO/messages/house.temperature/history -H "Authorization: Bearer $T
`GET /messages/` is the catalogue: every message any published flow declares,
with its last value. Publishing puts a value into the graph exactly as a
dashboard control does which means only a message some flow *declares* can be
dashboard control does, which means only a message some flow *declares* can be
published to. Flows own the namespace; everything else is a client of it.
## Runs
@@ -113,8 +112,8 @@ published to. Flows own the namespace; everything else is a client of it.
| `GET` | `/runs/{id}/metrics?name=&stride=` | one metric's series, in step order; every series of the run without `name` |
| `GET` | `/runs/series/compare?ids=a,b,c&metric=&x=` | that metric across several runs. `x` is what to plot against: nothing or `step`, `time` (seconds since each run's own first reading), or another metric's name, joined on the step the two share |
Submitting answers immediately with a `queued` run. Wrong parameters an
undeclared name, a value of the wrong type come back as a 422 naming the
Submitting answers immediately with a `queued` run. Wrong parameters (an
undeclared name, a value of the wrong type) come back as a 422 naming the
problem, before anything executes.
`?digest=` filters by the hash of the parameters, which is how you find "every
@@ -122,13 +121,13 @@ run that used exactly this configuration".
`?before=` is how a long history is paged: rows come newest first, so handing
back the last row's `created_at` reads the next page whatever landed
meanwhile which `?offset=` cannot, since a run submitted between two pages
meanwhile, which `?offset=` cannot, since a run submitted between two pages
shifts every row down one. `?since=` is inclusive and bounds the other end.
Both exports stream `csv` (the default) or `jsonl`, and take the selection the
history takes plus `?ids=a,b,c`, `?since=` and `?until=`. `export/runs` puts
the inputs that *vary* across the selection in `param.` columns the sweep
axis unless `?params=` names them, and the run's numbers in `metric.`
the inputs that *vary* across the selection in `param.` columns (the sweep
axis) unless `?params=` names them, and the run's numbers in `metric.`
columns. Both are dotted paths into whatever a node returned, so
`?metrics=final_metrics.train_loss` selects a field of a record and the
defaults reach every number inside one. The run id is on every row of both, which is what makes an exported
@@ -239,4 +238,4 @@ stored version moved past the one you were editing. Re-read, merge, save again.
The frontend's TypeScript client is generated from the OpenAPI schema
(`make generate-client`). Any OpenAPI generator will do the same for your
language point it at `/api/v1/openapi.json` on a non-production instance.
language: point it at `/api/v1/openapi.json` on a non-production instance.
+62 -65
View File
@@ -8,12 +8,12 @@ Installs the engine and the `fluksio` command. Python 3.12 or newer, Linux or
macOS.
The MQTT and InfluxDB connectors, outbound mail and error reporting are
`pip install 'fluksio[server]'` a deployment talking to devices wants them,
`pip install 'fluksio[server]'`: a deployment talking to devices wants them,
and a laptop waiting on them to install does not. Everything else, the engine
and every python node included, is in the plain install; a node type whose
library is missing says which extra to add when one is actually built.
There is a second, smaller distribution `fluksio-worker` for a machine that
There is a second, smaller distribution, `fluksio-worker`, for a machine that
should only *run nodes* for an engine elsewhere. It has none of the engine in
it. See [Remote workers](workers.md).
@@ -26,7 +26,7 @@ instance, while `login`, `sync`, `run`, `runs`, `artifacts`, `sweep` and
`.fluksio` beside your code, found the way `.git` is: from the working
directory, or any directory above it. Two repositories on one machine are
therefore two engines, with their own flows, runs and token. `fluksio serve`
makes one where there is none, and it ignores itself from within a
makes one where there is none, and it ignores itself from within with a
`.gitignore` of `*`, so a database and a credential cannot be committed by
accident.
@@ -47,17 +47,17 @@ Docker.
At a terminal this opens a dashboard with the engine running under it; see
[below](#the-dashboard). `--plain` prints the log stream instead, which is
also what happens with no terminal in a container, under systemd, or in CI.
also what happens with no terminal: in a container, under systemd, or in CI.
The default port moves out of the way when something already has it 8001,
8002, and so on and says which one it took; the URL written to
The default port moves out of the way when something already has it (8001,
8002, and so on) and says which one it took; the URL written to
`client.json` is the one it is actually on. A port you *asked* for is never
moved off: `--port 9000` on a taken 9000 fails, because something else is
there and you named it.
What it will *not* do is start a second engine for the same instance. If
the port is held by an engine already serving this directory, it says so and
stops one SQLite database wants one engine. Another instance's Fluksio
stops, since one SQLite database wants one engine. Another instance's Fluksio
on that port is named, and the move happens as usual.
| Option | Default | What it does |
@@ -76,8 +76,8 @@ on that port is named, and the move happens as usual.
| `--max-workers N` | 4 | python worker processes (`FLOW_MAX_WORKERS`) |
| `--gpus N` | 0 | GPUs on this machine a node may be given (`FLOW_GPUS`) |
Cards are declared rather than detected asking a vendor's tooling would make
one dependency two so a machine with a GPU reports none until `--gpus` says
Cards are declared rather than detected, since asking a vendor's tooling would
make one dependency two, so a machine with a GPU reports none until `--gpus` says
otherwise, and a node asking for one is clamped to zero and runs alongside
every other. `--gpus 1` is what serialises them.
@@ -87,7 +87,7 @@ restart. It is skipped if the instance is already enrolled.
!!! warning "One process"
`fluksio serve` holds the flow engine. A second one is a *second engine*
`fluksio serve` holds the flow engine. A second one is a *second engine*:
duplicated subscriptions, duplicated cron ticks, two webhooks answering the
same path. Run one, and distribute work with
[workers](workers.md) instead.
@@ -95,7 +95,7 @@ restart. It is skipped if the instance is already enrolled.
!!! note "`$HOME` on a cluster"
A login node's home directory is often NFS, where SQLite's write-ahead log
does not work the database would be locked or corrupt. `fluksio serve`
does not work, since the database would be locked or corrupt. `fluksio serve`
warns when it notices; point `--data-dir` at local disk.
### What it prints
@@ -116,8 +116,8 @@ dashboard is served from there rather than here.
### The dashboard
At a terminal, `serve` draws the health overview, the recent runs, and the
engine's own log in a pane below the output above is in there, not replaced
by it.
engine's own log in a pane below, so the output above is in there rather than
replaced by it.
| Key | What it does |
|---|---|
@@ -131,7 +131,7 @@ by it.
| `e` | pair with a portal, without leaving the screen |
The engine is a child process rather than a thread, which is what makes those
possible and what makes `q` a way out of the screen rather than a way to
possible, and what makes `q` a way out of the screen rather than a way to
stop the engine. Running `fluksio serve` again reattaches to it.
An engine started elsewhere is adopted rather than duplicated, and can be
@@ -140,30 +140,27 @@ beside the data and a token this directory's key signed have to agree. Another
instance's engine is named and left alone.
The screen subscribes to the engine's event bus over the same websocket a
browser uses, so a run appears the moment it starts rather than at the next
poll — which is what used to make a run that started and finished inside five
seconds visible only in the history. It reconnects on its own, quietly: an
engine stopped from this screen is a normal state, not an error to fill the
log pane with.
browser uses, so a run appears the moment it starts. It reconnects on its own,
and an engine stopped from this screen is a normal state rather than an error.
#### Comparing runs
`enter` opens a comparison of the ticked runs: one metric's curve for each,
over a table of what differs between them.
The curves are drawn in braille, five to a chart, in five distinct colours
a deliberate departure from the browser's chart palette, where the five series
The curves are drawn in braille, five to a chart, in five distinct colours: a
departure from the browser's chart palette, where the five series
step one hue by lightness. Hue is what survives a terminal. Every curve is
named in the legend either way.
Two pickers sit above the chart. The first is the metric, named as the run
records it `train.loss` rather than `loss`, since a name is qualified by the
node that published it. The second is what it is plotted against: the step,
records it (`train.loss` rather than `loss`, since a name is qualified by the
node that published it). The second is what it is plotted against: the step,
`time (s)` measured from each run's own first reading so runs started hours
apart lie on top of each other, or another metric of the same runs.
The table below keeps only what actually differs the inputs whose values are
not the same in every run, and the seed when it varies beside each run's
The table below keeps only what actually differs (the inputs whose values are
not the same in every run, and the seed when it varies) beside each run's
status, duration and final numbers. A parameter every run shared is not a
column worth scanning.
@@ -186,7 +183,7 @@ fluksio enroll ABCD-1234
Get the code from the portal under **Instances → Add instance**. It is
single-use and expires in fifteen minutes. `--as` matters when the instance
has several superusers — without it, enrolment refuses rather than guessing.
has several superusers. Without it, enrolment refuses rather than guessing.
Afterwards, `fluksio serve` dials the portal as it comes up, and keeps dialling:
a portal that restarts, a wifi that changes, a laptop that suspends and wakes
@@ -198,11 +195,11 @@ minute. See [Accounts and the portal](../interface/portal.md).
## `fluksio worker`
Runs nodes for an engine elsewhere. Everything after `worker` belongs to the
agent's own parser it is the same program `fluksio-worker` installs, so the
agent's own parser: it is the same program `fluksio-worker` installs, so the
two are interchangeable:
```sh
fluksio worker --url wss://api.example.com/api/v1/workers/attach \
fluksio worker --url wss://api.fluksio.com/api/v1/workers/attach \
--token "$FLUKSIO_WORKER_TOKEN" --labels gpu
```
@@ -211,21 +208,21 @@ See [Remote workers](workers.md).
## Talking to an engine
The commands below are the client half: they run wherever you work, and
address an engine over its API rather than being one except under `--local`,
address an engine over its API rather than being one, except under `--local`,
which boots one inside the command instead.
### `fluksio login`
```sh
fluksio login --url https://api.example.com
fluksio login --url https://api.fluksio.com
```
For an engine somewhere *else*. One you started yourself needs no login:
`fluksio serve` writes the token as it comes up and says where it put it.
The token goes in this project's `.fluksio/client.json`, or with `--global` in
`~/.fluksio/client.json`. Every command below reads it from there nearest
first, walking up from the working directory or from `FLUKSIO_URL` and
`~/.fluksio/client.json`. Every command below reads it from there, nearest
first walking up from the working directory, or from `FLUKSIO_URL` and
`FLUKSIO_TOKEN`, or from its own `--url` and `--token`. A token an older
version wrote to `~/.config/fluksio/client.json` is still read.
@@ -240,14 +237,14 @@ each one with a generated import shim per node. A directory that is a package
is walked; a dotted name is imported as it stands; nothing is loaded from a
file path, because the shim has to import the same way.
A plain directory is walked all the way down, so one folder per study
`fluksio sync dev` over `dev/s1_baseline/study.py` needs no naming. Hidden
A plain directory is walked all the way down, so one folder per study
(`fluksio sync dev` over `dev/s1_baseline/study.py`) needs no naming. Hidden
directories, `__pycache__`, `node_modules` and virtualenvs are left alone.
Each file is imported under the name its path spells beneath the directory
being synced, so `dev/s1/study.py` and `dev/s2/study.py` are `s1.study` and
`s2.study` and a `study.py` per study collides with nothing. No `__init__.py`
is needed the directories in between are namespace packages which leaves a
is needed, since the directories in between are namespace packages, which leaves a
bare `from study import ...` in a test beside it working. A file at the top of
what is synced keeps its plain name.
@@ -258,18 +255,18 @@ what is synced keeps its plain name.
| `--force` | overwrite a flow, or a node body, that was edited on the canvas |
Every sync retires the engine's workers, including one that had nothing to
upload a worker holds your package in memory, so an edit to it is invisible
upload: a worker holds your package in memory, so an edit to it is invisible
until the process goes.
It also records, per node, which of your modules that node's function imports
its way to, and what they hash to. That is what the
[stage cache](../concepts/runs.md#stage-caching) keys on, so editing a
helper a node calls into is reported as that node changing `train: updated
(flow, fit)` and re-runs it, while editing something the node never reaches
helper a node calls into is reported as that node changing (`train: updated
(flow, fit)`) and re-runs it, while editing something the node never reaches
is left alone.
A sync that changed nothing says `unchanged`, which is the answer worth
having. `published` and `draft` are said only when there was something to
having. `published` and `draft` are said only when there was something to
publish or a draft was genuinely left behind. An engine too old to store what
a node's code reaches says so in a line naming both versions; until it is
upgraded its cache is keyed on the whole repository, as it was before. See
@@ -281,7 +278,7 @@ upgraded its cache is keyed on the whole repository, as it was before. See
fluksio run train --lr 0.05 --seed 7 [--wait]
```
Syncs the working directory *and everything under it*, then submits a run so
Syncs the working directory *and everything under it*, then submits a run, so
the command after an edit is this one and nothing else, from the repository
root as readily as from the study's own folder. A study that will not import
is a warning rather than a stopped run; the upload is already a no-op for a
@@ -290,7 +287,7 @@ flow nothing changed in, so what the walk costs is importing the others.
is not free, and `--no-sync` skips it entirely.
Flags that are not its own are the flow's inputs, typed by what the flow
declares them as so a name the flow has not got is refused by name, and
declares them as, so a name the flow has not got is refused by name, and
`--param lr=0.002` is told that one value is `--lr 0.002` and several is a
[sweep](#fluksio-sweep). `--wait` blocks until the run
finishes and exits non-zero if it failed. `--follow` waits as well, and prints
@@ -309,13 +306,13 @@ timeout.
on the run, so what a result came from is answerable later. It goes into the
digest that identifies a run's inputs, so two runs of one configuration with
different seeds are different runs rather than a cache hit. And if the flow
declares an input named `seed`, that is what fills it so the number the run
declares an input named `seed`, that is what fills it, so the number the run
is labelled with is the number your code actually drew from, instead of merely
looking like it. A flow that declares no such input still records it, and
nothing reads it. Sweep over seeds with `--param seed=1,2,3`.
Any input takes what a previous run produced, named rather than typed out
a checkpoint, but equally a config object nobody wants to paste into a shell:
Any input takes what a previous run produced, named rather than typed out: a
checkpoint, but equally a config object nobody wants to paste into a shell:
```sh
fluksio run evaluate --dataset @run:1758042000123-9f2ab41c.dataset
@@ -327,7 +324,7 @@ fluksio run train --meta @run:1758042000123-9f2ab41c.dataset_meta
type; a bare digest is the content itself, resolved into a reference. Both
spellings are reserved on every input, `str` included, so an input that has to
carry one of them literally cannot. Passing the value as JSON still works and
is what a script that already holds one does the same thing
is what a script that already holds one does, the same thing
`flow.submit(dataset=run.result["dataset"])` does from Python.
Run a flow with no parameters at a terminal and it asks for them, one line per
@@ -346,15 +343,15 @@ the command, skips the questions, and `--defaults` skips them explicitly.
`--no-sync` runs what is already on the engine. Worth it in a tight loop where
you know nothing changed, since syncing retires the workers and the next call
pays its imports again. A directory that declares no flows syncs nothing and
says nothing a flow drawn on the canvas is run the same way.
says nothing: a flow drawn on the canvas is run the same way.
`--no-cache` executes every node, including one an earlier run already
answered. See [Stage caching](../concepts/runs.md#stage-caching).
`--local` boots the engine inside this process instead of talking to a served
one, so there is no `fluksio serve` terminal to keep open. It is the same
instance either way the same `.fluksio`, the same database, artifacts
and run history so a run made this way and a run made through a served
instance either way (the same `.fluksio`, the same database, artifacts and run
history) so a run made this way and a run made through a served
engine cache against each other. It always waits, because the engine it starts
lives exactly as long as the command. Starting one costs a few seconds of
worker pool and module reconcile, against the ~15 ms of submitting to an
@@ -378,11 +375,11 @@ engine that accounts for nothing.
The portal reads one of three ways. `no portal` means this instance was
never enrolled. `portal hub.fluksio.com` means the link is up. `portal
unreachable` names the error, and is the one worth acting on the dashboard is
unreachable` names the error, and is the one worth acting on: the dashboard is
served from the other end, so nobody can reach it while that is showing.
`--watch` keeps it on screen and refreshes every five seconds until Ctrl-C
the cadence the dashboard polls at, since nothing here moves faster. It needs a
`--watch` keeps it on screen and refreshes every five seconds until Ctrl-C, the
cadence the dashboard polls at, since nothing here moves faster. It needs a
terminal; without one, run it without `--watch` and the output pipes cleanly.
`--local` reads the flows and history out of this directory with no engine
@@ -398,10 +395,10 @@ fluksio runs [--flow train] [--limit 20]
The runs an engine has recorded, newest first: id, status, flow, duration, how
long ago it was submitted, the commit of the repository it came from, and the
inputs it was given. Statuses are coloured when a terminal is reading the
output `ok` green, `error` red, `cached` cyan.
output: `ok` green, `error` red, `cached` cyan.
Only the inputs that *differ from what the flow declares* are shown, and they
are clamped to what is left of the terminal's width — a run that took the
are clamped to what is left of the terminal's width. A run that took the
defaults lists none at all, and a flow taking a few kB of JSON does not push
everything else off the line. `Client.runs()` and
[`fluksio export runs`](#fluksio-export) are where the whole value is read.
@@ -414,7 +411,7 @@ to be served.
fluksio flavors
```
The named sizes a node can ask for `@node(resources="gpu-small")` with the
The named sizes a node can ask for, as in `@node(resources="gpu-small")`, with the
cores, memory and cards each stands for. Editing them is the Workers screen or
`POST /api/v1/flavors`; this is the read.
@@ -424,7 +421,7 @@ cores, memory and cards each stands for. Editing them is the Workers screen or
fluksio sweep train --param lr=0.1,0.01 --param epochs=10,50 --wait
```
Every combination of the parameter lists, submitted as one group four runs
Every combination of the parameter lists, submitted as one group: four runs
above, sharing a `group_id` and executing in parallel. Values are typed by the
flow's inputs, the same as `run`'s are, and `--seed`, `--no-sync`,
`--no-cache` and `--local` mean what they do there. `--wait` blocks until all
@@ -438,8 +435,8 @@ fluksio export metrics --flow train --name train.train_loss --stride 10 -o curve
fluksio export runs --flow train --status ok > arms.csv
```
The two tables an analysis reads. `export metrics` is the long one a row per
run, metric and step which is what a plotting library takes without
The two tables an analysis reads. `export metrics` is the long one, a row per
run, metric and step, which is what a plotting library takes without
reshaping; `--name` keeps the metrics it lists and `--stride` keeps every Nth
point of *each* curve. `export runs` is the wide one: a row per run with its
inputs as columns, its final numbers, its status, its duration and the commit
@@ -457,12 +454,12 @@ cell blank.
A node usually returns a record rather than a scalar, so both sides take
dotted paths into one: `--metrics final_metrics.train_loss,test_metrics.known.perfect`
selects three fields rather than two blobs, and `--params model.ansatz` does
the same for an input. The defaults reach the same depth every number a
the same for an input. The defaults reach the same depth: every number a
result carries becomes a column wherever it sits, and a record's inputs are
taken leaf by leaf rather than as one blob.
Metric names are flow-qualified a node of `train` writing `train_loss`
records `train.train_loss` so `--list` prints the names the selected runs
Metric names are flow-qualified (a node of `train` writing `train_loss` records
`train.train_loss`) so `--list` prints the names the selected runs
carry when the spelling is not obvious.
Both take `--flow`, `--run ID` (repeat it), `--group`, `--status`, `--since`,
@@ -470,8 +467,8 @@ Both take `--flow`, `--run ID` (repeat it), `--group`, `--status`, `--since`,
back to the run page and to what the run made.
`--format` is `csv` (the default), `jsonl` or `parquet`; output goes to stdout
unless `-o FILE` names somewhere. Parquet keeps the types and needs pyarrow
`pip install 'fluksio[parquet]'` — and a file to write, since it is not a
unless `-o FILE` names somewhere. Parquet keeps the types, needs pyarrow
(`pip install 'fluksio[parquet]'`) and needs a file to write, since it is not a
stream. In a notebook, `Client.export_metrics()` and `Client.export_runs()`
answer the same rows as a list of dicts, which `pandas.DataFrame` takes
directly.
@@ -483,7 +480,7 @@ fluksio artifacts 1758042000123-9f2ab41c
fluksio artifacts 1758042000123-9f2ab41c weights -o model.npz
```
The files a run produced what `fluksio.save_artifact(...)` wrote, and any
The files a run produced: what `fluksio.save_artifact(...)` wrote, and any
artifact a node returned. Named alone it lists them: the message each left on,
its size, its media type and the filename the node gave it. Name one and it is
written here, under that filename unless `-o` says otherwise.
@@ -496,9 +493,9 @@ The message name is the one to pass, since it is what addresses the bytes;
```text
.fluksio/ (or ~/.fluksio, with `--global`)
├── client.json the token `serve` wrote, mode 600
├── .gitignore `*` a database and a credential, ignored from within
├── .gitignore `*`, so a database and a credential are ignored from within
├── fluksio.db SQLite: users, runs, metrics, observability, agents
├── flows/ a git repository one directory per flow
├── flows/ a git repository, one directory per flow
│ ├── house/
│ │ ├── flow.json the published structure
│ │ ├── nodes/*.py the published node code
+6 -6
View File
@@ -36,8 +36,8 @@ uv add --editable ../../fluksio-connector-mydevice
```
Editing your connector's **code** then takes effect on the next engine restart,
with no reinstall. Changing its **metadata** the entry point, the version,
the package name needs `uv sync` to regenerate the installed metadata before
with no reinstall. Changing its **metadata** (the entry point, the version,
the package name) needs `uv sync` to regenerate the installed metadata before
a restart picks it up.
Two things to know about this:
@@ -84,7 +84,7 @@ class MyDevice(ConnectorNode):
Then, on the canvas: place the node, set `poll_interval` to how often the
device should be read, pick the API key from the secrets store, and add output
ports named `temperature` and `humidity`. The port names are the wiring any
ports named `temperature` and `humidity`. The port names are the wiring: any
node consuming those messages is downstream of this one.
## If the device is told rather than asked
@@ -102,7 +102,7 @@ message drives the device. Set `idempotent = False` on the class: a write is a
command, and a redelivery after a crash should not undo a newer one.
Keep the first version off the wire. A boolean setting the code checks before
it sends `artnet`'s `transmit` is the example lets a flow be built and
it sends (`artnet`'s `transmit` is the example) lets a flow be built and
watched in the logs before anything physically moves.
## If the reading is bytes
@@ -117,8 +117,8 @@ A camera or a microphone publishes a reference rather than the bytes:
Place the node with an `image`-typed output port and a Media widget draws each
frame as it lands. `fluksio-connector-test-media` publishes test frames and
tones this way, so the whole path can be wired up with no camera in the room
copy it if yours is a media device.
tones this way, so the whole path can be wired up with no camera in the room.
Copy it if yours is a media device.
## Try it without a device
+30 -31
View File
@@ -1,7 +1,7 @@
# Writing node code
A Function node is a Python file. That is all it is — no base class, no
decorator, no framework import unless you want one.
A Function node is a Python file: no base class, no decorator, no framework
import unless you want one.
```python
def process(temperature, setpoint=21.0):
@@ -22,18 +22,18 @@ See [Where a node's values come from](../concepts/values.md).
**The return value is a dict keyed by output ports.** Every value is checked
against the port's declared type before it is published. A key that is not a
declared port is an error, not a silent drop — nothing leaves a node except
declared port is an error, not a silent drop. Nothing leaves a node except
through a port it declared.
**Nothing else is importable from the engine.** Node code runs in a separate
process, on a separate interpreter, with none of Fluksio's own modules on its
path. What it can import is what the [Modules](../interface/operations.md)
screen installed which can include [your own project](#your-own-code-as-a-package),
screen installed, which can include [your own project](#your-own-code-as-a-package),
so a node need not be a self-contained file.
**A node is a pure function of its inputs.** No context object, no global
store, no handle to reach for. A running total or a debounce timer has a
specific shape see [Keeping state in a flow](../concepts/state.md).
specific shape; see [Keeping state in a flow](../concepts/state.md).
## Producing values over time
@@ -49,7 +49,7 @@ def process(lr, steps):
return {"final_loss": loss}
```
Whatever the generator `return`s at the end is the node's result what
Whatever the generator `return`s at the end is the node's result, which is what
downstream nodes read. If you never `return`, the last thing you yield is the
result instead.
@@ -60,15 +60,15 @@ Mark the port so the flow says what it does:
```
Two consequences. In a [run](../concepts/runs.md), the whole series is kept as
that run's metrics — this is why there is no `log_metric()` anywhere in the
that run's metrics, which is why there is no `log_metric()` anywhere in the
API. And **the node's timeout starts measuring silence rather than duration**:
each emission resets the deadline, so a node yielding every few seconds can run
for hours under a timeout of 300.
### `fluksio.emit`
Where a `yield` cannot reach the value comes from inside somebody else's
callback, and they call you rather than the other way round:
Where a `yield` cannot reach, because the value comes from inside somebody
else's callback and they call you rather than the other way round:
```python
import fluksio
@@ -103,19 +103,19 @@ def process(dataset):
```
`save_artifact` takes bytes or a path, stores them by their SHA-256 digest, and
returns a small reference digest, size, media type, name which is what an
returns a small reference (digest, size, media type, name) which is what an
`artifact`-typed port carries.
Because the address is the content's hash, a sweep whose fifty configs share
one preprocessed input stores it once, and a reference stays valid wherever the
store is reachable from including on another machine.
store is reachable from, including on another machine.
`fluksio artifacts RUN NAME` is how one comes back out at a terminal.
### `fluksio.logger`
A node's `print` is kept as that node's logs, and so is anything on
`fluksio.logger` the same logger the SDK exports at top level, so code that
`fluksio.logger` is the same logger the SDK exports at top level, so code that
runs both inside a node and outside one says it the same way:
```python
@@ -152,7 +152,7 @@ each carries and what rates are realistic.
Only what a node *returns* is recorded against its run. Frames yielded
along the way are replaced in state by the next one, and the artifact sweep
removes bytes nothing refers to any more which is what stops a camera
removes bytes nothing refers to any more, which is what stops a camera
filling the disk. If a particular frame matters, return it.
## Printing
@@ -161,14 +161,13 @@ each carries and what rates are realistic.
the flow editor's log panel and on the run's per-node record; the rest is
dropped, so a node printing in a loop cannot fill anything up.
Use it to debug. Do not use it to record results — a number worth keeping is an
output port, not a line of text.
Use it to debug, not to record results. A number worth keeping is an output
port, not a line of text.
## Errors
An exception fails that node's execution, not the flow. The message you see is
one line from the frame in *your* code, not a stack through the engine — that
is a deliberate choice about what is actionable.
one line from the frame in *your* code, not a stack through the engine.
The node keeps its last error visible after it recovers, so a failure that
fired an alert at 03:00 still says what it was at 09:00. It can also be
@@ -178,13 +177,13 @@ acknowledged from the canvas.
`timeout` on a node is how many seconds its code may be *silent* before it is
stopped. A yield or an `emit` resets the clock, and the first call's imports are
not charged to it a node importing torch is not being slow, it is loading.
not charged to it: a node importing torch is not being slow, it is loading.
There is no timeout by default. Training runs for hours and a node that reports
nothing is usually working, so the engine waits: what fails a call is the worker
dying, which arrives at once rather than after a deadline. Set a timeout on the
nodes where silence means stuck an HTTP call that should answer in seconds,
a loop that can spin either per node or, for all of them,
nodes where silence means stuck (an HTTP call that should answer in seconds, a
loop that can spin) either per node or, for all of them,
with `FLOW_NODE_TIMEOUT`. `timeout = 0` says explicitly that this node has none.
## Running a node somewhere else
@@ -197,7 +196,7 @@ A node declares the label of the machine it needs:
`require` (the default) waits for a worker carrying that label; `prefer` runs
locally when none is attached. A node bound to a device is compiled *on that
machine* a node importing `torch` is correct on the GPU box and a missing
machine*: a node importing `torch` is correct on the GPU box and a missing
module on the engine, so checking it here would fail something that is fine.
See [Remote workers](workers.md).
@@ -205,8 +204,8 @@ See [Remote workers](workers.md).
## Sharing code between flows
A node's source can be promoted to the shared library from its panel, and other
flows can then use it by reference. One copy, one place to edit and every
flow using it runs the edit, which is the point and also the caution.
flows can then use it by reference. There is one copy and one place to edit, and
every flow using it runs the edit.
Shared sources live in `_lib/` in the flow repository, so they are versioned
with everything else.
@@ -214,7 +213,7 @@ with everything else.
## Packages
Node code runs in a virtual environment of its own, on the instance's data
volume — deliberately separate from the one Fluksio itself runs on.
volume, separate from the one Fluksio itself runs on.
Declare what you import in [Modules](../interface/operations.md), or over the
API:
@@ -233,7 +232,7 @@ flows. An install takes effect immediately; nothing restarts.
!!! tip "If Fluksio is installed in the venv you work in, skip this"
Node code then runs on that environment, so your project and everything it
imports are already importable see
imports are already importable; see
[Getting started: data science](../getting-started/data-science.md). What
follows is for a Fluksio with a venv of its own, which is what a container
always has.
@@ -246,7 +245,7 @@ installable like any other dependency:
numpy>=2
```
A node body then imports it, and the logic stays where it already lives in
A node body then imports it, and the logic stays where it already lives: in
your repository, under your own version control, importing its own siblings:
```python
@@ -263,7 +262,7 @@ modules as it likes, and nothing was copied.
!!! tip "You can have those three lines written for you"
Decorate `fit` with `@node(...)` where it is defined, say which nodes make
a flow with `Flow(...)`, and `fluksio sync` generates the body above
a flow with `Flow(...)`, and `fluksio sync` generates the body above,
along with the flow document, so there is nothing to PUT by hand. The
declaration lives beside the function it describes and is checked against
its signature. See
@@ -271,7 +270,7 @@ modules as it likes, and nothing was copied.
!!! warning "Editable, but not live"
`-e` means edits reach the venv without reinstalling but a node's process
`-e` means edits reach the venv without reinstalling, but a node's process
already holds the imported module in memory. The engine's workers are
long-lived, so a change to your code is picked up when they are retired,
which is what **Apply** does. Pressing it after an edit is the loop.
@@ -279,20 +278,20 @@ modules as it likes, and nothing was copied.
A [worker](workers.md) you attach yourself is the exception: it starts a
process per call, so it reads your code fresh every run. If you are
iterating on the code many times an hour, point one at your own
interpreter `fluksio-worker --python "$(which python)"` and mark the
interpreter (`fluksio-worker --python "$(which python)"`) and mark the
node with its label.
!!! note "The path is a deployment detail"
It is resolved on whichever machine runs the node, and the manifest is
committed to the flow repository so an absolute path from your laptop
committed to the flow repository, so an absolute path from your laptop
means nothing inside a container or on a GPU box. Those need their own
install of the same project; a VCS requirement
(`myresearch @ git+ssh://…@a1b2c3d`) travels where a path does not.
## A worked example
The repository ships a small supervised fit as a seedable demo three nodes,
The repository ships a small supervised fit as a seedable demo: three nodes,
a batch flow, streaming metrics, artifacts between stages, and a GPU-labelled
node that falls back to the engine when no worker is attached. It is the
shortest complete thing to read:
+28 -30
View File
@@ -6,7 +6,7 @@ network as the other.
A **worker** is a process that runs the code of nodes marked for it. It dials
*out* to the engine over one authenticated websocket, so nothing on that
machine has to be reachable and nothing has to expose the engine's state
machine has to be reachable, and nothing has to expose the engine's state
backend across hosts, which it never should.
## Install and attach
@@ -15,13 +15,13 @@ backend across hosts, which it never should.
pip install fluksio-worker
fluksio-worker \
--url wss://api.example.com/api/v1/workers/attach \
--url wss://api.fluksio.com/api/v1/workers/attach \
--token "$FLUKSIO_WORKER_TOKEN" \
--labels gpu,cuda12 \
--python /opt/torch-venv/bin/python
```
`fluksio-worker` is its own distribution the agent, the node runner, and
`fluksio-worker` is its own distribution: the agent, the node runner, and
`websockets`. Nothing of the engine, so a GPU box does not install a database
driver in order to run a training step. An engine host already has it, and
`fluksio worker …` is the same program.
@@ -40,13 +40,13 @@ driver in order to run a training step. An engine host already has it, and
| `--ram-mb` | what the job or the machine has | memory to advertise, in MB |
| `--max-idle` | never | stop after this many seconds with nothing running |
`--python` is the important one. It is how this machine keeps its own wheels
the CUDA build, the vendor SDK, the thing that will not install anywhere else
`--python` is the important one. It is how this machine keeps its own wheels
(the CUDA build, the vendor SDK, the thing that will not install anywhere else)
without the engine ever installing them or knowing about them.
### What it says it has
A worker reports its inventory when it attaches cores, GPUs and memory and
A worker reports its inventory when it attaches (cores, GPUs and memory) and
the engine schedules against it: a node asking for two cores and a GPU goes to
a machine that has them free, not merely to one carrying the right label.
@@ -58,12 +58,12 @@ or `FLUKSIO_WORKER_GPUS`) or something you say with `--gpus`. A worker that
reports nothing still attaches and is scheduled by its label alone, as every
worker was before any of them reported anything.
The engine tells each call what it may use thread caps, and the devices it
The engine tells each call what it may use: thread caps, and the devices it
may see. The worker starts a process per call, so it applies them at the only
moment a numerical library still reads them: before the import.
`--max-idle` is for a worker something else started for one job a batch
scheduler, say. It exits when nothing has run for that long, so the allocation
`--max-idle` is for a worker something else started for one job, a batch
scheduler say. It exits when nothing has run for that long, so the allocation
goes back rather than idling until its walltime.
## Mint the token
@@ -75,15 +75,15 @@ curl -X POST $FLUKSIO/workers/tokens -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"name": "gpu-dev"}'
```
Shown once, valid for a year a worker is a machine somebody sets up and
Shown once, valid for a year, since a worker is a machine somebody sets up and
leaves running. It is signed with the same keypair agent tokens use, so
rotating that key revokes every worker along with them.
??? note "A host where pip is not an option"
The two files work copied into one directory and run with `python agent.py
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime`
it is the same module its own local workers run, deliberately standard
…`. The engine serves the runner itself at `GET /api/v1/workers/runtime`.
It is the same module its own local workers run, deliberately standard
library only.
## Send a node to it
@@ -107,10 +107,10 @@ A node declares the label of the machine it needs:
`prefer` is what makes a flow work before the GPU box exists. `require` is what
you want once it does.
!!! note "Set from the API"
!!! note "Not in the panel yet"
`device` and `device_policy` are not yet fields in the node panel. Set them
with `PUT /flows/{name}`.
`device` and `device_policy`, which machine a node's code runs on, are set
through the API rather than the panel, with `PUT /flows/{name}`.
## What follows from this
@@ -120,15 +120,15 @@ you want once it does.
`torch` is correct on the GPU box and a missing module on the engine, so
checking it here would fail something that is fine.
- **`import fluksio` inside a node is the worker's own reporter.** `emit`,
`save_artifact`, `load_artifact` installed before your code runs, so an
`save_artifact`, `load_artifact`, installed before your code runs, so an
installed `fluksio` package on that box never shadows it.
- **Cancelling a run kills what it is executing**, there or here, and leaves
other runs of the same node alone.
- **If the worker disappears mid-call**, the run fails in seconds with
`worker went away mid-call` rather than waiting out its timeout.
- **A worker sends a heartbeat every ten seconds while it executes**, so a long
node is distinguishable from a dead socket. Ninety seconds of nothing at all
not even a heartbeat fails the call as gone. A heartbeat says the *agent* is
node is distinguishable from a dead socket. Ninety seconds of nothing at all,
not even a heartbeat, fails the call as gone. A heartbeat says the *agent* is
alive and nothing about the node, so it never satisfies a node's own timeout:
one set to thirty seconds fires after thirty seconds of the node reporting
nothing, wherever it runs.
@@ -154,8 +154,8 @@ environment, and what it says it has: cores, GPUs and memory.
## Upgrading
The engine and the worker speak a version-matched protocol, and a worker
announcing anything else is refused rather than half-understood. Protocol 2
the one that carries inventory is `fluksio-worker` 0.2.0. An older agent is
announcing anything else is refused rather than half-understood. Protocol 2,
the one that carries inventory, is `fluksio-worker` 0.2.0. An older agent is
told so on the socket and stops, rather than retrying against an engine that
will never accept it; `pip install -U fluksio-worker` on that host is the whole
upgrade. Nothing changed in the runner served at `GET /api/v1/workers/runtime`,
@@ -163,7 +163,7 @@ so a host that copies its two files copies the same one as before.
## Machines from a batch scheduler
A cluster is not a machine that attaches and stays it is a queue somebody else
A cluster is not a machine that attaches and stays; it is a queue somebody else
owns. So Fluksio does not submit *nodes* to Slurm. It submits a job whose payload
is an ordinary worker dialling back in, and from there everything works the way
it already does: the same protocol, the same artifacts, the same cancellation.
@@ -176,7 +176,7 @@ Write the clusters into `provisioners.json` beside the flows:
"name": "hpc",
"login": "me@login.cluster",
"ssh_key": "/secrets/hpc_ed25519",
"engine_url": "wss://api.example.com/api/v1/workers/attach",
"engine_url": "wss://api.fluksio.com/api/v1/workers/attach",
"max_idle_s": 300,
"provision_timeout_s": 900,
"profiles": [{
@@ -192,12 +192,10 @@ Write the clusters into `provisioners.json` beside the flows:
A **profile** is what the scheduler is asked for, where a flavor is what a node
asks for. They are separate on purpose, and agree when you set them up to.
When a node needs a machine nothing attached can give, and a profile would fit,
the engine `sbatch`es one over ssh — the system `ssh`, so nothing new is
installed — and the run waits meanwhile, saying so. `prerun` owns the
environment: a `module load`, a venv with `fluksio-worker` already in it. There
is deliberately no `pip install` in the generated script, because what is
installed on a cluster is somebody's decision and not this program's.
When a node needs a machine nothing attached can give, and a profile fits, the
engine `sbatch`es one over the system `ssh`, and the run waits meanwhile, saying
so. `prerun` owns the environment: a `module load`, or a venv with
`fluksio-worker` already in it. The generated script runs no `pip install`.
One outstanding request per profile, however often it is asked for. A job that
never attaches within `provision_timeout_s` is `scancel`led, as is anything
@@ -219,8 +217,8 @@ queue does, which is what a cluster that queues overnight needs, and
## What a worker is not
It is not a second engine. Subscriptions, schedules, webhooks, the dashboards
and the run queue all stay in one process — that is what keeps a message having
one definition and a cron tick happening once. A worker executes node bodies.
and the run queue all stay in one process, which keeps a message having one
definition and a cron tick happening once. A worker executes node bodies.
Running two engines against one data directory is not supported. Distribute
work with workers.
+12 -14
View File
@@ -1,7 +1,6 @@
# Flows, nodes and messages
Three ideas hold the whole system up. They are worth twenty minutes, because
almost everything else follows from them.
Three ideas hold the system up. Almost everything else follows from them.
## A flow is a graph you did not draw
@@ -22,21 +21,20 @@ def process(temperature, setpoint=21.0):
`decide` is downstream of `read` because it needs `temperature` and `read`
produces it. Nobody drew a wire.
This is the one structural decision everything else rests on, so it is worth
being explicit about the consequences:
The consequences:
- **Fan-in is free.** Two nodes providing `temperature` are two producers of
one message. The consumer does not change.
- **A node runs when something it reads was published.** Not merely when
something upstream of it ran: a node that produced nothing this time held
back by a rate limit, say leaves what reads it on the value it already
something upstream of it ran: a node that produced nothing this time (held
back by a rate limit, say) leaves what reads it on the value it already
has, and so does everything behind that.
- **A wire cannot be wrong.** There is no wire. There is a name that either
matches or does not, and the canvas tells you at edit time which it is.
- **Layout is not a document.** The canvas computes the arrangement, so a flow
has no stored positions to maintain, merge or fight over.
- **Flows stay small.** A graph nobody can hand-arrange is one worth keeping
small which is the intent. Several atomic flows that name each other beat
small, which is the intent. Several atomic flows that name each other beat
one flow with sixty nodes in it.
### Message names are namespaced
@@ -65,7 +63,7 @@ def process(reading, unit="C"):
return {"shown": reading if unit == "C" else reading * 1.8 + 32}
```
`reading` is a port. `unit` is a **setting** a constant of this node's code,
`reading` is a port. `unit` is a **setting**: a constant of this node's code,
typed into its panel and stored with the flow. Both arrive as arguments, which
is why a setting may not share a name with a port. See
[Where a node's values come from](values.md).
@@ -86,7 +84,7 @@ Types are not decoration. They are what lets the dashboard editor offer you
only the messages a gauge can actually draw, and what lets the canvas refuse a
binding before anything runs. See [Payload types](../reference/payload-types.md).
Everything on the wire is JSON. Bytes a checkpoint, an image, a model
Everything on the wire is JSON. Bytes (a checkpoint, an image, a model)
travel as an `artifact`: the bytes go to a content-addressed store and the
message carries a small reference to them. The media types are that same
reference, saying what kind of bytes are behind it.
@@ -98,7 +96,7 @@ once at the end: a training loss, a progress fraction, a frame from a camera, a
second of speech. A node publishes on one by being a generator and yielding, or
by calling `fluksio.emit`.
Each value is delivered to the nodes reading it, in the order it was produced
Each value is delivered to the nodes reading it, in the order it was produced,
so a recogniser slower than the microphone in front of it still sees every
chunk rather than only the newest. What is in state remains the latest value,
which is what everything else reads, and what a run keeps is the whole series.
@@ -115,7 +113,7 @@ store, no handle to reach for.
That is deliberate: a node with hidden state cannot run twice in parallel,
cannot be replayed, and cannot be moved to another machine. Plenty of real
automations do need to remember something, and there is a specific way to say
so see [Keeping state in a flow](state.md).
so; see [Keeping state in a flow](state.md).
## Two shapes of flow
@@ -142,8 +140,8 @@ Every flow has a published version and, while you are working, a draft.
- **Publishing** promotes the draft. The engine reloads and picks it up.
- **Discarding** throws the draft away.
The store is a git repository `flow.json` for the structure, `nodes/*.py` for
the code and each save is a commit. So a flow's history is readable with
The store is a git repository (`flow.json` for the structure, `nodes/*.py` for
the code) and each save is a commit. So a flow's history is readable with
ordinary git tooling, and copying a flow between instances is copying a
directory.
@@ -166,7 +164,7 @@ to:
| `node_unhealthy` | the node loaded but is not working: a connector that cannot reach its device, or whose last publication failed |
A flow with any of these except the advisory one and `node_unhealthy` does not
run — a node reporting itself down is a live condition, not a build error, so
run. A node reporting itself down is a live condition, not a build error, so
the rest of the flow keeps going and the issue clears by itself once the node
reports well again. The health summary on Home counts them, so "why is nothing
happening?" has an answer that does not involve reading logs.
+49 -52
View File
@@ -1,6 +1,6 @@
# Runs: pipelines that finish
A flow that automates a house never ends a value arrives, nodes fire, and it
A flow that automates a house never ends: a value arrives, nodes fire, and it
waits for the next one. A research pipeline is the other shape: parameters go
in, stages execute in order, and at some point it is *done* and has produced
something worth keeping. Fluksio calls the second one a **run**, and it is the
@@ -8,7 +8,7 @@ same engine either way.
This is what makes Fluksio usable where Kedro, MLflow or ClearML would be: a
run has parameters that identify it, a result, per-step metrics, artifacts and
a place in a queryable history without a second server, and without paying a
a place in a queryable history, with no second server and no
project bootstrap on every execution.
## A batch flow
@@ -28,15 +28,15 @@ Set `mode: "batch"` on the flow and name the messages its result should hold:
}
```
A batch flow is built and validated like any other it appears on the canvas,
its ports are type-checked but it is never *activated*: no subscriptions, no
A batch flow is built and validated like any other (it appears on the canvas,
its ports are type-checked) but it is never *activated*: no subscriptions, no
schedules, no webhooks. It runs when a run asks it to, and not otherwise.
Its `inputs` are its parameters. A run supplies values for them; anything it
does not supply keeps the declared initial value.
One thing a batch flow may not do is rate-limit a port (`interval`). A rate
limit holds a value back for a timer to release, and a run has no timer the
limit holds a value back for a timer to release, and a run has no timer, so the
value would be dropped rather than delayed, so submitting is refused instead.
## Submitting
@@ -51,7 +51,7 @@ The answer is immediate and the run is `queued`; a training run is measured in
hours, so nothing waits for it. Poll `GET /api/v1/runs/{id}` for its status,
result, per-node record and artifacts.
Wrong parameters are refused before anything executes an undeclared name, or
Wrong parameters are refused before anything executes: an undeclared name, or
a value of the wrong type, comes back as a 422 naming the problem.
### Sweeps
@@ -98,20 +98,17 @@ Mark the port it streams on, so the flow says what it does:
{"name": "loss", "dtype": "float", "stream": true}
```
Every yield is published the instant it happens same port, same type check,
Every yield is published the instant it happens: same port, same type check,
same place on the canvas as any other value. Whatever the generator `return`s
is the node's result, and is what downstream nodes read. If you never
`return`, the last thing you yield is the result instead.
This is the whole reason the framework does not have a logging API. A metric
that escapes through `log_metric()` is undeclared: invisible to validation,
absent from the canvas, and stored somewhere the graph knows nothing about.
A metric that leaves through a port is a message — so a chart binds to it
directly, a downstream node can consume it, and the run keeps its series
without anyone asking.
There is no `log_metric()`. A metric leaves through a port, which makes it a
message: a chart binds to it directly, a downstream node can consume it, and the
run keeps its series.
**Where a yield cannot reach** — the value comes from inside somebody else's
callback, and they call you rather than the other way round `fluksio.emit`
callback, and they call you rather than the other way round. `fluksio.emit`
writes the same ports the same way:
```python
@@ -129,8 +126,8 @@ def process():
Every number a node emits is kept as the run's series, stepped by the count of
emissions on that message. Read one back with
`GET /api/v1/runs/{id}/metrics?name=<flow>.loss` or leave `name` off for
every series the run kept or compare runs:
`GET /api/v1/runs/{id}/metrics?name=<flow>.loss`, or leave `name` off for
every series the run kept, or compare runs:
```
GET /api/v1/runs/series/compare?ids=<a>,<b>,<c>&metric=<flow>.loss
@@ -141,15 +138,15 @@ training curves side by side is a widget binding. During a run the values also
arrive live on the flow socket, so a chart bound to the port fills in as the
training goes.
A streaming port may set `interval` to thin out what reaches the canvas — the
A streaming port may set `interval` to thin out what reaches the canvas. The
run's history still keeps every value, because the interval is asking for the
display not to be flooded, not for the curve to have holes in it.
Emitting has a second effect: **a node's timeout measures silence, not
duration**. A node that yields every few seconds can run for hours under a
`timeout` of 300; one that says nothing for longer than its timeout is killed.
There is no timeout unless one is set a training node that reports nothing is
usually working — so set `timeout` where going quiet means stuck, at how long
There is no timeout unless one is set, since a training node that reports
nothing is usually working. Set `timeout` where going quiet means stuck, at how long
the node may plausibly be.
In a *live* flow, an emission also wakes whatever is downstream of it, exactly
@@ -160,8 +157,8 @@ finished" with nothing to mean.
## Artifacts
Bytes never travel as a message. `save_artifact` writes them to a
content-addressed store and returns a small reference digest, size, media
type, name which is what an `artifact`-typed port carries:
content-addressed store and returns a small reference (digest, size, media
type, name) which is what an `artifact`-typed port carries:
```python
def process(weights): # requires: weights, dtype "artifact"
@@ -177,8 +174,8 @@ downloadable at `GET /api/v1/artifacts/{digest}`.
## Stage caching
A run mostly does not redo what an earlier one already did. Before a node
executes it is fingerprinted a sha256 over its source, its settings, the ports
it declares and the values it is about to read — and if some earlier run of that
executes it is fingerprinted with a sha256 over its source, its settings, the
ports it declares and the values it is about to read. If some earlier run of that
same fingerprint finished, what that one returned is restored into this run's
state and the node is skipped. It is recorded with the status `cached` and a
duration of zero, and its artifacts are listed on the new run as well, so they
@@ -192,14 +189,14 @@ on a run carries the `cache_key` it was looked up by.
For a [code-defined flow](../getting-started/data-science.md), "its source" is
the generated shim, which imports the real function and does not change when
that function does. So the key carries one thing more: a digest of the project
modules that node's function reaches, worked out by `fluksio sync` which is
the only side that imports your code and can see what it imports and read
modules that node's function reaches, worked out by `fluksio sync`, which is
the only side that imports your code and can see what it imports, and read
again from those files when the run starts.
Editing a helper three calls down from the node invalidates it, which is the
point: the alternative is a re-run answering with the previous code's numbers.
Editing something the node does not reach leaves the hit standing, which is
the other half a notebook two directories away is not a reason to retrain.
the other half: a notebook two directories away is not a reason to retrain.
The walk follows imports statically and stops at the standard library, at
anything installed, and at Fluksio itself; a module imported under a name the
code computes is not followed. An engine that cannot see the files keeps what
@@ -207,20 +204,20 @@ sync recorded instead of nothing, so a worker on another machine no longer
keys every run the same.
The run history *is* the cache; there is no second store. A node's returned
outputs are kept on its run record as canonical JSON, up to 256000 characters
a node returning more than that is simply not cacheable that run. An entry
outputs are kept on its run record as canonical JSON, up to 256000 characters.
A node returning more than that is not cacheable that run. An entry
whose artifact bytes have since left the store is a miss, not an error.
Only `python` nodes are cached, and by default all of them are. A built-in node
type or a connector node has side effects and no source to fingerprint, so
neither is ever a candidate. Turn it off for one node with
`@node(..., cache=False)` the flow document carries it as `cache`, so the
canvas and the API can change it too or for one run with
`@node(..., cache=False)` (the flow document carries it as `cache`, so the
canvas and the API can change it too) or for one run with
`fluksio run --no-cache`, `fluksio sweep --no-cache`, or `"no_cache": true` in
the submission body.
A cached node replays no emissions those values were the story of an
execution that is not happening this time so its series is not rewritten
A cached node replays no emissions, since those values were the story of an
execution that is not happening this time, so its series is not rewritten
either. The run it was restored from is recorded instead, and that is where the
curve is read back from: asking the reusing run for its metrics answers with
the same points, under its own flow's names. The one way to be left with a
@@ -229,7 +226,7 @@ deleting its flow does.
## Objects that cannot be serialized
A live model, a `DataLoader`, a JAX-compiled function — these do not cross a
A live model, a `DataLoader`, a JAX-compiled function: none of these cross a
node boundary, and no framework flag will make them. There are exactly two
patterns, and they are both deliberate:
@@ -249,7 +246,7 @@ A node that needs a GPU declares the label of the machine that has one:
```
A worker on that machine dials **out** to the engine, because the engine
generally cannot reach it different network, no inbound route and because
generally cannot reach it (different network, no inbound route) and because
nothing should expose Redis across hosts. Install it on the box, mint it a
token, and start it:
@@ -259,13 +256,13 @@ curl -X POST $FLUKSIO/workers/tokens -d '{"name": "gpu-dev"}' # once, as an ad
pip install fluksio-worker
fluksio-worker \
--url wss://api.example.com/api/v1/workers/attach \
--url wss://api.fluksio.com/api/v1/workers/attach \
--token "$FLUKSIO_WORKER_TOKEN" \
--labels gpu,cuda12 \
--python /opt/torch-venv/bin/python
```
`fluksio-worker` is its own distribution the agent, the node runner, and
`fluksio-worker` is its own distribution: the agent, the node runner, and
`websockets`. Nothing of the engine, so a GPU box does not install a database
driver to run a training step. Where pip is not an option, the two files still
work copied into one directory and run with `python agent.py …`; the engine
@@ -278,7 +275,7 @@ travels with every call, so nothing has to be deployed there.
A few consequences worth knowing:
* `import fluksio` inside a node is the worker's own reporter — `emit`,
`save_artifact`, `load_artifact` installed before the node's code runs, so
`save_artifact`, `load_artifact`, installed before the node's code runs, so
the installed `fluksio` package (if the box has one) never shadows it.
* A node bound to a device is **compiled on that machine**. A node importing
`torch` is correct on the GPU box and a missing module on the engine, so
@@ -295,32 +292,32 @@ A few consequences worth knowing:
## Durability
Submitting journals the run to a Redis stream of its own, separate from the
one the automations use a burst of five hundred sweep runs must not stand
one the automations use: a burst of five hundred sweep runs must not stand
between a house and its heating. An engine that is down when a run is
submitted picks it up when it starts.
From the moment a run is claimed, its database row is the record and the queue
is finished with it. Redelivering two hours of training because an
acknowledgement was late is not recovery; instead a running run refreshes a
lease, and one whose lease goes stale is marked `abandoned` which is what a
lease, and one whose lease goes stale is marked `abandoned`, which is what a
run whose engine was killed mid-training becomes.
## Looking at what ran
The **Runs** screen is the experiment log: every run newest-first, filtered by
flow, by status, or down to one sweep. A sweep is worth filtering to the
flow, by status, or down to one sweep. A sweep is worth filtering to, since the
table then draws a column per parameter that actually varied, which is what
makes fifty runs of one flow readable.
Tick two or more and their curves go side by side — shift-click to take a
Tick two or more and their curves go side by side. Shift-click to take a
range, or the header box to take everything on screen. That comparison is the
address, so a link to it is a link someone else can open.
The curves are drawn against the step by default. They can also be drawn
against elapsed seconds, which answers "which one got there sooner" and is
measured from each run's own first reading so that runs started hours apart
still lie on top of each other; or against another metric of the same runs
an epoch, or samples seen joined on the step the two share.
still lie on top of each other; or against another metric of the same runs (an
epoch, or samples seen) joined on the step the two share.
One run in full is params, the per-node record with its logs and traceback,
the artifacts it made, its metrics and its result.
@@ -328,8 +325,8 @@ the artifacts it made, its metrics and its result.
### Taking it into a dataframe
An analysis wants a table rather than a screen, and there are two it usually
wants. `fluksio export metrics` is the long one a row per run, metric and
step — and `fluksio export runs` is the wide one, a row per run with the
wants. `fluksio export metrics` is the long one, a row per run, metric and
step; `fluksio export runs` is the wide one, a row per run with the
inputs that varied as columns beside its final numbers:
```python
@@ -343,20 +340,20 @@ arms = pd.DataFrame(client.export_runs(flow="train", status="ok"))
The run id is on every row of both, so a curve joins to the arm it came from
and to the run page it was recorded on, and the wide table carries the commit
and the code digest an exported file says what produced its numbers.
and the code digest, so an exported file says what produced its numbers.
[The CLI](../code/cli.md#fluksio-export) writes the same rows as csv, jsonl or
parquet, which is where an export belongs: in the script beside the analysis.
A node's numbers usually arrive inside a record, and the wide table follows
them in: every number a result carries is a column of its own, named by its
path `metric.final_metrics.train_loss` and `metrics="final_metrics.train_loss"`
path (`metric.final_metrics.train_loss`) and `metrics="final_metrics.train_loss"`
selects one. Inputs work the same way, and are compared leaf by leaf, so two
configurations differing in one field give that field as the axis.
### A dashboard, read against runs
A run records values under the same names a dashboard binds to a run of
`study` writes `study.loss` so a dashboard is already a way of looking at
A run records values under the same names a dashboard binds to (a run of
`study` writes `study.loss`) so a dashboard is already a way of looking at
one. **Open in dashboard** from a run or a comparison opens any dashboard with
`?runs=a,b,c` on it, and the widgets resolve from those runs instead of from
the live engine: each chart draws a line per run, the tiles that show one
@@ -367,7 +364,7 @@ Nothing about the dashboard is specific to runs. The page built to watch a
training run happen is the page that shows the finished ones.
If a flow has no dashboard yet, the same menu offers to build one from what
the flow declares a chart per streaming port, a tile per output:
the flow declares: a chart per streaming port, a tile per output:
```
POST /api/v1/dashboards/from-flow/study
@@ -386,8 +383,8 @@ sweep, or specific runs. It re-reads on its own and whenever a run finishes.
### When a run draws nothing
A node restored from the [stage cache](#stage-caching) has its curve read back
from the run that recorded it. Delete that run deleting its flow does — and
the reusing run is left with a result and an empty curve, and the chart says so
from the run that recorded it. Delete that run, which deleting its flow does,
and the reusing run is left with a result and an empty curve, and the chart says so
rather than looking broken.
## What this costs, compared
+6 -6
View File
@@ -1,7 +1,7 @@
# Keeping state in a flow
Logic nodes are pure functions of their inputs. There is no `context`, no
`global`, no handle to a store — a node is called with the values of the
`global`, no handle to a store. A node is called with the values of the
messages it declares and returns the values of the messages it provides.
That is deliberate: a node with hidden state cannot be run twice in parallel,
@@ -10,7 +10,7 @@ automations need to remember something. This is how.
## State is a message the node both reads and writes
A running total, a debounce timer, a last-seen reading each is a value that
A running total, a debounce timer, a last-seen reading: each is a value that
survives between runs. Give it a message name, declare it as both an input and
an output, and it is state:
@@ -36,7 +36,7 @@ first call is well-formed even if the value is missing.
## Feeding a value back between two nodes
Sometimes the value comes from a different node a controller reading back what
Sometimes the value comes from a different node, such as a controller reading back what
an estimator computed from its own last output. Written plainly that is a cycle,
and the validator rejects it, because a graph where A waits for B and B waits
for A can never start.
@@ -51,7 +51,7 @@ Say so, by marking the input non-triggering:
A non-triggering input:
- creates no dependency, so it cannot form a cycle;
- never makes the node wait — if the message has no value yet, the port is
- never makes the node wait. If the message has no value yet, the port is
simply left out of the call, and the function's default applies;
- is read fresh from state whenever the node does run, for whatever reason.
@@ -60,8 +60,8 @@ input, and marking it non-triggering would mean the node never runs at all.
## What still holds engine-side state
Built-in nodes that are *about* time or change rate limiting, filter-on-change,
delay, cron keep what they need in the engine's own state backend, under keys
Built-in nodes that are *about* time or change (rate limiting, filter-on-change,
delay, cron) keep what they need in the engine's own state backend, under keys
that never appear as messages. That is the engine's business, not a flow's: the
line is that node *code* you write never reaches for state, while node *types*
the engine ships may.
+3 -3
View File
@@ -28,7 +28,7 @@ def process(reading, unit="C"):
return {"shown": reading if unit == "C" else reading * 1.8 + 32}
```
Because both arrive by name, a setting cannot share a name with a port — the
Because both arrive by name, a setting cannot share a name with a port. The
node reports it rather than picking one.
Settings are part of the flow document, so changing one is an edit that gets
@@ -41,7 +41,7 @@ consumer visibly downstream of it.
## Flow inputs: what arrives from outside
Some messages are not computed by any node in the flow a dashboard control
Some messages are not computed by any node in the flow, such as a dashboard control
writes them, the API publishes them, a batch run passes them in. Declare those
as the flow's **inputs**, with the value they start from:
@@ -54,7 +54,7 @@ and the canvas says so. With it, the flow starts from 21.0 and whatever writes
the message afterwards takes over.
The canvas draws each one as a labelled endpoint feeding the nodes that read
it, the same way it draws a dashboard tile or another flow so a value never
it, the same way it draws a dashboard tile or another flow, so a value never
appears from nowhere. They are edited in the flow's own panel.
A batch flow's inputs are also its run parameters, and its **outputs** name what
+65 -66
View File
@@ -32,8 +32,8 @@ Fluksio 0.1.0 — data in /home/you/my-research/.fluksio
```
Read the last two lines: you are already signed in. Signing in to your own
machine is a formality the password was printed by the same process that
would have checked it so `serve` writes the token itself and every command
machine is a formality, since the password was printed by the same process that
would have checked it, so `serve` writes the token itself and every command
below just works. `fluksio login` is for an engine somewhere *else*.
**Write that password down** anyway. It is shown once, and it is what the
@@ -41,7 +41,7 @@ dashboard asks for.
### One instance per project
`.fluksio` sits beside your code, and is found the way `.git` is from the
`.fluksio` sits beside your code, and is found the way `.git` is: from the
directory you are standing in, or any directory above it. So two repositories
on one machine are two engines: separate flows, separate run history,
separate token, and no chance of one experiment's graph turning up in the
@@ -59,7 +59,7 @@ other's.
climate/
```
It holds a database and a credential, so it ignores itself from within a
It holds a database and a credential, so it ignores itself from within with a
`.gitignore` of `*`, the same thing `uv` writes into `.venv`. Nothing to add
to your project's own.
@@ -67,7 +67,7 @@ Give them different ports (`--port`) if you want two running at once.
!!! tip "One engine for the machine instead"
`fluksio serve --global` uses `~/.fluksio` shared by every directory,
`fluksio serve --global` uses `~/.fluksio`, shared by every directory,
which is what you want for a personal server rather than a project. When
both exist, the banner says which one you are looking at and how to reach
the other. `--data-dir` still names any directory outright, which is worth
@@ -77,10 +77,9 @@ Give them different ports (`--port`) if you want two running at once.
!!! tip "Keep it running"
The engine is meant to be resident. That is the whole reason submitting a
run costs about 15 ms instead of the second a project-bootstrapping
orchestrator spends before it does anything. Leave it in a `tmux` window,
or write a small `systemd --user` unit for it.
The engine is meant to be resident, which is what makes submitting a run
cost about 15 ms. Leave it in a `tmux` window, or write a small
`systemd --user` unit for it.
## Talking to it over HTTP
@@ -96,13 +95,13 @@ While you are experimenting, the interactive schema at
<http://127.0.0.1:8000/docs> is the fastest way to see what is available.
For an engine on another machine, `fluksio login --url https://…` asks for a
password and stores the token the same way in this project's `.fluksio`, or
password and stores the token the same way: in this project's `.fluksio`, or
with `--global` in `~/.fluksio`.
## Your packages are already there
If you installed Fluksio into the environment you work in the venv that
already has torch or numpy in it that is the environment your nodes run on.
If you installed Fluksio into the environment you work in (the venv that
already has torch or numpy in it) that is the environment your nodes run on.
Nothing to declare, nothing to install twice:
```sh
@@ -119,8 +118,8 @@ fluksio serve
your environment, adopted. Add packages with pip.
```
That venv is yours. Add a package the way you added the rest `pip install
scikit-learn` and `fluksio sync` (or a restart) retires the workers so they
That venv is yours. Add a package the way you added the rest (`pip install
scikit-learn`) and `fluksio sync` (or a restart) retires the workers so they
pick it up. The Modules screen lists what is installed and stays read-only,
because the alternative would be Fluksio deciding what belongs in an
environment it did not make.
@@ -130,7 +129,7 @@ environment it did not make.
The cost of not having two environments: a package the engine depends on
is one you can now upgrade out from under it. In practice this is what
everybody does with every other tool in the venv, and the answer when it
bites is the same pin it back, or keep Fluksio somewhere separate with
bites is the same: pin it back, or keep Fluksio somewhere separate with
the venv of its own below.
### A venv of Fluksio's own
@@ -144,7 +143,7 @@ or an environment too precious to let a node's dependency near. Set
a venv of its own; the Modules screen installs into it.
```
Then the Modules screen is how packages get in a pip manifest, installed
Then the Modules screen is how packages get in: a pip manifest, installed
with `uv pip sync` and versioned alongside your flows, so what a run imported
is recorded with what it ran:
@@ -156,7 +155,7 @@ curl -X POST $FLUKSIO/modules/apply -H "Authorization: Bearer $TOKEN" \
Adding a package takes effect immediately; nothing restarts. The Docker image
sets `NODE_VENV=managed` for itself, because the venv in it holds the app and
nothing of yours so a container is always this case.
nothing of yours, so a container is always this case.
!!! tip "A GPU box works the same way"
@@ -169,13 +168,13 @@ nothing of yours — so a container is always this case.
--token "$WORKER_TOKEN" --labels gpu
```
Then give the node `device="gpu"` see [Remote workers](../code/workers.md).
Then give the node `device="gpu"`; see [Remote workers](../code/workers.md).
## Say which functions are nodes
A **flow** is a graph of nodes. A **batch flow** is one that runs on demand
from parameters to a result, which is what an experiment is. A node is one of
your own functions — it stays in your repository, imported by its siblings as
your own functions. It stays in your repository, imported by its siblings as
it always was.
Say your project looks roughly like this:
@@ -222,13 +221,13 @@ def fit(dataset, lr, epochs=25):
```
`yield` replaces `print`. Each one publishes on the node's `loss` port the
instant it happens, and the run keeps every value as a series which is why
instant it happens, and the run keeps every value as a series, which is why
Fluksio has no `log_metric()` call. A metric that leaves through a port is a
message like any other: a chart can bind to it, a downstream node can consume
it, and it shows up on the canvas. A metric that escapes through a logging
function is invisible to all three.
`fluksio.save_artifact` handles the things too big to be messages a
`fluksio.save_artifact` handles the things too big to be messages: a
checkpoint, a dataset, a plot. It stores the bytes by their hash and returns a
small reference. Nothing changes about how you write the file.
@@ -241,24 +240,24 @@ stops your own code from running.
- **Ports are declared, never inferred.** `requires` names the messages the
node reads; each arrives as the parameter of the same name. A bare string is
the shorthand `"dataset"` takes its type from whoever provides it, so a
the shorthand: `"dataset"` takes its type from whoever provides it, so a
message's type is written down exactly once.
- **Parameters with a default that are not ports become settings.** `epochs`
above is stored with the flow and tunable on the canvas without touching
this file.
- **The declaration is checked against the function.** A port with no matching
parameter, or a parameter that is neither port nor setting, is an error when
the module is imported not when the node is first called.
the module is imported, not when the node is first called.
- **The decorator hands the function back untouched.** `fit(dataset, lr=0.05)`
is an ordinary call, and the module is an ordinary script so a
is an ordinary call, and the module is an ordinary script, so a
`if __name__ == "__main__":` check beside the nodes runs. Syncing one
defined in a script *run as* `__main__` is what cannot work, since the
generated body would have no name to import it by.
!!! note "Where a `yield` cannot reach"
If the number comes from inside somebody else's callback Keras, Lightning,
HuggingFace `Trainer` `fluksio.emit(loss=...)` writes the same port the
If the number comes from inside somebody else's callback (Keras, Lightning,
HuggingFace `Trainer`) `fluksio.emit(loss=...)` writes the same port the
same way:
```python
@@ -271,7 +270,7 @@ stops your own code from running.
A training node is not like the rest of a flow. NumPy, JAX and PyTorch each
size their thread pool to every core they can see, so a few of them at once
oversubscribe the machine badly enough to starve the engine itself the API
oversubscribe the machine badly enough to starve the engine itself: the API
stops answering and every client waiting on a run dies with it. On a GPU the
same shape deadlocks: two processes each preallocating most of the card sit at
zero utilisation with nothing failing and nothing to read.
@@ -289,7 +288,7 @@ def finetune(checkpoint):
```
Two things follow. The node waits its turn rather than starting alongside more
of itself than the machine has room for the same waiting it already does for
of itself than the machine has room for, the same waiting it already does for
a worker. And the worker it runs in is *started* with thread limits matching
what it was given, because that is the only moment a numerical library reads
them; a GPU node is told which card is its through `CUDA_VISIBLE_DEVICES`, and
@@ -299,8 +298,8 @@ nothing else is given that card while it runs.
string whose contents depend on the version you have installed, so writing it
for you would silently replace whatever you had put there.
The engine has to be told how many cards it has `fluksio serve --gpus 1`, or
`FLOW_GPUS` because detecting them would mean depending on a vendor's
The engine has to be told how many cards it has (`fluksio serve --gpus 1`, or
`FLOW_GPUS`) because detecting them would mean depending on a vendor's
tooling. Until it is, a node asking for one is quietly given zero and runs
alongside every other; the log says so the first time it happens.
@@ -313,9 +312,9 @@ the next GPU run paying for its imports again.
imports jax without `resources={"gpus": 1}` runs on the shared pool, is never
given `CUDA_VISIBLE_DEVICES`, and leaves a warm worker holding whatever it
preallocated. Declaring the card is what makes it a worker the engine knows
to retire and what stops two such nodes running at once in the first place.
to retire, and what stops two such nodes running at once in the first place.
Declaring nothing is the default and is right for most nodes a poll, a
Declaring nothing is the default and is right for most nodes: a poll, a
threshold, a message on its way somewhere. Those share the engine's worker pool
and are given a fair share of `FLOW_CPUS` as a thread cap, which is what stops
several of them at once from each sizing to the whole box.
@@ -339,13 +338,13 @@ The flavor is read every time the node is built, so editing it changes what the
next run gets. `fluksio flavors` lists them; the Workers screen edits them.
`ram` takes `"2G"` and `duration_s` takes `"30m"`, and a flavor already says how
much pass one or the numbers, not both. `duration_s` is a statement about the
much, so pass one or the numbers, not both. `duration_s` is a statement about the
node for whoever is planning around it, not a limit; the limit is `timeout`.
### Where it runs
This is one decision, not two. Of every machine attached this engine and each
worker the node goes to one that can grant what it asked for and has it free.
This is one decision, not two. Of every machine attached (this engine and each
worker) the node goes to one that can grant what it asked for and has it free.
So a node asking for a card finds the box that has one, without naming it:
```python
@@ -354,7 +353,7 @@ So a node asking for a card finds the box that has one, without naming it:
```
If nothing attached can ever grant the ask, it is cut down to what is here and
runs anyway a flow written on a cluster still has to work on a laptop. If
runs anyway, since a flow written on a cluster still has to work on a laptop. If
something *could* but is busy, the node waits and says so.
Ask what is free, and what is waiting for it, at
@@ -363,8 +362,8 @@ Ask what is free, and what is waiting for it, at
!!! note "Accounted, not enforced"
Nothing stops a node that ignores its declaration; the numbers are
bookkeeping plus the environment its libraries read. Real enforcement
cgroups, rlimits is a next step, not this one.
bookkeeping plus the environment its libraries read. Real enforcement
(cgroups, rlimits) is a next step, not this one.
## Say which nodes make a flow
@@ -390,7 +389,7 @@ train = Flow(
`inputs` are the run's parameters, with the value a run gets when it names
none; `outputs` are what a run reports as its result.
Nodes are connected because one **provides** a message another **requires**
Nodes are connected because one **provides** a message another **requires**,
never because one imported the other. Importing `fit` into a second flow means
"the same code", not "wired to it".
@@ -439,7 +438,7 @@ def process(dataset, lr, **settings):
return (yield from fit(dataset=dataset, lr=lr, **settings))
```
So the flow store still holds a complete, runnable, git-versioned definition
So the flow store still holds a complete, runnable, git-versioned definition;
the body simply happens to import rather than duplicate. Your code stays in
your repository, under your version control.
@@ -450,7 +449,7 @@ alongside the store's own. `--dry-run` prints all of this and uploads nothing;
!!! warning "Sync after you edit"
The engine's workers are long-lived and hold your imported modules in
memory, so an edit to `myresearch` is invisible until they are retired
memory, so an edit to `myresearch` is invisible until they are retired,
which is what every `fluksio sync` does, including one that has nothing to
upload. If you are editing many times an hour, attach your own interpreter
as a worker instead: it starts a process per call and therefore reads your
@@ -467,7 +466,7 @@ alongside the store's own. `--dry-run` prints all of this and uploads nothing;
The generated body puts your repository on `sys.path` by absolute path,
which is right for the machine you synced from and meaningless in a
container. For anything else, install the package where the node runs
container. For anything else, install the package where the node runs:
`-e /home/you/my-research` in the module manifest, or
`myresearch @ git+ssh://…@a1b2c3d`, which travels where a path does not.
@@ -488,7 +487,7 @@ the GPU box is not switched on. See [Remote workers](../code/workers.md).
### What the canvas does with a synced flow
Its node bodies are generated, so the editor shows them read-only and says
where the real code is. Everything else behaves as usual but ports and
where the real code is. Everything else behaves as usual, but ports and
settings changed there are overwritten by the next sync, which is the point of
the repository being the source of truth. A node whose code you edit on the
canvas makes the next sync stop and say so rather than discarding your edit;
@@ -501,19 +500,19 @@ fluksio run train --lr 0.003 --seed 7
fluksio runs --flow train
```
`run` syncs first, so after an edit the command is just `fluksio run` — there
is no step to forget. `--no-sync` skips it for a tight loop where nothing
`run` syncs first, so after an edit the command is just `fluksio run`, with no
step to forget. `--no-sync` skips it for a tight loop where nothing
changed, since syncing retires the workers and the next call pays its imports
again.
`--lr` is typed by the flow's own inputs, so `0.003` arrives as a float. A
parameter you did not declare, or one of the wrong type, is refused before
anything executes. It answers immediately with a queued run training is
anything executes. It answers immediately with a queued run, since training is
measured in hours, so nothing waits for it unless you pass `--wait`, or
`--follow`, which waits and prints the numbers as the run reports them.
No engine has to be up for any of this: `fluksio run --local` boots the
engine inside the command instead, on the same `.fluksio` the same database,
engine inside the command instead, on the same `.fluksio`: the same database,
artifacts and history a served one would use. It costs a few seconds of
startup per invocation against the ~15 ms of submitting to an engine that is
already up, so it is for the run you want now rather than the loop you are
@@ -521,10 +520,10 @@ iterating in.
A second run of a flow you did not change mostly does not execute. Each of
your nodes is fingerprinted by its settings, the ports it declares, the values
it reads, and a digest of your repository's Python files so editing anything
it reads, and a digest of your repository's Python files, so editing anything
your node calls into invalidates it, not only the decorated function itself.
One an earlier run already answered is restored from that run rather than run
again reported as `cached`, so a flow with nothing left to do finishes as
again, reported as `cached`, so a flow with nothing left to do finishes as
`(3/3 cached)`. Change `--lr` and only the nodes downstream of it run. The
digest is blunt on purpose: an edit anywhere in the repository re-runs
everything. `--no-cache` turns caching off for one run, `@node(...,
@@ -555,7 +554,7 @@ curl -s "$FLUKSIO/runs/<id>/metrics?name=train.loss" -H "Authorization: Bearer $
The run carries its parameters, a digest of them, the seed, its result, how
long each node took, what it logged, every artifact it produced, and both
commits the flow store's and your repository's. That is the answer to "what
commits: the flow store's and your repository's. That is the answer to "what
was the learning rate on the run that got 94%?". When a run failed,
`run.failures` is the node that did it, with its traceback and its logs.
@@ -563,9 +562,9 @@ Beside your commit is `code_digest`, and `fluksio runs` prints the pair as
`a1b2c3d-dirty+9f0e1a2`. The commit alone cannot identify what ran: your node
bodies are imports, so the engine executes whatever is on disk when the worker
starts, and an uncommitted tree stamps `-dirty` for every run it ever produces.
The digest is over the modules this flow's nodes actually reach not the
The digest is over the modules this flow's nodes actually reach, not the
whole repository, so it moves when the code behind a number moves and stays
put when a notebook beside it changes — and it is read at the moment the run
put when a notebook beside it changes. It is read at the moment the run
starts, so in a sweep whose runs queue for hours each one records the code
that actually executed it rather than the code that was there when you
submitted. It is what makes the column worth joining an
@@ -594,7 +593,7 @@ and a blip is seconds:
A command that still cannot reach the engine says
`fluksio: engine not answering (…)` and names the run it started, which is
running on the engine regardless `fluksio runs` finds it again.
running on the engine regardless, and `fluksio runs` finds it again.
## Sweep it
@@ -618,7 +617,7 @@ curl -X POST $FLUKSIO/runs/flows/train/sweep -H "Authorization: Bearer $TOKEN" \
```
They share a `group_id`, so the sweep is `GET /runs?group=…`. Running them
concurrently is safe because each run gets a state backend of its own two
concurrently is safe because each run gets a state backend of its own: two
runs of one flow cannot overwrite each other's values.
Compare the curves in one call:
@@ -633,7 +632,7 @@ which answers in exactly the shape a chart widget draws.
## Take it into a dataframe
The analysis itself wants a table, and `export` writes the two an analysis
asks for the curves long, and one row per run with the parameters that
asks for: the curves long, and one row per run with the parameters that
varied beside its final numbers:
```sh
@@ -643,8 +642,8 @@ fluksio export runs --flow train --status ok -o arms.csv
Both carry the run id on every row, and `arms.csv` carries the commit and the
code digest beside it, so an exported file still says what produced its
numbers. Numbers inside a record are columns of their own
`metric.final_metrics.train_loss` and `--metrics` and `--params` take those
numbers. Numbers inside a record are columns of their own
(`metric.final_metrics.train_loss`) and `--metrics` and `--params` take those
dotted paths to narrow the table. Metric names are flow-qualified, so
`--list` prints the ones a selection carries. `Client.export_metrics()` and
`Client.export_runs()` answer the same rows to a notebook, ready for
@@ -653,8 +652,8 @@ dotted paths to narrow the table. Metric names are flow-qualified, so
## Looking at them in the portal
None of the above needs a browser, but there is one. **Runs** is the table of
everything that has run: filter it to a flow, to a status, or to one sweep
where it grows a column per parameter that actually varied and tick two runs
everything that has run: filter it to a flow, to a status, or to one sweep,
where it grows a column per parameter that actually varied, and tick two runs
to put their curves side by side.
A run records its metrics under the same names a dashboard binds to, so a
@@ -721,13 +720,13 @@ api.post("/flows/train/publish", json={"version": version}).raise_for_status()
```
Two things worth noticing, because the decorators only move where they are
said. Ports are declared, not inferred `process(lr, epochs)` gets its
said. Ports are declared, not inferred: `process(lr, epochs)` gets its
arguments from the ports of the same name, and the types are checked on every
value. And saving writes a *draft*; `publish` is what the engine picks up.
That separation is what lets you edit a flow that is running.
A flow uploaded this way carries no `origin`, which is what tells the canvas
and the next `fluksio sync` that it was not generated. See
A flow uploaded this way carries no `origin`, which is what tells the canvas,
and the next `fluksio sync`, that it was not generated. See
[The API](../code/api.md).
## Small scripts you are just playing with
@@ -741,7 +740,7 @@ cd .fluksio/flows
git log --oneline
```
Every save is a commit the node source, the parameters, the graph. A run
Every save is a commit: the node source, the parameters, the graph. A run
records the commit it ran at, so `git show` on that hash is literally the code
that produced the number. You get the versioning without adopting anything.
@@ -759,8 +758,8 @@ A submit is around 15 ms, so calling that in a loop is a reasonable thing to do.
## Get a browser onto it
The pip install gives you the engine and the API, not a web interface a
machine with no inbound route cannot serve one usefully anyway. To see the
The pip install gives you the engine and the API, not a web interface, since a
machine with no inbound route cannot serve one usefully. To see the
canvas, the run history and live loss curves, pair the instance with a
portal, which serves the dashboard from its side:
@@ -773,7 +772,7 @@ Nothing needs to be exposed: your machine dials out and holds the connection
open. See [Accounts and the portal](../interface/portal.md).
If you would rather stay entirely offline, run the dashboard SPA yourself from
the app's Docker image see the [facility path](facility-automation.md), which
the app's Docker image; see the [facility path](facility-automation.md), which
is the same stack.
## Where to go next
+25 -27
View File
@@ -20,7 +20,7 @@ in this interface a lot, so it is worth learning it properly.
└── workers on other boxes (optional)
```
One engine holds every flow. Nothing here is a plugin you install separately
One engine holds every flow. Nothing here is a plugin you install separately:
the broker client, the time-series writer, the dashboards and the alerting are
all part of the same process, editing the same graph.
@@ -28,7 +28,7 @@ all part of the same process, editing the same graph.
- Docker and Compose v2 on the host
- The address you will type into a browser. On a home network that is the
server's own address `192.168.1.50` and nothing else is needed. A
server's own address, `192.168.1.50`, and nothing else is needed. A
hostname is optional, and covered below
- Optionally: an MQTT broker and an InfluxDB you already run. If not, the stack
can start both for you
@@ -53,15 +53,15 @@ $EDITOR .env # FIRST_SUPERUSER, ENVIRONMENT=production
| `FIRST_SUPERUSER_PASSWORD` | leave it as `changethis` and one is generated for you |
| `ENVIRONMENT` | `production` closes the interactive API schema; `local` leaves it open |
Everything the instance owns the database, your flows, secrets,
artifacts, the packages your node code imports is on one Docker volume.
Everything the instance owns (the database, your flows, secrets, artifacts,
the packages your node code imports) is on one Docker volume.
Backing that volume up is backing up the instance.
??? note "Even smaller: no Docker at all"
`pip install fluksio && fluksio serve` gives you the same engine with no
containers, keeping its data in `~/.fluksio`. What it does not give you is
the web interface, which the SPA container serves so you would drive it
the web interface, which the SPA container serves, so you would drive it
from the API, or [pair it with a portal](../interface/portal.md) that
serves the dashboard for you. Good for a Raspberry Pi that only runs flows;
less good as your main instance.
@@ -82,7 +82,7 @@ APP_PORT=8080 docker compose -p fluksio-app --env-file .env \
-f docker/compose.yml -f docker/compose.lan.yml up -d --build
```
Open `http://192.168.1.50:8080`. That one address is the whole application
Open `http://192.168.1.50:8080`. That one address is the whole application:
the interface, the API under `/api`, and the websocket that puts live values on
the canvas.
@@ -96,13 +96,13 @@ maintaining: the browser is talking to one origin.
A published port is plain HTTP with no certificate. Between your own
machines that is fine, and it is not something to forward from a router. To
reach the house from outside, put a reverse proxy with TLS in front of it
reach the house from outside, put a reverse proxy with TLS in front of it,
or [pair it with a portal](../interface/portal.md), which needs no inbound
route at all.
### With a hostname, behind a reverse proxy
If you have DNS or want certificates set `DOMAIN` and run `make up`
If you have DNS, or want certificates, set `DOMAIN` and run `make up`
instead. The interface is then served on `app.${DOMAIN}` and the API on
`api.${DOMAIN}`, and nothing is published on a host port.
@@ -113,7 +113,7 @@ attach it to the `proxy` network and forward `app.${DOMAIN}` →
!!! tip "Developing on the same machine"
`make dev-lan` is the same idea for the integrated dev stack the app on
`make dev-lan` is the same idea for the integrated dev stack: the app on
`http://<host-ip>:8080`, `APP_PORT` to move it. `make dev` rebuilds without
the published port, so re-run `make dev-lan` after one.
@@ -128,9 +128,8 @@ Go to **Flows → New flow** and call it `house`. You land on the canvas.
A flow is a set of nodes that talk to each other through **named messages**.
You do not draw wires: a node says which messages it needs and which it
produces, and the canvas draws the graph that follows from those names. That
sounds like a small difference and turns out to be a large one — renaming is
safe, fan-in is free, and two flows can share a value by naming it.
produces, and the canvas draws the graph that follows. Renaming is safe, fan-in
is free, and two flows share a value by naming it.
### Read a sensor
@@ -150,8 +149,8 @@ broker sends one. Click the wire to see the last payload and its history.
### Do something with it
Add a **Function** node. This is a Python node the code editor opens in its
panel:
Add a **Function** node. This is a Python node, and the code editor opens in
its panel:
```python
def process(living_temperature, comfortable=21.0):
@@ -160,7 +159,7 @@ def process(living_temperature, comfortable=21.0):
```
Declare `living_temperature` as an input (type `float`) and `heat_wanted` as an
output (type `bool`). `comfortable` is not a port — it is a **setting**,
output (type `bool`). `comfortable` is not a port but a **setting**,
because it is a constant of this node rather than something the graph carries.
It shows up as a field in the node's Settings section.
@@ -178,8 +177,8 @@ Publish the flow. You have a thermostat.
!!! warning "Test before it touches a relay"
A flow can be **paused** (it holds messages instead of running them) and
**stepped** (release exactly one). Together with the run button which
injects a value by hand that is how you convince yourself the logic is
**stepped** (release exactly one). Together with the run button, which
injects a value by hand, that is how you convince yourself the logic is
right before the contactor finds out. Both live on the dock at the bottom
of the canvas.
@@ -212,11 +211,10 @@ point.
Note the token. Credentials never sit in a flow: `{"$secret": "influx-token"}`
is a reference into an encrypted store, and the editor renders those fields as
a secret picker. Add the actual value once under **Secrets**. Flows are a git
repository you might well push somewhere this is what keeps a password out of
it.
repository you might well push somewhere, so this keeps a password out of it.
Reading back is the same node with `queries` instead of `writes`, or for
anything a chart asks for a pair of small Python nodes on either side that
Reading back is the same node with `queries` instead of `writes`, or, for
anything a chart asks for, a pair of small Python nodes on either side that
build a Flux query and shape its rows. That indirection is deliberate: the
database node holds the connection and nothing else, so a dashboard widget
never learns which database answered it.
@@ -231,7 +229,7 @@ Widgets are typed the same way ports are: a switch binds to a `bool`, a gauge
to a number, an agenda to a `list`. Bind it wrong and the editor says so rather
than drawing nothing.
Controls work in the other direction a switch on a dashboard publishes the
Controls work in the other direction: a switch on a dashboard publishes the
message it is bound to, exactly as a node would. The canvas draws it as a
labelled endpoint feeding the nodes that read it, so a value never appears from
nowhere.
@@ -247,7 +245,7 @@ lives on a different network, or something needs hardware the engine's host
does not have.
The unit of distribution is the **worker**. It runs the code of nodes you mark
for it, and it dials *out* to the engine so the Pi in the shed does not need
for it, and it dials *out* to the engine, so the Pi in the shed does not need
an inbound route, and the engine does not need to reach it.
On the engine, mint a token:
@@ -276,7 +274,7 @@ flow before the hardware arrives.
!!! note "What a worker is not"
It is not a second engine. Subscriptions, schedules, webhooks and the
dashboards all stay in one process — that is what keeps a value having one
dashboards all stay in one process, which keeps a value having one
definition. A worker executes node bodies, nothing else. Scaling *the
engine* to several processes is not supported: run one.
@@ -291,9 +289,9 @@ What you probably want on day one is everything, to ntfy:
- add an ntfy channel with your server and topic
- add a rule with no events ticked, which means *all of them*
The engine deduplicates aggressively the same node failing every second is
one alert, not thirty-six thousand and caps the total at ten an hour however
bad it gets. See [Secrets, modules and alerts](../interface/operations.md).
The engine deduplicates aggressively, so the same node failing every second is
one alert rather than thirty-six thousand, and caps the total at ten an hour
however bad it gets. See [Secrets, modules and alerts](../interface/operations.md).
## Keep it alive
+3 -5
View File
@@ -1,8 +1,6 @@
# Pick your starting point
People arrive at Fluksio from two directions, and the honest answer to "how do
I set this up?" is different for each — not just in the commands, but in how
much of an afternoon it is reasonable to spend.
Setup differs depending on what you are here for. Pick the path that matches.
Pick the one that sounds like you. Everything past this section is the same for
both.
@@ -52,8 +50,8 @@ Some rough tells:
| **Data lives in** | SQLite beside the flows | InfluxDB, usually |
| **The thing you look at** | run history and loss curves | a dashboard, maybe on a wall |
If both describe you a lab with instruments to drive *and* models to
fit — start with the data-science path. It is the smaller instance, and it
If both describe you (a lab with instruments to drive *and* models to fit),
start with the data-science path. It is the smaller instance, and it
grows into the other one without being reinstalled: the same engine, the same
flows, just more of them running all the time.
+18 -23
View File
@@ -1,45 +1,40 @@
# Fluksio
Fluksio is a node-based automation engine. You describe what should happen as a
graph of small pieces of logic, and it keeps that graph running reacting to
graph of small pieces of logic, and it keeps that graph running: reacting to
what arrives, or executing once from parameters to a result.
Two very different jobs turn out to be the same shape, which is why the same
engine does both:
It runs two kinds of graph:
- **A house, a lab or a plant** produces values forever. A sensor publishes, a
rule fires, a relay closes, a dashboard on the wall shows what happened. The
flow never ends.
- **An experiment** produces a value once. Parameters go in, stages execute,
and at some point it is *done* and has left behind metrics, artifacts and a
record you can compare against last month's.
- A **live flow** produces values forever. A sensor publishes, a rule fires, a
relay closes, a dashboard on the wall shows what happened. It never ends.
- A **run** produces a value once. Parameters go in, stages execute, and it
finishes, leaving metrics, artifacts and a record to compare against.
The first is a **live flow**, the second is a **run**. Both are the same nodes,
the same type checking, the same editor and the same API.
Both use the same nodes, the same type checking, the same editor and the same
API.
## What you actually get
## What you get
- A **flow engine** that owns your graph, checks the types on every edge, and
keeps running when a node fails rather than taking the rest down with it.
keeps running when a node fails.
- A **canvas** that lays flows out for you and an editor for the Python inside
each node with the running values drawn on the wires while you work.
each node, with the running values drawn on the wires while you work.
- **Dashboards** built next to the logic that feeds them, including ones you
can hang on a wall tablet that has no keyboard.
- **Runs**: parameters, metrics, artifacts, sweeps, and a queryable history
without a second server and without paying a project bootstrap per execution.
- **Runs**: parameters, metrics, artifacts, sweeps and a queryable history, with
no second server and no project bootstrap per execution.
- **Distributed workers**: a node marked `device: gpu` runs on the machine that
has one, which dials out to the engine rather than needing to be reachable.
- An **HTTP API** that came first — everything the browser does, you can do from
a script — plus an MCP endpoint for agents.
- An **HTTP API** covering everything the browser does, plus an MCP endpoint for
agents.
## Start here
The setup is genuinely different depending on what you are here for, so
[Getting started](getting-started/index.md) splits in two: one path installs a
Python package and gets out of your way, the other stands up a server you will
be running for years. Everything after that is shared.
[Getting started](getting-started/index.md) has two paths: one installs a Python
package, the other stands up a server. Everything after that is shared.
If you would rather look before installing, the hosted demo at
To look before installing, the hosted demo at
[fluksio.com](https://fluksio.com) runs a real instance with a small-house panel
and a training pipeline on it.
+33 -35
View File
@@ -1,6 +1,6 @@
# Dashboards and panels
A dashboard is a grid of widgets bound to message names the same names that
A dashboard is a grid of widgets bound to message names, the same names that
wire the graph. It is its own document rather than a set of nodes, so it reads
across flows without belonging to any of them, and a flow stays the logic it
was.
@@ -39,7 +39,7 @@ and dragging is off. Picking a widget and editing its settings still works.
| **Clock** | — | the time, in a size a wall can read |
Every widget carries a **title**, and **Show title** decides whether the panel
draws it. Turned off, the tile is just the reading which is what a row of
draws it. Turned off, the tile is just the reading, which is what a row of
gauges under one heading wants. The title is still the widget's name: what a
screen reader calls its controls, and what a published value is labelled with.
@@ -54,8 +54,8 @@ screen reader calls its controls, and what a published value is labelled with.
| **Selector** | one of a list | a mode, a scene, a preset — as a menu, or as a row of choices with the active one held |
| **Colour** | `[h, s, v]`, `[r, g, b]` or `"#rrggbb"` | a hue wheel with saturation and brightness, for an RGB fixture |
The colour wheel sends `[h, s, v]` by default hue 0-360 degrees, saturation
and value 0-100 percent, which is what a DMX encoder expects and its
The colour wheel sends `[h, s, v]` by default (hue 0-360 degrees, saturation
and value 0-100 percent, which is what a DMX encoder expects) and its
**Sends** setting switches that to `[r, g, b]` (0-255 each) or to a
`"#rrggbb"` string, because fixtures differ. The first two bind a `list`
message, the third a `str`.
@@ -82,21 +82,20 @@ buffer of recent values, extended over the socket as new ones land. Nothing
else is involved. This is what you want for "the last few hours of the living
room".
**Querying.** The chart publishes a *request* the window and resolution it
wants exactly as a slider publishes a value, and draws the `series` some flow
**Querying.** The chart publishes a *request* (the window and resolution it
wants) exactly as a slider publishes a value, and draws the `series` some flow
answers with. What serves that request is the flow's business: typically a
small Python node that builds a Flux query, an InfluxDB node that runs it, and
another Python node that shapes the rows.
That indirection is the point. The widget never learns which database answered
it, so swapping the store is a change to one flow and nothing else. The answer
also states what it was computed for, so an answer to a different question is
ignored rather than two charts overwriting each other's picture.
The widget never learns which database answered it, so swapping the store
changes one flow and nothing else. An answer states what it was computed for,
so an answer to a different question is ignored.
!!! note "Drop the bucket that is still filling"
The request carries a window and an interval, and the binning is the flow's
own work so the newest bucket only ever holds the part of an interval
own work, so the newest bucket only ever holds the part of an interval
that has elapsed. Drawn, it reads as a fall that never happened. The
engine's own [Activity charts](index.md#activity) end on the last closed
bin for that reason; a flow answering a chart has to drop or hold back its
@@ -105,8 +104,8 @@ ignored rather than two charts overwriting each other's picture.
## Media tiles
A media widget draws what its message points at: a picture, a clip with
controls, a video. Media does not travel as a message a reference to it does
— so the tile fetches the bytes behind whichever reference the message holds,
controls, a video. Media does not travel as a message; a reference to it does.
The tile fetches the bytes behind whichever reference the message holds,
and redraws when a new one arrives.
**Crop or fit** decides how a picture fills the tile. **Play as it arrives**
@@ -115,7 +114,7 @@ touched the page, so a screen nobody has tapped stays silent.
Rate is the thing to get right. A frame every second or two is a glance at a
door, and works; through the portal, make that every few seconds. Live video is
not something to push through the message plane at all — put the camera's own
not something to push through the message plane at all. Put the camera's own
address in **Live stream** and the browser plays it from source, leaving the
messages to carry the occasional still that a flow can actually react to.
@@ -136,7 +135,7 @@ carrying what to do about it:
The buttons and the bar publish words: `toggle`, `next`, `prev` and
`seek:<seconds>`. Those are a streamer's own vocabulary rather than this app's,
which is what lets one tile drive whatever is on the other end the node that
which is what lets one tile drive whatever is on the other end: the node that
receives them decides what they mean for its device.
The position counts forward in the browser between readings, so the bar moves
@@ -147,7 +146,7 @@ reading that arrives is taken as the truth and the count restarts from it.
Most of what a dashboard carries is a widget: a tile bound to a message. A
handful of things are not, because they belong to the whole surface rather than
to any tile on it and a screen bolted to a wall has nobody standing at it to
to any tile on it, and a screen bolted to a wall has nobody standing at it to
set them.
| Setting | Is | Driven by |
@@ -169,25 +168,24 @@ They all work the same way, and both halves are optional:
fallback: what the dashboard uses before the first message arrives, and
whenever the message is silent.
A bound setting is type-checked like a widget binding a Theme pointed at a
`float` is refused by the editor and by the server — and the panel's credential
A bound setting is type-checked like a widget binding, so a Theme pointed at a
`float` is refused by the editor and by the server. The panel's credential
is extended to it, so a paired screen may read its own theme message and
nothing further.
There is no schedule field, on purpose. **A schedule is a node publishing to
the bound message**: an `inject` with a cron expression, feeding a `change`
node that maps the hour onto a palette, is the whole of "warmer after sunset"
and the same channel then serves anything else you want to drive, including
locking a panel down remotely.
There is no schedule field. **A schedule is a node publishing to the bound
message**: an `inject` with a cron expression feeding a `change` node that maps
the hour onto a palette is the whole of "warmer after sunset". The same channel
drives anything else, including locking a panel down remotely.
### Look
**Fluksio** is the app's own design `--card` surfaces separated by a hairline
and a low shadow, pill controls, one slate-blue accent and is what a dashboard
**Fluksio** is the app's own design (`--card` surfaces separated by a hairline
and a low shadow, pill controls, one slate-blue accent) and is what a dashboard
wears until you choose otherwise. **Material** lays flat, tonal cards on a plain
ground. **Glass** floats translucent tiles over a soft, slowly moving one.
They are three complete sets of components, not three stylesheets but they are
They are three complete sets of components, not three stylesheets, but they are
the same dashboard: every widget, every control and every keystroke behaves
identically, so switching look never changes what a panel can do.
@@ -205,7 +203,7 @@ A palette is an ordered list of colours, and **position is the role**:
| 6+ | further chart colours |
Paste a [coolors.co](https://coolors.co) link (or a list of hex colours) and the
whole dashboard is recoloured every widget, the rail and the charts. Leave the
whole dashboard is recoloured: every widget, the rail and the charts. Leave the
later roles off and they are worked out from the ones you gave, so **three
colours are a whole dashboard**. **Rotate** turns the list when the roles landed
in the wrong order. Text that could not be read on a surface is replaced with
@@ -213,8 +211,8 @@ black or white there, so no palette can produce a line nobody can see.
While a palette is set, **Theme is idle**: the first colour is the ground, so
whether the panel reads light or dark is already decided by the palette itself.
Fault and success keep their own colours in every palette a failure must never
be paintable as a reading.
Fault and success keep their own colours in every palette, so a failure is never
paintable as a reading.
### Background
@@ -249,25 +247,25 @@ A wall tablet has no keyboard, so it pairs.
2. Point the device's browser at the link the dialog shows. The device then
displays a six-character code.
3. Type that code into the same panel's **Pair device** field. The line under
it names what is holding the code — check it is the screen you just hung,
it names what is holding the code. Check it is the screen you just hung,
because approving adopts whatever answered. The screen picks the credential
up within a few seconds and never asks again.
What the screen holds is not a login. It reaches that panel's published
dashboards and the messages its own widgets read or publish, and nothing else
a message no tile on it draws is refused in both directions.
dashboards and the messages its own widgets read or publish, and nothing else.
A message no tile on it draws is refused in both directions.
The *credential* cannot be made strictly read-only, and that is honest rather
than an oversight: a querying chart publishes its request, and a control on a
panel is the reason you put one there. The dashboard's own **Lock** setting
above stops its controls publishing and can be driven by a flow, which is how
you quieten a screen remotely but that is the surface behaving, not the
you quieten a screen remotely, but that is the surface behaving, not the
credential being narrowed.
Deleting the panel revokes the credential and the assignment together, which is
how you retire a device and what it showed. Unpairing revokes only the
credential: the panel, its dashboards and their arrangement stay exactly where
they are, and the screen falls back to asking for a new code which is how you
they are, and the screen falls back to asking for a new code, which is how you
swap the device out without rebuilding what hangs there.
!!! note "If the link is wrong"
@@ -290,7 +288,7 @@ The portal names the panel and nothing else. What the panel may read is decided
on the instance, on every call, by the same check a locally paired screen
passes. Two differences: it acts as the account the instance was enrolled
with rather than as whoever approved it, and deleting the panel stops it here
immediately while the portal's copy of the token expires on its own which is
immediately while the portal's copy of the token expires on its own, which is
also why unpairing, which works on a credential this instance signed, does
not reach a remote screen. Revoke that one at the hub.
+12 -14
View File
@@ -5,14 +5,12 @@ full-bleed graph, so what you are looking at is always the flow.
## The canvas
Nodes are laid out for you left to right on a desktop, top to bottom on a
phone and the arrangement is recomputed as the graph changes. There is
Nodes are laid out for you, left to right on a desktop and top to bottom on a
phone, and the arrangement is recomputed as the graph changes. There is
nothing to drag and nothing to tidy.
That is on purpose. The edges come from message names, so a stored layout would
be a second, weaker description of the same thing, and something to merge when
two people edit the same flow. It also keeps flows honest: a graph nobody can
hand-arrange is one worth splitting.
The edges come from message names, so there is no layout to store or to merge
when two people edit the same flow.
What is drawn:
@@ -38,7 +36,7 @@ the canvas with its panel open.
Floats on the right; a full-screen sheet on a phone.
**Consumes** and **Provides** are the ports. Each row is a message name, a
payload type, and on an output a *streaming* toggle for ports that publish
payload type, and, on an output, a *streaming* toggle for ports that publish
repeatedly during one execution. The name field suggests messages already in
the flow, which is usually how you wire something: type the name that already
exists.
@@ -47,7 +45,7 @@ Renaming an output renames the message everywhere it is read. Inputs are not
renamed, because an input is more often re-pointed at a different message than
renamed.
**Settings** is generated from the node type's parameter schema so an MQTT
**Settings** is generated from the node type's parameter schema, so an MQTT
node shows broker, topic and QoS, and an ntfy node shows server, topic and
priority. A field marked as a credential renders as a picker over your stored
[secrets](operations.md) instead of a text box.
@@ -59,7 +57,7 @@ ports. See [Where a node's values come from](../concepts/values.md).
**Timeout** (Function nodes) is how many seconds the code may be *silent*
before it is stopped: it measures silence, not duration, so a node that yields
every few seconds can run for hours under a timeout of 300. Left empty there is
no limit, which is the default — set one where a node going quiet means it is
no limit, which is the default. Set one where a node going quiet means it is
stuck rather than working. A `0` says so explicitly.
**Code** is the editor. It saves as you type and applies on ⌘S; the node
@@ -71,10 +69,10 @@ instead.
Editing a shared node edits the copy every flow using it runs, which is the
point and also the caution.
!!! note "Not in the panel (yet)"
!!! note "Not in the panel yet"
`device` and `device_policy` which machine a node's code runs on are set
through the API rather than the panel. See
`device` and `device_policy`, which machine a node's code runs on, are set
through the API rather than the panel, with `PUT /flows/{name}`. See
[Remote workers](../code/workers.md).
## The flow panel
@@ -95,8 +93,8 @@ Click the canvas title to open the flow's own settings.
Click a wire. You get the last payload that went down it, when, and a sparkline
of its recent numeric history. You can also republish that value from the node
that produced it, which re-runs everything downstream the fastest way to test
a change without waiting for the real sensor.
that produced it, which re-runs everything downstream. It is the fastest way to
test a change without waiting for the real sensor.
## The dock
+5 -5
View File
@@ -1,6 +1,6 @@
# The dashboard app
The web interface is a single-page app served at `app.${DOMAIN}` or, for an
The web interface is a single-page app served at `app.${DOMAIN}`, or, for an
instance reached through a portal, at `${DOMAIN}/i/{instance-id}`.
Either way it is the same application, and it is a client of the same REST API
you can script against.
@@ -54,7 +54,7 @@ Each flow also has a switch beside it in the list, which starts and stops it.
Two columns under the graph, exactly as tall as each other, most recently
worked on first. The flows column is the list with the switches; the dashboards
column is a mosaic, each tile a schematic of that dashboard's layout blocks
column is a mosaic, each tile a schematic of that dashboard's layout: blocks
where its widgets sit, shaded by what kind of widget each one is. Neither
column grows past about six rows: past that it scrolls in place rather than
pushing the health block down the page.
@@ -98,7 +98,7 @@ Opening one takes you to [the flow editor](flow-editor.md).
### Deleting several at once
Press and hold a card or ctrl-click it to pick it, then tap the rest. While
Press and hold a card, or ctrl-click it, to pick it, then tap the rest. While
anything is picked, **New flow** in the toolbar becomes a trash button, and it
asks once before deleting the lot. Unpicking the last one puts the list back;
so does Escape. The Dashboards screen works the same way.
@@ -117,5 +117,5 @@ so does Escape. The Dashboards screen works the same way.
## Appearance
Light and dark follow your system by default; **Settings → Appearance**
overrides it. Both themes are first-class the wall-panel view in particular is
designed to be legible in dark from across a room.
overrides it. Both themes are first-class, and the wall-panel view is designed
to be legible in dark from across a room.
+6 -6
View File
@@ -23,8 +23,8 @@ They are encrypted at rest with a key derived from the instance's
flows are a git repository you may well push somewhere: what gets committed and
shared never contains a password.
A name that no longer resolves is a clear error on the node that wanted it
"No secret named 'x' add it under Secrets" rather than a mysterious
A name that no longer resolves is a clear error on the node that wanted it
("No secret named 'x'; add it under Secrets") rather than a mysterious
authentication failure.
!!! warning "Rotating `SECRET_KEY`"
@@ -36,7 +36,7 @@ authentication failure.
## Modules
Node code runs in a virtual environment of its own, on the instance's data
volume separate from the packages Fluksio itself runs on. A pin of yours can
volume, separate from the packages Fluksio itself runs on. A pin of yours can
never shadow one of ours, and vice versa.
**Modules** is a pip manifest and an **Apply** button. It is installed with
@@ -49,7 +49,7 @@ and whether the manifest and the environment agree. A failed resolve comes back
as `uv`'s own output, in the pane, because that is the only thing anyone can
act on.
An install takes effect immediately the worker processes pick up the new
An install takes effect immediately: the worker processes pick up the new
interpreter state without the engine restarting.
!!! note "Big scientific stacks"
@@ -79,7 +79,7 @@ literal, so a webhook URL with a token in it does not sit in the config.
### Rules
A rule is a set of events and the channels they go to. **A rule with no events
ticked covers all of them**, including ones added in later versions which is
ticked covers all of them**, including ones added in later versions, which is
usually what you want for the first rule you write.
| Event | Fires when |
@@ -113,7 +113,7 @@ Superusers get an **Admin** screen: add, edit and remove accounts, and see who
is a superuser.
Accounts that arrived through a portal are badged *Portal*. They have no
password they sign in at the portal, not here and they are never
password (they sign in at the portal, not here) and they are never
superusers. Deleting one cuts their access immediately and independently of the
portal, so it works even if the portal cannot be reached. See
[Accounts and the portal](portal.md).
+9 -10
View File
@@ -13,13 +13,13 @@ The **portal** is optional, and it exists to solve two specific problems:
An enrolled instance dials *out* to the portal and holds one websocket
open. The portal serves the dashboard from its own side, and only the API calls
travel down the tunnel so the interface loads at portal speed and your
travel down the tunnel, so the interface loads at portal speed and your
machine stays unreachable from the internet.
## Enrolling
Two halves, deliberately: whoever performs the second step decides what the
instance's owner gets.
Two halves: whoever performs the second step decides what the instance's owner
gets.
**On the portal** (`hub.${DOMAIN}`, or [fluksio.com](https://fluksio.com) for
the hosted one): **Instances → Add instance**, give it a name, and copy
@@ -39,10 +39,9 @@ fluksio enroll ABCD-1234 --portal https://hub.fluksio.com
The code is single-use and expires in fifteen minutes.
A portal session then arrives as *that local account* the settings screen
states this plainly, because it is the whole security model in one sentence.
Use `--as someone@example.com` to enrol as a specific local account when the
instance has several superusers.
A portal session then arrives as *that local account*, which the settings screen
states plainly. Use `--as someone@example.com` to enrol as a specific local
account when the instance has several superusers.
Once enrolled, the instance appears under **Instances** with its
status, when it was last seen and its version. **Open** takes you to its
@@ -55,7 +54,7 @@ the tunnel. What those requests may do is decided **on the instance**, by
the same checks a local session passes.
The trust anchor is a signing keypair on the portal. Every instance pins
its public half at enrolment and rejects anything else which is what stops a
its public half at enrolment and rejects anything else, which is what stops a
hijacked DNS entry or a mis-issued certificate from impersonating the portal.
## Letting someone else in
@@ -72,7 +71,7 @@ admits them, and they arrive as a local user of their own rather than as you.
you.
The instance redeems that code against the portal using its own credential.
A portal session cannot do this which is what stops somebody you let in from
A portal session cannot do this, which is what stops somebody you let in from
letting others in.
On the instance they appear under **Admin → Users**, badged *Portal*, never
@@ -95,7 +94,7 @@ unsure: it does not need the portal's cooperation.
## Running your own portal
The portal is the `index` stack's `hub` service accounts, the registry of
The portal is the `index` stack's `hub` service: accounts, the registry of
connected instances, and the websocket each one dials in on. Two things
about it are load-bearing:
+2
View File
@@ -31,8 +31,10 @@
- [The HTTP API](/code/api/): auth, flows, messages, runs, observability, websocket
- [Remote workers](/code/workers/): fluksio-worker, device labels, artifacts across hosts
- [Agents over MCP](/code/agents/): the tool list and the security model
- [Writing a connector](/code/connectors/): packaging a device integration as a node type
## Reference
- [Node types](/reference/node-types/): every built-in type and its settings
- [Payload types](/reference/payload-types/): float, int, str, bool, json, record, list, series, artifact
- [Configuration](/reference/configuration/): every environment variable
- [The connector contract](/reference/connector-contract/): what a connector package must provide
+10 -10
View File
@@ -17,7 +17,7 @@ Anything already exported wins over the file.
| `DATA_DIR` | `flow-data` (`./.fluksio` via the CLI; `~/.fluksio` with `--global`) | everything below derives from this |
| `DATABASE_URL` | SQLite in `DATA_DIR` | any SQLAlchemy URL |
| `FLOWS_DIR` | `$DATA_DIR/flows` | the git repository holding flows |
| `SECRETS_FILE` | `$DATA_DIR/secrets.enc` | encrypted credentials, deliberately outside the repo |
| `SECRETS_FILE` | `$DATA_DIR/secrets.enc` | encrypted credentials, kept outside the repo |
| `ALERTS_FILE` | `$DATA_DIR/alerts.json` | channels and rules |
| `PROVISIONERS_FILE` | `$DATA_DIR/provisioners.json` | clusters a machine can be started from; absent means none |
| `PANELS_FILE` | `$DATA_DIR/panels.json` | wall-panel pairings |
@@ -25,7 +25,7 @@ Anything already exported wins over the file.
| `CLOUD_CONFIG_FILE` | `$DATA_DIR/cloud.json` | the portal enrolment, if any |
| `NODE_VENV` | `auto` | which interpreter node code runs on — see below |
Set `DATA_DIR` and the rest follow. Set one explicitly and it wins which is
Set `DATA_DIR` and the rest follow. Set one explicitly and it wins, which is
what the container images do to pin everything onto `/data`.
`NODE_VENV` is the exception, being about an environment rather than a path:
@@ -43,7 +43,7 @@ An instance that already has a managed venv keeps it on upgrade under
`secrets.enc`, `alerts.json`, `panels.json` and `oauth-key.pem` are written
at runtime. In a container, anything not on a volume lands in the writable
layer and is lost on the next rebuild un-pairing every screen and
layer and is lost on the next rebuild, un-pairing every screen and
revoking every agent.
## State
@@ -54,8 +54,7 @@ An instance that already has a managed venv keeps it on upgrade under
| `REDIS_PORT` | `6379` | |
Flow state is the last value of every message, node memory, and the run queue.
Redis here is persistence, not a cache — run it with append-only persistence
on.
Redis here is persistence, not a cache. Run it with append-only persistence on.
## Identity and access
@@ -85,7 +84,7 @@ there to satisfy.
Empty is the useful value: the interface then addresses the API relative to
whichever origin served the page, so one image answers on a hostname, on a
`http://<host-ip>:<port>`, and through an ssh tunnel alike and no origin has
`http://<host-ip>:<port>`, and through an ssh tunnel alike, and no origin has
to be added to `BACKEND_CORS_ORIGINS`, because there is only one.
`docker/compose.lan.yml` builds it that way and puts an `/api` proxy in front
of the backend to complete it; see
@@ -115,7 +114,7 @@ o'clock nor the same hour in summer as in winter. Set it to where the
instance is.
`production` closes `/docs`, `/redoc` and the OpenAPI document, because the
schema enumerates every endpoint the instance serves including the paths
schema enumerates every endpoint the instance serves, including the paths
webhook nodes mounted at runtime. It also turns a `changethis` secret from a
warning into a refusal to start.
@@ -133,8 +132,9 @@ warning into a refusal to start.
| `ARTIFACT_GC_INTERVAL_S` | `3600` | how often artifact bytes nothing refers to are swept away; 0 never sweeps |
| `ARTIFACT_GC_GRACE_S` | `3600` | how long a freshly written artifact is spared, whatever refers to it |
The three concurrency limits are also flags on `fluksio serve``--max-workers`,
`--max-cascades`, `--max-runs` as is the card count, `--gpus`. The flags
The three concurrency limits are also flags on `fluksio serve`
(`--max-workers`, `--max-cascades`, `--max-runs`), as is the card count,
`--gpus`. The flags
outrank the file, and the engine says which numbers it started with in its
first lines. Each pool size must be at least 1 and the card count at least 0:
a number below that is refused as a flag error naming it, rather than read as
@@ -153,7 +153,7 @@ several at once cannot each size themselves to the whole machine. Setting
that default.
These two are this machine's figures. An attached worker reports its own when
it dials in, and a node goes to whichever machine can grant what it asked for
it dials in, and a node goes to whichever machine can grant what it asked for,
so a GPU on a worker needs no `FLOW_GPUS` here. What every machine has free, and
which nodes are queued, is `GET /api/v1/workers/resources` and the Workers
screen. Named sizes live in the database and are `GET /api/v1/flavors`;
+10 -10
View File
@@ -1,7 +1,7 @@
# The connector contract
A connector is the device-facing node class. It talks to something outside the
engine a device, a service, a protocol and publishes what it finds as
engine (a device, a service, a protocol) and publishes what it finds as
ordinary typed messages, so the rest of a flow cannot tell the difference
between a reading from a heat pump and one from a `print` statement.
@@ -75,7 +75,7 @@ class Params(ConnectorNode.Params):
```
- Booleans render as switches, numbers and strings as inputs. Objects and
arrays are not rendered keep settings flat.
arrays are not rendered, so keep settings flat.
- A field carrying **`x-secret`** renders as a picker over the stored secrets
and writes a reference, `{"$secret": "name"}`, rather than the value. The
engine resolves it when the node is built, so a credential never lands in
@@ -98,13 +98,13 @@ optional `interval`.
at most that often. A connector should poll at the rate the device is
comfortable with and leave delivery rates to whoever wires it up.
- A key no port declares is an error, and it fails the whole reading rather
than the one value — a mistyped metric name is how a training curve goes
than the one value. A mistyped metric name is how a training curve goes
missing.
### When the device decides what the ports are
A connector for a device whose readings vary by model which components a
relay has, which entities were flashed onto a board cannot know its ports in
A connector for a device whose readings vary by model (which components a
relay has, which entities were flashed onto a board) cannot know its ports in
advance, and returning everything the device reports would fail on the first
value nobody bound. Narrow the reading to the ports that were declared:
@@ -115,7 +115,7 @@ return {port: value for port, value in reading.items() if port in declared}
`spec.port` is the local, unqualified name and stays that way for the node's
whole life, so the set can be taken fresh each time. Say something in the log
when a declared port is not one the device has once, not once per poll: from
when a declared port is not one the device has, once rather than once per poll: from
the canvas a renamed entity and a typo look the same, and both leave a port
silent forever.
@@ -131,7 +131,7 @@ async def poll(self) -> dict[str, Any] | None:
- Return `None` when there is nothing new.
- **Only changed values are published.** A device polled every few seconds
usually says the same thing, and every publication wakes everything
downstream, so the loop compares against what it last published what it
downstream, so the loop compares against what it last published, meaning what it
actually published, so a publication that failed is retried next tick rather
than counting as said.
- Raising is not fatal: it is reported as a health problem and retried on the
@@ -182,7 +182,7 @@ async def poll(self) -> dict[str, Any] | None:
```
`save_artifact` stores the bytes and returns what an `image`, `audio` or
`video` port carries the media type has to match the port's type. It only
`video` port carries: the media type has to match the port's type. It only
works once the node has started, since the store is the engine's and is handed
over then.
@@ -200,7 +200,7 @@ async def stop(self, app=None) -> None: ...
- `start` is called when the connector's flow starts, and after every rebuild.
- `stop` is called before a rebuild and when the flow is stopped. **It must be
idempotent** it is called whether or not `start` succeeded.
idempotent**: it is called whether or not `start` succeeded.
- Overriding either means calling `super()` if the polling loop is also wanted.
- The `app` argument is the FastAPI application, for the rare connector that
needs to mount a route. Most ignore it.
@@ -216,7 +216,7 @@ self.report_health("down", str(exc))
```
Three values, `ok`, `degraded` and `down`, plus an optional detail string.
Reporting the same status twice is free only changes are published. A node
Reporting the same status twice is free, since only changes are published. A node
reporting `down` is named among its flow's issues and counted on the health
summary on Home; `degraded` means still working, and is not. The polling loop
already reports around `poll()`; a connector managing its own connection should
+17 -20
View File
@@ -3,10 +3,8 @@
Every type the canvas can place. Each is configured by filling in a form the
editor generates from its parameter schema, so they all behave the same way.
Anything here could be written as a **Function** node — that is what the
function node is for. These exist because the same handful of shapes account
for most of a real instance, and a rule you fill in is easier to read on a
canvas, and to change, than five lines of code repeated eighty times.
Anything here can also be written as a **Function** node. These types cover the
shapes that come up most often, as a form to fill in rather than repeated code.
`GET /flows/node-types` returns this list with each type's full parameter
schema, including any [connectors](#connectors) installed on your instance.
@@ -40,13 +38,15 @@ A node with *outputs only* subscribes; a node with *inputs* publishes.
| `qos` | `0` | 0, 1 or 2 |
| `retain` | `false` | on published messages |
| `keepalive` | `60` | seconds |
| `timeout` | `10.0` | seconds before a broker operation is given up on |
| `publish_queue_size` | `256` | payloads that may wait for the broker; past this the oldest is dropped and the node reports degraded |
| `json_key` | — | key to lift out of an object payload; one for every port, or `{"port": "key"}` |
A topic may be a filter: `+` matches one level, `#` the rest. Everything a
filter matches lands on the same port, so use one port per topic where the
difference matters.
`json_key` is for a device that wraps its reading Victron publishes
`json_key` is for a device that wraps its reading: Victron publishes
`{"value": 47}` on every path. Without it, a payload object is unwrapped only
when it happens to carry the port's own name as a key.
@@ -100,10 +100,8 @@ Outputs only makes it a **webhook**: the engine mounts a route at
written, and the rows come back on the first output port as
`{"rows": [{ts, value, field, measurement, tags}], ...}`.
That is what keeps a database node a database node: it holds the credentials
and the connection and nothing else. Building a query and shaping its rows are
ordinary Function nodes on either side, so a dashboard widget never learns which
database answered it.
A database node holds the credentials and the connection, nothing else.
Building a query and shaping its rows are Function nodes on either side.
### Notification
@@ -126,7 +124,7 @@ thing configured elsewhere.
### Inject
**`inject`** — emit a value on request, on a schedule, or when the flow starts.
**`inject`** — emit a value on request, on a timer, or when the flow starts.
| Setting | Default | Notes |
|---|---|---|
@@ -138,8 +136,8 @@ thing configured elsewhere.
| `start_delay` | `1.0` | how long to wait before that first emission |
The scheduler: a `cron` expression here is what makes a flow run by the clock.
It is also the most-placed node in a real instance mostly as a button
someone presses.
It is also the most-placed node in a real instance, mostly as a button someone
presses.
### Delay & schedule
@@ -165,7 +163,7 @@ needs a busy cascade slot waits for one.
!!! note "Not in a batch flow"
A rate limit holds a value back for a timer to release, and a run has no
timer the value would be dropped rather than delayed. Submitting a batch
timer, since the value would be dropped rather than delayed. Submitting a batch
flow with a rate-limited port is refused instead.
### Trigger
@@ -183,13 +181,13 @@ starts from a value arriving. For a cron tick, see [`inject`](#inject).
| `passthrough` | `false` | send the incoming value instead of `first` |
| `wait_port` | — | an input carrying the wait, when it differs per message |
The shape this exists for: *the door opened turn the light on, and off again
in two minutes unless it opens again.*
The shape this exists for: *the door opened, so turn the light on, and off
again in two minutes unless it opens again.*
`wait_port` covers the case where how long to wait is itself a value: a
rollershutter takes 26 seconds up and 28 down, so the node that decides the
direction says how long to run for as well. A wait of zero or less sends
nothing afterwards and still cancels whatever the last message scheduled,
nothing afterwards, and still cancels whatever the last message scheduled,
which is how a *stop* is commanded exactly once.
## Logic
@@ -289,13 +287,12 @@ reloads identically.
|---|---|---|
| `seed` | `0` | |
Kept as a worked example of numeric logic rather than as a modelling tool. If
you are training something, that is a [batch flow](../concepts/runs.md) and a
Function node.
A worked example of numeric logic, not a modelling tool. To train something,
use a [batch flow](../concepts/runs.md) and a Function node.
## Connectors
Anything else in the palette came from an installed **connector** package a
Anything else in the palette came from an installed **connector** package: a
node type written against a published contract and discovered through the
`fluksio.node_types` entry point group. The editor shows which package supplied
it.
+13 -14
View File
@@ -20,11 +20,10 @@ the state backend, the work queue and the worker protocol unchanged.
| `str` | a string |
| `bool` | exactly `true` or `false` |
`bool` is an `int` subclass in Python and deliberately not a number here: a flag
is not a measurement, and a switch bound to a temperature is a mistake worth
catching.
`bool` is an `int` subclass in Python but is not a number here, so a switch
cannot bind to a temperature.
`NaN` and infinity are refused, wherever they sit including inside a `json`,
`NaN` and infinity are refused, wherever they sit, including inside a `json`,
`record`, `series` or `list`. JSON cannot spell either, so one that travelled
would come back as a response nobody can parse and a row the database rejects,
a long way from the node that made it. An empty subset or a division with no
@@ -43,8 +42,8 @@ Flat named scalars.
{"title": "Boiler", "body": "Pressure low", "severity": "warning"}
```
Nesting is deliberately out: a record that can contain a record is a schema
language, and the shape stops being readable from the declaration alone.
A record cannot contain a record, so its shape is readable from the declaration
alone.
Read by the **Notification** widget. It is also what an alert channel of kind
*dashboard* writes.
@@ -64,7 +63,7 @@ point.
### `series`
Labelled lines of `(timestamp, value)` pairs what a chart draws.
Labelled lines of `(timestamp, value)` pairs, which is what a chart draws.
```json
{
@@ -91,7 +90,7 @@ A reference to stored bytes.
{"digest": "sha256:…", "size": 4194304, "media_type": "application/octet-stream", "name": "weights.pt"}
```
Binary payloads tensors, checkpoints, images never travel as a message. The
Binary payloads (tensors, checkpoints, images) never travel as a message. The
bytes go to a content-addressed store and the message carries this. A
thirty-megabyte checkpoint never sits in the state backend, and the reference
stays valid wherever the store is reachable from, including on another machine.
@@ -110,14 +109,14 @@ The same reference, narrowed to a kind of media by its `media_type`.
An `audio` port takes `audio/*` and refuses anything else, so a speech
recogniser declares what it eats rather than taking any bytes at all and
finding out. An `artifact` port still accepts all three media narrows
finding out. An `artifact` port still accepts all three: media narrows
artifact, not the other way round.
`meta` is optional and nothing here reads it: sample rates, dimensions and
sequence numbers are for whoever consumes the media.
Bytes still never travel as a message. A camera publishes one reference per
frame and a microphone one per chunk which makes a media stream an ordinary
frame and a microphone one per chunk, which makes a media stream an ordinary
[streaming port](../concepts/flows.md#streaming-ports), and each frame an
artifact. What that costs is worth knowing before pointing a camera at it:
@@ -143,8 +142,8 @@ next author nothing.
## Naming a run's output
Any *run parameter* is also accepted as text, since nobody wants to paste an
object into a shell. `@run:<id>.<output>` names what an earlier run produced
whatever its type, an `artifact` reference or a `json` config alike — and a bare
object into a shell. `@run:<id>.<output>` names what an earlier run produced,
whatever its type, an `artifact` reference or a `json` config alike. A bare
`sha256:…` digest names content in the artifact store. Both resolve before the
run starts, so the CLI, the run dialog and a python caller all mean the same
thing by the same string.
@@ -171,8 +170,8 @@ Enforced on the server as well as in the editor.
A value that does not match its port's declaration raises on the node that
published it, naming the port and what arrived. It does not get published, and
it does not reach anything downstream — a wrong value stopping at its source is
much easier to diagnose than one propagating.
it does not reach anything downstream. A wrong value stopping at its source is
easier to diagnose than one propagating.
## See also