Files
app/backend
stroblmeandClaude Opus 5 608d30d884 Let a node say how much of the machine it takes
Five concurrent training nodes, each sizing its thread pool to every core,
left the engine's own event loop unscheduled: the API stopped answering
within 10 s and every client died. The same shape on a GPU deadlocked a run
for 21 minutes at 0% utilisation with nothing failing and nothing to read --
it just sat in `running`.

@node(resources={"cpus": 2}) is the declaration. The engine holds that much
for the length of the execution, so more of them than the machine has room
for wait their turn rather than oversubscribing it, and a `gpus` node holds
its card exclusively. FLOW_CPUS defaults to every core but two, and those two
are what keeps the engine answering.

Because a thread cap is read when the process imports the library, a warm
worker cannot be told a different one -- so an environment gets a pool of its
own and nodes deriving the same one share it, rather than paying a cold start
per call on exactly the nodes whose imports are slowest. XLA_FLAGS is never
derived: it is a composed, version-dependent string, so it travels in
resources.env where it is visible.

A node that declares nothing is not accounted for and behaves as it always
did -- it just gets FLOW_CPUS/FLOW_MAX_WORKERS as a thread cap, which is the
half of this that fixes the reported incident without anybody declaring
anything. An operator who set OMP_NUM_THREADS themselves still wins.

Resources are claimed strictly before a worker slot, so the two blocking
waits cannot deadlock. A node queued for them publishes node_queued and shows
on GET /workers/resources, because waiting and hanging looked identical.

Accounted, not enforced: no cgroups, no rlimits. Scheduling across machines,
flavours and enforcement are the next steps.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 21:36:59 +02:00
..
gc
2026-08-24 19:06:54 +02:00

Fluksio

Fluksio is a node-based automation software that brings trust and reliability to your flow. It just works and looks good. Get started by running

pip install fluksio
fluksio serve

and you're ready to go.

For data science

You can turn your existing data science project into a flow by decorating your functions with @node ...

# myresearch/train.py
import fluksio
from fluksio import Port, node

@node(
    requires=["dataset", Port("lr", "float")],
    provides=[Port("loss", "float", stream=True), Port("weights", "artifact")],
    device="gpu", device_policy="prefer",
)
def fit(dataset, lr, epochs=25):
    for epoch in range(epochs):
        loss = step(...)
        yield {"loss": loss}          # published as it happens, kept as a series
    return {"weights": fluksio.save_artifact("weights.pt")}

... and passing them to a Flow:

# myresearch/pipeline.py
from fluksio import Flow, Port
from myresearch.data import prepare
from myresearch.evaluate import evaluate
from myresearch.train import fit

train = Flow("train", nodes=[prepare, fit, evaluate],
             inputs=[Port("lr", "float", initial=0.01)], outputs=["score"])

Fluksio will automatically infer the order of nodes based on the inputs and outputs you defined. When everything is set, you can launch your first run as follows:

fluksio run train --lr 0.05 --wait

Checkout our documentation for more infos.

Some other features

  • Flows: typed messages between nodes, wired by name, edited on a canvas or declared in code. Every change is a commit in a git repository you own.
  • Runs: an experiment and a CI-style job are the same entity. Parameters, seed, result, per-node timings, artifacts and the commit it ran at.
  • Dashboards: charts and controls bound to the same messages the flows carry, with no separate metrics pipeline.
  • Remote workers: pip install fluksio-worker on the GPU box; it dials out over one websocket, so nothing there has to be reachable.

Fluksio can also be used for facility automation. Visit us on fluksio.com or go straight to our documentation.

License

Copyright (C) 2026 Melvin Strobl - GNU Affero General Public License v3.0 or later. Running a modified version over a network obliges you to offer its users the corresponding source (AGPL §13).