Execution was fire-and-forget: an MQTT message or webhook ran a cascade on a ThreadPoolExecutor built for that one wave, and an engine that died halfway through simply lost whatever was in flight. Concurrent triggers each built their own pool, so load meant unbounded threads. Every external trigger is now journaled to a Redis Streams queue before anything runs, and acknowledged only once its cascade finishes. A consumer thread drives cascades on one long-lived pool while node bodies run on another, so a cascade cannot starve the nodes it is waiting for. A reaper reclaims what a dead consumer never acknowledged — verified end to end: work journaled while the engine was stopped runs on restart, and work abandoned mid-cascade comes back as a second delivery. At-least-once needs a guard, so nodes that reach outside are marked non-idempotent and skipped on a redelivery they already completed. Without Redis the queue degrades to an in-memory one that does not pretend to be durable, and interactive callers still run inline. Also fixes two things this turned up: a delay node was sleeping on a worker thread, where a handful of them could occupy the whole pool, and webhooks 404'd whenever MCP was enabled because the app mounted at / answered first for every path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011LF61rxW1FG5YCD2J9YqjY
8.8 KiB
Roadmap
Component-level breakdown. The milestone-level master (M1–M5, with the vision
decisions behind it) is docs/private/roadmap.md in the docs submodule.
Implementation strategy and record of existing/planned features. Completed items are
terse checklists — the requirement detail lives in docs/private/vision.md (goals,
requirements, decisions) and docs/architecture/structure.canvas (the four-way component split).
Remaining tasks keep enough scope to be actionable.
Legend: [x] done · [ ] planned · sub-lists split done vs. remaining for partial items.
Within each phase, remaining [ ] items are listed in rough priority order: making the
existing flow engine reachable and persistent precedes new feature breadth.
Phase 0 — Workspace and platform
- Root orchestrator repo with
app,indexanddocsas submodules make initbootstrap: secrets generation, per-stack.envpropagation, sharedproxydocker network- Layered compose (
compose.yml→compose.dev.yml→compose.local.yml) for both stacks, one Traefik serving${DOMAIN},app.${DOMAIN},api.${DOMAIN} - Design token contract: root
DESIGN-GUIDELINES.md, per-repoDESIGN.md, byte-identical token blocks verified bymake design-check - CI on Codeberg (Forgejo Actions): pre-commit, backend tests, Playwright, compose smoke
Phase 1 — Backend: management
Python, optimised for development speed. Owns the graph structure, persistence and the
external interfaces. See docs/architecture/structure.canvas → Backend – Management.
- FastAPI + SQLModel + Alembic + Postgres base with JWT auth and user management
- Flow engine in
backend/app/flow/:Node/Pipeline/StateBackend(memory + Redis) /FlowController - Node types: HTTP, MQTT, InfluxDB, Delay, MLP
app/flowis an importable package with absoluteapp.flow.*imports- Typed, serializable node I/O: every port declares a
DType, messages are JSON on the wire and in Redis, no pickle anywhere. Binary codecs are still open —DType.JSONcarries everything non-scalar for now - Message namespacing per flow (
flow.message), with several producers per message resolving to real fan-in - Secrets/credentials store for node integrations managed via the API/UI
(encrypted at rest, referenced from node params as
{"$secret": "name"});.envbootstrap-only - Connector node contract:
ConnectorNodewith a declared contract version, a polling coordinator that deduplicates,x-secretparameters the editor renders as a secret picker, and health reporting. Connectors are installed packages found through thefluksio.node_typesentry point group; the contract is documented indocs/connectors/with a working skeleton atconnector-skeleton/. The registry follows later - Node lifecycle as a protocol (
start/stop/report_healthonNode), replacing the controller's per-type isinstance chains — the same hooks a connector implements, validated on the built-in nodes first - Flow persistence:
flow.jsonplus node sources per flow, replacing the watch-directory prototype - REST + WebSocket API over the engine: create/read/update flows, edit node source, run, and stream values, node status and execution events
- Dependency-loop detection and graph validation surfaced as API errors
- Per-flow start/stop, stored in a
runtime.jsonbeside the flow so it survives a restart and stays out of the autosaved document; pause/resume holds a flow's nodes while its values keep arriving - Node log streaming: what a node prints, and the traceback of one that
fails, reach the editor as
node_logevents - MQTT broker / InfluxDB compose services for local development
- Git-based versioning of the flow store (one commit per saved change)
- Draft/publish split: edits autosave to
flow.draft.json/nodes.draft/, the engine runs only the published files, and publishing promotes the draft. Saves carry the version they were based on, so a second client editing the same flow is refused rather than overwritten - Import/export of a flow as human-readable code plus a JSON structure
- Per-input/-output discretization interval setting: a port publishes, or wakes its node, at most every n seconds. State keeps the latest value, so only the delivery is skipped
- Alert / notification handler
- Deep health check (
GET /utils/health/): reports event-loop lag and state-backend reachability and fails the container healthcheck, so a wedged engine is restarted rather than counted as up. One engine per deployment — the API image runs a single worker, because a second one would be a second engine - Supervised background tasks: a node's subscription, schedule or poll loop is restarted with growing delay when it dies, and a flow that spends its failure budget is quarantined and surfaced rather than left crash-looping. The loops themselves no longer carry private retry logic
- Durable work queue: every external trigger is journaled to Redis Streams before anything runs and acknowledged once its cascade finishes, so an engine that dies mid-cascade picks the work up again instead of losing it. A reaper reclaims what a dead consumer never acknowledged; nodes that reach outside are skipped on a redelivery they already ran. Long-lived worker pools replace the per-wave executors, and a delay now waits in the queue rather than on a worker thread
- Test nodes: a small node dragged onto an existing one, smoke or unit, blocking deployment on failure
- User management scoped per flow and per data set
- MCP server over the same API: agents authenticate through a built-in
OAuth 2.1 authorization server (dynamic registration, PKCE, rotating
refresh tokens) and drive the flow API through 20 tools. Tokens are
RS256, signed with their own keypair, so the set can be revoked on its
own — and an additional issuer is one branch in
deps.decode_token, which is the seam remote access needs later - LLM interface for natural-language flow authoring beyond the MCP tools
Phase 2 — Backend: processing
Rust, optimised for throughput. Executes nodes and distributes them across workers. See
docs/architecture/structure.canvas → Backend – Processing.
- Parallel invocation of stateless nodes over independent input sets, to
keep I/O delay minimal (stateful I/O nodes keep serializing via the
synchronousmechanism) - Extract node execution from the Python prototype into a Rust engine
- Worker distribution and load balancing across capable devices
- Input/output validation at the node boundary
- Data aggregation and discretization
Phase 3 — Frontend: admin view
React + Vite, primarily desktop but usable on mobile. See docs/architecture/structure.canvas →
Frontend – Admin View.
- Dashboard SPA shell: TanStack Router, floating frosted sidebar, auth flows, generated OpenAPI SDK
- Node canvas (
@xyflow/react) showing nodes and their connections, which are derived from message names rather than stored - Tab-style view of atomic flows, with a floating dock
- Embedded code editor (Monaco) for node source
- Live values on the edges, with the last payload and its time on click
- Validation shown on the node it belongs to, and summarised in the dock
- Publish control and draft markers in the flow bar, discard in the flow panel, and a conflict dialog when another client got there first
- Marking a node reusable, and placing a shared one from the palette
- Secret picker for credential parameters, so a password never lands in
flow.json - Dashboard showing which flows run, which are stopped and which have errors, with a switch per flow
- Logs panel in the canvas dock, pause/resume beside Run, and replaying an edge's last message from the inspector
- Device assignment per node, selectable from compatible devices
- Test-node affordance on the canvas
- User management screens
- Mobile-friendly canvas: touch connect, full-screen node panel
- Installable as a PWA (
vite-plugin-pwa)
Phase 4 — Frontend: dashboard view
Shares components with the admin view. See docs/architecture/structure.canvas →
Frontend – Dashboard View.
- User-defined dashboard layout with edit and view modes
- Responsive layout targeting wall panels, mobile and desktop
- Per-device view
Phase 5 — Website and docs
- Marketing site with a live node-graph demo, shared design system
- Published documentation site fed from the
docssubmodule - Umami analytics configured (the site still ships the placeholder script)