39 KiB
This file captures tasks which derive from roadmap tasks (unfinished, deferred), bugs encountered during usage and feature requests/improvements which are not fitting directly in the roadmap. Always sort by priority and put tasks blocked by other tasks/features at the dedicated section. When working on a task, check for other, similar tasks that could be resolved on the way. Use following pattern to classify tasks: TYPE/SCOPE Where TYPE could be BUG, FEAT, PERF, CHORE and SCOPE could be UX, UI, FLOW, NODE, API, INFRA, DOCS appended by MOBILE if only for mobile use case. Don't write temporary reasons for deferring a task in the task description (only strategical reasons should be noted). Deferring because out of scope is fine, but don't mention deferring than.
Deferred holds what stays open on purpose, each with the condition that should reopen it.
Open
To be sorted
-
CHORE/PKG the SPA is not in the wheel:
fluksio serveserves no UI, on the assumption that a pip install is paired with a portal. Bundlingdist/and mounting it withStaticFileswould give a local dashboard — it needs a hatch build hook running bun,VITE_API_URL=""for the same-origin case, and a decision about the MCPmount("/")it would collide with. -
CHORE/PKG the Docker image still starts with
fastapi run;fluksio servenow does the same thing plus the bootstrap. Switching would give the container and a pip install one code path. -
CHORE/DEPS
sentry-sdkwent to 2.x and therequires-pythoncap came off with it. Nothing exercises Python 3.13/3.14 in CI — the matrix is one version. -
FEAT/UI add a list of the dashboards next to the list of flows in the home view. Use a mosaic like structure with previews of the dashboards; Then make the list of flows capped at a certain number; everything above that should be scrollable; sort by recently modified. The dashboard mosaic should have the exact same height as the flow list both capped at a lower limit to allow showing 1 flow and 1 dashboard
-
BUG/UI assimilate the design of the settings in the app to mirror the design of the settings in the portal
-
BUG/UI the "Installation Offline" warning (and notification in general) should be centered w.r.t. the viewport (currently it is a bit left, discarding the width of the sidebar). Furthermore, the notification does not seem to disappear on its own. Reloading the page solves it
-
CHORE/INFRA: the
generate-frontend-sdkpre-commit hook runsscripts/generate-client.shon anybackend/**change, and that script ends by formatting the whole frontend tree — whilebiome.jsonexcludessrc/client, so it formats nothing the generator wrote. Every backend commit therefore rewrites files it never touched. Dropping the trailing format call is the fix; it was kept this wave only to preserve behaviour. -
CHORE/UI: acknowledging a node's failure clears it on the engine but publishes no event, so another browser watching the same flow keeps the marker until its next snapshot or rebuild. One event on the bus would close it, the way
dashboard_changeddoes for panels. -
CHORE/API:
save_dashboardstill catchesDashboardNotFoundfromwrite_draft, which can no longer raise it. Harmless, and the same shapesaveFlowhas: a PUT to an unknown name now creates that dashboard's first draft rather than answering 404. -
CHORE/UI: the editor-side
ModePickerandStylePickerinpanels.tsxare still the older flex row with a jumpingbg-accentfill, while the widget-side segmented control andRangePickernow slide one thumb over equal grid tracks. Two shapes for one control. -
CHORE/UI: an icon set on a dashboard cannot be cleared from the UI — Radix forbids an empty
SelectItemvalue, so neither the rail-icon select nor the existing "Otherwise" select offers a "none". Both would need the same affordance. -
FEAT/UX add an option to the settings of an installation to configure automatic updates. If enabled, the installation would send a request e.g. every 1h to the hub at fluksio.com and the hub then checks if a new version is available. The settings should include a second toggle for automatically installing an update (which might cause a short outage). Later this mechanism should be extended to check if updating would cause things to break.
-
FEAT/UI allow setting icons for multi-page dashboard (when configuring a panel, we could simply add an icon picker there)
-
BUG/UI hovering the sidebar where we can switch dashboards on a multi-page dashboard shows a horizontal scrollbar. We should remove that; no scrollbars at all should be shown in this type of sidebar
-
BUG/UI the "Installation Offline" warning (and notification in general) should be centered w.r.t. the viewport (currently it is a bit left, discarding the width of the sidebar). Furthermore, the notification does not seem to disappear on its own. Reloading the page solves it
-
CHORE/UI:
.u-legendis styled in the dashboard's own CSS chunk, so a uPlot legend on a Health or Home page renders unstyled until a dashboard has been visited in that session.RangePickeralready side-effect-importsdashboard.cssfor the segmented thumb; the legend rules want the same treatment, or a shared chart stylesheet. -
CHORE/UI: the bar widget's readout swaps between an inline
rightand an inlineleftas the value crosses 30%, and the side it stops setting resets toauto, which does not interpolate — so the label jumps once at that threshold while everything else animates. Positioning it always byleftplus atranslateX(-100%)would put the whole travel on one property. -
CHORE/UI: two of the mobile-overflow floors are over-determined. Removing
.widget-grid { min-width: 0 }, or the segmented fieldset'smin-w-0, leaves the mobile suite green — the grid tracks are alreadyminmax(0, 1fr)and the fieldset became a grid. The uPlot legend is the one offender the assertion actually catches. Both are cheap insurance for a future widget, but nothing would notice if they regressed. -
CHORE/API: artifact blobs are content-addressed and have no GC, so deleting a flow drops its
run_artifactrows and leaves the bytes on the data volume. -
CHORE/API:
DELETE /flows/{name}does not refuse while the flow has arunningorqueuedRun.RunService._record_nodecan then insertrun_noderows for a run that no longer exists;_finishis an UPDATE, so it degrades to a harmless 0-row no-op.make test-backendthen fails at the coverage HTML step after every test has passed — which reads like a test failure and is not one. (tests/flows.spec.ts, tests/admin.spec.ts)". There are nine. -
FEAT/UI add (multi-)select to the flows and dashboards view to allow deleting (multiple) items; long press to select -> "Add" button should change into "Trash" icon button
-
BUG/UI in the brain view: make the chasing circle animation running entirely in the gap between the ring and the node (using the full width)
-
BUG when clicking "edit" in the "Home" dashboard of the demo on hub.fluksio.com, most of the panels disappear (only a handfull is left for actual edit)
-
CHORE/UI: the edge popover shows the same value twice —
MessageSparklinefalls through to a collapsedValuePreviewfor a non-numeric value, andEdgeInspectorthen renders its ownValuePreview defaultOpenbelow it. Cosmetic; one of the two is redundant. -
FEAT/UI: a settings-and-inputs overview page, so what every node of an installation is configured with can be read and searched in one place rather than one panel at a time.
-
FEAT/UI: an input endpoint opens the flow panel, which is right for editing but not for reading one value. A panel of its own — the declaration, the current value, its history — is what clicking a label wants to give.
-
FEAT/UI: sync between the header of the python function and the node configuration. The config→header half exists for ports and settings —
scaffoldForwritesdef process(<ports>, <settings>)andeditNodekeeps it in step — but only while the source is still exactly the generated scaffold (SCAFFOLD_SHAPE), and never for shared code. What is missing is the same for code someone has edited, and the reverse direction: nothing parses adef process(...)header back into ports and settings. -
BUG/UI on flows like "House history" where the widget sets the range for the "draw the window" node to generate some data, the edges overlap the nodes. We should adjust the flow visualization to account for these cyclic behaviors
-
CHORE/UI: loop lag on Home reads a real number with no flows, and that is right —
LoopWatchdogtimes how lateasyncio.sleep(1.0)wakes on the API's event loop and is started unconditionally, so it measures the engine process rather than any flow, and it is what turns the health badgedegraded. Nothing to fix; recorded so it is not reopened. -
BUG/UI auto node placement on flows should be improved in regards to least crossing edges and a more vertical layout on mobile devices
-
INFRA: ensure that all the packages/ dependencies needed to run fluksio are available on arm to make this software runnable on e.g. raspbian
-
INFRA: merge the philosophy statement at the beginning of vision.md into the rest of the document. Dissolve the decision dates and fold the decisions into a clean structure
-
FEAT/UI add a loading animation for the initial app load and when loading individual pages; make sure that elements e.g. in the home dashboard load independently to ensure a fast loading of the initial site but figures charts, tables, graph etc. follow after that
-
FEAT/UI introduce a graph panel which renders at the top right next to the graph view (to make more use of the horizontal space) and which allows (de-) selecting flows to be excluded from the graph view or search for individual nodes where only the flows containing this node should be shown (like slicing the brain)
-
FEAT/UI labels in flows (indicating dashboard widget connections) naturally can't pulse. Instead add an animation (enlightning fade) from either ltr or rtl depending if the label is in- or outbound
-
CHORE/UI:
layoutGraphtreats every node as 220×56 rather than measuring, because feeding a measurement back into the layout oscillates. A node wider than that crowds its neighbours; take the sizes fromnode.measuredonce they have settled if it shows. -
FEAT/UI/MOBILE: a rank of many nodes — a connector feeding eight dashboard tiles — is thousands of pixels wide however the graph is turned, so on a phone the fit shrinks it past reading. The layout is right and the flow is simply too big for the screen; a "one rank at a time" reading mode, or wrapping a wide rank, is what would make it legible.
-
CHORE/UI: an edge's value chip sits at the bezier midpoint while the layout reserves its room at dagre's label rank. The two agree closely enough today; if chips ever pile up, take the position from the layout instead.
-
FEAT/UI (deferred until MCP lands): add a "bot" icon button to the home view (graph panel) which opens a chat window (reuse general concept of a side panel like in flows/nodes to make it a chat panel which can open on any screen (stacks below any other existing panel -> introduce stacking) to give support on errors/write code, generate dashboards etc) to explain the error(s)
-
FEAT/UI make the header (Fluksio - YEAR) and the logo in the sidebar link to the main page (fluksio.com)
-
FEAT/UI consider adding a diagram to the Home view which shows a histogram of the different classes of nodes and which time it takes to execute (logarithmic scale); this should give a hint on the load and help to detect bottle necks/hotspots
Persistence and databases
From the 2026-08 database review. Verdict recorded under Deferred: the Postgres + Redis + git-files split stays; the actionable part is durability.
- CHORE/INFRA: Redis AOF runs at
appendfsync everysec, so up to ~1 s of journaled work-queue entries can vanish on a crash — softer than "journaled before it runs" reads. Queue write volume is low, soappendfsync alwaysis likely affordable; otherwise document the loss window. - CHORE/INFRA: Redis has no auth (
requirepassunset). Fine on the compose-internal network; a blocker for M5 remote workers, which turn Redis into a network-exposed shared bus.
Connector write paths
Needs someone watching the real hardware, so it is not a background task. This is what M4 still waits on, together with porting the flows.
Art-Net can write now: ConnectorNode.write carries a node's input ports, a
per-port channels map places each on its own DMX channel, and transmit
still gates the socket. Verified on the wire against a listener (channel 33 =
255, channel 31 = 60, nothing else set) and the MQTT half was driven end to
end against the house broker. The rig is the house_control flow and its
dashboard, seeded by make -C app seed-house.
- FEAT/NODE: Art-Net against the real fixtures is still untried. The house's own dmxnet sender re-emits universe 1 every 1000 ms, so fluksio and Node-RED overwrite each other; the test needs Node-RED's Art-Net sender stopped, and while it is stopped every channel fluksio does not set is dark.
- CHORE/NODE: the operatorId worry was unfounded — the reference Node-RED
setstatnode for this unit is configured with an empty operatorId and deviceId, so a command needs no registration. The second unit may still differ. - PERF/NODE: a
wfraccommand is two round trips (read, then set) on the scheduler's thread, so at the default timeout a command can hold a cascade for several seconds. Fine for a person pressing a button; a flow commanding it on a schedule would want the work off that thread. - CHORE/NODE:
wfracwrites carry the unit's whole state, so two flows commanding one unit will each undo whatever the other set between their read and their write. One writer per unit, the same rule Art-Net has for a universe. - FEAT/NODE: the second WF-RAC unit (the one Node-RED addresses with operatorId "0") closes the connection on an anonymous read. It likely wants an account registered; the first unit answers without one.
- CHORE/NODE:
wfracsometimes reportsmodeas "unknown" while the unit is off, because the mode bits hold a value outside the known set. Narrower than it first looked: a unit switched off by a command keeps its last mode in those bits and reads back correctly, so this is about however the remote turns it off. "off" would still read better than "unknown". - PERF/NODE:
ArtNetOut.writesends one frame per input port, so a node with two ports emits two frames per run. The last one carries both channels, so the end state is right; folding them into one send would halve the traffic. - CHORE/NODE: the Art-Net node starts from an all-zero universe and has no way to learn what the fixtures are currently at — Art-Net has no read-back. Taking over a universe therefore blanks everything the flow does not drive. A baseline setting, or driving every channel, is what a real cutover needs.
Porting the Node-RED flows
What the reference actually does, extracted while building the write-path rig.
One Art-Net Out node in 865 covers every physical device: mqtt in <topic>
-> change (msg.topic = DMX channel) -> an nCH encoder function -> Art-Net,
to the Art-Net node's address universe 1. Payloads are bare: ON/OFF, a number, [h,s,v],
UP/DOWN. Nothing is retained, so state lives only in Node-RED globals and
is lost on its restart.
- CHORE/FLOW: three DMX channel collisions in the reference — ch 9 (
light/bathRoomLightvslight/bathRoomSinkLight), ch 28 (actor/windowOpenerStorage28-29 vslight/traverseAmbientLight28-30), ch 129 (actor/canopy129-130 vs an orphaned 1CH mapping). Decide these deliberately rather than porting them. - CHORE/FLOW: 1CH values are not scaled.
light/traverseSpotLightandlight/kitchenDirectLightreceive 0-100 and that number lands on DMX as-is, so those fixtures never go above 100/255. The 4CH "A" channel getsvraw for the same reason. Faithful is ugly; deliberate is better. - BUG/FLOW: the reference's
function 9/function 10publish the string"undefined"for unselected zones (var a, b, c = [0,0,0]only initialisesc), which reaches the 3CH/4CH encoders and producesNaNDMX values. A port should emit an explicitOFF. - CHORE/FLOW: dead in the reference and not worth porting — the
AmbientModeToLightchain, the four Dashboard toggle chains,light/generalLight(written, no consumer),light/outdoorPavillonLight(no wiring),Color Adapt.
Bugs found while building the screens
- CHORE/API: revoking an OAuth client does not invalidate access tokens already issued; they are stateless JWTs valid up to
MCP_TOKEN_EXPIRE_MINUTES. Immediate revocation meansapp/mcp/http.pychecking the client row still exists. - CHORE/FLOW:
Pipeline.trigger's docstring says a paused flow still publishes so the value shows on the canvas. True only without a queue; with one the item parks beforeapply_outputsand nothing shows. Docstring and behaviour disagree.
Out-of-process nodes and modules
- CHORE/FLOW:
compile_checksends the draft source under the running node's cache key, so the worker recompiles the published source on its next call. Correct, but one wasted compile per save on a busy node. - FEAT/API:
POST /modules/applyrebuilds the whole pipeline so a node that could not import its package stops being red. That resubscribes every MQTT node in the deployment; a targeted rebuild of the flows that actually failed to load would be gentler. - CHORE/FLOW: a node's return value now round-trips through JSON, so tuples arrive downstream as lists and anything non-JSON is an explicit error. That is the message contract, but flows written before this may notice.
- BUG/FLOW: a node whose cold-start imports plus body exceed its timeout can never succeed. The timeout covers the first call's imports, a timeout kills the worker so the next attempt is cold again, and
compile()only ever warms one of the N workers. Broadcastingcompileto every worker is the candidate fix, at the cost of N module executions per reload. - CHORE/FLOW: worker protocol loose ends — the request
idis echoed but never checked,json.dumpsruns twice per result (once to prove it is JSON, once to send it),_remote_typesis an unbounded cache keyed on class names that user code chooses, andPythonWorkerPool._lockguards less than its name suggests.
Engine history
- CHORE/FLOW: a rate-limit flush gets no run record — it is the tail of the run that scheduled it, and there is no id linking the two. A flush that fails therefore shows as a failure with no run beside it.
- CHORE/FLOW:
Pipeline.flushreleasing a held value runs its cascade without a run id, so those executions land in the minute rollups but in no run. Threading the scheduling run's id through the queue item would close it. - CHORE/FLOW:
EventBus.emitscounts a node's publishes so a reconnecting client can restore what it missed. Two deliberate shortcuts: the increment is a read-modify-write, so two threads emitting from one node can lose a count —publishis documented as never blocking, and a dropped increment is invisible in an animation — and the dict is never pruned, so a deleted flow's node ids sit there until restart. Bounded by distinct ids seen in the process, and orphans are never read, since lookups go throughbrain_graphmembers. - CHORE/API: the metrics collector is a bus subscriber, so a storm that overflows the bus queue undercounts. The events dropped are the same ones the websocket drops; exact accounting would need the collector to be fed from the engine rather than the bus.
- CHORE/UI: the Home block's "Changes" list is the newest 15 audit rows whatever range is selected. Deliberate — an audit trail is worth reading past the window — but it sits under a control that governs everything else on the screen.
- CHORE/FLOW: run records for a deleted flow stay until the retention window passes, so a flow that no longer exists keeps appearing in the history. Deliberate — it is a record of what ran — but
forget_flowcould offer to clear it. - CHORE/API: nothing can ask the collector to flush now, so anything needing the tables to be current has to wait out
FLUSH_INTERVAL_S— which is what the soak harness does before clearing its own rows. - CHORE/FLOW:
RedisWorkQueue.clear_flowdeletes onlypipeline:__parked__:{flow}, so a deleted or renamed flow's__queue__stream entries,__delayed__zset members and__done__:*markers stay behind. The stream is capped and the entries are dropped when they reach a node that no longer exists, so it costs work rather than correctness. - CHORE/FLOW:
MemoryWorkQueue's in-flight count is a counter around claim/ack, and claiming already removed the item — so an item a handler leaves unacknowledged (no pipeline bound) counts as in flight until the process ends. Nothing can hand it back either way, which is what the memory queue is. - CHORE/API: audit rows ride the same drop-oldest bus as telemetry, so a storm can lose one. Writing a node's source is not audited either; publishing is.
- PERF/API:
queue.stats()does a keyspacescan_iteron every call while two endpoints poll it. - CHORE/INFRA: dev only — memory-queue ids (
mem-{seq}) restart at 0 each boot andFlowRun.idis the primary key, so a restart without Redis upserts over the previous boot's run rows.
Wall-panel parity with the current home dashboard
What a fluksio dashboard still lacks to replace geli-dash (Dash/Plotly, e-ink
panel: clock and nav chrome, indoor climate, weather forecast strip, calendar
agenda, room light groups, sliders, power/battery bars, and three pages of
InfluxDB time series). Component-level only; the arrangement and the styling are
this design system's business, not that one's.
Decisions taken up front, because most items below depend on them:
-
Structured data reaches a widget as a declared shape, not as opaque JSON with a path per binding. A path would leave the picker with nothing to offer and
widgetIssueunable to judge a tile from the document alone. -
BUG/UI double check that this aligns with the new data-science pipeline feature
-
A chart asks a flow for its series the way every other input widget speaks: it publishes a request message and reads the answer. No query API, no database knowledge in the widget.
-
Database nodes are transport and credentials only. The InfluxDB node runs the Flux it is handed and echoes back every other field of the request; building the query and shaping the answer are Python nodes either side of it. That is what keeps a widget ignorant of the database, and it is also what a series read mode inside the node would have prevented. A "grouped nodes" concept could later package the standard chart→build→db→parse→chart quintet so a dashboard is not five nodes of wiring each time.
-
Nothing e-ink-specific in the widgets. Panel access is a credential problem (see below); the display's demands are a rendering profile, deferred.
-
CHORE/UI: identical in-flight chart requests are deduplicated per browser tab, so two wall panels showing the same tile still run the query twice. An
intervalon the request port is the backstop, and it belongs to the flow serving the request rather than to the widget asking. -
CHORE/FLOW: one request/answer pair per InfluxDB node — the first input carrying a
fluxkey is the request and the answer leaves on the first output port. A second query stream through one node needs a second node. -
FEAT/UI: the slider offers
stepnow, but no tick labels — thedatalistmarks are unlabelled and drop out past fifty steps. -
FEAT/UI: per-dashboard theme — forced light, forced dark, or switched on a schedule. View mode inherits localStorage and the OS preference today, which a panel in a room has no way to set. NOTE: to solve this, we could introduce a general message sending to the overall dashboard (so far we only treat widgets in a dashboard as a receiver). We could e.g. have a toggle in the dashboard settings which says "propagate theme" which enables a field for defining a consume input (identical to a standard node input) and then a node can connect to this property by producing a corresponding message. This would nicely generalize to other dashboard settings later. This could later also serve as a security mechanism, i.e. the possibility to lock down dashboards remotely
-
CHORE/FLOW: porting the controls needs a declared writable message per control, since an input widget can only target what a flow declares. Declaring one is no longer API-only — the flow panel edits a flow's inputs and the canvas draws each as a label — so the dashboard-input node this asked for has largely been answered by flow inputs. What is left is the naming: an input a panel writes looks the same as one a run passes in.
Deliberately not ported: the local-state/timestamp reconciliation the old dashboard does per widget — publishing on release and reading the value back covers it — and its demo mode, since an unbound or silent message already renders as an em dash.
Dashboard follow-ups
- FEAT/NODE: the hosted demo places six of the fifteen built-in node types (
python,inject,change,join,rbe,trigger); it does cover all fifteen dashboard widget types.switchanddelayare the awkward ones — aswitchbranch needs either a dead-end port or trivial nodes to turn a branch back into a label, and neither reads as something a person would hang — while the I/O types (mqtt,http,influxdb,exec,file,ntfy,mlp) are unplaced because the demo has nothing real to talk to. Worth revisiting when the demo grows a second page. - CHORE/UI:
ROW_HEIGHTis a fixed 80px while column width follows the canvas, so a 1920-wide panel at 12 columns has 160×80 cells. If that reads too wide, the row height could derive from the canvas too. - CHORE/UI: multi-page and multi-section dashboards still have no UI, and now need none — a panel carries several whole dashboards instead, each with its own canvas and its own publish.
PageDef/SectionDefstay in the schema and the editor still editssectionsOf(page)[0], so the pageTabsinDashboardEditorare dead until something writes a second page through the API. - CHORE/API: a panel credential may publish any message, not only the ones its own widgets bind to — the allowlist is the
/messages/prefix rather than a walk of the panel's widgets. The walk now exists:panels.messages_for()is what bounds the live socket. Pointing_panel_mayat it would close this too, but it tightens what already-paired screens may do, so it wants a deliberate look at the query-chart request path first. - CHORE/API: unpairing a device means deleting the panel. A per-panel nonce in the token, bumped on demand, would let one screen be re-paired without disturbing the assignment.
- CHORE/API: a panel paired through the portal is revoked here the moment the panel is deleted —
_panel_mayfinds nothing and answers 401 — but the hub's copy of the token stays valid until it expires or the installation's generation counter is bumped ("New code"). The hub has no per-panel revocation, and giving it one means telling it which panels exist, which is exactly what this design avoids. The generation bump is the lever; it is blunt, cutting every credential the portal minted for the installation. - CHORE/UI: the device line under a pairing code is the raw user agent plus the address the request came from. Both are self-reported and neither is proof; it is there so an admin can tell the screen they just hung from one they were not expecting, not to authenticate anything.
- CHORE/API:
POST /panels/pairis reachable from the internet once an installation is enrolled — the hub forwards it without a session, since a device with no credential is the point of it. Bounded three ways (the hub's per-installation and per-address limits, and the fifty-code cap here), but it is the first unauthenticated surface this installation exposes outward. - CHORE/API:
POST /panels/pairis unauthenticated and capped at fifty pending codes in one process. A second API worker would each keep their own dictionary, so pairing would work only when the poll lands on the process that minted the code. The same holds for a screen pairing through the portal, which lands on whichever worker holds the tunnel. - CHORE/UI: only
layout.lgis ever written, andmd/smstay unwritten by decision — a phone stacks the widgets (.widget-stacked) rather than carrying an arrangement of its own, since arranging is not a phone feature. The keys stay in the schema for a panel that one day wants a second size. - PERF/UI:
ChartWidget's cost per live value is theuPlot.joininUplotChart, not the tail append — the fetched half comes from React Query and is replaced wholesale on every refetch, so a ring buffer over the live tail would leave the dominant cost untouched. If this is ever profiled and fixed, thesetDataeffect must stay dependency-free: a mutable buffer's identity never changes, so keying the effect on it reintroduces the staleness that the point-count dependency used to cause, and more quietly. - CHORE/UI: opening edit mode on a dashboard whose widgets predate placement writes the migrated positions immediately, bumping the version once.
Flow editor follow-ups
- PERF/FLOW: every save rebuilds the whole pipeline. Fine at the current flow count; rebuild only the touched flow when it starts to show.
- CHORE/API:
POST /flows/{name}/renameis no longer reachable from the UI. A flow's title is what the panel edits, matching how nodes work; the canonical name is fixed at creation, so either the endpoint goes or renaming comes back deliberately. - CHORE/UI: ⌘C/⌘V
preventDefaulton the canvas blocks the native clipboard there (fields are guarded). The node clipboard islocalStorage, so it does not cross browsers or profiles. - PERF/UI:
useParamSuggestionsfetches every flow's detail to build the suggestion list. An aggregate endpoint if an installation ever has many flows. - CHORE/UX: free-form params (python nodes) get no suggestions, since there is no schema to key them off.
- PERF/UI:
BrainViewruns 300 force-layout ticks synchronously inside auseMemo, so the graph is laid out on the render thread. - FEAT/UI: the brain is a band on a scrolling page now, so it neither pans nor zooms — the fit keeps the whole graph in view instead. An installation with enough flows to make the labels unreadable at that fit needs a way to open the graph larger.
- CHORE/UI: React Flow measures a node's handle bounds out of the DOM once and never again, and in the brain that one measurement falls inside the graph's
scaleInentrance — so everysourceX/targetXit hands an edge there is the entrance's 4% short of the centre, permanently.BrainEdgetakes both ends from the layout instead (position + radius). Any future view that mounts a canvas inside a transform and reads node internals meets the same thing.
Infrastructure
- CHORE/INFRA:
make soak's redis scenario stops the container the whole stack shares, so every flow briefly fails to journal, not just the soak fixtures. They recover on their own — nothing was dead-lettered or quarantined in the run this note comes from — but it is not a thing to run against a stack someone is relying on. - CHORE/INFRA: the soak harness's engine kill only catches a couple of items unacknowledged, because a cascade finishes in about four milliseconds. Redelivery is proven but barely stressed; a fixture node with a deliberate sleep would widen the window enough to test it properly.
Deferred
Open on purpose. Each names what should bring it back.
- PERF/UI: the app's entry chunk exceeds the warning threshold. React Flow and Monaco are already lazy; a manualChunks split measured no better, so this needs route-level work on the shell rather than chunking config.
- PERF/UI: the Monaco chunk is 2.6 MB. It only loads when a node panel opens, but the editor could be trimmed further or swapped for CodeMirror if that becomes a problem. NOTE: switch to codemirror; loading speed is definitely an issue.
- CHORE/API: node source saves carry no version precondition, so two clients editing the same node's code are last-writer-wins. The flow document is what the optimistic lock protects; code files would need their own, and an exact-match one produces false conflicts against a single client's own interleaved flow and source saves. Revisit with the M5 multi-user work.
- CHORE/FLOW: shared node sources bypass the draft/publish split. Editing one writes the library copy and reloads immediately, since the code is not any single flow's to hold back. Deliberate, but it means a shared node is the one thing publish does not gate.
- CHORE/INFRA:
requires-pythonis capped below 3.14 because the MCP SDK wants a newer starlette there than the pinnedsentry-sdk<2allows. Lift the cap when sentry-sdk moves to 2.x. - FEAT/UI: the node-panel and edge trend curves take no range, unlike the health block. They are drawn from a Redis ring of the last 120 values per message, which has no window to ask for — a hover caption names what the curve covers instead of a picker promising a span nothing can serve. Reopen if per-message history ever gains a time window.
- FEAT/UI: an e-ink rendering profile for a dashboard — motion off, hover-only affordances resolved to something visible, high-contrast palette, thick strokes, and a repaint cadence low enough for a display that takes a second to settle. Reopen when a panel with such a display is actually hung.
- CHORE/INFRA: Postgres stays. The 2026-08 review rejected YugabyteDB/CockroachDB (multi-node cluster systems, ~4 GB+ RAM per node, against the small-server target — the scaling story is remote workers, not a distributed DB) and found merging Postgres into Redis or vice versa buys little: the stores hold disjoint data and both sit behind abstractions. SQLite would fit the single-instance design and drop a container; reopen if the home-install footprint becomes a product concern.
- CHORE/INFRA: NATS JetStream as the work-queue backend — durable streams whose consumer semantics match the
WorkQueueinterface, in one small binary. Reopen with M5 remote workers, when the queue crosses hosts. NOTE: remote workers landed without it — a worker dials the engine's own socket and never touches Redis, so the queue still does not cross a host. Reopen if a second engine ever pulls from the same stream. - FEAT/RUNS: stage caching.
run_node.cache_keyis written on every run and the artifact store is content-addressed, so the pieces are in place; what is missing is computing the key from the node's source digest plus its input values and skipping a node whose key already has anokrow with its artifacts still present. The two research repos want this more than they want resume — neither persists checkpoints, and both re-run unchanged preprocessing every time. - FEAT/RUNS: per-label requirements overlays (
requirements-gpu.txt) synced into a remote worker's venv, with drift surfaced against the engine's manifest. Today a worker's environment is whatever--pythonpoints at, which is fine for one hand-managed GPU box and not for several.venv_digestalready arrives at attach and is shown on/workers, so the reporting half exists. - FEAT/UI: a dashboard shows a run's curve only while it is running. Emissions reach the socket live, but a run's values live in its own state namespace, so reloading the panel afterwards leaves the chart empty — the durable series is on the run (
/runs/{id}/metrics) and nothing binds a widget to it. A chart variant that reads a run's series, or the existing querying chart pointed at/runs/series/compare, is what would close it. This is also what a demo needs to show a finished experiment rather than only a live one. - FEAT/UI: launching a sweep is API-only. Pressing Run on a batch flow asks for its parameters, but the many-runs-at-once shape has no UI; a run detail screen is what it wants to land next to.
- FEAT/RUNS: a run detail screen. The API answers everything — params, per-node status with logs and tracebacks, artifacts, metrics, and
/runs/series/comparein the chart widget's ownseriesshape — but nothing in the dashboard reads it yet, so a run is inspected over HTTP. Comparing curves is a widget binding once someone builds the page around it. - FEAT/RUNS: a thin client CLI (
fluksio run/runs/sweep/worker) over the same API. The engine being resident is what makes runs cheap; a CLI is ergonomics on top, andcurlcovers it until someone is running sweeps daily. - FEAT/RUNS: the step on a run's series is the count of emissions on that message, so a node yielding every tenth training step records steps 0, 1, 2 rather than 0, 10, 20 — a faithful x-axis of its own emissions, not of the loop inside it. If a real step number ever matters, a
record-typed streaming port carrying its ownstepis the shape to read it from; the column is already there. - CHORE/RUNS: an emission publishes on the node's port and, in a live flow, enqueues a cascade with no payload of its own — the value is already in state, and an item carrying it would re-apply that value whenever it was claimed, which is how a mid-node emission overwrites the one the node returned at the end. Downstream therefore reads what is current rather than the value that caused it to run. Right for a curve; worth revisiting if something ever needs every intermediate value delivered rather than sampled.
- CHORE/RUNS:
run_metrichas no retention. Deliberately outsideOBS_RETENTION_DAYS— an experiment nobody deleted should not vanish on a rollup window — but a few thousand runs at 3000 steps will want a policy eventually, probably per-flow rather than global. - CHORE/RUNS: a run holds one worker slot per node for its whole duration, and
MAX_PARALLELrun drivers bound how many graphs are in flight. A sweep of 500 therefore queues behind the pool rather than the driver count. Fine — the GPU is the scarce thing — but the two limits are unrelated numbers that read as if they were one.
Blocked
-
CHORE/INFRA:
bun installinside the frontend Docker build intermittently fails with "Fail extracting tarball" for several packages at once, and succeeds on a plain rebuild. It looks like concurrent extraction under memory pressure. Pin down or retry in the Dockerfile if it starts costing CI time. NOTE: memory lifted; retry and close if stale -
CHORE/DOCS:
docs/(the published site) documentsdeviceanddevice_policyas API-only, because the node panel has no field for either. That is the one place the public docs have to say "use the API instead of the UI". Adding a Device section toNodePanel.tsx— a label field plus a require/prefer toggle — would close it. -
CHORE/DOCS: the site's node-type reference is hand-written from
NODE_TYPESand each node'sParams. It will drift.GET /flows/node-typesalready returns the whole thing with its schemas, so a generator (the way n3xd generates its command catalog) is the obvious fix once the type list stops moving. -
make update/make up(production overlay) runsup -d --remove-orphansunder the samefluksio-appproject asmake dev, so it silently deletes the local Traefik containerfluksio-app-proxy-1. WithREVERSE_PROXY=externaland no host proxy, the whole stack then has nothing bound to :80 and*.localhoststops resolving. Either drop--remove-orphansor makeupdaterefuse whenENVIRONMENT=local.