Skip to content

feat: redesign the visual pipelines module as a validated DAG (atomized nodes, typed contracts) - #910

Open
pip3alfar0 wants to merge 97 commits into
developfrom
dag-pipelines
Open

pip3alfar0 wants to merge 97 commits into
developfrom
dag-pipelines

Conversation

@pip3alfar0

Copy link
Copy Markdown
Collaborator

Summary

Redesigns the visual pipelines module: from a strictly linear flow with a monolithic training node and frontend-side validation, to an expressive DAG-based system with atomized nodes, typed contracts, and backend validation. This is the full body of work developed on dag-pipelines (88 commits) and is intended as the integration of the module into develop.


Motivation

The previous pipelines module had three limitations:

  • Flows were strictly linear (no branching or merging).
  • Training was a single monolithic node (data split + task/model + metrics).
  • Structural/compatibility validation lived in the frontend.

What's included

1. DAG-based flows

  • Pipelines are now a directed acyclic graph with branching (fan-out) and merging (fan-in). This enables model comparison within a single flow, either several models on the same data or one model over several datasets.
  • Multiple typed output handles per node.

2. Atomized nodes

  • The monolithic Train Model node is split into Split Data, Task and Model, and Metrics.
  • Legacy Train Model is removed from both the backend and the node palette.
  • Retrieve Model is adapted to load models trained via Task and Model.

3. Typed contracts + backend validation

  • New node contract system (back/pipeline/contracts.py): each node declares typed input/output ports and their cardinality.
  • Node compatibility is derived from port types. Acyclicity is guaranteed by construction: the derivation excludes same-type pairs, and the remaining port types only progress forward, so no valid edge can close a cycle.
  • PipelineValidator validates graph structure on the backend (moved off the frontend) and returns per-node/edge diagnostics.

4. Execution & orchestration

  • New data model that separates the pipeline definition from its executions (PipelineRun / NodeRun).
  • PipelineJob is now a per-node orchestrator. It executes nodes in topological order and runs independent branches concurrently, with training serialized by a lock.
  • Live node coloring on the canvas during execution. Status is exposed via the API and polled by the frontend.
  • Concurrency fix: the orchestrator owns NodeRun tracking in a single serialized session. This eliminates the SQLite write-lock contention that previously dropped node states and inflated run time.

5. Frontend (React Flow)

  • New pipeline editor with a node palette, a right-side configuration panel, history and templates sidebars, and a branched-flow canvas.
  • Results view with per-branch comparison metrics and outputs from all nodes.

Breaking changes / compatibility

  • Train Model node removed. The atomization migration adds new nullable columns to the pipeline table and performs no destructive data migration. Pipelines authored with the old linear node are not auto-converted; as the redesigned module was not yet exposed in the UI, no existing user pipelines need migrating.
  • New Alembic migrations for pipeline-run tracking and node atomization, plus merge migrations that unify heads with develop. alembic heads resolves to a single head, and the full backend suite (which upgrades a fresh SQLite DB to head) passes.

How to try it

  1. Upload any tabular dataset.
  2. Open /app/pipelines (hidden from navigation, see below).
  3. Load a template from the Templates sidebar (e.g. Train Model), or build Data Selector → Split Data → 2× Task and Model → Metrics.
  4. Run the pipeline and watch the live node coloring, then compare branches in the results view.

Testing

  • The full backend suite passes (3030 passed).
  • No new automated backend tests were added for the new pipeline modules (contracts, validator, orchestrator); they are currently covered by manual verification and are a natural follow-up.
  • The frontend has no automated tests; UI behavior was verified manually.
  • Manually verified: building and running branched flows, per-branch comparison, rejection of invalid flows, and live node coloring.
  • This version was also evaluated in a usability test with 12 participants (93% task success).

Exposure / integration note

  • The /app/pipelines route is fully functional but intentionally hidden from the home page and the top navigation. It is reachable by URL. To enable it, revert commit 19eba5cd ("chore: remove pipelines from home and navbar"), which removed the card in DashAI/front/src/pages/home/Home.jsx and the nav entry in DashAI/front/src/components/ResponsiveAppBar.jsx.

Notes for reviewers

  • tests/back/migrations/test_add_preprocessing_to_model_session.py: I changed upgrade("head") → upgrade("f4a91c62d8e7") so that the relative downgrade("-1") stays unambiguous. This test (new in develop) downgrades one step from head. Once any branch is merged in, head becomes a merge revision with two parents and -1 is ambiguous. This affects any branch merged after that migration, not just this one. Targeting the revision tests exactly that migration's up/down behavior.
  • DataGrid: the pipeline history grid is migrated to the @mui/x-data-grid v8 selection model ({ type, ids: Set }).
  • Ruff / Prettier formatting is applied per the repo's pinned config.

feat: enhance PipelineJob with logging and error handling in run method
…e enqueuing job and add validation error handling in TrainNode
…in DataSelectorNode and PipelineResults components
pip3alfar0 and others added 30 commits June 26, 2026 17:24
Add Data Selector, Split Data and Retrieve Model sections to the pipeline
results, alongside the existing Exploration/Train/Prediction:

- Data Selector: dataset name, row/column counts, task and a column
  preview (reuses DatasetSummaryTable / the dataset sample endpoint).
- Split Data: input/output columns and per-partition ratios and sizes.
- Retrieve Model: model, task, input columns and model path.

Partial results now render progressively with a "running" indicator.
All data is derived from what getPipelineById already returns (steps
config and split_data); no backend or schema changes required.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant