feat: redesign the visual pipelines module as a validated DAG (atomized nodes, typed contracts) - #910
Open
pip3alfar0 wants to merge 97 commits into
Open
pip3alfar0 wants to merge 97 commits into
pip3alfar0 wants to merge 97 commits into
Conversation
feat: enhance PipelineJob with logging and error handling in run method
…assification and regression data types
… output handling in ModelFactory
…e enqueuing job and add validation error handling in TrainNode
…epth controlnet model
…_depth_controlnet.py
…to dag-pipelines
…d load() methods in DistilBertTransformer
…r for load() method
…hyperopt_optimizer.py
…in DataSelectorNode and PipelineResults components
…ed metrics handling and display
…peline creation and update
…ults, default name)
Add Data Selector, Split Data and Retrieve Model sections to the pipeline results, alongside the existing Exploration/Train/Prediction: - Data Selector: dataset name, row/column counts, task and a column preview (reuses DatasetSummaryTable / the dataset sample endpoint). - Split Data: input/output columns and per-partition ratios and sizes. - Retrieve Model: model, task, input columns and model path. Partial results now render progressively with a "running" indicator. All data is derived from what getPipelineById already returns (steps config and split_data); no backend or schema changes required. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Redesigns the visual pipelines module: from a strictly linear flow with a monolithic training node and frontend-side validation, to an expressive DAG-based system with atomized nodes, typed contracts, and backend validation. This is the full body of work developed on
dag-pipelines(88 commits) and is intended as the integration of the module intodevelop.Motivation
The previous pipelines module had three limitations:
What's included
1. DAG-based flows
2. Atomized nodes
Train Modelnode is split intoSplit Data,Task and Model, andMetrics.Train Modelis removed from both the backend and the node palette.Retrieve Modelis adapted to load models trained viaTask and Model.3. Typed contracts + backend validation
back/pipeline/contracts.py): each node declares typed input/output ports and their cardinality.PipelineValidatorvalidates graph structure on the backend (moved off the frontend) and returns per-node/edge diagnostics.4. Execution & orchestration
PipelineRun/NodeRun).PipelineJobis now a per-node orchestrator. It executes nodes in topological order and runs independent branches concurrently, with training serialized by a lock.NodeRuntracking in a single serialized session. This eliminates the SQLite write-lock contention that previously dropped node states and inflated run time.5. Frontend (React Flow)
Breaking changes / compatibility
Train Modelnode removed. The atomization migration adds new nullable columns to thepipelinetable and performs no destructive data migration. Pipelines authored with the old linear node are not auto-converted; as the redesigned module was not yet exposed in the UI, no existing user pipelines need migrating.develop.alembic headsresolves to a single head, and the full backend suite (which upgrades a fresh SQLite DB tohead) passes.How to try it
/app/pipelines(hidden from navigation, see below).Data Selector → Split Data → 2× Task and Model → Metrics.Testing
Exposure / integration note
/app/pipelinesroute is fully functional but intentionally hidden from the home page and the top navigation. It is reachable by URL. To enable it, revert commit19eba5cd("chore: remove pipelines from home and navbar"), which removed the card inDashAI/front/src/pages/home/Home.jsxand the nav entry inDashAI/front/src/components/ResponsiveAppBar.jsx.Notes for reviewers
tests/back/migrations/test_add_preprocessing_to_model_session.py: I changedupgrade("head")→upgrade("f4a91c62d8e7")so that the relativedowngrade("-1")stays unambiguous. This test (new indevelop) downgrades one step fromhead. Once any branch is merged in,headbecomes a merge revision with two parents and-1is ambiguous. This affects any branch merged after that migration, not just this one. Targeting the revision tests exactly that migration's up/down behavior.@mui/x-data-gridv8 selection model ({ type, ids: Set }).