Skip to content

Feat session preprocessing structure - #911

Open
Creylay wants to merge 41 commits into
developfrom
feat/session-preprocessing-structure
Open

Creylay wants to merge 41 commits into
developfrom
feat/session-preprocessing-structure

Conversation

@Creylay

@Creylay Creylay commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Summary

Redesigns the preprocessing step of the model session wizard so users build converter chains knowing, at every step, which columns exist and what type they have.

Before this change the wizard asked for converters first and inputs/output last, and the frontend guessed each converter's output type on its own. It never removed columns a converter consumed (after PCA on age, weight, both were still offered) nor updated columns replaced in place, so a chain could look valid and only fail when the session trained.

Now:

  • The user picks the output column and the candidate columns first, then builds the chain, then picks the model inputs from what the chain leaves. The output column is never part of any converter scope.
  • Every converter declares what it does to its columns (COLUMN_OPERATION: replace, add, expand, select, rows) and can describe its exact output with infer_output_columns.
  • A new backend function, infer_structure, estimates the dataset state after every step without fitting anything. The wizard calls it on every change; session creation uses the same function and rejects invalid chains with a 422 before creating anything.
  • Output columns whose names are only known after fit (e.g. one-hot columns) are shown as symbolic blocks with their size, or "N" when unknown.

Type of Change

Check all that apply like this [x]:

  • Backend change
  • Frontend change
  • CI / Workflow change
  • Build / Packaging change
  • Bug fix
  • Documentation

Changes (by file)

Structure estimation (backend)

  • DashAI/back/preprocessing/structure_types.py (new): state items (ColumnItem, BlockItem), StructureDelta, StepStructure, StructureResult and the type_fields helper.
  • DashAI/back/preprocessing/structure.py (new): infer_structure. Walks the chain resolving each scope against the current state, checks types and cardinality like the frontend does, applies each converter's delta with the same column renaming as the runtime, and marks steps after an invalid one as blocked. Also resolve_state_refs and StructureError.
  • DashAI/back/converters/base_converter.py: COLUMN_OPERATION attribute, default infer_output_columns per operation, and column_operation in the metadata.
  • DashAI/back/converters/structure_mixins.py (new): ComponentsOutputMixin, whose output size follows n_components (PCA family and kernel approximations).
  • DashAI/back/converters/category/*.py: each category declares its COLUMN_OPERATION.
  • DashAI/back/converters/** (sklearn, simple, Hugging Face and SAM3 converters): exact infer_output_columns overrides where the output follows from params, and operation overrides where a converter differs from its category.

Deterministic output types

A converter's output type now follows from its params and input types, never from the values, so the estimate always matches the real output:

  • scikit_learn/simple_imputer.py: mean/median always produce Float.
  • simple_converters/column_arithmetic.py, numeric_expansion.py: integer results stay Integer even with nulls (computed with pyarrow.compute).
  • simple_converters/character_replacer.py: keeps its input type instead of turning all-digit text into Integer, a decision that was made per batch (also affects notebooks).
  • simple_converters/type_cast.py: on_error="skip" adds a cast_may_skip warning.

Runtime fixes

  • DashAI/back/preprocessing/session_preprocessor.py: supervised converters now receive the output column as y (they previously got none and failed); a new column renamed on a name clash is recorded under its final name.
  • DashAI/back/preprocessing/column_ref.py: GroupColumnRef gains an optional name to reference one generated column whose name is known before fit (e.g. date_month). No migration needed.
  • DashAI/back/converters/dataset_columns.py: the clash renaming moves to plan_new_column_names, shared by the runtime and the estimate.
  • DashAI/back/job/preprocessing_job.py: passes the output columns to the preprocessor.

API

  • DashAI/back/api/api_v1/endpoints/model_sessions.py:
    • New POST /model-session/preprocessing/structure (read only).
    • Session creation validates the chain and the input refs and returns a 422 with the details.
    • /validation types every input ref from the estimated structure; the old frontend computed converter_output_types is removed.
  • DashAI/back/api/api_v1/schemas/model_sessions_params.py: PreprocessingStructureParams; ColumnsValidationParams takes preprocessing.

Frontend

  • components/models/CreateSessionSteps.jsx: with preprocessing on, the order is prepare, base columns, preprocessing, final inputs. Without it the wizard is unchanged. Preprocessing steps are no longer sent if the switch was turned off after adding them.
  • components/models/modelSession/BaseColumnsStep.jsx (new): output and candidate columns.
  • components/models/modelSession/usePreprocessingStructure.js (new): debounced call to the structure endpoint that discards stale responses.
  • components/models/modelSession/PreprocessingStep.jsx, SessionConvertersRightBar.jsx, ScopeStepSessionConverter.jsx, FormSessionConverterSection.jsx: the catalog and the scope picker only offer what exists at the end of the chain; row-changing converters are hidden.
  • components/models/modelSession/AppliedConvertersView.jsx:
    • Each card shows what its step produces.
    • Invalid steps are outlined in red with a translated reason, and later steps are dimmed.
    • A new "Dataset after preprocessing" panel lists the final state.
  • components/models/modelSession/SelectColumnsStep.jsx: final inputs come from the estimated final state (all selected by default); the output is fixed.
  • components/models/modelSession/sessionColumnRefs.js: keys for named refs, itemToRef, stateToOptions, labelForRef; the frontend type inference (buildColumnKeysAndTypes, resolveDeclaredOutputSlots) is removed.
  • components/models/modelSession/structureMessages.js (new): turns backend structure messages into translated text.
  • components/models/modelSession/DivideDatasetColumns.jsx: optional inputLabel and outputDisabled.
  • components/models/SessionInfoContent.jsx: scope labels via labelForRef.
  • components/notebooks/converterCreation/ParameterStepConverter.jsx: the data leakage warning only shows in notebooks, since session preprocessing already fits on the training split.
  • api/modelSession.ts, types/modelSession.ts: structure endpoint client and types.
  • utils/i18n/locales/*/models.json: new strings in the 5 languages.

Tests

  • tests/back/preprocessing/test_structure_matches_runtime.py (new): for every registered converter, compares the estimated structure with a real fit and transform; also fails if a new converter has no case.
  • tests/back/preprocessing/test_structure.py (new), tests/back/converters/test_infer_output_columns_defaults.py (new), tests/back/converters/test_download_converters_structure.py (new).
  • tests/back/preprocessing/test_session_preprocessor.py, test_column_ref.py, tests/back/api/test_model_session_api.py, test_run_preprocessing_guard.py: updated and extended.
  • Frontend: sessionColumnRefs.test.js, AppliedConvertersView.test.jsx, usePreprocessingStructure.test.js (new).

Testing (optional)

  1. Create a model session with "Apply preprocessing (advanced)" on and pick the output and candidate columns.
  2. Add PCA with n_components=2 over two numeric columns: both disappear from the "Dataset after preprocessing" panel and are no longer offered to later converters.
  3. Create a session.

Fits a session's converter sequence once at creation time (per fold for
Cross-Validation, once for Holdout) instead of applying converters on the
full dataset before any split exists, which leaked validation/test data
into the fit. Resolves the session's input columns from the fitted
sequence and persists preprocessing status/errors on the session.
…nation

Training reuses the SessionPreprocessor already fitted by PreprocessingJob
instead of ever re-fitting on new data. Prediction and explanation apply
the final fitted preprocessor to raw/manual input before running. Run
creation is blocked with a clear error while a session's preprocessing is
still pending or failed.
…ut group

BagOfWordsConverter (and any converter with the same shape, e.g. Binarizer)
keeps its scope column unchanged and appends new derived columns instead
of replacing it. SessionPreprocessor was recording every column the
converter's transform() returned as the step's output group, so the
untouched original column leaked into it too — selecting only "this
converter's output" as a session's input still smuggled in the raw column,
which fails task validation when it isn't an allowed input type (e.g. Text).
output_type already reported the semantic type name (e.g. "Integer") from
a default-constructed instance; output_dtype adds the concrete storage
dtype (e.g. "int64") the same way, so a converter's not-yet-materialized
output group can show a real dtype instead of falling back to unknown.
…nverter picker

A Models-module session can now optionally configure a sequence of
converters as part of creation. The wizard reuses Notebooks' own tool
picker (search, list/grid, drag-and-drop) and column selector so the
scope/output-group UX matches exactly, and represents a converter's
not-yet-materialized output group as a synthetic column key so it can be
picked as an input before any real fit exists (and chained into a later
converter's own scope).
…echanism

Uses useJobTracker (the same mechanism every other job-backed indicator in
the app already relies on: RunnerDialog, ComponentDownloadControl,
prediction/explainer panels) instead of an independent timer polling
preprocessing_status, so the session's processing/failed views never drift
out of sync with what the Job Queue widget shows for the same job.
handleCreateRun always showed the same generic message regardless of
cause, so a session blocked by the preprocessing-not-ready guard (or any
other backend rejection) never told the user why. Uses the project's
existing getApiErrorMessage helper (already used by RAG) to show the
backend's actual detail, falling back to the generic message only when
there isn't one.
… column types

- Introduced `_classify_by_type` method in `SessionPreprocessor` to categorize output columns by their types.
- Updated the fitting process to resolve output slots for each step, allowing converters to handle mixed types.
- Modified `AppliedConvertersView` to display output entries as chips for each declared slot.
- Adjusted tests to cover new functionality, including validation of slotted group references and output slot resolution.
- Enhanced `resolveDeclaredOutputSlots` to determine output types based on the converter's scope and type preservation.
- Updated frontend components to accommodate changes in output structure and ensure proper rendering of converter outputs.
…nce SessionInfoContent with converter metadata
…imation

- Added infer_output_columns method to ImageEmbeddingConverter to estimate output structure based on input StateItems.
- Implemented infer_output_columns in SAM3SegmentConverter to provide output structure based on segmentation masks.
- Updated CharacterReplacer to preserve input types and added infer_output_columns method for output estimation.
- Enhanced ColumnArithmetic with infer_output_columns to estimate output based on selected columns and operations.
- Introduced infer_output_columns in ColumnConcat to estimate output structure based on concatenated columns.
- Added infer_output_columns to ColumnRemover to indicate no columns are kept after removal.
- Implemented infer_output_columns in DateFeaturesConverter to estimate output based on date features.
- Enhanced NumericExpansion with infer_output_columns to estimate output structure based on numeric operations.
- Added infer_output_columns in TypeCast to provide warnings for potential casting issues.
- Introduced tests for structure estimation in converters that download pretrained models.
- Updated test_structure_matches_runtime to include new converters and ensure output structure matches runtime behavior.
…tion

- Updated FormSessionConverterSection to utilize final state and step display names for converter scoping.
- Enhanced PreprocessingStep to integrate preprocessing structure estimation and error handling.
- Modified ScopeStepSessionConverter to derive column options from the estimated dataset state.
- Revamped SessionConvertersRightBar to align with new structure handling and converter validation.
- Removed obsolete resolveDeclaredOutputSlots function and related tests.
- Introduced structureMessages for user-friendly error messages based on backend responses.
- Added translations for new structure-related messages in multiple languages.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant