Conversation
Fits a session's converter sequence once at creation time (per fold for Cross-Validation, once for Holdout) instead of applying converters on the full dataset before any split exists, which leaked validation/test data into the fit. Resolves the session's input columns from the fitted sequence and persists preprocessing status/errors on the session.
…nation Training reuses the SessionPreprocessor already fitted by PreprocessingJob instead of ever re-fitting on new data. Prediction and explanation apply the final fitted preprocessor to raw/manual input before running. Run creation is blocked with a clear error while a session's preprocessing is still pending or failed.
…ut group BagOfWordsConverter (and any converter with the same shape, e.g. Binarizer) keeps its scope column unchanged and appends new derived columns instead of replacing it. SessionPreprocessor was recording every column the converter's transform() returned as the step's output group, so the untouched original column leaked into it too — selecting only "this converter's output" as a session's input still smuggled in the raw column, which fails task validation when it isn't an allowed input type (e.g. Text).
output_type already reported the semantic type name (e.g. "Integer") from a default-constructed instance; output_dtype adds the concrete storage dtype (e.g. "int64") the same way, so a converter's not-yet-materialized output group can show a real dtype instead of falling back to unknown.
…nverter picker A Models-module session can now optionally configure a sequence of converters as part of creation. The wizard reuses Notebooks' own tool picker (search, list/grid, drag-and-drop) and column selector so the scope/output-group UX matches exactly, and represents a converter's not-yet-materialized output group as a synthetic column key so it can be picked as an input before any real fit exists (and chained into a later converter's own scope).
…echanism Uses useJobTracker (the same mechanism every other job-backed indicator in the app already relies on: RunnerDialog, ComponentDownloadControl, prediction/explainer panels) instead of an independent timer polling preprocessing_status, so the session's processing/failed views never drift out of sync with what the Job Queue widget shows for the same job.
handleCreateRun always showed the same generic message regardless of cause, so a session blocked by the preprocessing-not-ready guard (or any other backend rejection) never told the user why. Uses the project's existing getApiErrorMessage helper (already used by RAG) to show the backend's actual detail, falling back to the generic message only when there isn't one.
… column types - Introduced `_classify_by_type` method in `SessionPreprocessor` to categorize output columns by their types. - Updated the fitting process to resolve output slots for each step, allowing converters to handle mixed types. - Modified `AppliedConvertersView` to display output entries as chips for each declared slot. - Adjusted tests to cover new functionality, including validation of slotted group references and output slot resolution. - Enhanced `resolveDeclaredOutputSlots` to determine output types based on the converter's scope and type preservation. - Updated frontend components to accommodate changes in output structure and ensure proper rendering of converter outputs.
…n references for manual predictions
…ups in SelectColumnsStep
…nce SessionInfoContent with converter metadata
… messages in multiple languages
…videDatasetColumns
…put fields on mount
…ed processing capabilities
…new structure mixins and utility functions
…imation - Added infer_output_columns method to ImageEmbeddingConverter to estimate output structure based on input StateItems. - Implemented infer_output_columns in SAM3SegmentConverter to provide output structure based on segmentation masks. - Updated CharacterReplacer to preserve input types and added infer_output_columns method for output estimation. - Enhanced ColumnArithmetic with infer_output_columns to estimate output based on selected columns and operations. - Introduced infer_output_columns in ColumnConcat to estimate output structure based on concatenated columns. - Added infer_output_columns to ColumnRemover to indicate no columns are kept after removal. - Implemented infer_output_columns in DateFeaturesConverter to estimate output based on date features. - Enhanced NumericExpansion with infer_output_columns to estimate output structure based on numeric operations. - Added infer_output_columns in TypeCast to provide warnings for potential casting issues. - Introduced tests for structure estimation in converters that download pretrained models. - Updated test_structure_matches_runtime to include new converters and ensure output structure matches runtime behavior.
…g with target columns
…tion - Updated FormSessionConverterSection to utilize final state and step display names for converter scoping. - Enhanced PreprocessingStep to integrate preprocessing structure estimation and error handling. - Modified ScopeStepSessionConverter to derive column options from the estimated dataset state. - Revamped SessionConvertersRightBar to align with new structure handling and converter validation. - Removed obsolete resolveDeclaredOutputSlots function and related tests. - Introduced structureMessages for user-friendly error messages based on backend responses. - Added translations for new structure-related messages in multiple languages.
… FormSessionConverterSection
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Redesigns the preprocessing step of the model session wizard so users build converter chains knowing, at every step, which columns exist and what type they have.
Before this change the wizard asked for converters first and inputs/output last, and the frontend guessed each converter's output type on its own. It never removed columns a converter consumed (after PCA on
age, weight, both were still offered) nor updated columns replaced in place, so a chain could look valid and only fail when the session trained.Now:
COLUMN_OPERATION:replace,add,expand,select,rows) and can describe its exact output withinfer_output_columns.infer_structure, estimates the dataset state after every step without fitting anything. The wizard calls it on every change; session creation uses the same function and rejects invalid chains with a 422 before creating anything.Type of Change
Check all that apply like this [x]:
Changes (by file)
Structure estimation (backend)
DashAI/back/preprocessing/structure_types.py(new): state items (ColumnItem,BlockItem),StructureDelta,StepStructure,StructureResultand thetype_fieldshelper.DashAI/back/preprocessing/structure.py(new):infer_structure. Walks the chain resolving each scope against the current state, checks types and cardinality like the frontend does, applies each converter's delta with the same column renaming as the runtime, and marks steps after an invalid one as blocked. Alsoresolve_state_refsandStructureError.DashAI/back/converters/base_converter.py:COLUMN_OPERATIONattribute, defaultinfer_output_columnsper operation, andcolumn_operationin the metadata.DashAI/back/converters/structure_mixins.py(new):ComponentsOutputMixin, whose output size followsn_components(PCA family and kernel approximations).DashAI/back/converters/category/*.py: each category declares itsCOLUMN_OPERATION.DashAI/back/converters/**(sklearn, simple, Hugging Face and SAM3 converters): exactinfer_output_columnsoverrides where the output follows from params, and operation overrides where a converter differs from its category.Deterministic output types
A converter's output type now follows from its params and input types, never from the values, so the estimate always matches the real output:
scikit_learn/simple_imputer.py:mean/medianalways produce Float.simple_converters/column_arithmetic.py,numeric_expansion.py: integer results stay Integer even with nulls (computed withpyarrow.compute).simple_converters/character_replacer.py: keeps its input type instead of turning all-digit text into Integer, a decision that was made per batch (also affects notebooks).simple_converters/type_cast.py:on_error="skip"adds acast_may_skipwarning.Runtime fixes
DashAI/back/preprocessing/session_preprocessor.py: supervised converters now receive the output column asy(they previously got none and failed); a new column renamed on a name clash is recorded under its final name.DashAI/back/preprocessing/column_ref.py:GroupColumnRefgains an optionalnameto reference one generated column whose name is known before fit (e.g.date_month). No migration needed.DashAI/back/converters/dataset_columns.py: the clash renaming moves toplan_new_column_names, shared by the runtime and the estimate.DashAI/back/job/preprocessing_job.py: passes the output columns to the preprocessor.API
DashAI/back/api/api_v1/endpoints/model_sessions.py:POST /model-session/preprocessing/structure(read only)./validationtypes every input ref from the estimated structure; the old frontend computedconverter_output_typesis removed.DashAI/back/api/api_v1/schemas/model_sessions_params.py:PreprocessingStructureParams;ColumnsValidationParamstakespreprocessing.Frontend
components/models/CreateSessionSteps.jsx: with preprocessing on, the order is prepare, base columns, preprocessing, final inputs. Without it the wizard is unchanged. Preprocessing steps are no longer sent if the switch was turned off after adding them.components/models/modelSession/BaseColumnsStep.jsx(new): output and candidate columns.components/models/modelSession/usePreprocessingStructure.js(new): debounced call to the structure endpoint that discards stale responses.components/models/modelSession/PreprocessingStep.jsx,SessionConvertersRightBar.jsx,ScopeStepSessionConverter.jsx,FormSessionConverterSection.jsx: the catalog and the scope picker only offer what exists at the end of the chain; row-changing converters are hidden.components/models/modelSession/AppliedConvertersView.jsx:components/models/modelSession/SelectColumnsStep.jsx: final inputs come from the estimated final state (all selected by default); the output is fixed.components/models/modelSession/sessionColumnRefs.js: keys for named refs,itemToRef,stateToOptions,labelForRef; the frontend type inference (buildColumnKeysAndTypes,resolveDeclaredOutputSlots) is removed.components/models/modelSession/structureMessages.js(new): turns backend structure messages into translated text.components/models/modelSession/DivideDatasetColumns.jsx: optionalinputLabelandoutputDisabled.components/models/SessionInfoContent.jsx: scope labels vialabelForRef.components/notebooks/converterCreation/ParameterStepConverter.jsx: the data leakage warning only shows in notebooks, since session preprocessing already fits on the training split.api/modelSession.ts,types/modelSession.ts: structure endpoint client and types.utils/i18n/locales/*/models.json: new strings in the 5 languages.Tests
tests/back/preprocessing/test_structure_matches_runtime.py(new): for every registered converter, compares the estimated structure with a real fit and transform; also fails if a new converter has no case.tests/back/preprocessing/test_structure.py(new),tests/back/converters/test_infer_output_columns_defaults.py(new),tests/back/converters/test_download_converters_structure.py(new).tests/back/preprocessing/test_session_preprocessor.py,test_column_ref.py,tests/back/api/test_model_session_api.py,test_run_preprocessing_guard.py: updated and extended.sessionColumnRefs.test.js,AppliedConvertersView.test.jsx,usePreprocessingStructure.test.js(new).Testing (optional)
n_components=2over two numeric columns: both disappear from the "Dataset after preprocessing" panel and are no longer offered to later converters.