Skip to content

One implementation of model training: composable units, the engine that runs them, and the cross-validation reconciliation - #875

Open
Felipedino wants to merge 50 commits into
developfrom
reconcile/cv-units
Open

Felipedino wants to merge 50 commits into
developfrom
reconcile/cv-units

Conversation

@Felipedino

Copy link
Copy Markdown
Collaborator

What this brings

Three bodies of work, together because the third depends on the first two:

  1. The jobs, atomized into composable units (DashAI/back/units/).
  2. The DAG engine that runs them (DashAI/back/dag/).
  3. The reconciliation with cross-validation, which is what is new on this
    branch.

It can be reviewed in two parts instead: merge the units and the engine first,
then rebase the reconciliation on top. Say the word and I will split it.

Why

ModelJob.run() was 400 lines. Atomizing it cut it into units that declare what
they read and what they produce, and that work the same inside a job or as a
node of a graph.

While that was happening, develop gained cross-validation — and it did not
add a branch to the monolith. It replaced it with a second layer of
abstraction
: BaseSplitter (10 splitters) and BaseEvaluationStrategy (4
strategies). The result was two complete implementations of model training,
each passing its own tests, and each able to give a different answer for the
same model.

This PR leaves one.

Where it landed

ModelJob composes units on all three paths — holdout, cross-validation and
nested CV — and the strategy classes are left only declaring how a run is
carved and which partitions it records.

LoadDatasetUnit
  → PrepareAndSplitUnit | PrepareAndFoldUnit      (chosen by the splitter)
  → BuildModelUnit
  → FitModelUnit | FitModelOverFoldsUnit | FitModelOverNestedFoldsUnit
  → EvaluateModelUnit
  → SaveModelUnit
DashAI/back/evaluation/ 881 → 153 lines
Registered units 33
Tests 3692 passing

The strategies stay registered even though they are emptied

That is not deference to dead code. The frontend reads these classes in four
places, and only one of them is about metrics:

  • the session wizard lists them so the user can pick one, and
    ModelSession.evaluation_strategy is NOT NULL — without that listing a
    session cannot be created at all
    ;
  • it starts on the first one whose kind is holdout;
  • kind decides the shape of the splits payload and which controls are shown;
  • scored_splits tells the metric charts which partitions exist to plot.

There is precedent in this codebase for a class that declares and does not
execute: BaseSplitter.PARTITIONING and explainable_partitions are read the
same way.

Where to start reviewing

Suggested order, most to least load-bearing:

  1. DashAI/back/models/base_model.py — the sharpest edge, below.
  2. DashAI/back/job/model_job.py — the three paths composed.
  3. DashAI/back/units/fit_scope.py and the three fitting units.
  4. tests/back/api/test_model_job_cross_validation.py — the net that pins
    the observable behaviour.
  5. The rest of units/ and dag/, which is the earlier atomization.

The conflict git does not mark

base_model.py auto-merges without a conflict and produces a class with
two compute_metrics: both branches had added a method under that name, in
different parts of the file, with different semantics.

atomization CV
default split VALIDATION TEST
nothing to score None {}
non-finite scores dropped kept

Python keeps the last one. The only thing that noticed was ruff F811, by
accident — no test did, because the CV path barely had any.

Resolved by hand: the CV copy is deleted, the extracted one is kept with its
filter for non-finite values
, and fold_index, inner_fold_index and
_epoch_reporter are kept. A test parses BaseModel and asserts a single
definition, so the next merge cannot repeat this quietly.

Decisions worth a close look

  • A trial may not score the test partition. FitModelUnit.trial_splits does
    not offer it — not as a default, but as a value that cannot be chosen.
    Scoring it once per trial lets the search see the test set, and a model picked
    that way has no honest score left to report.
  • Whether validation reaches the fit is policy, not shape. A holdout run
    hands it to the model to watch training and stop early; a fold must not,
    because those are the rows it will be scored on. Nothing raises when it goes
    wrong: the score simply comes out better than the model deserves.
  • The model is pointed at its data when it is fitted, not when it is built.
    Binding the splits at construction only worked while a model was fitted once.
  • The job aggregates the fold metrics, not the unit. A summary row carries
    std_value, and a unit does not write domain rows —
    BaseModel._save_metrics, the one sanctioned write, has nowhere to put a
    deviation.
  • Nested CV is a separate unit rather than a flag, because its inner
    splitter is a required component field, and a component field cannot be made
    optional without leaving the user with no selector.

What is deliberately missing

  • Per-fold progress reporting is gone. The bar sits at 20% for the whole CV
    loop. Restoring it needs a callback in a unit's contract, and a runtime
    parameter the engine cannot supply makes that unit unusable as a graph node.
    This is the one visible regression and it is unresolved on purpose — if
    the progress matters more than that property, it can be implemented.
  • Prediction.split / predicting on a split of a run (PR Enable predictions on dataset partitions and remove refitting on validation for forecasting #858) fell out
    while resolving the merge, because a coherent file was preferred to a
    half-applied one. It needs a unit that selects rows; that was not merge work.
  • FitModelOverNestedFoldsUnit has no contract tests of its own yet (it is
    covered end to end).

Other changes worth naming

  • An Alembic merge migration (82fb7a6b8ac2). The two branches left
    different heads, which made every test that builds the app fail. It is an
    empty revision: the two sides touch disjoint tables.
  • Saving a model is atomic now. A save that died partway left a truncated
    artifact that the row went on pointing at.
  • PredictJob raises JobError rather than HTTPException on the two
    branches that fail while predicting: an HTTPException stores nothing in
    args, so dill brings it back from the worker without its message. Six
    pre-flight raises still have this shape and are noted in the code.
  • DatasetJob clears file_path when a failure removes the folder it
    created.

Verification

uv run pytest tests/

3692 passing. One test was deselected during development:
tests/back/test_main.py::test_app_front builds the app against the real
~/.DashAI rather than a temporary directory, and if that database is stamped
with an Alembic revision from another branch, the migration code backs up and
recreates the user's database
. Worth fixing separately.

The new regression net — tests/back/api/test_model_job_cross_validation.py,
34 tests — was written and verified against develop with nothing modified,
before any of this. It pins the shape of the splits, which rows reach which
partition, how many fits happen and on what, the fold metrics and their
aggregation, and the verbatim text of every error branch. That it still
passes without a single assertion changed is what says the observable behaviour
did not move.

Felipedino and others added 29 commits August 13, 2026 20:35
- Introduced tests for atomic units including registration, execution, and validation.
- Added a new scratch.py module to manage temporary directories for dataset caching during tests.
- Updated existing tests to utilize the new scratch directory management to avoid cluttering the repository.
- Ensured that unit tests cover various scenarios including validation failures and context management.
- Updated PrepareAndSplitUnit to require dataset_id and provide task_name.
- Introduced SaveDatasetUnit for persisting datasets to disk.
- Added comprehensive tests for SaveDatasetUnit and LoadDatasetUnit to ensure correct functionality.
- Implemented contract tests for ApplyConverterUnit to validate context handling and converter behavior.
- Enhanced unit contract tests to ensure all context keys are declared and properly managed.
…asetUnit

- Introduced FitConverterUnit to fit a converter on a dataset without transforming it.
- Introduced TransformDatasetUnit to apply an already fitted converter to a dataset.
- Updated ApplyConverterUnit to utilize the new units for fitting and transforming.
- Added ConverterScopeMixin for shared scope resolution logic between converter units.
- Updated initial_components.py to include new units.
- Added tests to ensure correct functionality of the new units and their interactions.
Brings the fourth contract check from feat/atom-expl-predict-explor, where it
was written, so the file is byte-identical on both branches and the two PRs
cannot diverge on it.

The check matters on its own: __call__ demands every key in REQUIRES
unconditionally, so a key that is declared but never read is not harmless
documentation — it rejects any upstream that does not happen to publish it.

It also matters that this file specifically stays in sync. A mismerge here is
the one that fails silently: the audit keeps passing, it just audits less.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Introduced `test_exploration_units.py` to validate the functionality of exploration units, ensuring proper handling of datasets, explorers, and saving results.
- Added `test_prediction_units.py` to test prediction units, focusing on model loading, dataset handling, and prediction saving.
- Enhanced `test_unit_contracts.py` with a new test to ensure units do not require keys they do not read, preventing potential composability issues.
as_shap_predictor handed SHAP a closure over model.predict. The callers
(KernelShap, RegressionKernelShap, ContrastiveShap) first move the
background into the model's feature space with prepare_model_input, so
going through predict ran the model's input preparation a second time
over an already prepared matrix. SHAP then perturbs that matrix into
plain arrays, which the preparation cannot consume at all, and the
failure surfaced far from its cause as

    AttributeError: 'numpy.ndarray' object has no attribute 'types'

The module docstring and a comment above each caller already said the
model had to be queried through predict_prepared; only the call was
left behind.

test_the_wrapped_predictor_never_routes_through_predict pins it: its
stub raises if predict is reached, so a future regression fails at the
wrapper instead of five frames away inside SHAP.
- Introduced new tests for `LoadUploadedDatasetUnit`, `LoadDatafileDatasetUnit`,
  `InferDatasetTypesUnit`, `ApplyDatasetSchemaUnit`, `ComputeDatasetMetadataUnit`,
  and `SaveDatasetToPathUnit` in `test_dataset_ingest_units.py`.
- Updated expected unit schemas in `test_units_api.py` to include new dataset ingestion units.
- Enhanced validation checks and error handling in the dataset processing workflow.
…saving

- Introduced tests for tracking DAG execution in `test_tracking.py`, ensuring proper recording of node runs and artifacts.
- Updated `test_dag_engine_spike.py` to reflect changes in run_id handling, ensuring it is no longer an orphan input.
- Enhanced `test_build_model_unit.py` with tests for handling run_id and model evaluation metrics.
- Added tests in `test_evaluate_model_to_artifact_unit.py` to verify model evaluation and metric publication.
- Implemented validation checks in `test_evaluate_model_unit.py` to ensure run_id presence during evaluation.
- Created `test_save_model_unit.py` to validate model saving behavior and artifact prefix handling.
- Updated `test_fit_model_unit.py` to ensure proper handling of model state and artifact naming conventions.
- Enhanced unit contract tests in `test_unit_contracts.py` to ensure configuration keys are properly declared and classified.
Cross-validation landed on develop as a second abstraction layer for training
(BaseSplitter + BaseEvaluationStrategy) rather than as a branch inside the
monolith, so this merge reconciles two complete implementations of the same
work rather than combining two feature sets.

Six files conflicted. Where a file had been rewritten end to end by one side,
it is taken whole -- resolving those hunk by hunk produced files that belonged
to neither side and did not run:

  model_job.py      develop's, plus the guard for a run id that does not exist.
                    The units rebuild this file in a later slice.
  pipeline_job.py   this branch's; develop still carries the dead subsystem.
  predict_job.py    this branch's. Reverts develop's HTTPException -> JobError
  explainer_job.py  fix, which is restored below, and leaves two develop
                    features out; see the note at the end.
  dataset_job.py    this branch's, plus develop's on_cancel, which merged
                    cleanly on its own.
  initial_components.py   both sides: 330 components, no duplicates.

base_model.py did NOT conflict, and that is the sharp edge of this merge. Both
sides had added a method named compute_metrics to BaseModel -- this branch by
extracting the scoring body out of calculate_metrics so a caller with no Run
row could reuse it, develop by copying that body for the fold loop. They do not
overlap textually, so git merged them into one class with two definitions and
Python kept the second. That silently changed what every caller of
calculate_metrics computes, including dropping the filter for non-finite
scores. Only ruff's F811 noticed.

Resolved deliberately: develop's copy is deleted, this branch's is kept with
its non-finite filter, and develop's fold_index, inner_fold_index and
_epoch_reporter are kept on calculate_metrics. cv.py is adapted to the unified
contract -- compute_metrics returns None when there was nothing to score, which
is not the same as scoring nothing, so a fold with no validation data now
raises instead of contributing to the mean. A test parses BaseModel and asserts
a single definition, so the next merge cannot repeat this quietly.

Alembic had two heads (pipeline run tracking, and the cross-validation and
reports line), which made every test that builds the app fail. Joined by an
empty merge revision: the two sides touch disjoint tables.

Also restored or fixed while verifying:

  - DatasetJob clears file_path when a failure removes the folder it created.
    develop's on_cancel writes the path as soon as the folder exists so a
    cancelled job can clean up; without this the row survives a failure
    pointing at a folder that is gone. folder_is_ours is now bound before the
    try that reads it.
  - PredictJob raises JobError rather than HTTPException on the two branches
    that fail while predicting. An HTTPException stores nothing in args, so
    dill brings it back from the worker without its message. Six pre-flight
    raises still have this shape and are left for the slice that rebuilds this
    job.
  - Three fixtures that build a ModelSession now name their evaluation
    strategy, which is NOT NULL since cross-validation, and their splitter.

Test suite: 2 failed, 3628 passed. The two are named debts, not surprises:
test_cross_validation_run_is_explained_on_its_reserved_rows needs develop's
CV-aware explanation indexes ported into the explanation units, and
test_app_front needs a frontend build that no checkout here has.

New: tests/back/api/test_model_job_cross_validation.py, 34 tests pinning the
observable contract of a cross-validated run -- split shapes, which rows reach
which partition, how many fits happen and on what, the fold metrics and their
aggregation, and the verbatim text of every error branch. Written and verified
against develop untouched, before any of this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PrepareAndSplitUnit carried the partitioning policy inside itself, reached
through prepare_for_model_session, and took its configuration as an untyped
`splits` dictionary -- the kind of field the atomization notes admit only
because there was nothing better to declare. develop meanwhile grew ten
splitters, each a registered component with its own schema, multilingual
labels and compatibility per task. Those are the better thing: the unit now
picks one with a component field, the same way BuildModelUnit picks a model and
FitModelUnit picks an optimizer.

Two units rather than one with a flag, for two reasons that both bite:

  A component field carries a single `parent` and the front reads it directly
  off the property, so a field offering both families would leave the user
  without a selector at all -- the same wall that made the two explainer units
  siblings.

  What comes back has a different *type*. A holdout splitter returns one
  DatasetDict per side; a fold splitter returns a list of them, plus a trailing
  entry that is not a fold. Publishing that list as `x` would give one key two
  shapes, which a contract comparing key names cannot express: a graph would
  validate and then fail at run time, or quietly train on the wrong thing. So
  PrepareAndFoldUnit publishes `x_folds` and `y_folds`.

The two families are told apart with no renaming and no registry change:
component_parent matches any ancestor by name, and the hierarchy already
partitions the ten exactly -- PartitionSplitter covers the two holdout
splitters, FoldSplitter the eight fold ones.

The shared body lives in splitter_scope.py, next to converter_scope.py and for
the same reason: one implementation of resolving the task, preparing the
dataset and selecting the columns, so the siblings cannot drift into two
answers for the same dataset. It takes and returns plain values and never
touches the context -- a ctx.put hidden in a helper is invisible to the audit
that parses each unit's own source, so a broken PROVIDES would pass it.

Two details worth naming:

  BaseSplitter.__init__ takes a single `splits_data` mapping rather than
  keyword arguments, so this is the one component field in the units that is
  not expanded with **params.

  The instance state is declared on each unit and not in the mixin's __init__.
  BaseUnit.__init__ comes first in the MRO and does not chain, so a mixin
  __init__ never runs -- which surfaced as the resolved task being missing
  rather than as anything about construction. ApplyConverterUnit already does
  it this way.

The splitter's own refusal passes through undecorated: it already names the
numbers that explain it, and the caller that knows which run this was frames it
from outside, which is how the message the user reads is built today.

Contract tests build the context by hand rather than going through a job,
including that two of these units in one context do not share a resolved task.
The spike is untouched in substance: it only ever used this unit for static
validation, which reads REQUIRES and PROVIDES and never constructs anything.

969 passed in units, dag, spike and api; the one failure is the CV-aware
explanation indexes still to be ported, which is a later slice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…of it

Two changes, both policy rather than shape, and both measured by the
cross-validation net before they were made.

BuildModelUnit no longer takes the data. ModelFactory attached the splits to
the model instance at construction, which worked only because a model was
fitted once on one split. Fitted over folds it sees different data on every
iteration, so binding one partition at build time would leave the metrics
describing whichever fold happened to be built with. The unit now needs only
the label count -- a property of the dataset, not of a split -- and whoever
fits the model points it at what it is being fitted on. REQUIRES loses `x` and
`y`, which is a relaxation: every graph that fed it still validates, with two
fewer wires. The measurement in the graph test moves from fifteen to thirteen
and says why.

FitModelUnit gained `validation_during_fit`. It handed the validation partition
to `train` unconditionally, and both halves of that are wrong for folds:

  A model uses validation data to watch the fit and stop early, which is what
  an ordinary holdout run wants and exactly what a fold does not -- a fold is
  scored on the rows it held back, so letting the fit watch them measures it on
  data it was allowed to see. Nothing raises; the score just comes out better
  than the model deserves. The net recorded four fits in a cross-validated run
  and none of them receiving validation data, which is the behaviour this field
  now expresses.

  The trailing entry a fold splitter produces has no validation partition at
  all -- it holds the pooled rows and the reserved ones -- so reading
  x["validation"] is a plain KeyError on the very partition set that fits the
  model which gets kept.

Read with `.get` and the schema's own placeholder, the way the other units read
a declared optional field: a caller that builds this unit by hand should not
have to name a policy it is happy to leave alone.

Also moved the runs directory out of the top of execute and into the branch
that needs it. It is only used to name the plots a search produces, so a fit
without a search had been requiring a service of its caller for nothing --
which is what made these tests need a container before they could watch a fit.

973 passed across units, dag, spike and api; the one failure is the CV-aware
explanation indexes still to be ported.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BaseOptimizer.optimize took the task as its sixth argument and did the fitting
and the scoring inline. develop replaced that argument with a callable, because
cross-validation needs a trial to mean k fits rather than one, and the
optimizer has no business knowing which. FitModelUnit was still passing the
task into that position -- a silent mismatch, since a task is not callable, so
it would have surfaced from inside a trial rather than from the call.

The objective is now built by the unit that fits: one fit of the training
partition and one score of the validation partition. That is what makes the
same search reusable over anything that can be fitted and scored, which is the
whole point of the inversion -- the fold sibling will hand it a loop instead,
and nothing in the optimizer changes.

The trial metrics move with it, and that is the part worth noticing. They were
written by the optimizer, which meant the search decided what counted as a
scored partition. It is a property of the thing being fitted: a partition with
no metrics configured writes nothing, because calculate_metrics finds nothing
to score and returns. So the objective writes them.

`task` leaves FitModelUnit.REQUIRES, since nothing reads it there any more --
the audit would have caught it otherwise. It is still produced and still
consumed, by the local explanation unit. The graph measurement drops from
thirteen wires to twelve and says which one went and why.

Covered directly rather than through a job: the graph test's model declares no
optimizable parameters, so the search branch never runs there, and the
orchestration net exercises develop's ModelJob rather than these units. Three
tests pin what the objective computes, what it logs, and -- separately -- that
it is what reaches the optimizer, because passing the wrong sixth argument is
invisible until something calls it.

976 passed across units, dag, spike and api; the one failure is the CV-aware
explanation indexes still to be ported.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ExplainerJob read train_indexes, test_indexes and val_indexes straight off
Run.split_indexes. That is the shape a holdout run stores. A cross-validated
one stores an entry per fold plus the pooled rows and the reserved ones, so
explaining such a run raised KeyError inside the wrapper that reports a
preparation failure -- the user was told the dataset could not be prepared,
which is true and is not the reason.

develop had already built what this needs: explainable_indexes asks the
splitter that produced the run which partitions it has and what they are
called, and maps whichever answer it gives onto the three slots the explainers
are built from. A splitter added later needs no change here, and a fold run is
explained on the rows no fold ever saw.

Resolved in the job rather than in the unit. Unpacking the JSON column of a row
is an artifact of how the column is stored rather than part of the
transformation, and deciding which splitter wrote it is the same kind of
unpacking. The unit's contract does not change: it still requires
split_indexes, and still gets the three lists it always did.

A payload that does not match its splitter now says so instead of reaching the
generic wrapper. The old message read as a problem with the dataset rather than
with the run's own record of how it was split, so the test that pinned it is
updated with the reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SaveModelUnit called model.save straight at the destination. A save that died
partway left a truncated artifact there, and the row went on pointing at it as
if it were a model -- the failure is only visible later, when something tries
to load it.

develop had already fixed this in the job it kept, with atomic_save_path: the
model is handed a temporary sibling path, and what it leaves there is moved
into place once it returns. The temporary path is handed over rather than
derived here because only the model knows whether it writes a single file or a
directory of weights.

Two consequences worth having in the tests. The model no longer sees the final
path, so what is asserted is where the artifact ended up rather than what the
model was told -- which is what the caller and the row care about anyway. And
a double that recorded a path without writing anything now fails, correctly:
leaving nothing to move is the same thing a model that silently saved nothing
would do, and it should be reported rather than hidden.

978 passed across units, dag, spike and api.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ModelJob had its own copy of everything before the model is fitted: loading the
dataset, resolving the task, validating the dataset against it, counting the
labels, separating features from targets, resolving the metrics and the model
class, checking the downloads, and calling the splitter. The units had the same
steps. That was most of the duplication this reconciliation exists to remove,
and it goes in one piece rather than one path at a time -- holdout and
cross-validation differ only in which unit prepares the data.

_prepare_dataset_and_components now does what only it can: read the rows, unpack
the JSON columns stored on them, and choose which unit prepares the data. That
choice follows from how the splitter carves the dataset, which the splitter
declares -- it is a choice of unit rather than a flag on one, because the two
publish different shapes. The file loses eighty-four lines.

The evaluation strategy is built in run() now rather than in the helper,
because it takes the factory the build unit produced.

Two orderings are deliberate and were not obvious:

  BuildModelUnit.validate runs before the data is partitioned, and the unit
  itself after. The download gate lives in validate, and a model that cannot be
  trained should be reported as that rather than surfacing later as a splitting
  failure -- the same reason ModelJob has always resolved the optimizer before
  changing the run's status.

  The splitter class is resolved in the helper and again inside the unit. The
  helper needs it to know which unit to build, and the message a missing one
  produces belongs to the preparation step. The unit resolves its own because
  it must work for a caller that is not this job.

Both regression nets pass unchanged -- 46 tests, including the verbatim text of
every error branch, which is what says the messages did not drift. 978 across
units, dag, spike and api.

The evaluation strategy still owns the training loop. That is the next piece.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
HoldoutEvaluationStrategy.execute was FitModelUnit, EvaluateModelUnit and
SaveModelUnit in sequence, written a second time. ModelJob now composes those
three for a holdout run, and the strategy is no longer called for one. Fold
runs still train through it; their loop is the next piece.

Saving moved for both paths at once. The strategy hands the model back rather
than leaving it in the context, so the job puts it there and one SaveModelUnit
serves whichever path produced it -- which also gets the fold path the atomic
replacement it did not have separately.

FitModelUnit gained trial_splits, which is the SCORED_SPLITS question answered
for the search. Which partitions a run records a score for is declared by the
strategy the session chose, and it is not the same as which ones have metrics
configured: a forecaster has training metrics and still must not be judged on
the dates it was fitted on, because an in-sample fit statistic is not
comparable with a forecast. The job reads that declaration and passes it on, so
the strategy classes keep deciding it while the units do the work.

The test partition is deliberately not an option in that field. Scoring it once
per trial would let the search see it, and a model chosen with the test set in
view has no honest score left to report -- so a trial may score the partition
it fitted on and the one it is measured against, and nothing else. It was a
default before, which is a weaker statement than a value that cannot be chosen.

The runs directory leaves the job: naming the artifact is the saving unit's, and
the job had been resolving it only to build a path the unit builds itself.

Both nets pass unchanged, 46 tests. 992 across units, dag, spike, api and
evaluation -- the strategies' own tests included, since the classes still stand
and are still what declares the scored partitions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The cross-validation sibling needs the same optimizer resolution, the same
search, the same check that the optimizer gave back the model it was handed,
the same recording of the best parameters and the same plot writing. Copying
that is how the two implementations this branch is removing came to exist, so
it moves to fit_scope.py first and the sibling is written against it.

The contract audit caught a real mistake in the first attempt, and it is worth
recording because the rule reads like tidiness until it bites. The helper took
the context and did its own require and put. The audit parses each unit's own
source, so moving the reads out of the unit made four declared keys look
unread: FitModelUnit was suddenly requiring `factory` and `model_parameters`
and declaring `run_id` and `artifact_prefix` while appearing to touch none of
them. A caller reading only the declarations -- the graph validator among them
-- would have been told the truth by the declarations and contradicted by the
audit, or worse, the declarations would have been trimmed to match.

So the helper takes and returns plain values and never touches the context.
Every require and put stays in the unit, and so does reading the runtime
parameters. It is the mirror of the rule already written down for a helper that
publishes: a ctx.put hidden in one makes a broken PROVIDES pass.

992 passed across units, dag, spike, api and evaluation. No behaviour changed:
this is the same fit, moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FitModelOverFoldsUnit is FitModelUnit's sibling: it takes a list of partition
sets instead of one, which is a different REQUIRES and so a different unit.
Everything around the fit is the shared mixin; what is written here is the
objective the search measures -- the whole fold loop, so one trial costs k fits
-- and what happens once the search is over.

ModelJob composes it for a fold run that is not nested. Nested cross-validation
still trains through the strategy: its inner splitter is a required component
field, and a component field cannot be made optional without leaving the user
without a selector, so it is a further sibling rather than a flag on this one.

Three decisions worth naming.

The per-fold scores are published rather than aggregated in the unit. A summary
row carries a standard deviation, and a unit may not write domain rows -- the
one sanctioned write in the domain layer has nowhere to put one. So the unit
hands the numbers over and the job does the arithmetic and the writing, where
every other row it persists is written. A single fold gets a deviation of zero
rather than none, because none is what the reserved-rows measurement carries
and the two say different things.

A trial records one row per split holding the mean over its folds, not one per
fold: the folds of a trial measure a hyperparameter setting rather than the
model that gets kept, and recording each would bury the rows that describe it.
That write is guarded on the run, which _save_metrics does not guard for
itself -- a caller with no run would write rows against a foreign key pointing
at nothing, and they insert without complaint because nothing enforces it.

Scoring the reserved rows is not the unit's. It is an ordinary LAST metric, so
it is EvaluateModelUnit, the same one a holdout run uses, and whether there is
anything to score is the caller's to know: a session that reserved nothing
leaves that partition empty rather than absent.

calculate_metrics now returns what it wrote, so a caller that wants both the
row and the number scores the split once instead of twice.

The assertion that the optimizer gave back the model it was handed caught a
real gap while this was written: with the data attached at fit time rather than
at build time, nothing had pointed the model at anything during a fold search.
It does now, per fold, the same as the scoring loop.

1001 passed across units, dag, spike, api and evaluation; both nets unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tegy

FitModelOverNestedFoldsUnit is FitModelOverFoldsUnit plus one measurement taken
before it. Inheritance rather than a shared mixin because that is the actual
relationship: everything the sibling does still happens, and this adds a step
in front. Two units and not one with a flag because its inner splitter is a
required component field, and a component field cannot be made optional -- the
front reads `parent` straight off the property, and an anyOf buries it where it
does not look.

What the nested loop is for, since the code alone does not say it: in an
ordinary cross-validated search the same folds choose the hyperparameters and
report the score, so the score is optimistic by however much the search managed
to fit them. The nested loop measures that honestly -- for each outer fold a
search runs on folds carved out of that fold's training rows alone, and what it
chooses is scored on the outer fold's validation rows, which it never saw.

What it does not do is choose the hyperparameters: each outer fold picks its
own and they generally differ, so there is no single model to keep out of that
loop. The ordinary search still runs afterwards and produces the model that
gets saved. The nested numbers describe the procedure, not the artifact, which
is why they are kept at their own level -- LAST_OUTER against LAST -- and why
the inner trials record nothing at all.

The two fold branches in the job became one, choosing a unit rather than
repeating a body.

**The evaluation strategies are no longer called.** The job reads SCORED_SPLITS
and KIND off the class it resolves and never touches execute() on any path. The
classes are now what they always were underneath -- a declaration of how a run
is carved and what it records -- and emptying them of the code that is now
unreachable is the last piece.

1014 passed across units, dag, spike, api and evaluation, both nets unchanged,
plus thirteen contract tests for the fold unit built on a hand-made context:
the end-to-end net runs it inside a real job, which cannot show what it reads,
promises and refuses on its own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
They ran the training: execute took the run row and the database session and
did the fitting, the search, the scoring, the aggregation and the persistence
behind one method. Every piece of that is now a unit, and nothing has called
execute since the fold paths moved -- the job reads SCORED_SPLITS and KIND off
the class it resolves and never touches it otherwise.

So the code goes. base_evaluation_strategy, cv and holdout drop from 881 lines
to 153, and what is left is what was underneath all along: how a run is carved,
and which partitions it records a score for.

They stay registered. That is not deference to dead code -- the frontend reads
these classes in four places, and only one is about metrics. The session wizard
lists them so the user can choose one, and ModelSession.evaluation_strategy is
NOT NULL, so without that listing a session cannot be created at all. It starts
on the first one whose kind is holdout. `kind` decides the shape of the splits
payload and which controls are shown. Only `scored_splits` is about the charts.
Removing the classes would not have cost two screens; it would have cost the
way sessions are made.

There is precedent for a class here that declares and does not execute:
BaseSplitter.PARTITIONING and explainable_partitions are read exactly this way,
by the backend and by the frontend, and nothing calls them to do work.

Five of the forecasting tests exercised behaviour rather than declarations --
the final fit, and which partitions a trial scores. That behaviour moved rather
than disappeared, so they are pointed at the units that carry it out now. They
stay in the same file, next to the declarations, because that is the pair that
has to stay consistent: a strategy that says it does not score the training
partition, and a fit that then does not.

1014 passed across units, dag, spike, api and evaluation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two optimizer test files built the objective they measure by reaching for
HoldoutEvaluationStrategy.evaluate. That method is gone: the strategies declare
how a run is evaluated and the fitting unit carries it out, so the objective
comes from there now. The tests themselves are unchanged -- they still check
that a bad trial is pruned, that disabling the pruner completes every trial,
and that a real model reports each epoch to its trial.

Which partitions a trial records is still read off the strategy class, the same
way the job reads it, so the declaration stays connected to what it produces.

These four failures were not caught earlier because the verification runs had
been narrowed to the directories this work was touching -- units, dag, spike,
api and evaluation -- after the full suite was dropped for containing a test
that builds the app against the real ~/.DashAI. Deselecting that one test was
the right answer; shrinking the suite to what seemed relevant was not, and it
is precisely the change that removes a caller elsewhere that this hides.

Whole suite: 3682 passed, one test deselected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five real findings, one of them hiding the others.

**A search needs a tuner, not only a target.** Both evaluation strategies
guarded on `self.optimizer and self.run_optimizable_parameters`; the units kept
only the second half. `Run.optimizer_name` is a plain string and the wizard
leaves it empty when no search is asked for, while the model may still declare
a parameter optimizable -- a combination that has always meant "fit it once
with the values given". It had become a lookup of the empty string in the
registry, surfacing as "Metric is not compatible with the Task. ''", a message
with nothing to do with what happened. Reproduced, fixed with a shared
`_will_search`, and pinned by a test.

**The nested unit was not being audited at all.** `_unit_class` matched only
classes whose direct base is literally `BaseUnit`, so a unit that extends
another unit fell out of every contract check -- and would have failed them,
because its PROVIDES are written by the parent's body rather than its own. It
is the same blindness a shared helper causes, arriving by inheritance instead:
the audit reads one class's source. It now follows the lineage for
declarations, context calls and config reads. 32 audited units became 33.

**The inner splitter was resolved unconditionally**, so a run still carrying a
nested configuration it no longer uses failed on a splitter it would never have
touched.

**The fold branch hardcoded `splits=["TEST"]`** where the holdout branch derives
it from SCORED_SPLITS. Latent today -- no strategy excludes TEST -- but it is
exactly the coupling this work exists to remove.

And a docstring describing `{split: [scores]}` for something shaped
`{split: {metric: [scores]}}`.

Two findings were left alone, deliberately. `best_parameters` is published
without being in PROVIDES, which is the already-declared limitation that there
is no way to express an optional output; the new units repeat it rather than
inventing an exception to it. And per-fold progress reporting is gone, which is
a real regression: restoring it needs a callback in a unit's contract, and a
runtime parameter the engine cannot supply makes the unit unusable as a node --
the static validator rejects it. Both are recorded rather than patched over.

Whole suite: 3692 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 11, 2026 02:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Felipedino
Felipedino marked this pull request as draft September 11, 2026 02:44
@Felipedino
Felipedino marked this pull request as ready for review September 21, 2026 02:16
A search hands back the model with the best values written onto it as
attributes, but its weights are whatever the last trial left behind. The
LAST metrics and the serialized model therefore described the last trial
rather than the best one, and a pruned last trial left a half-trained
model. The merge-base's holdout strategy refitted after the search, and
so does FitModelOverFoldsUnit on its pooled rows; FitModelUnit now does
the same on the training partition.
The endpoint stored the split, the splits endpoint listed the partitions
and the front offered the selector, but the job loaded every row of the
dataset and never read prediction.split, so asking for the test partition
silently returned train, validation and test alike, and a partition the
run does not have no longer failed.

Restore what v0.10.0 did: when a split is named and the dataset is the
one the model was trained on, resolve its rows with run_split_indexes and
narrow the loaded dataset to them before predicting. A ValueError from the
lookup ends the prediction in error with the same message as before.
A regression predicts continuous values whatever type its target was
trained as, but SavePredictionUnit inherited the training dataset's type
for the predicted column. With an integer-typed target the save cast the
predictions to int64 and Arrow refused the truncation, so every
prediction of such a run ended in error. v0.10.0 overrode the type for
regression tasks and the unit lost it when the job was decomposed.

The unit now receives the task name as configuration, the same way
PredictUnit does, resolves its class from the registry and declares the
predicted column a float when the task is a regression.
Resolve the nine conflicts the way the PR 875 review handoff prescribes.

- model_job.py, predict_job.py, explainer_job.py: PR side kept as is.
  develop's only changes there are the session preprocessing commits
  (1c4b9fb, 63c5d87); porting that into the units is deferred to
  follow-up work branching off this commit.
- base_model.py: one compute_metrics, final, returning _score_split;
  calculate_metrics keeps the getattr run_id guard, scores through
  _score_split and returns what it wrote. None still means nothing to
  evaluate, an empty mapping means nothing finite.
- base_evaluation_strategy.py, cv.py, holdout.py: the PR's declarative
  strategies, plus develop's DESCRIPTION on CrossValidation and Holdout.
- base_optimizer.py: develop's contour and importance fixes for
  categorical parameters together with the PR's artifact_prefix.
- converter_job.py: PR side; the PR's private copy of
  rebuild_dataset_with_transformed_columns is dropped in favour of
  develop's shared DashAI/back/converters/dataset_columns.py.

Ported alongside, from the same handoff:

- fit_scope._search keeps a None slot for a plot the optimizer skipped,
  so the contour column of the run stays empty instead of pointing at a
  pickle of the text "None".
- FitModelOverFoldsUnit._score_folds raises RuntimeError on a non-finite
  fold objective, naming the metric and the fold.
- tests/back/evaluation/test_cv_goal_metric.py exercises that guard
  through the unit, since FoldEvaluationStrategy.evaluate no longer
  exists; test_metric_scoring pins compute_metrics to None when no
  metrics are declared.
The merge left two heads, 2e1b3462553d from the pipeline tracking chain
and f4a91c62d8e7 from session preprocessing, and startup refuses to
migrate past "Multiple head revisions" (and then backs up and recreates
the database). An empty merge revision joins them; upgrading an empty
database, or one stamped at either old head, now ends at 2f0172d883dc.
Three tests that arrived with develop assumed things the merge changed.

- test_mlp_regression drove Optuna through
  HoldoutEvaluationStrategy.evaluate, which the strategies no longer
  carry; the objective now comes from FitModelUnit._score_one_trial, the
  same port tests/back/optimizers/test_optuna_real_model.py already made.
- test_evaluate_model_to_artifact_unit borrows BaseModel's scoring
  methods onto a stub; both of them now score through _score_split, so
  the stub borrows that one too.
- test_add_preprocessing_to_model_session upgraded to head and then
  downgraded one step, but head is now a merge revision with two
  parents and a relative downgrade from it is an ambiguous walk. It
  upgrades to the revision under test instead.
A contour plot needs two numeric axes, and an optimizer with fewer
returns None in its place. The search loop keeps that None at its index
rather than normalizing it, since normalize_artifacts would turn it into
a text artifact and the run's contour column would point at a pickle of
the word "None". The test drives FitModelUnit with an optimizer that
skips the contour and checks the third slot stays empty on disk and in
the published paths.
ExplainerJob hands PrepareExplanationDataUnit and GenerateLocalExplanationUnit the session's preprocessing_artifacts_path, as runtime configuration and only when the session declared steps. The prepare unit transforms the whole dataset with the persisted final.pkl before replaying the run's split; the local unit transforms manual input before selecting columns and stored rows before the task prepares them, mirroring develop's explainer_job.
Unit tests fit a real Binarizer through SessionPreprocessor, pickle it as final.pkl and check both units over it; an API test trains a scaler session and asserts the stored explained instance is the transformed one.
A session that declares preprocessing steps is trained on the columns the
converters PreprocessingJob fitted produce, and its input_columns are
rewritten with those names. The prediction path selected them from the
raw rows, so a converter that derives a column ended every prediction in
"Model prediction failed", and one that keeps the names handed the model
raw data while develop applied final.pkl. Port develop's PredictJob
behaviour: unpickle the final preprocessor, never refit it, transform the
rows before select_columns and predict, and save the raw rows plus y_pred.

ApplySessionPreprocessingUnit does the transform and publishes the model
input under its own key, so the raw rows under "dataset" survive for
SavePredictionUnit. It always runs, passing the rows through when the
session has no artifacts path, so PredictUnit reads one key in both
cases and the graph keeps one shape. The preview narrows its rows from
the model input, since with preprocessing the session's input columns
only exist there. Hand-typed rows go through the same transform.
develop's test reaches the preview through a run trained by ModelJob,
and training on a derived column is being fixed separately. Fit the run
by hand on the column the persisted Binarizer produces and drive the job
on the whole dataset, on one partition of it and on hand-typed rows, plus
the preview, comparing what was saved with the labels the model gives
rows transformed by the fitted preprocessor directly. Without the port
every one of them dies with 'Field "bin_SepalLengthCm" does not exist'.
ModelJob selected the resolved input columns from the raw dataset, so a
converter that produces new columns failed with 'Can not prepare Dataset'
and one that keeps the names trained on raw data while predict and explain
applied final.pkl. This mirrors develop's apply_persisted_preprocessing in
the units.

PrepareAndSplitUnit and PrepareAndFoldUnit take two optional fields as a
pair, input_column_refs and preprocessing_artifacts_path, which ModelJob
sets only when the session has preprocessing steps. With them the task is
validated on the raw refs alone, the split carries every raw column, and
each entry is transformed with its own fit (fold_{i}.pkl, final.pkl) and
narrowed to what the refs resolve to before it is published; y is left as
the splitter narrowed it.

They are schema fields rather than runtime params because the DAG
validator demands every runtime param of every node, which would leave the
two units unusable on a canvas without preprocessing. The path is
classified in the widget registry and the property sets are re-pinned.
Pickled stand-ins for the SessionPreprocessor that PreprocessingJob
persists, each producing a column named after the artifact it was loaded
from, so a test can read off every entry which fit transformed it: the one
holdout entry with final.pkl, fold i with fold_{i}.pkl and the trailing
entry with final.pkl.

Also pins that the task is validated on the raw refs alone, that the split
carries a raw column the refs never name so a converter can read it, that
y is untouched, that one field without the other is refused, and that a
missing artifact is reported by the name of the entry it was fitted for.
…t fitted

A session that declares preprocessing steps but has no artifacts path yet
(its PreprocessingJob is pending or failed) was predicted and explained on
raw rows without a word, while training already refused that state. The
runs endpoint keeps it out of reach with a 409, so nothing reached it
today; the jobs now refuse it themselves with their own message, and the
preview answers 409 like the runs endpoint does.
@cristian-tamblay

Copy link
Copy Markdown
Member

Diseño propuesto: forecasting y clustering sobre las units

Propuesta de diseño, no implementada, escrita el 2026-09-28 sobre el árbol de #908
(esta rama más develop). Las referencias archivo:línea son sobre ese head. Prior art:
origin/clustering (PR #841, draft de constanzaru). Forecasting se des-registró en develop
en #895 porque su abstracción no quedó bien integrada; este documento dice qué tiene que ser
cierto para volver a registrarlo, y qué hace falta para clustering.

Respuesta corta

  • No hacen falta sub-jobs. Lo que difiere entre familias es qué units se componen y qué
    filas Metric se escriben. Lo que comparten (estado del Run, progreso, errores, plots,
    guardado del modelo) es exactamente lo que un job hace. Un SupervisedJob sería el
    ModelJob de hoy con otro nombre y un ForecastingJob no tendría cuerpo.
  • Forecasting corre hoy como familia supervisada con otras declaraciones (la task
    ordena por fecha, los splitters temporales ignoran y, las estrategias declaran
    SCORED_SPLITS sin TRAIN, PREDICTS_FORWARD_ONLY filtra las particiones predecibles),
    y para que corra sin fugas le falta una declaración y un test end to end. Pero el
    concepto es incorrecto
    : la fecha va como feature y la serie como target, cuando la
    fecha es el índice y la serie es a la vez historia y salida (autorregresión). Eso no se
    arregla con declaraciones; necesita roles de columna y una familia temporal propia en
    la tabla (ver "El problema de fondo" más abajo). Es la razón real de que se haya quitado
    del registro y es más complejo que clustering.
  • Clustering sí es una familia nueva, porque no tiene y, no tiene particiones y
    puntúa con score(X, labels). Necesita dos units nuevas, una estrategia declarativa
    full, un splitter declarativo full, y los contratos de task, modelo y métrica que feat: add clustering support across models and notebooks #841
    ya escribió.
  • La elección de familia vive en ModelJob como una tabla explícita indexada por
    (splitter.PARTITIONING, task.REQUIRES_TARGET), no como if/else ni como jerarquía de
    jobs. Hoy run ya elige por strategy.KIND y run.nested, y _prepare_dataset_and_components
    por PARTITIONING: son dos fuentes para una decisión. La tabla las une.

Forecasting: qué significa "bien integrada" sobre las units

Lo que ya está resuelto por declaraciones (verificado en el árbol mergeado):

Diferencia con regresión Dónde se declara Quién la lee
Orden temporal de las filas ForecastingTask.prepare_for_task ordena por fecha SplitterScopeMixin._prepare selecciona del dataset preparado
Cortes temporales, sin y TemporalHoldoutSplitter (holdout) y RollingOriginSplitter (folds), COMPATIBLE_COMPONENTS=["ForecastingTask"] PrepareAndSplitUnit / PrepareAndFoldUnit por PARTITIONING
No se puntúa TRAIN ForecastingHoldout/CV.SCORED_SPLITS=(VALIDATION, TEST) ModelJob arma trial_splits y EvaluateModelUnit(splits=...)
Solo se predice hacia adelante ForecastingTask.PREDICTS_FORWARD_ONLY + strategy.FINAL_FIT_PARTITIONS splits_payload.predictable_splits
Métricas familia de regresión más MAPE/SMAPE BaseModel._score_split

Lo que NO está resuelto y explica el "no bien integrada":

  1. Fuga latente de validación en el fit. FitModelUnit entrega x["validation"] al
    train del modelo (validation_during_fit=True por defecto,
    fit_model_unit.py:127) y ModelJob nunca lo cambia. Hoy funciona solo porque los
    forecasters ignoran ese argumento (arima.py, "Unused"). Un forecaster con early
    stopping vería la partición con que se le puntúa. Arreglo: atributo
    VALIDATION_DURING_FIT = True en BaseEvaluationStrategy, False en las dos estrategias
    de forecasting, y ModelJob lo pasa a FitModelUnit. Es un atributo aparte, no una
    derivación de FINAL_FIT_PARTITIONS: que el modelo mire validación no es lo mismo que
    ajustarse con ella (holdout.py:29-31).
  2. PreprocessingJob parte el dataset crudo y ModelJob el preparado (problema heredado de
    develop). Con una serie desordenada, fold_{i}.pkl y final.pkl se ajustan sobre filas
    distintas a las de las folds de entrenamiento. Arreglo: PreprocessingJob llama
    task.prepare_for_task(loaded_dataset, nombres raw, output_columns) antes de
    splitter.split, igual que SplitterScopeMixin._prepare. Mientras no esté, rechazar
    con 409 una sesión de forecasting con pasos de preprocesamiento.
  3. No existe ningún test end to end de ModelJob con ForecastingTask (los tests de
    orquestación usan tasks dummy). Ese test es la condición para re-registrar: holdout
    temporal y rolling origin llegan a FINISHED, escriben LAST solo para VALIDATION y TEST,
    split_indexes en orden temporal, y un stub que registre x_validation recibe None.
  4. Re-registrar las 12 entradas que quitó 6c3b7c741 en initial_components.py. Los
    splitters temporales y la compatibilidad de Optuna con ForecastingTask nunca se
    quitaron.

Tamaño: S a M (unas 30 líneas de registro, 40 en PreprocessingJob, 200 de tests).
Sin cambios de front ni de dispatch.

El problema de fondo: el concepto autorregresivo no está en la abstracción

Los cuatro puntos anteriores hacen que forecasting corra sin fugas sobre las units, pero no
corrigen el modelo mental que hoy usa el software y que es el motivo real de que la
abstracción no haya madurado:

  • Hoy la fecha se declara como input column (feature) y la serie como output column.
    ForecastingTask.metadata exige exactamente un input Date y un output numérico, y por
    debajo ForecastingModel.train(x_train, y_train, ...) ignora las features: lee la serie
    desde y_train y usa x_train solo para recordar fechas, y predict(x) convierte las
    fechas de x en un número de pasos desde _last_train_date. La fecha no es una feature,
    es el índice temporal. Las features de un modelo autorregresivo son los valores pasados
    de la propia columna objetivo (rezagos, ventanas), y la salida es esa misma columna hacia
    adelante.
  • Consecuencias en la abstracción actual: select_columns(x, y) y splitter.split(x, y)
    reparten "features" y "target" cuando en realidad reparten "índice" y "serie";
    PredictJob pide "fechas a predecir" cuando lo natural es un horizonte; el
    preprocesamiento de sesión (converters por columna) no tiene forma de expresar rezagos o
    ventanas sobre la serie; las variables exógenas (ramas
    origin/feat/exogenous-forecasting-*) no caben porque el único input permitido es la
    fecha; y n_labels, output_columns[0] en predict y la validación de cardinalidades
    están todas pensadas para tabular supervisado.
  • Lo que una abstracción madura tendría que declarar (a resolver antes de re-registrar
    forecasting como familia de primera clase, no en PR 875):
    1. Roles de columna en la task: index (temporal), target (serie, que es a la vez
      historia y salida) y exogenous (features de verdad, opcionales). El wizard pide
      roles, no "inputs y outputs".
    2. Una unit de preparación propia (PrepareSeriesUnit) que publique la serie ordenada por
      el índice más las exógenas, y que sea donde se construyan rezagos/ventanas si el modelo
      los necesita (o que los modelos los construyan, pero con un contrato explícito).
    3. Splitters temporales que corten por el índice, sin y en la firma.
    4. Contrato del modelo: train(series, index, exogenous=None) y predict(horizon, exogenous_future=None); la lógica "fechas a pasos" desaparece del modelo.
    5. Métricas y SCORED_SPLITS sobre el horizonte (validation y test como ventanas), lo
      que hoy ForecastingHoldout ya documenta pero sin apoyo del contrato.
      Esto es más complejo que clustering en el sentido de que cambia qué significan x e
      y en las units, no solo si hay y. Por eso conviene que la tabla de familias tenga
      desde el inicio la fila ("temporal", ...) como familia propia y no como caso de la
      supervisada; con las declaraciones de hoy forecasting corre, pero no es correcto.

Clustering, en cambio, tiene otra dificultad: también se puede ver como converter (agregar
una columna de etiquetas al dataset, que es lo que #841 hace en la ruta de notebooks). Las
dos vistas pueden coexistir: como familia de entrenamiento produce un run con métricas
internas y un modelo guardado; como converter produce columnas. La decisión es cuál se
ofrece primero y si comparten el mismo ClusteringModel.

Clustering: qué hace falta

Declaraciones (reutilizando #841)

  • BaseTask.REQUIRES_TARGET = True; UnsupervisedTask lo pone en False;
    ClusteringTask con outputs_cardinality 0 y num_labels() que devuelve None
    (BaseTask.num_labels sigue siendo abstracto en la PR). Se expone como
    requires_target en get_metadata para que el wizard oculte el selector de salida.
  • FullDatasetEvaluationStrategy (nueva, solo declara, como sus hermanas):
    KIND="full", SCORED_SPLITS=(SplitEnum.FULL,), FINAL_FIT_PARTITIONS=("full",),
    COMPATIBLE_COMPONENTS=["ClusteringTask"], DESCRIPTION. Satisface
    ModelSession.evaluation_strategy NOT NULL con un nombre real, el wizard la autoselecciona
    (única compatible) y los charts leen scored_splits=["full"] por el mismo camino que ya
    usan para forecasting. Reemplaza a SESSION_CONFIG_SCHEMA={"split_strategy":"none"} de feat: add clustering support across models and notebooks #841
    y al evaluation_strategy: "" que mandaba su front.
  • FullDatasetSplitter (nuevo, declarativo): PARTITIONING="full",
    TRAINING_PARTITION="full", schema vacío, split(x, y=None) devuelve
    ({"full": x}, None, {"full_indexes": [...]}). Existe para que _validate_splits,
    _prepare_dataset_and_components, run_splits, predictable_splits y PreprocessingJob
    no ganen una rama None cada uno. Ojo: normalize_splits_payload rellena un
    splitter_name ausente con HoldoutSplitter, así que un payload {"splitType":"none"}
    hoy pasa la validación como holdout por accidente; con el splitter real eso desaparece.
  • SplitEnum.FULL = "full". No requiere migración: Metric.split es Enum(SplitEnum) sin
    create_constraint, SQLite lo guarda como VARCHAR (feat: add clustering support across models and notebooks #841 lo agregó sin migración y con la
    suite verde). runs.py itera SplitEnum, así que full_metrics aparece en la respuesta
    del run sin tocar el endpoint. Agregar un test que escriba y lea una fila FULL sobre una
    DB migrada por alembic, no solo sobre el metadata en memoria.
  • ClusteringModel(BaseModel) de feat: add clustering support across models and notebooks #841: train(x), get_cluster_labels(x=None),
    COMPATIBLE_COMPONENTS=["ClusteringTask"]. Sobrescribe _score_split (que en la PR NO es
    @final; solo compute_metrics, calculate_metrics y _save_metrics lo son) para puntuar
    con metric.score(x, labels). Con eso EvaluateModelUnit y calculate_metrics funcionan
    sin cambios y escriben FULL/LAST por el camino normal. No hace falta el SupervisedModel
    de feat: add clustering support across models and notebooks #841
    ni reparentar todos los modelos: ese split existía solo para obtener el hook que
    _score_split ya es.
  • ClusteringMetric(BaseMetric).score(X, labels) y las tres métricas de feat: add clustering support across models and notebooks #841, sin cambios.

Units nuevas

LoadDatasetUnit
  → PrepareFeaturesUnit (NUEVA, BaseUnit + SplitterScopeMixin)
      SCHEMA: task_name, input_columns, splitter(parent="FullDatasetSplitter")  (sin output_columns)
      REQUIRES (dataset, dataset_id)   PROVIDES (x, n_labels, task, split_indexes, task_name)
      prepare_for_task(..., output_columns=[]) → select inputs → splitter.split(x, None)
      → _apply_preprocessing([x], ["final"]) (el helper puro que ya aplica final.pkl en #908) → publica x={"full": ...}
  → BuildModelUnit (existente; gana full_metrics junto a train/validation/test_metrics)
  → FitUnsupervisedModelUnit (NUEVA, BaseUnit + ModelFitScopeMixin)
      REQUIRES (model, factory, optimizable_parameters, model_parameters, x)
      PROVIDES (model, plot_paths)   RUNTIME (run_id, artifact_prefix)
      model.train(x["full"]); labels = get_cluster_labels(x["full"]);
      error de #841 si quedan < 2 clusters tras quitar ruido (-1);
      búsqueda opcional con objetivo = índice interno sobre x["full"] (misma coreografía que FitModelUnit)
  → EvaluateModelUnit (existente; su enum de splits gana "FULL")
  → SaveModelUnit

Publicar x con la misma clave y forma (dict de particiones) que PrepareAndSplitUnit
respeta la regla "una clave, un tipo" de la PR. n_labels=None ya viaja hoy para
forecasting. Cada unit nueva debe pasar la auditoría de test_unit_contracts.py y entrar en
initial_components.py y en los sets fijados de test_units_api.py.

Métricas FULL sin columna nueva

En v1 ModelJob llena full_metrics con todas las métricas compatibles con la task (la
regla de #841) y se lo pasa a BuildModelUnit. Así no hay columna ModelSession.full_metrics,
ni migración, ni campo nuevo en ModelSessionParams. Si después el front quiere un selector,
se agrega la columna encadenada tras la migración de merge 2f0172d883dc.

Preprocesamiento en vez de _standardise_features

#841 escalaba las features dentro del job, invisible para predict y explain. Sobre develop
el escalado es un paso StandardScaler de la sesión (stack de leakage, PR 876) que
PrepareFeaturesUnit aplica con final.pkl. Dos cambios pequeños: PreprocessingJob con
KIND="full" ajusta sobre {"full": dataset} y persiste solo final.pkl;
SessionPreprocessor ajusta sobre split[splitter.TRAINING_PARTITION] en vez del literal
"train". Costo: DBSCAN con eps por defecto sobre unidades crudas etiqueta todo como
ruido y el run termina en error, así que el wizard debería proponer el scaler para
ClusteringTask.

DB, API y front

  • DB: sin migración. La sesión guarda output_columns=[],
    evaluation_strategy="FullDatasetEvaluationStrategy", splits con
    splitter_name="FullDatasetSplitter". Run.split_indexes guarda {"full_indexes": ...}.
  • API: _validate_splits valida el schema vacío del splitter; /model_session/validation
    ya acepta outputs=[] con cardinalidad 0. predict.py y los endpoints de particiones del
    explainer devuelven [] porque explainable_splits y predictable_splits no ofrecen la
    partición de entrenamiento.
  • Front (re-encajando lo que feat: add clustering support across models and notebooks #841 ya hizo, indexado por kind y scored_splits en vez de
    splitType): STRATEGY_KINDS.FULL, resolveSplitterName → FullDatasetSplitter, ocultar el
    selector de salida cuando requires_target es falso (y no autoseleccionar la última
    columna, SelectColumnsStep.jsx:104-109), pestaña FULL en LiveMetricsChart,
    ResultsTabsHeader, RunResults, ModelComparisonTable, types/run.ts; ocultar
    Predict para tasks sin target hasta que exista PredictClustersUnit. api/job.ts sigue
    mandando ModelJob.
  • Optimizadores: registrar ClusteringTask en COMPATIBLE_COMPONENTS de Optuna/Hyperopt
    solo cuando se quiera HPO por índice interno; sin eso la búsqueda queda fuera sin flag.

De #841 se descarta

SupervisedModel y el reparenting de modelos y tasks; _standardise_features;
_prepare_without_target/_train_without_target inline en ModelJob; escritura directa
de filas Metric desde el job; SESSION_CONFIG_SCHEMA; evaluation_strategy: ""; el
vocabulario de "task executors". Los converters y explorers de clustering para notebooks
son otra PR.

Dónde vive la elección: tabla de familias en ModelJob

# DashAI/back/job/training_family.py (puro: sin ctx, sin db)
@dataclass(frozen=True)
class Family:
    prepare: type          # PrepareAndSplitUnit | PrepareAndFoldUnit | PrepareFeaturesUnit
    fit: type              # FitModelUnit | FitModelOverFoldsUnit | FitUnsupervisedModelUnit
    nested_fit: type | None
    partitions_key: str    # "x" | "x_folds"
    prepare_extra: Callable[[ModelSession], dict]
    fit_extra: Callable[[type, Run], dict]        # trial_splits, scored_splits, inner_splitter, validation_during_fit
    final_splits: Callable[[list, object], tuple]  # qué LAST puntúa EvaluateModelUnit
    aggregate: tuple       # (("outer_fold_metrics", LAST_OUTER), ("fold_metrics", LAST)) o ()

FAMILIES = {
    ("holdout", True): ...,   # hoy
    ("folds",   True): ...,   # hoy
    ("full",    False): ...,  # clustering
}

def resolve_family(task_cls, strategy_cls, splitter_cls) -> Family:
    # exige que strategy.KIND corresponda a splitter.PARTITIONING (holdout/holdout, cv/folds, full/full)
    # y levanta JobError antes de set_status_as_started si la combinación no existe.

ModelJob.run queda lineal para todas las familias: load, prepare, build, fit,
aggregate, evaluate, plots, save. Lo compartido no cambia (transiciones, progreso, textos
de JobError que los tests comparan, best_parameters, columnas de plots, run_path,
_aggregate_fold_metrics). La misma tabla es la plantilla de grafo por familia para el
motor DAG (test_model_job_as_a_graph.py): solo cambian los nombres de la unit de prepare y
de fit, y connect() cablea por intersección REQUIRES/PROVIDES.

Por qué no sub-jobs (comparación honesta)

Jerarquía TaskJob → SupervisedJob / UnsupervisedJob (+ ModelJob fachada) Tabla de familias en un ModelJob
Superficie de dispatch Nueva: from_request en jobs.py, dos jobs registrados, task_type nuevo en el widget de cola, dos lecturas del Run por job Ninguna: el front sigue mandando ModelJob
Forecasting ForecastingJob vacío; la jerarquía no le sirve Nada que pagar
Localidad Cada familia se lee de arriba abajo en su run() Tabla más run genérico: dos lugares
Extensión por plugin Un plugin registra su job Un plugin edita la tabla del core (o la task declara sus units, que es A con otra ropa)
Riesgo de regresión Aislado por familia run compartido: un refactor toca las tres
Costo hoy ~500 líneas movidas con --color-moved mientras las units siguen cambiando ~150 líneas, sin mover archivos, tests de orquestación y CV pasan sin cambios

Con la evidencia de hoy (una familia de cada tipo y forecasting absorbido por
declaraciones) la tabla es la codificación más liviana. Si aparece una tercera familia que
difiera en qué persiste (no solo en qué units corre), es el momento de convertir las filas
de la tabla en clases; la tabla no lo impide.

Fases y tamaños

  1. PR 875 (esta PR): nada de esto. Solo devolver 0.10.0 (hecho) y mergear. Opcional y
    pequeño, si el autor acepta: VALIDATION_DURING_FIT en las estrategias más su test
    (XS), porque cierra una fuga que las units de esta PR introducen.
  2. F1, forecasting re-registrado (S a M): los 4 puntos de la sección de forecasting.
  3. C1, clustering sobre las units (M a L, dos tercios son archivos de feat: add clustering support across models and notebooks #841 reutilizados):
    training_family.py con las dos filas actuales y la de clustering, las dos units nuevas,
    FullDatasetSplitter, FullDatasetEvaluationStrategy, SplitEnum.FULL, task/modelo/
    métricas de feat: add clustering support across models and notebooks #841, PreprocessingJob para full, front de feat: add clustering support across models and notebooks #841 re-encajado. Tests: run
    end to end con KMeans (tres filas FULL/LAST), DBSCAN todo ruido termina en error sin
    filas, un paso StandardScaler cambia el resultado de DBSCAN, auditoría de contratos y
    una variante del test de grafo con la familia full.
  4. C2, predicción de clustering (S a M, opcional): PredictClustersUnit y un camino
    de guardado que escriba una columna cluster en vez de output_columns[0]. Solo para
    algoritmos inductivos (KMeans, GMM); DBSCAN y Agglomerative devuelven las etiquetas del
    ajuste sea cual sea x.

Riesgos

  • Regla del registro: ningún ancestro nuevo puede llamarse Base* (por eso la PR usa
    *ScopeMixin). Vale para las units nuevas y para cualquier intermedio de jobs.
  • La tabla tiene huecos ((holdout, False), (folds, False), (full, True)): resolve_family
    falla con mensaje claro antes de STARTED, y COMPATIBLE_COMPONENTS de splitters y
    estrategias mantienen disjuntas las listas del wizard por task.
  • SplitEnum.FULL sin migración depende de que SQLAlchemy no cree CHECK sobre SQLite; un
    backend Postgres futuro necesitaría ALTER TYPE.
  • Las filas FULL no son comparables con VALIDATION de tasks supervisadas; el front no debe
    mezclarlas en una tabla (la comparación ya es por task).
  • PrepareFeaturesUnit depende de que el helper _apply_preprocessing(entries, names) de
    Sync reconcile/cv-units with develop and restore v0.10.0 behaviour #908 sea solo de x; si ese helper empaqueta y en las entradas, la unit necesita una variante.
  • ClusteringModel._score_split llama get_cluster_labels(x), que en algoritmos no
    inductivos devuelve las etiquetas del ajuste ignorando x. Correcto para la partición
    full; incorrecto si algún día se evalúa clustering sobre filas reservadas.

Felipedino and others added 2 commits September 30, 2026 08:57
PreprocessingJob split the dataset as loaded, while ModelJob splits what
task.prepare_for_task hands back (SplitterScopeMixin._prepare). The fitted
fold_{i}.pkl and final.pkl are applied to ModelJob's entries by position,
so the two sides must carve the same rows. A task that reorders its rows,
as forecasting does by sorting on the date, made each preprocessor fit on
rows that training uses as validation or test, and nothing noticed: the
entries still came in the same number.

PreprocessingJob now prepares the dataset for the task before splitting,
validating only the raw refs as the unit does, and resolves the task
before anything is fitted. For every registered task the prepared dataset
is the loaded one, so their partitions are unchanged.

The test registers a task that reverses the rows and checks, for holdout
and K-fold, that every preprocessor is fitted on the training rows of the
entry it is applied to.
…evelop

Sync reconcile/cv-units with develop and restore v0.10.0 behaviour

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants