Skip to content

About

Python library to train and use Khiops decision-tree-driven models with a core and scikit-learn compatible API.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

dtkhiops

Python library to train and use Khiops decision-tree-driven models with a scikit-learn compatible API.

What Is Included

  • Core API: dtkhiops.core
    • train_dtpredictor
    • train_dtrecoder
  • Estimators: dtkhiops.estimators
    • DTKhiopsClassifier
    • DTKhiopsRegressor
    • DTKhiopsEncoder
  • Ensemble methods: dtkhiops.ensemble
    • StackedDTKhiopsClassifier
    • GradientBoostingKhiopsEncoder
    • GradientBoostingSNBClassifier

Installation

From a local checkout:

pip install .

From a Git repository:

pip install "git+https://github.com/KhiopsLab/decision-tree-khiops.git"

Minimum Khiops dependency:

  • khiops>=11.0.1.0

Core API Reference

train_dtpredictor(...)

Trains a Khiops predictor using a decision-tree recoding step, then launches predictor training on the enriched dictionary.

Main purpose:

  • Build tree-based engineered variables from the target.
  • Inject these variables into a dictionary used by Khiops predictor training.
  • Produce both a report and a model dictionary.

train_dtrecoder(...)

Trains a Khiops recoder configured for decision-tree variable generation.

Main purpose:

  • Generate tree-derived variables only (without full predictor training).
  • Export a recoder model dictionary usable for downstream pipelines.

Note:

  • leaves_groupping is the historical parameter name kept for backward compatibility (typo preserved intentionally).

Class Reference

DTKhiopsClassifier

Scikit-learn style classifier powered by Khiops.

Main behavior:

  • fit(X, y) trains a classification model.
  • predict(X) returns predicted class labels.
  • predict_proba(X) returns class probabilities.

DTKhiopsRegressor

Scikit-learn style regressor powered by Khiops.

Main behavior:

  • fit(X, y) trains a regression model.
  • predict(X) returns numerical predictions.

DTKhiopsEncoder

Tree-based feature encoder / transformer.

Main behavior:

  • fit(X, y) learns tree-based encodings.
  • transform(X) returns transformed features.
  • fit_transform(X, y) fits and transforms in one step.

StackedDTKhiopsClassifier

Multi-layer stacking classifier based on repeated DTKhiopsEncoder transformations.

Main behavior:

  • Supports one or multiple stacking layers (nb_stack).
  • Aggregates probabilities from all layers or only the last (aggregation_mode).
  • Can optionally train a final KhiopsClassifier on transformed data.

GradientBoostingKhiopsEncoder

Boosting-style encoder that sequentially trains tree-based regressors and appends learned outputs as new features.

GradientBoostingSNBClassifier

Classifier that combines boosted tree encodings with a Naive-Bayes-like final modeling strategy.

Quickstart: train_dtpredictor

Use this when you already have a Khiops dictionary (.kdic) and a training table.

from dtkhiops import train_dtpredictor

report_path = "./artifacts/adult_predictor.khj"

report_output, model_output = train_dtpredictor(
    dictionary_file_path_or_domain="/path/to/Adult.kdic",  # Input dictionary
    dictionary_name="Adult",                               # Dictionary name in .kdic
    data_table_path="/path/to/Adult.txt",                  # Training table
    target_variable="class",                               # Target column name
    report_file_path=report_path,                           # Output report path
    detect_format=False,                                    # Explicit table format
    header_line=True,                                       # Input file has header
    field_separator="\t",                                  # TSV format
    output_dir="./artifacts/predictor_run",                # Artifacts/debug folder

    # Tree-related controls
    n_random_seed=1,                                        # Reproducibility
    n_depth_max=3,                                          # Max depth, must be >= 2 if bounded
    best_tree=True,                                         # Best-tree split strategy
    leaves_groupping=False,                                 # Use tree leaves directly

    # Predictor complexity controls
    max_trees=10,                                           # Max number of trees
    max_constructed_variables=100,                          # Feature engineering budget
    max_text_features=10000,                                # Text feature cap
    max_pairs=0                                             # Pair features cap
)

print("Report:", report_output)
print("Model:", model_output)

Key parameters explained:

  • dictionary_file_path_or_domain: path to the input .kdic or a loaded DictionaryDomain.
  • dictionary_name: main dictionary to train from.
  • data_table_path: training table path.
  • target_variable: target column name.
  • report_file_path: output report file.
  • output_dir: optional run artifacts directory.
  • n_depth_max: controls tree complexity; 0 means no explicit limit, and bounded values must be at least 2 because Khiops generates trees with at least 2 internal nodes.
  • best_tree: toggles a best-tree split selection mode.
  • max_trees: upper bound on generated tree variables.

Parameter Cheat Sheet (train_dtpredictor)

Parameter Typical value Why it matters
dictionary_file_path_or_domain "/path/to/file.kdic" Defines the schema used by Khiops.
dictionary_name "Adult" Selects which dictionary inside .kdic is used.
data_table_path "/path/to/data.txt" Input training table.
target_variable "class" Target column to predict.
report_file_path "./artifacts/run.khj" Location of the generated report.
output_dir "./artifacts/run" Stores intermediate/debug artifacts.
n_depth_max 3 Controls tree depth/complexity; bounded values must be >= 2.
best_tree True Uses best-tree split search mode.
max_trees 10 Caps number of tree variables generated.

Quickstart: DTKhiopsClassifier

Use this when you want a scikit-learn style workflow.

from sklearn.model_selection import train_test_split
from dtkhiops import DTKhiopsClassifier

# X: features DataFrame, y: target Series
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=42, stratify=y
)

clf = DTKhiopsClassifier(
    output_dir="./artifacts/classifier",  # Folder for Khiops artifacts
    n_trees=1,                             # Number of DT-based variables to inject
    n_random_seed=1,                       # Reproducibility
    n_depth_max=0,                         # Max depth (0 = no limit; bounded >= 2)
    best_tree=True,                        # Best-tree split selection mode
    n_features=100,                        # Feature construction budget
    n_pairs=0                              # Pair feature budget
)

clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
y_proba = clf.predict_proba(X_test)

Key parameters explained:

  • output_dir: where model artifacts and reports are written.
  • n_trees: how many decision-tree-derived variables to include.
  • n_depth_max: tree depth limit; 0 means no explicit limit, and 1 would only be a simple binarization rather than a generated decision tree.
  • best_tree: enables best-tree selection strategy.
  • n_features: maximum number of constructed features.
  • n_pairs: number of pairwise constructed features.

Parameter Cheat Sheet (DTKhiopsClassifier)

Parameter Typical value Why it matters
output_dir "./artifacts/classifier" Where reports/models are written.
n_trees 1 to 5 Number of DT-derived variables used.
n_random_seed 1 Reproducibility of tree generation.
n_depth_max 0 or any value >= 2 Depth budget for generated trees; 1 is not a generated decision tree.
best_tree True Activates best-tree split strategy.
n_features 100 Feature construction budget.
n_pairs 0 Pairwise feature construction budget.

Development

pip install -r requirements-dev.txt
pytest -q

See also:

  • examples/starter.py
  • docs/classes.md
  • docs/usage.md

Visualization with Khiops Visualization

The reports generated by dtkhiops can be opened and explored with the Khiops Visualization application.

Khiops Visualization provides an interactive interface for inspecting Khiops reports, including:

  • the global target distribution;
  • the evaluated, informative, and selected variables;
  • the distribution of variable values;
  • the generated decision tree;
  • the distribution of observations in each leaf;
  • the target distribution for each tree leaf;
  • the decision rules associated with each leaf;
  • the tree structure and its relationships through an interactive graph;
  • leaf population and purity information.

The following screenshot illustrates the visualization of a Khiops decision-tree report:

Khiops Visualization application showing a MODL decision tree

Opening a Khiops report

After training a model with dtkhiops, a Khiops report is generated. For example:

from dtkhiops import train_dtpredictor

report_output, model_output = train_dtpredictor(
    dictionary_file_path_or_domain="/path/to/Adult.kdic",
    dictionary_name="Adult",
    data_table_path="/path/to/Adult.txt",
    target_variable="class",
    report_file_path="./artifacts/adult_predictor.khj",
    output_dir="./artifacts/adult_predictor",
)

print("Report:", report_output)
print("Model:", model_output)

The generated report can then be opened with Khiops Visualization.

Depending on the selected model and report type, the report may contain an interactive representation of the MODL decision tree. Internal nodes represent segmentation rules, while leaves represent groups of observations with their associated target-class distributions.

Interactive tree exploration

The tree visualization allows users to:

  1. Navigate through the complete tree structure.
  2. Select an internal node or a leaf.
  3. Inspect the values and frequencies associated with the selected node.
  4. View the target-class distribution of a leaf.
  5. Examine the population and purity of a leaf.
  6. Display the rules leading from the root to the selected leaf.
  7. Explore the tree as a graph using the hyper-tree visualization.
  8. Adjust the displayed scales and visualization options.

The tree view is especially useful for understanding how the MODL model segments the input variables and how these segmentations contribute to the classification task.

Installing Khiops Visualization

Installation instructions for Khiops Visualization are available on the official Khiops website:

Install Khiops Visualization

For more information about Khiops and its visualization tools, see the official Khiops documentation.

Foundations: MODL Decision Trees

What is a decision tree?

A decision tree is a classification model that predicts a categorical output variable from a set of numerical or categorical input variables. It is expressed as a hierarchical partition of the learning space, represented by connected nodes:

  • An internal node has children and is defined by a segmentation rule.
  • A leaf has no children and represents a decision by assigning the distribution of output values (typically the majority class) to every instance reaching it.

The following diagram illustrates the general structure of a decision tree, as described in Voisine et al. (2010):

flowchart TD
    N1["Internal Node 1<br/>Variable X1<br/>I1 children"]
    N2["Internal Node 2<br/>Variable X2<br/>I2 children"]
    N3["Internal Node 3<br/>Variable X3<br/>I3 children"]
    L4["Leaf 4<br/>Class distribution N4.j"]
    L5["Leaf 5<br/>Class distribution N5.j"]
    L6["Leaf 6<br/>Class distribution N6.j"]
    L7["Leaf 7<br/>Class distribution N7.j"]

    N1 --> N2
    N1 --> N3
    N2 --> L4
    N2 --> L5
    N3 --> L6
    N3 --> L7
Loading

Building a decision tree is a difficult trade-off problem: a tree must be predictive enough while remaining as simple as possible, in order to generalize well to unseen data. Finding an optimal decision tree from a dataset is known to be NP-hard, which is why practical algorithms rely on heuristics (top-down pre-pruning or post-pruning strategies), rather than exhaustive search.


What is a MODL decision tree?

MODL (for Modeling Optimal Discovery from data) is a parameter-free Bayesian approach originally developed for supervised discretization and value grouping, and later extended to decision trees by Voisine, Boullé and Hue (2010).

Instead of relying on ad-hoc local splitting criteria (information gain, gain ratio, Gini index, chi-squared statistic, etc.) evaluated independently at each node, the MODL approach selects the tree with the highest posterior probability given the data, evaluated globally over the whole tree structure.

Formally, given:

  • Data: the training dataset,
  • Tree: a candidate decision tree model,

MODL searches for the tree maximizing:

p(Tree | Data) ∝ p(Tree) × p(Data | Tree)

where p(Tree) is the prior probability of the tree model and p(Data | Tree) is the likelihood of the data given the tree. This is the Maximum A Posteriori (MAP) tree.

A MODL decision tree model T is entirely defined by:

  • the subset of K_T input variables used by the tree (among K available variables),
  • for every node, its number of child nodes I_s (I_s = 1 for a leaf, I_s > 1 for an internal node),
  • for every internal node s: the segmentation variable X_s, the number of parts I_s (intervals for numerical variables, groups of values for categorical variables), and the distribution of instances across the child nodes {N_si.},
  • for every leaf l: the distribution of output values {N_l.j}.

The MODL decision tree criterion

The evaluation criterion is the negative logarithm of the posterior tree probability:

c(Tree) = -log( p(Tree) × p(Data | Tree) )

The prior p(Tree) is built hierarchically, describing the tree recursively from the root to the leaves:

p(Tree) = p(K_T)
         × Π_{s ∈ S_T} p(I_s) p(X_s | K_T) p(N_si. | K_T, X_s, N_s., I_s)
         × Π_{l ∈ L_T} p(I_l) p(N_l.j | K_T, N_l.)

where S_T is the set of internal nodes and L_T the set of leaves of tree T.

Putting all terms together, and using S_T^n / S_T^c for internal nodes segmented on numerical / categorical variables respectively, the optimal tree cost is given by:

C_opt(T) =
      log(K + 1) + log( C(K + K_T - 1, K_T) )

    + Σ_{s ∈ S_T^n} [ log(K_T) + CR_is(I_s)·log(2)
                       + log( C(N_s. + I_s - 1, I_s - 1) ) ]

    + Σ_{s ∈ S_T^c} [ log(K_T) + CR_is(I_s)·log(2)
                       + log( B(V_Xs, I_s) ) ]

    + Σ_{l ∈ L_T} [ CR_is(1)·log(2)
                       + log( C(N_l. + J - 1, J - 1) ) ]

    + Σ_{l ∈ L_T} log( N_l.! / (N_l.1! N_l.2! ... N_l.J!) )

Notation:

Symbol Meaning
N Number of instances
J Number of output classes
K Number of available input variables
K_T Number of variables actually used by the tree
S_T Set of internal nodes of the tree
L_T Set of leaves of the tree
X_s Segmentation variable of internal node s
I_s Number of child nodes of node s
N_s. Number of instances in node s
N_si. Number of instances in the i-th child of node s
V_Xs Number of distinct values of categorical variable X_s
N_l. Number of instances in leaf l
N_l.j Number of instances of class j in leaf l
C(n, k) Binomial coefficient "n choose k"
B(X, Y) Number of ways to partition X values into Y groups (Bell-type number, sum of Stirling numbers of the second kind)
CR_is(I) Universal (Rissanen) prior code length for encoding integer I

Each term of the criterion has a direct interpretation:

  1. Variable selection cost log(K+1) + log(C(K+K_T-1, K_T)) Penalizes the number of variables used by the tree and the specific subset chosen among all available variables.

  2. Internal node cost (numerical variables) Accounts for choosing the segmentation variable, the number of intervals I_s (via the universal integer prior CR_is), and the distribution of instances across intervals — similar to the univariate MODL discretization criterion.

  3. Internal node cost (categorical variables) Same principle, but the segmentation corresponds to grouping categorical values into I_s groups, whose number of possible partitions is B(V_Xs, I_s).

  4. Leaf cost Accounts for the number of possible multinomial class distributions C(N_l. + J - 1, J - 1) in each leaf.

  5. Likelihood term Σ log( N_l.! / (N_l.1! ... N_l.J!) ) measures how well the leaf class distributions fit the observed data. Using Stirling's approximation, this term is asymptotically equivalent to the target entropy in the tree leaves — connecting MODL to classical entropy-based impurity measures, while adding an explicit penalty for tree complexity.

The tree minimizing C_opt(T) is the selected MODL decision tree. Because this criterion derives entirely from a Bayesian formulation, it is parameter-free: there is no manually tuned pruning threshold or confidence level, unlike C4.5 or CART.


Optimization: pre-pruning and post-pruning

Since finding the exact global optimum is NP-hard, two deterministic top-down heuristics are used to search for a tree minimizing C_opt(T):

  • Pre-pruning: starting from the root, the algorithm searches, for every current leaf and every candidate variable, the best univariate MODL segmentation. A split is only applied if it strictly improves the global tree cost. The process stops when no split can further reduce the criterion. This is fast but can suffer from the horizon effect: a locally unpromising split may prevent discovering a more informative deeper structure.

  • Post-pruning: first builds the deepest possible tree by applying the best univariate MODL partition at every leaf, even without immediate improvement of the global criterion (the best tree encountered is memorized along the way). Then, starting from this deep tree, internal nodes whose children are all leaves are iteratively pruned back whenever this improves the global criterion. Post-pruning is guaranteed to find a tree at least as good as pre-pruning, at the cost of a longer training time.

Both binary trees (I_s ≤ 2) and N-ary trees (I_s > 2 allowed) can be built with either strategy. Experiments on UCI and WCCI 2006 challenge datasets show that MODL trees (in particular the binary, post-pruned variant) reach predictive performance comparable to J48 (C4.5) and SimpleCART (CART), while producing trees that are 2 to 4 times smaller on average.

Reference

The MODL decision trees included in Khiops are based on the foundational paper:

N. Voisine, M. Boullé, C. Hue. A Bayes Evaluation Criterion for Decision Trees. Advances in Knowledge Discovery and Management (AKDM-1), 292:21-38, 2010.

About

Python library to train and use Khiops decision-tree-driven models with a core and scikit-learn compatible API.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages