Python library to train and use Khiops decision-tree-driven models with a scikit-learn compatible API.
- Core API:
dtkhiops.coretrain_dtpredictortrain_dtrecoder
- Estimators:
dtkhiops.estimatorsDTKhiopsClassifierDTKhiopsRegressorDTKhiopsEncoder
- Ensemble methods:
dtkhiops.ensembleStackedDTKhiopsClassifierGradientBoostingKhiopsEncoderGradientBoostingSNBClassifier
From a local checkout:
pip install .From a Git repository:
pip install "git+https://github.com/KhiopsLab/decision-tree-khiops.git"Minimum Khiops dependency:
khiops>=11.0.1.0
Trains a Khiops predictor using a decision-tree recoding step, then launches predictor training on the enriched dictionary.
Main purpose:
- Build tree-based engineered variables from the target.
- Inject these variables into a dictionary used by Khiops predictor training.
- Produce both a report and a model dictionary.
Trains a Khiops recoder configured for decision-tree variable generation.
Main purpose:
- Generate tree-derived variables only (without full predictor training).
- Export a recoder model dictionary usable for downstream pipelines.
Note:
leaves_grouppingis the historical parameter name kept for backward compatibility (typo preserved intentionally).
Scikit-learn style classifier powered by Khiops.
Main behavior:
fit(X, y)trains a classification model.predict(X)returns predicted class labels.predict_proba(X)returns class probabilities.
Scikit-learn style regressor powered by Khiops.
Main behavior:
fit(X, y)trains a regression model.predict(X)returns numerical predictions.
Tree-based feature encoder / transformer.
Main behavior:
fit(X, y)learns tree-based encodings.transform(X)returns transformed features.fit_transform(X, y)fits and transforms in one step.
Multi-layer stacking classifier based on repeated DTKhiopsEncoder transformations.
Main behavior:
- Supports one or multiple stacking layers (
nb_stack). - Aggregates probabilities from all layers or only the last (
aggregation_mode). - Can optionally train a final
KhiopsClassifieron transformed data.
Boosting-style encoder that sequentially trains tree-based regressors and appends learned outputs as new features.
Classifier that combines boosted tree encodings with a Naive-Bayes-like final modeling strategy.
Use this when you already have a Khiops dictionary (.kdic) and a training table.
from dtkhiops import train_dtpredictor
report_path = "./artifacts/adult_predictor.khj"
report_output, model_output = train_dtpredictor(
dictionary_file_path_or_domain="/path/to/Adult.kdic", # Input dictionary
dictionary_name="Adult", # Dictionary name in .kdic
data_table_path="/path/to/Adult.txt", # Training table
target_variable="class", # Target column name
report_file_path=report_path, # Output report path
detect_format=False, # Explicit table format
header_line=True, # Input file has header
field_separator="\t", # TSV format
output_dir="./artifacts/predictor_run", # Artifacts/debug folder
# Tree-related controls
n_random_seed=1, # Reproducibility
n_depth_max=3, # Max depth, must be >= 2 if bounded
best_tree=True, # Best-tree split strategy
leaves_groupping=False, # Use tree leaves directly
# Predictor complexity controls
max_trees=10, # Max number of trees
max_constructed_variables=100, # Feature engineering budget
max_text_features=10000, # Text feature cap
max_pairs=0 # Pair features cap
)
print("Report:", report_output)
print("Model:", model_output)Key parameters explained:
dictionary_file_path_or_domain: path to the input.kdicor a loadedDictionaryDomain.dictionary_name: main dictionary to train from.data_table_path: training table path.target_variable: target column name.report_file_path: output report file.output_dir: optional run artifacts directory.n_depth_max: controls tree complexity;0means no explicit limit, and bounded values must be at least2because Khiops generates trees with at least 2 internal nodes.best_tree: toggles a best-tree split selection mode.max_trees: upper bound on generated tree variables.
| Parameter | Typical value | Why it matters |
|---|---|---|
dictionary_file_path_or_domain |
"/path/to/file.kdic" |
Defines the schema used by Khiops. |
dictionary_name |
"Adult" |
Selects which dictionary inside .kdic is used. |
data_table_path |
"/path/to/data.txt" |
Input training table. |
target_variable |
"class" |
Target column to predict. |
report_file_path |
"./artifacts/run.khj" |
Location of the generated report. |
output_dir |
"./artifacts/run" |
Stores intermediate/debug artifacts. |
n_depth_max |
3 |
Controls tree depth/complexity; bounded values must be >= 2. |
best_tree |
True |
Uses best-tree split search mode. |
max_trees |
10 |
Caps number of tree variables generated. |
Use this when you want a scikit-learn style workflow.
from sklearn.model_selection import train_test_split
from dtkhiops import DTKhiopsClassifier
# X: features DataFrame, y: target Series
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42, stratify=y
)
clf = DTKhiopsClassifier(
output_dir="./artifacts/classifier", # Folder for Khiops artifacts
n_trees=1, # Number of DT-based variables to inject
n_random_seed=1, # Reproducibility
n_depth_max=0, # Max depth (0 = no limit; bounded >= 2)
best_tree=True, # Best-tree split selection mode
n_features=100, # Feature construction budget
n_pairs=0 # Pair feature budget
)
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
y_proba = clf.predict_proba(X_test)Key parameters explained:
output_dir: where model artifacts and reports are written.n_trees: how many decision-tree-derived variables to include.n_depth_max: tree depth limit;0means no explicit limit, and1would only be a simple binarization rather than a generated decision tree.best_tree: enables best-tree selection strategy.n_features: maximum number of constructed features.n_pairs: number of pairwise constructed features.
| Parameter | Typical value | Why it matters |
|---|---|---|
output_dir |
"./artifacts/classifier" |
Where reports/models are written. |
n_trees |
1 to 5 |
Number of DT-derived variables used. |
n_random_seed |
1 |
Reproducibility of tree generation. |
n_depth_max |
0 or any value >= 2 |
Depth budget for generated trees; 1 is not a generated decision tree. |
best_tree |
True |
Activates best-tree split strategy. |
n_features |
100 |
Feature construction budget. |
n_pairs |
0 |
Pairwise feature construction budget. |
pip install -r requirements-dev.txt
pytest -qSee also:
examples/starter.pydocs/classes.mddocs/usage.md
The reports generated by dtkhiops can be opened and explored with the
Khiops Visualization application.
Khiops Visualization provides an interactive interface for inspecting Khiops reports, including:
- the global target distribution;
- the evaluated, informative, and selected variables;
- the distribution of variable values;
- the generated decision tree;
- the distribution of observations in each leaf;
- the target distribution for each tree leaf;
- the decision rules associated with each leaf;
- the tree structure and its relationships through an interactive graph;
- leaf population and purity information.
The following screenshot illustrates the visualization of a Khiops decision-tree report:
After training a model with dtkhiops, a Khiops report is generated. For
example:
from dtkhiops import train_dtpredictor
report_output, model_output = train_dtpredictor(
dictionary_file_path_or_domain="/path/to/Adult.kdic",
dictionary_name="Adult",
data_table_path="/path/to/Adult.txt",
target_variable="class",
report_file_path="./artifacts/adult_predictor.khj",
output_dir="./artifacts/adult_predictor",
)
print("Report:", report_output)
print("Model:", model_output)The generated report can then be opened with Khiops Visualization.
Depending on the selected model and report type, the report may contain an interactive representation of the MODL decision tree. Internal nodes represent segmentation rules, while leaves represent groups of observations with their associated target-class distributions.
The tree visualization allows users to:
- Navigate through the complete tree structure.
- Select an internal node or a leaf.
- Inspect the values and frequencies associated with the selected node.
- View the target-class distribution of a leaf.
- Examine the population and purity of a leaf.
- Display the rules leading from the root to the selected leaf.
- Explore the tree as a graph using the hyper-tree visualization.
- Adjust the displayed scales and visualization options.
The tree view is especially useful for understanding how the MODL model segments the input variables and how these segmentations contribute to the classification task.
Installation instructions for Khiops Visualization are available on the official Khiops website:
For more information about Khiops and its visualization tools, see the official Khiops documentation.
A decision tree is a classification model that predicts a categorical output variable from a set of numerical or categorical input variables. It is expressed as a hierarchical partition of the learning space, represented by connected nodes:
- An internal node has children and is defined by a segmentation rule.
- A leaf has no children and represents a decision by assigning the distribution of output values (typically the majority class) to every instance reaching it.
The following diagram illustrates the general structure of a decision tree, as described in Voisine et al. (2010):
flowchart TD
N1["Internal Node 1<br/>Variable X1<br/>I1 children"]
N2["Internal Node 2<br/>Variable X2<br/>I2 children"]
N3["Internal Node 3<br/>Variable X3<br/>I3 children"]
L4["Leaf 4<br/>Class distribution N4.j"]
L5["Leaf 5<br/>Class distribution N5.j"]
L6["Leaf 6<br/>Class distribution N6.j"]
L7["Leaf 7<br/>Class distribution N7.j"]
N1 --> N2
N1 --> N3
N2 --> L4
N2 --> L5
N3 --> L6
N3 --> L7
Building a decision tree is a difficult trade-off problem: a tree must be predictive enough while remaining as simple as possible, in order to generalize well to unseen data. Finding an optimal decision tree from a dataset is known to be NP-hard, which is why practical algorithms rely on heuristics (top-down pre-pruning or post-pruning strategies), rather than exhaustive search.
MODL (for Modeling Optimal Discovery from data) is a parameter-free Bayesian approach originally developed for supervised discretization and value grouping, and later extended to decision trees by Voisine, Boullé and Hue (2010).
Instead of relying on ad-hoc local splitting criteria (information gain, gain ratio, Gini index, chi-squared statistic, etc.) evaluated independently at each node, the MODL approach selects the tree with the highest posterior probability given the data, evaluated globally over the whole tree structure.
Formally, given:
Data: the training dataset,Tree: a candidate decision tree model,
MODL searches for the tree maximizing:
p(Tree | Data) ∝ p(Tree) × p(Data | Tree)
where p(Tree) is the prior probability of the tree model and
p(Data | Tree) is the likelihood of the data given the tree. This is the
Maximum A Posteriori (MAP) tree.
A MODL decision tree model T is entirely defined by:
- the subset of
K_Tinput variables used by the tree (amongKavailable variables), - for every node, its number of child nodes
I_s(I_s = 1for a leaf,I_s > 1for an internal node), - for every internal node
s: the segmentation variableX_s, the number of partsI_s(intervals for numerical variables, groups of values for categorical variables), and the distribution of instances across the child nodes{N_si.}, - for every leaf
l: the distribution of output values{N_l.j}.
The evaluation criterion is the negative logarithm of the posterior tree probability:
c(Tree) = -log( p(Tree) × p(Data | Tree) )
The prior p(Tree) is built hierarchically, describing the tree recursively
from the root to the leaves:
p(Tree) = p(K_T)
× Π_{s ∈ S_T} p(I_s) p(X_s | K_T) p(N_si. | K_T, X_s, N_s., I_s)
× Π_{l ∈ L_T} p(I_l) p(N_l.j | K_T, N_l.)
where S_T is the set of internal nodes and L_T the set of leaves of tree
T.
Putting all terms together, and using S_T^n / S_T^c for internal nodes
segmented on numerical / categorical variables respectively, the optimal
tree cost is given by:
C_opt(T) =
log(K + 1) + log( C(K + K_T - 1, K_T) )
+ Σ_{s ∈ S_T^n} [ log(K_T) + CR_is(I_s)·log(2)
+ log( C(N_s. + I_s - 1, I_s - 1) ) ]
+ Σ_{s ∈ S_T^c} [ log(K_T) + CR_is(I_s)·log(2)
+ log( B(V_Xs, I_s) ) ]
+ Σ_{l ∈ L_T} [ CR_is(1)·log(2)
+ log( C(N_l. + J - 1, J - 1) ) ]
+ Σ_{l ∈ L_T} log( N_l.! / (N_l.1! N_l.2! ... N_l.J!) )
Notation:
| Symbol | Meaning |
|---|---|
N |
Number of instances |
J |
Number of output classes |
K |
Number of available input variables |
K_T |
Number of variables actually used by the tree |
S_T |
Set of internal nodes of the tree |
L_T |
Set of leaves of the tree |
X_s |
Segmentation variable of internal node s |
I_s |
Number of child nodes of node s |
N_s. |
Number of instances in node s |
N_si. |
Number of instances in the i-th child of node s |
V_Xs |
Number of distinct values of categorical variable X_s |
N_l. |
Number of instances in leaf l |
N_l.j |
Number of instances of class j in leaf l |
C(n, k) |
Binomial coefficient "n choose k" |
B(X, Y) |
Number of ways to partition X values into Y groups (Bell-type number, sum of Stirling numbers of the second kind) |
CR_is(I) |
Universal (Rissanen) prior code length for encoding integer I |
Each term of the criterion has a direct interpretation:
-
Variable selection cost
log(K+1) + log(C(K+K_T-1, K_T))Penalizes the number of variables used by the tree and the specific subset chosen among all available variables. -
Internal node cost (numerical variables) Accounts for choosing the segmentation variable, the number of intervals
I_s(via the universal integer priorCR_is), and the distribution of instances across intervals — similar to the univariate MODL discretization criterion. -
Internal node cost (categorical variables) Same principle, but the segmentation corresponds to grouping categorical values into
I_sgroups, whose number of possible partitions isB(V_Xs, I_s). -
Leaf cost Accounts for the number of possible multinomial class distributions
C(N_l. + J - 1, J - 1)in each leaf. -
Likelihood term
Σ log( N_l.! / (N_l.1! ... N_l.J!) )measures how well the leaf class distributions fit the observed data. Using Stirling's approximation, this term is asymptotically equivalent to the target entropy in the tree leaves — connecting MODL to classical entropy-based impurity measures, while adding an explicit penalty for tree complexity.
The tree minimizing C_opt(T) is the selected MODL decision tree. Because
this criterion derives entirely from a Bayesian formulation, it is
parameter-free: there is no manually tuned pruning threshold or confidence
level, unlike C4.5 or CART.
Since finding the exact global optimum is NP-hard, two deterministic top-down
heuristics are used to search for a tree minimizing C_opt(T):
-
Pre-pruning: starting from the root, the algorithm searches, for every current leaf and every candidate variable, the best univariate MODL segmentation. A split is only applied if it strictly improves the global tree cost. The process stops when no split can further reduce the criterion. This is fast but can suffer from the horizon effect: a locally unpromising split may prevent discovering a more informative deeper structure.
-
Post-pruning: first builds the deepest possible tree by applying the best univariate MODL partition at every leaf, even without immediate improvement of the global criterion (the best tree encountered is memorized along the way). Then, starting from this deep tree, internal nodes whose children are all leaves are iteratively pruned back whenever this improves the global criterion. Post-pruning is guaranteed to find a tree at least as good as pre-pruning, at the cost of a longer training time.
Both binary trees (I_s ≤ 2) and N-ary trees (I_s > 2 allowed) can be built
with either strategy. Experiments on UCI and WCCI 2006 challenge datasets show
that MODL trees (in particular the binary, post-pruned variant) reach
predictive performance comparable to J48 (C4.5) and SimpleCART (CART), while
producing trees that are 2 to 4 times smaller on average.
The MODL decision trees included in Khiops are based on the foundational paper:
N. Voisine, M. Boullé, C. Hue. A Bayes Evaluation Criterion for Decision Trees. Advances in Knowledge Discovery and Management (AKDM-1), 292:21-38, 2010.
