INF264 – Project 1
Implementing Decision Trees
Mathias Fuglum Skjelvik
September 20, 2026
Contents
1 Implement decision tree learning algorithm from scratch
2
1.1
Node class . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2
1.2
Model construction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2
1.3
Entropy and Gini impurity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3
1.4
Information gain . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
1.5
Splitting dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
1.6
Selecting the best split . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5
1.7 Tree construction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6
1.8
Model fitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
1.9
Prediction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
1.10 Pruning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
8
2 Model selection
9
2.1
Data handling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
9
2.2
Best model selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
9
2.3
Model evaluation and comparing to Sklearn’s DecisionTreeClassifier . . . . . . .
10
3 Feature Importance
13
3.1
Permutation Importance Function . . . . . . . . . . . . . . . . . . . . . . . . . .
13
3.2
Evaluation of Importances . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
14
1
1
Implement decision tree learning algorithm from scratch
The first part of the project is implementing a classification decision tree from scratch in Python
using the ID3 learning algorithm. The purpose of the implementation learn to understand and
build machine learning models, instead of relying on complete implementations from a library
such as scikit-learn.
The model receives a feature matrix X and a label vector y. It divides the observations into
smaller and smaller groups by selecting a feature and using mean as a numerical treshold. Each
node asks a binary question of the form
xj ≤ t,
(1)
where xj is the value of feature j and t is the threshold stored in the node. If the condition
is true, the branch continues to the left. Else, it goes to the right branch. This continues until
the leaf node which contains a predicted class label.
The implementation accepts entropy and Gini impurity as splitting criteria. Tree growth is
controlled using maximum depth, minimum samples in a leaf, and minimum samples required to
split a node. The implementation is divided into a Node class and a DecisionTree class.
1.1
Node class
The tree is represented using connected Node objects. An node has a feature index, a threshold,
and references to its left and right children. A leaf node stores only a class label.
Listing 1: Node class.
1
2
3
4
5
6
7
8
class Node :
def __init__ ( self , feature = None , threshold = None ,
left = None , right = None , label = None ) :
self . feature = feature
self . threshold = threshold
self . left = left
self . right = right
self . label = label
If label is None, the node contains another decision. Otherwise, the node is a leaf. This
makes the same class fit for both decision nodes and leaves.
1.2
Model construction
The DecisionTree constructor stores the parameters controlling the learning algorithm:
• max_depth: the maximum permitted tree depth;
• min_samples_leaf: the minimum number of samples allowed in either child;
• min_samples_split: the minimum number of samples required to split a node;
• criterion: the impurity criterion, either entropy or Gini impurity.
2
The criterion is checked in the constructor.
Listing 2: Decision tree configuration.
1
2
3
4
5
6
def __init__ ( self , max_depth = None , min_samples_leaf =1 ,
mi n_sam ples_s plit =2 , criterion = " entropy " ) :
if criterion not in [ " entropy " , " gini " ]:
raise ValueError (
" Criterion must be either ’ entropy ’ or ’ gini ’. "
)
7
self . max_depth = max_depth
self . min_samples_leaf = min_samples_leaf
self . mi n_samp les_s plit = m in_sam ples_ split
self . criterion = criterion
self . root = None
8
9
10
11
12
Multiple hyperparameters were created to possibly create a more versitile model that give better
predictions.
1.3
Entropy and Gini impurity
To measure how mixed the class labels are in a node entropy and Gini impurity. Entropy is
calculated as
H(y) = −
K
∑︂
pk log2 (pk ),
(2)
k=1
while Gini impurity is calculated as
G(y) = 1 −
K
∑︂
p2k .
(3)
k=1
Here, pk is the proportion of labels belonging to class k. Both measures equal zero for a
pure node. The Counter class is used to count the labels and their amounts, and the impurity
method selects the calculation dependent by criterion.
Listing 3: Impurity calculations.
1
2
3
4
5
6
7
8
def entropy ( self , y ) :
counts = Counter ( y )
total = len ( y )
result = 0
for count in counts . values () :
p = count / total
result -= p * np . log2 ( p )
return result
9
10
11
12
def gini ( self , y ) :
counts = Counter ( y )
3
total = len ( y )
result = 1
for count in counts . values () :
p = count / total
result -= p ** 2
return result
13
14
15
16
17
18
19
20
21
22
23
24
def impurity ( self , y ) :
if self . criterion == " gini " :
return self . gini ( y )
return self . entropy ( y )
Keeping the criterion selection in one method allows the remaining algorithm to work with
either measure without duplicated code. By default the entropy method is used.
1.4
Information gain
Information gain measures how much a split will reduce impurity. A parent node P divided into
child nodes Ci , gives the gain by:
Gain(P ) = I(P ) −
∑︂ |Ci |
i
|P |
I(Ci ),
(4)
where I is the selected impurity measure. A higher value indicates that the split creates
purer child groups.
Listing 4: Information gain.
1
2
3
4
def info_gain ( self , parent_y , children_y ) :
parent_impurity = self . impurity ( parent_y )
total = len ( parent_y )
conditional_impurity = 0
5
for child_y in children_y :
p = len ( child_y ) / total
c o n d i t i o n a l _ i m p u r i t y += p * self . impurity ( child_y )
6
7
8
9
return parent_impurity - c o n d i t i o n a l _ i m p u r i t y
10
The method only receives label arrays because its responsibility is the mathematical calculation,
not creating the split. This was done to reduce complexity in the method and make it easier to
implement.
1.5
Splitting dataset
For each feature, the implementation uses its mean value in the current node as threshold. This
creates a binary split between left and right mathematically as:
4
L = {i | xij ≤ tj },
(5)
R = {i | xij > tj }.
(6)
NumPy Boolean masks select the corresponding rows from both X and y. This implementation
by code is shown below.
Listing 5: Mean-based binary split.
1
2
def split_dataset ( self , X , y , feature_index ) :
threshold = np . mean ( X [: , feature_index ])
3
left_mask = X [: , feature_index ] <= threshold
right_mask = X [: , feature_index ] > threshold
4
5
6
return {
" threshold " : threshold ,
" left " : { " X " : X [ left_mask ] , " y " : y [ left_mask ]} ,
" right " : { " X " : X [ right_mask ] , " y " : y [ right_mask ]}
}
7
8
9
10
11
The subsets are needed to calculate the information gain, while both the feature and target
subsets are required when recursively constructing the child branches. Categorical features are
treated as ordered numerical values by this implementation.
1.6
Selecting the best split
The best_split method evaluates the features in the current node. Each features get splitted
and checked that both children satisfy min_samples_leaf.
Listing 6: Best-split selection.
1
2
3
4
def best_split ( self , X , y ) :
best_feature = None
best_gain = -1
best_split = None
5
6
7
8
9
10
11
12
for feature_index in range ( X . shape [1]) :
split = self . split_dataset (X , y , feature_index )
left_child = split [ " left " ][ " y " ]
right_child = split [ " right " ][ " y " ]
left_size = len ( left_child )
right_size = len ( right_child )
inf_gain = self . info_gain (y , [ left_child , right_child ])
13
14
15
16
17
if self . min_samples_leaf is not None :
if ( left_size < self . min_samples_leaf or
right_size < self . min_samples_leaf ) :
continue
5
18
gain = self . info_gain (y , [ left_y , right_y ])
19
20
if gain > best_gain :
best_feature = feature_index
best_gain = gain
best_split = split
21
22
23
24
25
return best_feature , best_gain , best_split
26
Lastly the gain is updated if it gets larger, and returns the best split to build_tree later.
1.7
Tree construction
The build_tree method is the main learning algorithm. It is called one time for each node. The
method checks if a node should become a leaf and the building stops when:
• all labels in the node are identical;
• the node has fewer observations than min_samples_split;
• the maximum depth has been reached;
• no split satisfies min_samples_leaf;
• the best split has no positive information gain.
When a node must become a leaf, the largest size of class is used as its label. Else, the
method builds a left and right branch.
Listing 7: Recursive tree construction.
1
2
3
def build_tree ( self , X , y , depth =0) :
if len ( y ) == 0:
return None
4
5
largest_class = Counter ( y ) . most_common (1) [0][0]
6
7
8
if len ( np . unique ( y ) ) == 1:
return Node ( label = y [0])
9
10
11
if len ( y ) < self . m in_sam ples_s plit :
return Node ( label = largest_class )
12
13
14
if self . max_depth is not None and depth >= self . max_depth :
return Node ( label = largest_class )
15
16
feature , gain , split = self . best_split (X , y )
17
18
19
if split is None or gain <= 0:
return Node ( label = largest_class )
6
20
left_branch = self . build_tree (
split [ " left " ][ " X " ] , split [ " left " ][ " y " ] , depth + 1
)
right_branch = self . build_tree (
split [ " right " ][ " X " ] , split [ " right " ][ " y " ] , depth + 1
)
21
22
23
24
25
26
27
return Node (
feature = feature ,
threshold = split [ " threshold " ] ,
left = left_branch ,
right = right_branch
)
28
29
30
31
32
33
Recursion is fitting because every child branch is itself a smaller decision tree.
1.8
Model fitting
The fit method checks that the training target isn’t empty and starts the tree construction.
The root gets returned as a node.
Listing 8: Fitting the tree.
1
2
3
4
5
def fit ( self , X , y ) :
if len ( y ) == 0:
raise ValueError (
" Cannot fit DecisionTree on an empty dataset . "
)
6
self . root = self . build_tree (X , y )
return self
7
8
1.9
Prediction
To predict a class, the model starts at the root and follows the already created decisions. A value
less than or equal to the threshold follows the left branch, whilst a larger value follows the right
branch. The model stops at a leaf, where the label is not None and returns it.
Listing 9: Prediction by tree traversal.
1
2
3
4
5
6
7
8
def predict ( self , X ) :
if self . root is None :
raise ValueError (
" Decision Tree has not been trained yet , use ’. fit (X , y ) ’
before predicting ! "
)
predictions = []
for x in X :
node = self . root
7
while node . label is None :
if x [ node . feature ] <= node . threshold :
node = node . left
else :
node = node . right
predictions . append ( node . label )
9
10
11
12
13
14
15
return np . array ( predictions )
16
The prediction gets restarted at the root for every value in y.
1.10
Pruning
Lastly in the Decision Tree - class, a prune method is added. This is a way to remove unnecessary
branches after the tree has been created, called post-pruning. Doing this ensures that useful
branches don’t get removed early in pre-pruning.
Listing 10: Main logic of pruning found in prune_node()-method
1
2
s ub t r ee _ p re d i ct i o ns = np . array ([ self . predic t_from _node (x , node )
for x in X_val ])
subtree_correct = np . sum ( s u bt r e e_ p r ed i c ti o n s == y_val )
3
4
5
largest_label = Counter ( y_val ) . most_common (1) [0][0]
leaf_correct = np . sum ( largest_label == y_val )
6
7
8
9
10
11
12
if leaf_correct >= subtree_correct :
node . feature = None
node . threshold = None
node . left = None
node . right = None
node . label = largest_label
Now both pre-pruning and post-pruning is incorporated to the decision tree. Pre-pruning comes
from the hyperparameters min_samples_leaf and min_samples_split, whilst post-pruning is
done with the prune() method. Pre-pruning can save time and stop the tree from overfitting by
becoming too large. Post-pruning can remove noise after the tree has been created. Using both
pre- and post-pruning makes them both more effective.
In my particular case post-pruning was not done because it would require the data to be split
into a separate dataset, to avoid data-leakage by pruning on the same data being used for model
validation. The reason for this is to save data for training.
8
2
Model selection
2.1
Data handling
Data retrieving was done exactly as suggested in INF264_project1.pdf, which contains the project
task and description. The hold-out validation method was used in the model selection. The
data was splitted as follows:
• 70% training data
• 15% validation data
• 15% test data
Listing 11: Splitting data
X_train , X_test_val , y_train , y_test_val = train_test_split (X , y ,
test_size =0.3 , random_state =42)
X_val , X_test , y_val , y_test = train_test_split ( X_test_val ,
y_test_val , test_size =0.5 , random_state =42)
1
2
Most of the data goes to training instead of validation and testing. This makes it possible for
the model to find more patterns and possibly become more complex. Validation data makes it
possible to test many hyperparameters and prevent overfitting before it reaches the test data.
Having completely unseen test data that is not part of the training process shows how the model
generalizes.
2.2
Best model selection
Finding the best model was done by testing multiple values for every hyperparameter in the
decision tree.
Listing 12: Hyperparameter values
1
2
3
4
depths = [3 , 5 , 7 , 10 , None ]
min_splits = [2 , 10 , 50 , 100]
min_leafs = [1 , 10 , 50 , 100]
criteria = [ " entropy " , " gini " ]
Only a few values were used per hyperparameter as it would quickly become time consuming to
test thousands of combinations. In total 160 combinations were tested.
Listing 13: Main logic for parameter testing
1
2
3
4
for criterion in criteria :
for depth in depths :
for min_split in min_splits :
for min_leaf in min_leafs :
5
6
tree = DecisionTree (
9
max_depth = depth ,
m in_sam ples_s plit = min_split ,
min_samples_leaf = min_leaf ,
criterion = criterion
7
8
9
10
)
11
12
tree . fit ( X_train , y_train )
prediction = tree . predict ( X_val )
13
14
Listing 13 shows how the best parameter search was done in practice. Testing every parameter
is important to give more/less complexity to the trained model and finding the one that fits the
best. The hyperparameters and F1-score from the best model was as follows:
• criterion: ’gini’
• max_depth: 7
• min_samples_split: 100
• min_samples_leaf: 10
To evaluate the best model, a macro-averaged F1-score was used to individually calculate the
F1-score for each class and average the result. Choosing this performance measure ensures the
models value each class equally and doesn’t hide flaws for a smaller class. The given formula
shows how the score is calculated for this exact dataset with only two classes.
F 1macro =
F 10 + F 11
2
The heavy class imbalance of 5163/1869 is some of the reason for this method. This will still
punish models only guessing one single class, whilst acknowledging a model predicting well for
every class. I also value both classes as much with this method. Only F1-score is going to be
used every time when training, evaluation and comparing models to ensure a fair comparison.
A normal F1-score could also have been used if the main goal is finding out which customers
churn. This is because a model guessing only no-churn would have 0 in F1-score because of 0
true positives.
To ensure reproducibility, random_state is always given the value 42 for every run. Splitting
the data in 70/15/15 ratio and ensuring no data leakage as well as using F1-score on the skewed
contributes with ensuring a fair model selection. The test data was only used once to ensure no
bias and not to change model to get a better generalization ability.
2.3
Model evaluation and comparing to Sklearn’s DecisionTreeClassifier
When we have decided which model is going to be used to generalize with test data, we can
combine both training and validation data to give the model more data to train on and increase
learning possibility. Afterwards the model predicts X_test once and returns F1-score.
10
Listing 14: Generalizing the DecisionTree model
1
2
3
4
5
def d e c i s i o n _ t r e e _ g e n e r a l i z a t i o n ( best_criterion , best_max_depth ,
best_min_samples_split , b e s t _ m i n _ s a m p l e s _ l e a f ) :
tree = DecisionTree ( max_depth = best_max_depth ,
min_samples_leaf = best_min_samples_leaf ,
mi n_sam ples_s plit = best_min_samples_split ,
criterion = best_criterion )
6
7
8
tree . fit ( X_full_train , y_full_train )
prediction = tree . predict ( X_test )
9
10
11
f1score = f1_score ( y_test , prediction , average = " macro " )
print ( " F1 - score on test data = " , round ( f1score ,4) )
12
13
return f1score , prediction , y_test , tree
Baseline models are necessary to verify the learning of our own model. The DummyClassifier
is therefore implemented with strategy ’most_frequent’. For the given churn-data.csv, this means
that the baseline model always predicts 0 (no-churn) as its the most common class. As seen
further below the model didn’t score 0 as it would have with a normal F1-score as opposed to
Macro F1 Score.
Traning and model selection for sklearn’s DecisionTreeClassifier is the exact same as for
DecisionTree except one added criteria, log-loss.
Below shows the comparison of the best selected parameters for DecisionTreeClassifier
and DecisionTree. To ensure reproducibility, random_state is given the value 42 for the sklearn
model.
Table 1: Best model hyperparameters
DecisionTree DecisionTreeClassifier
criterion
gini
entropy
max_depth
7
5
max_samples_split
100
2
max_samples_leaf
10
100
The differences in hyperparameters are not too large. The DecisionTree has a bit deeper tree
but DecisionTreeClassifier has much smaller requirement for sample size to allow splitting
and could therefore have many branches and more complexity. Next, F1-score will be used
evaluate under/overfitting and model complexity vs error.
11
Table 2: Summary of F1-scores for every model
Model
F1-score (average="macro")
DecisionTree (Validation)
0.74
DecisionTree (Generalization)
0.71
DecisionTreeClassifier (Validation)
0.72
DecisionTreeClassifier (Generalization)
0.66
DummyClassifier (Baseline)
0.43
Both models seem to have been learning from the churn_data.csv dataset and gets a
much better score than the dummy model. The DecisionTreeClassifier has a bigger drop
in F1-score which reflects possibly larger overfitting to the data compared to the other model.
Something that would be expected from the given hyperparameters in table 1.
Looking at a confusion matrix gives more insight in how the two models perform in each
class, and why they have their given F1-score.
Figure 1: Confusion matrix of self created DecisionTree
12
Figure 2: Confusion matrix of Sklearn’s DecisionTreeClassifier
DecisionTreeClassifier is better at predicting the larger class 0 (no-churn) than the
DecisionTree. Whilst the opposite applies to class 1 (churn). The probability of the models
predicting churn given that the customer churns are quite different.
P (ChurnDecisionT reeClassif ier |ChurnActual ) =
P (ChurnDecisionT ree |ChurnActual ) =
111
≈ 41%
111 + 163
154
≈ 56%
154 + 120
Which model is best is dependent on what class is the most important to predict correctly. If
predicting no-churn is the most valuable, choosing the DecisionTreeClassifier is the best
option. Predicting if a customer churns seems to be the most hard to predict for both models
and could mean that the features don’t give enough information and that it is mostly random.
3
Feature Importance
To find out what features are the most important in a decision tree, a feature importance algorithm can be implemented. This is useful to find out which features the model are using making
accurate predictions. Furthermore, we will go through how this algorithm was implemented,
what features are most important and discuss the weaknesses of the algorithm.
3.1
Permutation Importance Function
The function, permutation_importance, implements the permutation importance algorithm
into the main run_experiments.py file. A code snippet from the algorithm is shown below.
13
Listing 15: Main logic in permutation_importance
for j in range ( n_features ) :
permuted_score = []
for k in range ( n_repeats ) :
org_col = X_arr [: , j ]. copy ()
1
2
3
4
5
randomize . shuffle ( X_arr [: , j ])
6
7
y_pred = model . predict ( X_arr )
score = metric ( y_arr , y_pred )
permuted_score . append ( score )
8
9
10
11
X_arr [: , j ] = org_col
12
13
avg_score = np . mean ( permuted_score )
importances [ j ] = baseline_score - avg_score
return importances
14
15
16
Here the randomized shuffle lets the user input a fixed seed to ensure reproducibility. Score is
saved n_repeats times and averaged before finding one single importance value. The process is
repeated for every feature until we get the importance of every n_features.
3.2
Evaluation of Importances
Matplotlib is utilized to visualize the importance values for the features found by permutation_importance().
The parameters used in the function are the best fitted DecisionTree, test data, F1-score and a
seed of 42. Figure 3. shows the importance in increasing order from the bottom.
14
Figure 3: Feature importance of DecisionTree model
From the figure we see the top 5 most important features:
• Contract
• MonthlyCharges
• InternetService
• tenure
• OnlineSecurity
This is quite as expected. The type of contract a customer has is going to be very important
in deciding churn. A month-to-month contract is going to be much more likely to churn than
year-to-year. How long a customer has been at the company, meaning tenure, is also going to
be extremely in prediction. Newer customers are definitely more likely of churning than loyal
customers who’s been at the company for many years.
Feature importance scores tell us how much the accuracy decreases when we randomly
shuffling the values of a column. By doing this we’re breaking the relationship between the
specific feature and the label. If the feature is completely unrelated and random in relation to
the label, shuffling will most likely neither increase nor decrease the accuracy.
One weakness with this algorithm can be correlating features. Consider to features correlated
1:1, shuffling only one of the features at a time does not create less model accuracy because we get
15
the exact same information from the other correlated feature. Because of this dependency, both
the features could get negligible levels of importance even though they are both very important.
Another weakness of the algorithm is getting very unrealistic data. By shuffling we could
get feature combinations which don’t exist in reality, which again give importance scores that
doesn’t reflect the real importance.
16
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )