Data Science

AdaBoost Machine Learning Algorithm: How to Improve Performance of Hard to Predict Cases

An intuitive explanation of the Adaptive Boosting algorithm and its difference from other Decision Tree based Machine Learning algorithms

Saul Dobilas
February 21, 202110 min read

Machine Learning

Adaptive Boosting (AdaBoost). Image by author.
Adaptive Boosting (AdaBoost). Image by author.

Intro

The number of Machine Learning algorithms keeps increasing with time. If you want to be a successful Data Scientist, you must understand how they differ from one another.

This story is part of a series where I provide an in-depth look at different algorithms, how they work, and how to build them in Python.

The story covers the following topics:

  • The category of algorithms AdaBoost belongs to

  • Visual comparison of model predictions between a Single Decision Tree, Random Forest, and AdaBoost.

  • Explanation of how AdaBoost differs from other algorithms

  • Python code examples

What category of algorithms does AdaBoost belong to?

AdaBoost, same as its other tree-based algorithm "friends," belongs to the supervised branch of Machine Learning. While it can be used for both classification and regression problems, I focus on the classification side in this story.

Side note, I have put Neural Networks in a category of their own due to their unique approach to Machine Learning. However, they can be used to solve a wide range of problems, including but not limited to classification and regression. The below chart is interactive so make sure to click👇 on different categories to enlarge and reveal more.

If you enjoy Data Science and Machine Learning, please subscribe to get an email whenever I publish a new story.

Decision Tree vs. Random Forest vs. AdaBoost

Let us start with a comparison of the prediction probability surface made by each of the three models. They all use the same Australian weather data:

  • Target (a.k.a dependent variable): 'Rain Tomorrow'. Possible values: 1 (yes, it rains) and 0 (no, it does not rain);

  • Features (a.k.a. independent variables): 'Humidity at 3 pm' today, 'Wind Gust Speed' today. Note, we use only two features to enable us to visualize the results easily.

Here are the model prediction planes. Note that you can use Python code in the last section of this story if you would like to replicate these graphs.

1.Decision Tree (1 tree, max_depth=3); 2.Random Forest (500 trees, max_depth=3); 3.AdaBoost (50 trees, max_depth=1). Images by author.
1.Decision Tree (1 tree, max_depth=3); 2.Random Forest (500 trees, max_depth=3); 3.AdaBoost (50 trees, max_depth=1). Images by author.

Interpretation

Let's interpret visualizations and see what they tell us about the differences in these algorithms. The z-axis in the above graphs is the probability of rain tomorrow. Meanwhile, the thin white line is the decision boundary, i.e., where the probability of rain = 0.5.

We will now quickly review CART and Random Forest before diving into AdaBoost.

1. Single Decision Tree (CART)

CART is an acronym for a standard decision tree algorithm that stands for Classification and Regression Trees. In this example, we have built a single decision tree, which goes 3 levels deep and has 8 leaves. Here is the exact tree for reference:

CART decision tree. Image by author.
CART decision tree. Image by author.

Every leaf within this tree gives us a probability of rain tomorrow, which is a ratio between the number of cases where it rained and the total number of observations inside that leaf. E.g., the bottom left leaf gives us a probability of rain tomorrow of 6% (2,806 / 46,684).

Each leaf corresponds to a flat surface in the 3D prediction graph, while the splits in the tree are the step changes. See the below illustration.

CART decision tree and its corresponding predictions. Image by author.
CART decision tree and its corresponding predictions. Image by author.

As you can see, this is a pretty simple tree that gives us only 8 different probabilities. 5 of those leaves lead to a prediction of no rain tomorrow (probability < 0.5), while the remainder 3 leaves suggest it will rain tomorrow (probability > 0.5).

If you would like to go deeper into the mechanics of the CART algorithm, you can refer to my earlier story here:

CART: Classification and Regression Trees for Clean but Powerful Models

2. Random Forest

The first thing to note is that in our case, Random Forest uses the same CART algorithm for its base estimator. However, there are a few major differences:

  • It builds many random trees and combines predictions from each individual tree to generate the final prediction.

  • It uses bootstrapping (sampling with replacement) to create many samples from the original data. These samples maintain the same size but have different distributions of observations.

  • Finally, it uses feature randomness to minimize the correlation between the trees. This is done by only making a random subset of features available to the algorithm at each node split.

In the end, Random Forest creates many trees (500 in our example) and calculates the overall probability based on the predictions by each tree. This is why the prediction plane surface is smoother compared to a single tree (i.e., it has many little steps instead of a few big steps).

Random Forest prediction plane. Image by author.
Random Forest prediction plane. Image by author.

You can find more details on Random Forests in my separate story here:

Random Forest Models: Why Are They Better Than Single Decision Trees?

3. AdaBoost

Finally, we arrive at the main topic of this story.

Like Random Forest, we use CART as a base estimator inside the Adaptive Boosting algorithm. However, AdaBoost can also use other estimators if required.

The core principle of AdaBoost is to fit a sequence of weak learners, such as decision stumps, on repeatedly modified versions of data. A decision stump is a decision tree that is only one level deep, i.e., it consists of only a root node and two (or more) leaves. Here is the example of 3 separate decision stumps from our AdaBoost model:

Three decision stumps. Image by author.
Three decision stumps. Image by author.

Similar to Radom Forest, the predictions from all weak learners (in this case, stumps) are combined through a weighted majority vote to produce the final prediction. The major difference, though, is how these weak learners are generated.

Boosting iterations consist of applying weights to each of the training samples (observations). Initially, those weights are equal across all observations, so the first step trains a weak learner on the original data.

Linking this to our example, it means that the first decision stump (first split) is the same as what you get using a single decision tree approach:

Single Decision Tree vs. the first decision stump in AdaBoost. Image by author.
Single Decision Tree vs. the first decision stump in AdaBoost. Image by author.

For each successive iteration, the sample weights are individually modified, and the learning algorithm is reapplied to the reweighted data. Those training examples that were incorrectly predicted at the previous step have their weights increased. Meanwhile, the ones that were correctly predicted have their weights decreased. Hence, each subsequent weak learner is thereby forced to concentrate on the examples missed by the previous ones.

In the end, combining all of the weak learners into a final prediction generates a much "flatter" distribution of predictions. This is because the algorithm deliberately reduced the weights on the most confident examples and shifted focus towards harder to classify ones. As a result, we have the model predictions displayed in the graph below.

AdaBoost prediction plane. Image by author.
AdaBoost prediction plane. Image by author.

Note, the more weak learners (stumps) you add, the "flatter" the prediction distribution becomes.

Prediction distribution

Another way to visualize the prediction distribution is with a simple frequency line chart. As expected,

  • The single decision tree has few spaced-out predictions.

  • Random forest shows a much more even distribution of predictions.

  • AdaBoost has all of its predictions located close to the decision boundary (at 0.5).

Performance

While the three approaches produced very different probability distributions, the final classification results were very similar. The performance can be evaluated in many ways, but for the sake of simplicity, I am only showing accuracy (on test sample) here:

  • Single Decision Tree: 82.843%

  • Random Forest: 83.202%

  • AdaBoost: 83.033%

While the performance can be slightly improved with some additional hyperparameter optimization, the similarity of the results above tells us that we are fairly close to extracting the maximum information contained within the features used.

AdaBoost Limitation

The resulting "flat" probability distribution of AdaBoost is its main limitation. Depending on your use case, it may not be an issue for you. Say if you only care about assigning the correct class, then the prediction probability is less relevant.

However, if you care more about the probability itself, you may want to use Random Forest, which provides you with probability predictions such as 9% or 78%, as shown in the rain prediction modeling above. This is in contrast to AdaBoost, where all predictions are much closer to 50%.

Python section

Now that we know how AdaBoost works and understand its differences from other tree-based modeling approaches, let's build a model.

Setup

We will use the following data and libraries:

Let's import all the libraries:

python
import pandas as pd # for data manipulationimport numpy as np # for data manipulationfrom sklearn.model_selection import train_test_split # for splitting the data into train and test samplesfrom sklearn.metrics import classification_report # for model evaluation metricsfrom sklearn.ensemble import AdaBoostClassifier # for AdaBoost modelfrom sklearn.tree import DecisionTreeClassifier # to use as base estimatorimport plotly.express as px  # for data visualizationimport plotly.graph_objects as go # for data visualization

Then we get the Australian weather data from Kaggle, which you can download following this link: https://www.kaggle.com/jsphyg/weather-dataset-rattle-package.

We ingest the data and derive a few new variables for usage in the models.

python
# Set Pandas options to display more columnspd.options.display.max_columns=50# Read in the weather data csvdf=pd.read_csv('weatherAUS.csv', encoding='utf-8')# Drop records where target RainTomorrow=NaNdf=df[pd.isnull(df['RainTomorrow'])==False]# For other columns with missing values, fill them in with column meandf=df.fillna(df.mean())# Create a flag for RainToday and RainTomorrow, note RainTomorrowFlag will be our target variabledf['RainTodayFlag']=df['RainToday'].apply(lambda x: 1 if x=='Yes' else 0)df['RainTomorrowFlag']=df['RainTomorrow'].apply(lambda x: 1 if x=='Yes' else 0)# Show a snaphsot of datadf
A snippet of Kaggle&#039;s Australian weather data with some modifications. Image by author.
A snippet of Kaggle's Australian weather data with some modifications. Image by author.

Next, let's build a model by following these steps:

  • Step 1 - select model features (independent variables) and model target (dependent variable)

  • Step 2 - split data into train and test samples

  • Step 3 - set model parameters and train (fit) the model

  • Step 4 - predict class labels on train and test data using our model

  • Step 5 - generate model summary statistics

python
##### Step 1 - Select data for modelingX=df[['WindGustSpeed', 'Humidity3pm']]y=df['RainTomorrowFlag'].values##### Step 2 - Create training and testing samplesX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)##### Step 3# Set model and its parametersmodel = AdaBoostClassifier(base_estimator=DecisionTreeClassifier(min_samples_leaf=1000, max_depth=1),                           n_estimators=50, # default=50                           learning_rate=1.0, # default=1.                            algorithm='SAMME', # SAMME' - discreate, 'SAMME.R' - real                           random_state=0, # random state for reproducibility                          )# Fit the modelclf = model.fit(X_train, y_train)##### Step 4# Predict class labels on training datapred_labels_tr = model.predict(X_train)# Predict class labels on a test datapred_labels_te = model.predict(X_test)##### Step 5 - Model summary# Basic info about the modelprint('*************** Tree Summary ***************')print('No. of classes: ', clf.n_classes_)print('Classes: ', clf.classes_)print('No. of Estimators: ', len(clf.estimators_))print('Base Estimator: ', clf.base_estimator_)print('--------------------------------------------------------')print("")print('*************** Evaluation on Test Data ***************')score_te = model.score(X_test, y_test)print('Accuracy Score: ', score_te)# Look at classification report to evaluate the modelprint(classification_report(y_test, pred_labels_te))print('--------------------------------------------------------')print("")print('*************** Evaluation on Training Data ***************')score_tr = model.score(X_train, y_train)print('Accuracy Score: ', score_tr)# Look at classification report to evaluate the modelprint(classification_report(y_train, pred_labels_tr))print('--------------------------------------------------------')

The above code generates the following output that summarizes the model performance.

AdaBoost model performance. Image by author.
AdaBoost model performance. Image by author.

The model generalized quite well with similar performance on test data compared to training data. The number of estimators was 50, meaning the final model is comprised of 50 different decision stumps.

Finally, as promised, here is the code to generate the 3D prediction plane graph:

python
def Plot_3D(X, X_test, y_test, clf, x1, x2, mesh_size, margin):    # Specify a size of the mesh to be used    mesh_size=mesh_size    margin=margin    # Create a mesh grid on which we will run our model    x_min, x_max = X.iloc[:, 0].fillna(X.mean()).min() - margin, X.iloc[:, 0].fillna(X.mean()).max() + margin    y_min, y_max = X.iloc[:, 1].fillna(X.mean()).min() - margin, X.iloc[:, 1].fillna(X.mean()).max() + margin    xrange = np.arange(x_min, x_max, mesh_size)    yrange = np.arange(y_min, y_max, mesh_size)    xx, yy = np.meshgrid(xrange, yrange)    # Calculate predictions on grid    Z = clf.predict_proba(np.c_[xx.ravel(), yy.ravel()])[:, 1]    Z = Z.reshape(xx.shape)    # Create a 3D scatter plot with predictions    #fig = px.scatter_3d(x=X_test[x1], y=X_test[x2], z=y_test,    fig = px.scatter_3d(x=[], y=[], z=[],                     opacity=0.8, color_discrete_sequence=['black'])    # Set figure title and colors    fig.update_layout(#title_text="Scatter 3D Plot with Adaboost Prediction Surface",                      paper_bgcolor = 'white',                      scene = dict(xaxis=dict(title=x1,                                              backgroundcolor='white',                                              color='black',                                              gridcolor='#f0f0f0'),                                   yaxis=dict(title=x2,                                              backgroundcolor='white',                                              color='black',                                              gridcolor='#f0f0f0'                                              ),                                   zaxis=dict(title='Probability of Rain Tomorrow',                                              backgroundcolor='lightgrey',                                              color='black',                                               gridcolor='#f0f0f0',                                              range=[0, 1],                                              tickmode = 'linear',                                              tick0 = 0,                                              dtick = 0.2                                              )))    # Update marker size    fig.update_traces(marker=dict(size=1))    # Add prediction plane    fig.add_traces(go.Surface(x=xrange, y=yrange, z=Z, name='AdaBoost Prediction',                              colorscale='Bluered',                              reversescale=True,                              showscale=False,                               contours = {"z": {"show": True, "start": 0.5, "end": 0.9,                                                 "size": 0.5, "color":"white"}}))    fig.show()    return fig# Call the above fucntion to generate a graphfig = Plot_3D(X, X_test, y_test, clf, x1='WindGustSpeed', x2='Humidity3pm', mesh_size=1, margin=1)

Conclusion

I sincerely hope that this story will help you choose the right algorithm for your use case. Thanks for reading, and feel free to use the above code and materials in your own Data Science projects.

Cheers! 👏 Saul Dobilas


Related stories you may like:

Gradient Boosted Trees for Classification - One of the Best Machine Learning Algorithms

XGBoost: Extreme Gradient Boosting - How to Improve on Regular Gradient Boosting?

Related Articles