Exploring the Perceptron Algorithm, Using Python
From theory to practice, here is everything that you need to know about this simple yet interesting and powerful method.

Ok, classic Machine Learning situation. You have a tabular dataset and you have to classify it. How do you do it?
Well, first thing first, you need to know very well the tools that you may use. A very well-known algorithm that you may try to use is the Perceptron.
From theory to practice, we will examine this Machine Learning method starting from a brief theoretical introduction and then showing a practical implementation.
At the end of this blog post, you will be able to understand when and how to use this Machine Learning algorithm, having a clear idea of all its pros and cons.
1. The Theory
1.1 Introduction
The perceptron has a biological reason to exist. Our neurons constantly receive energy from other neurons but they decide to be "activated" and emit their own signal only after the quantity of the energy they receive is greater or equal to a certain amount.
Let's start with the final product. At the very end, given a 4-dimensional input, this input is processed with 4 different weights, the sum gets into the activation function, and you get the result. Nothing more complicated than this. :)

Let's make it more clear. Imagine you have this table of features (columns) X1, X2, X3 and X4. These features are 4 different values that characterize a single instance (row) of your dataset.
This instance needs to be binary classified, so that you will have an additional value t, which is the target, that can be -1 or 1.
The Perceptron algorithm multiplies X1, X2, X3 and X4 by a set of 4 weights. For this reason, we consider the Perceptron to be a linear algorithm (more on this later).
Then, an activation function will be applied on the result of this multiplication (again, more about the activation function later).
Here is the whole process in an equation:

Where a is the so called activation function.
Of course, the input can be N dimensional (N does not have to be four) so that you may use N weights + 1 bias as well. Nonetheless, the pure Perceptron algorithm is meant to be used for binary classification (more on this later).
Of course, the result of y=a(w_1x_1+...+w_4x_4) needs to be between -1 and 1. In other words, at the end of the day, the so called activation function needs to be able to give you a classification.
So what is this so often discussed activation function? Well, nothing more than a step function. What does it mean?
The product of your N dimensional input with N dimensional weights will give you a single number. Then if this number is greater than 0 your algorithm will say "1", otherwise it will say "-1".

This is the final product. This is how it works, and this is how it takes the decision. Nothing really mysterious here :).
Let's move on.
1.2 Loss Function
We all know that Machine Learning algorithms come with a Loss Function. Well, in this case, the Loss function is nothing more but a weighted sum of the wrong classified points.
Let's make it easier. Let's say you have a point which is not well classified. It means that, for example, multiplying your parameters and your input you will get a final result of -0.87.

Ok, but the point is again, wrongly classified, remember? So it means that the target is indeed "1" for that point (t=1). So it means that if you do this multiplication:

You actually get a quantity that tells you how much you are wrong and you should change your weights and bias to do a better classification job.
In general, the loss function is the negative sum for all the wrongly classified points:

Where S is the set of the wrong classified points. The idea is that we will start optimizing this loss function, that of course we want to minimize.

The equation you see above is known as gradient descent. It means that we follow the direction where the loss goes to its minimum value and we update the parameters following this direction.
As the loss function is dependent on the number of wrongly classified points, it means that we will slowly start to correct the instances up to a point where, if the dataset is linearly separable (more on this later), there will be no more target to "correct" and our classification task will be just perfect. :)
2. The Implementation
Of course, the SkLearn Perceptron is a well known and ready implementation. Nonetheless, in order to understand it better, let's create this Perceptron from scratch.
Let's start with the libraries:
Let's define the decision function:
2.1 Linearly Separable Dataset
Let's create a linearly separable dataset using SkLearn.
2.2 The Perceptron Function
Using this function, all the ideas that have been explained before are actually implemented:
Then we can plot the decision boundaries using the following code:
So let's see what happens in our toy dataset:
As it is possible to see, all the points are well classified (even the small red triangle).
Let's see the Loss function plot:
It means that the dataset is perfectly classified now.
2.2 Non Linearly Separable Dataset
Let's consider a dataset which is harder to consider to be "linearly separable".
Let's run the algorithm:
Ok, right now we probably need a little bit of work to have our best classification.
Let's run different number of epochs and different learning rate (the so called hyperparameter tuning) to get the best version of the Perceptron:
So these are the optimal number of epochs and learning rate:
3. More Considerations
These are some things to consider:
The perceptron algorithm is fast. In fact, it is nothing but a linear multiplication + a step function application. It is super straightforward and easy to use.
The algorithm doesn't converge in terms of the loss function when the dataset is not linearly separable. It means that this perceptron is meant to (perfectly) work on linearly separable dataset only. Nonetheless, we can apply a transformation on the dataset and apply the perceptron algorithm on the transformed dataset
An hyperparameter tuning part could drastically increase the performance of the algorithm.
4. Conclusions
If you liked the article and you want to know more about Machine Learning, or you just want to ask me something you can:
A. Follow me on Linkedin, where I publish all my stories B. Subscribe to my newsletter. It will keep you updated about new stories and give you the chance to text me to receive all the corrections or doubts you may have. C. Become a referred member, so you won't have any "maximum number of stories for the month" and you can read whatever I (and thousands of other Machine Learning and Data Science top writer) write about the newest technology available.








