Shahrukh Khan Face Recognizer - Your 4th CNN
Fourth CNN?
Face Recognizer - Your 4th CNN

Fourth CNN?
This is my fourth post in the series of Do-it-yourself CNN models. As usual, the format will remain the same. I'll give you a Colab file (which works perfectly fine), which you have to run without any pre-conditions before moving forward. The reason being, once you run it yourself you become much more attached to the problem and you tend to get deeper insights.
You can check my previous posts in this series here:
In this post, we will be solving and understanding one of the most important problems of CNN - Face Recognizer. That is, given a photo of a person, find out its name. We will take a photo of Bollywood's most famous actor - Shah rukh khan and will be determining him.

So, without further ado, let me handover to you the Colab file 🚀 . This is not my work but has been copied from various places. The only thing I did here was to make sure it runs well in a Colab setup (which was surprisingly very difficult to do). But, nevertheless, it works just fine now, phew!
Feel free to upload your images and detect them! Now, the file can be a bit daunting to look at but is actually not so. As usual, let's start with the theory and will cover the implementation details later. Also, as per my previous articles will try to keep it short and sweet.
Face Detection Vs Verification vs Recognizer
First thing first, what do we actually mean by recognizer?
Face Detection - Its entirely similar to Object Detection. Detect all faces in an image and find their positions
Face Verification - Give two faces to determine whether they belong to the same person or not.
Face Recognizer -Recognize the name of the detected person.
Face Detection
This is exactly in lines of Object Detection where the aim is to detect specifically the face (rather than any random object). There are a few famous models out there which one can choose from. Will not go deeper into it since it's out of the scope of this blog.
Dlib: Works on Histogram of Oriented Gradients (HOG) and linear SVM. This can detect only front-facing faces though. Read more.
Haar Cascade: It's an optimization technique on CNN. All kernels aren't applied altogether. Instead, they are broken into multiple smaller groups. If nothing is detected in the first group it won't go further.
MTCNN: Its CNN with multiple stages. The first state detects the bounding boxes and in later stages removes the false positives and chooses the face's final bounding landmarks.

Face Recognizer
Now, once the face is detected, its entirely a different problem statement to recognize the person's name. Before moving forward, let me introduce a few concepts here.
Concept 1: Face Embedding
This is exactly in terms of word embedding. The only difference being it represents faces on a euclidean space.
Each face is converted to a vector which somehow represents all the characteristics of that face.
All such vectors are then plotted as points on the euclidean space.
The nearer the points, the more similar are the faces.
Concept 2: Siamese Network
Have you heard of Siamese twins? These are identical looking, conjoined twins. And so is our network.
So, a Siamese network is an architecture with two parallel neural networks, each taking a different input. Both individual outputs are then combined to provide a prediction of the similarity of the two images.
Siamese Network Training Techniques
Siamese networks can be used to predict the name of the person via two techniques based on the amount of training data we had.
Technique 1: One-Shot Learning
When to use this: When we don't have enough learning data.
How does it work:
Training data: 100 different images of 100 different people.
Scenario: Check if a picture matches any of the above 100 people.
Solution: Figure out the similarity score of the given picture with all the above 100 images.
Result: The one with the highest image similarity score among the 100 (above cutoff), will be marked as a match.
Technique 2: Triplet loss
When to use this: When we have enough learning data to train the network well.
How does it work:
Training data: 1000 different images of 100 different people.
Training Step: Take an anchor image and train it with a matching image and non-matching image. And calculate the loss. This loss is called triplet loss.
Scenario: Check if a picture matches any of the above 100 people.
Solution: Just like other CNNs, send it as an input to an already trained system and get the relevant output.

For implementing face detection and recognition, have taken the example of Shahrukh Khan. As you can see, how face-detection and recognition works and gives the result so accurately. Over here MTCNN based face-detection and Triplet Loss based Face recognizer have been used.
Feel free to upload pictures of your favorite actor and play around with it.








