This episode explores two powerful shallow learning algorithms: support vector machines (SVMs) and naive Bayes classifiers. SVMs create a wide-margin decision boundary between classes to prevent overfitting, using support vectors as key data points. The kernel trick allows them to handle non-linear data by transforming it into higher dimensions. Naive Bayes, rooted in Bayesian probability and conditional logic, is fast and memory-efficient, particularly effective in text classification like spam detection, due to its assumption of feature independence. Both algorithms excel with categorical data and handle missing values well. Choosing between them depends on practical constraints: data size, memory, speed, and the nature of input features. A structured approach—narrowing options based on domain knowledge, evaluating data characteristics, and testing multiple algorithms—helps determine the best fit. While deep learning models offer greater flexibility, SVMs and naive Bayes remain valuable for specific use cases where speed, simplicity, and efficiency are critical. The episode emphasizes that no single algorithm is universally best; instead, a systematic evaluation and experimentation process is essential for successful model selection. Tyler also highlights resources like Andrew Ng’s course and the scikit-learn decision tree to guide algorithm selection.
Welcome back to Machine Learning Guide. I'm your host, Tyler Renelli. MLG teaches
the fundamentals of machine learning and artificial intelligence. It covers
intuition, models, math, languages, frameworks, and more. Where your other
machine learning resources provide the trees, I provide the forest. Visual is the
best primary learning modality, but audio is a great supplement during exercise
commute and chores. Consider MLG or syllabus with highly curated resources for
each episode's details at ocdevelop.com/mlg. Speaking of curation, I'm a curator of
life hacks. My favorite hack being treadmill desks. While you study machine
learning or work on your machine learning projects, walk. This helps improve
focus by increasing blood flow and endorphins. This maintains consistency and
energy, alertness, focus, and mood. Get your CDC recommended 10,000 steps
while studying or working. I get about 20,000 steps per day, walking just two
miles per hour, which is sustainable without instability at the mouse or
keyboard. Save time and money on your fitness goals. See a link to my favorite
walking desk setup in the show notes. This is episode 13. Shallow Learning
Algorithms Part 2. Support vector machines and naive bays. In this episode, I'm
going to be talking about support vector machines and the naive bays classifier.
These are two very powerful machine learning techniques, shallow learning algorithms.
These are kind of those power machine learning algorithms, power tools. I
consider decision trees, support vector machines, and naive bays. I consider them all
three to be sort of power tools. A lot of the other shallow learning algorithms that
you'll learn are sort of dedicated to particular tasks, or even if they're
multipurpose, they shine under specific circumstances. But decision trees, support
vector machines and naive bays are sort of these power tools that can be applied
across a very wide spectrum of machine learning applications. They're primarily built
for classification. All three of these algorithms are primarily built for
classification, but can be used for regression. Before we get too far into this
episode, I want to talk about the fact that we're talking about so many
machine learning algorithms. In the last episode, I dropped a bunch on you. In this
episode, we're going to be talking about two more. And in the next episode, I
actually decided to stretch the two parter into a three parter. In the next
episode, I'm going to be talking about even more machine learning algorithms.
If there's so many machine learning algorithms, how are you supposed to decide
what to use when? Well, there's a multi-part approach to deciding which
machine learning algorithm to use given your circumstances. We've talked in the
past that sort of deep learning can be seen as a silver bullet that can be used
across a wide spectrum of machine learning problems. But right now, we're taking a
diversion. We're talking about shallow learning algorithms. And with so many
shallow learning algorithms, you have to know specifically which algorithms are
supposed to be used under which circumstances. So this kind of this multi-tier
approach to deciding which algorithm to use under your circumstance.
At a top level, certain machine learning algorithms only handle specific tasks.
So, for example, in the next episode, I'm going to talk about an
anomaly detection algorithm. Well, you only use one or a handful of algorithms
for anomaly detection. You don't use things like linear regression or
logistic regression or any of the algorithms that we're going to be talking
about in this episode. So there is a level of domain knowledge that will
filter down the types of algorithms that you're going to be using for your
specific purposes. It's apples to oranges. This situation only calls for a
handful of machine learning algorithms that can be applied whatsoever. And that
just takes familiarity with a lot of the machine learning algorithms and what
they're specifically built for. I think listening to this podcast series and
taking the Android course and some other follow-up material, you'll start
to get a feel for what algorithms are built for what purposes. But then we go
down a level. We've decided that we're going to be working in supervised
learning classification. And so that only includes a handful of specific
algorithms. We've excluded a whole bunch of algorithms, like all the
unsupervised learning algorithms, any regression algorithms, etc. We're not
going to use linear regression for this. We're not going to use k-means. We're
going to do classification. But now we have a whole bunch of classification
algorithms to work with. So in the last episode I talked about decision
trees. In this episode, I'm going to talk about support vector machines and naive
bays. All of these are classifier algorithms. We also have logistic regression.
We have neural networks. How the heck am I going to choose from amongst all
these classifiers? Well, at this point, what a lot of people in machine learning
do is they look at their data and they look at their environment, their system,
the computer, how much RAM does it have, and how flexible are we with time,
as far as running the algorithms, making the inferences and predictions.
Certain algorithms work better with certain types of data. Maybe logistic
regression works really well with numerical input, but it's also very
sensitive to missing features and such like this.
Now you've bays works better with categorical input, but is not sensitive
at all to missing data. Now you've bays classifiers are very fast to run
and very memory efficient, but maybe a little bit less precise than
something like a neural network. A neural network takes a lot longer to run
and train, but it's very precise and can represent very complex models.
So once we know what category of machine learning algorithms we're going to be
using for us specific purposes, then we consider the situation at hand.
Memory restrictions, time restrictions, what does our data look like?
How many samples do we have? Certain machine learning algorithms work
really well with only a handful of samples while other machine learning
algorithms require tons and tons of examples, your training data set.
Neural networks, for example, are vexed by the fact that they need lots
and lots and lots of data, whereas a naive bays classifier, for example,
doesn't need that much data to get up and running. So you look at your data,
you plot it, you chart it, you graph it. You decide if you're missing any
information. Is it numerical? Is it categorical?
How much memory do I have on my machine? How much time do I have to work with?
How many examples do I have? And then you decide. So that part is a little bit tough.
It's tough to explain in podcast format. I'm going to link in the resources
section to a table of pros and cons for specific algorithms,
given situations like memory constraints, time constraints, etc.
There's also a decision tree put out by Scikit-learn,
a Python library for shallow learning machine learning algorithms.
They have a decision tree, a picture, like a flow chart for helping you decide
which algorithm to use, given the circumstances. It'll ask you some yes-no
questions. Things like, do I have greater than 50,000 examples in my training data?
Yes, go this way. If not, go that way. Okay, is this text-based? Go this way.
Am I missing any data? Go that way. And it'll help narrow down what
machine learning algorithm you're supposed to use. And the reason I'm explaining
all this is that it can seem so overwhelming at first when I just throw a million
machine learning algorithms at you. And you're thinking, I have all these algorithms.
How am I supposed to know what to use when and where? And there's a system.
There's a system to deciding what to use. And the final approach,
finally, once we've decided a handful of algorithms that we can use,
given the circumstances, is to actually just try them all.
You'll often see a lot of machine learning engineers, what they'll do is they'll
import from scikit-learn or TensorFlow, just all the algorithms that can possibly
be applied to their circumstance. They'll clean up the data, they'll visualize the data,
they'll do some stuff with the training data, they'll split it into training,
validation and test set will get into that stuff in another episode.
And then they'll just throw all their data through 10 machine learning algorithms.
Run them all, parallel or serial, whatever. And then at the very end of the file,
you'll see they'll evaluate the performance of all the machine learning algorithms.
They'll write some code that determines how well each algorithm did,
compare them all to each other and find the champion or champions, throw away the losers,
and maybe keep the top three on hand as they continue in their programming.
Eventually, they'll sort of come to a conclusion that one is clearly the winner.
This is the right algorithm for the job we're going to roll with that.
So it's not necessarily very clear what algorithm to use when.
It's like a three-part approach.
We start at the top where we decide what algorithms are even grossly applicable.
Okay, is this a supervised or an unsupervised learning situation?
Do I have the labels?
Okay, if it's supervised, is it regression or classification?
Okay, it's classification.
Now we have maybe 30 algorithms that we can choose from.
Let's plot our data.
Let's look at the situation, let's look at the environment.
What kind of constraints are we up against?
And at this point, it's a little bit difficult to memorize which algorithms work best,
given constraints and data.
What you usually do then is you look at this reference table
or that scikit-learn flow chart to help you pick a handful of algorithms to try.
And then you just throw them all against the wall.
You just shock an approach, all these algorithms,
and the evaluation metric at the end of your script will tell you which one did best.
So you're going to learn a lot of machine learning algorithms
and use this approach to determine which algorithm to use, given your circumstance.
With that out of the way, let's jump into the first of these two algorithms called support vector machines,
SVM.
It's a very weird word.
I'll tell you why it's called that in a bit.
But let's just try to understand an intuition of what it does.
So like I said, these power tool machine learning algorithms,
like decision tree support vector machines and naive bays,
they can all be used for both classification and regression.
So they're supervised learning machine learning algorithms that can be used
both for classification and regression.
Also, neural networks can be used for classification and regression as well.
You'll find that the primary use case of all of these algorithms is classification.
I don't know if this is true or not, but it seems to me from my experience
that classification is kind of the majority use case of machine learning that you'll see in the wild.
I don't know if this is true, don't quote me on it.
But you'll see that these machine learning algorithms, these power tools,
are primarily built for classification but can be used for regression.
But because their primary use case is classification,
you'll see the examples or the
tutorials, they'll all be showing you how to use them for classification.
And that's what I will be doing in this episode.
And then you'll have to look up how to use them for regression on your own.
So support vector machines can be used for classification and regression.
When you use a support vector machine for classification,
it's called a support vector classifier SVC.
And if you use it for regression,
it's called a support vector regressor SVR.
And the broad category of these is called support vector machines.
So we're going to go with the classification examples, like I mentioned.
How it works is it determines a decision boundary,
a decision boundary between your things over here and your things over there.
That sounds a lot like logistic regression.
It's very similar to logistic regression.
But it's got some perks over logistic regression that we'll give you to in a minute.
But let's remember what a decision boundary is.
Let's say that you have all the cats on the left and all the dogs on the right.
You have a graph of cats and dogs.
They're just dots on a graph, right?
Imagine blue dots and red dots.
These are your data points. These are your training examples.
You have all the cats on the left and all the dogs on the right.
And what you want to do is come up with a line.
You're going to draw a line in the sand between the cats and the dogs.
So no dogs allowed over here on the left, say the cats.
So that line that separates your cats from your dogs is called your decision boundary.
And now if you add a new animal into the mix,
based on some features about the animal, whether it has whiskers, does it bark?
How many lives does it have, etc?
These are all the features will be used to determine where,
locationally, the object gets placed in 3D or 4D space.
And if it's on the right side of the line, then it's a dog.
And if it's on the left side of the line, then it's a cat.
That's your decision boundary.
And that looks a lot like a logistic regression situation.
Now, what makes a support vector machine different from logistic regression
in categorizing things over here and over there is this decision boundary specifically.
A support vector machine, it doesn't use a line.
It doesn't draw a hair thin line between the two sets, like logistic regression does.
Instead, it tries to make that line as fat as it possibly can.
It makes a wall.
It doesn't use a one point line.
It uses a 16 point brush stroke.
How fat is this wall?
Well, the borders of the wall bump up against the innermost cats and dogs.
So, the rightmost edge of this decision boundary is going to bump up against the leftmost dogs
and the leftmost edge of the decision boundary will bump up against the rightmost cats.
Okay, so that makes sense.
It's just, you're just trying to fill a river or make a wall between the two things
these over here and those over there, as wide as you can before you touch them.
So, you do that with your training set of examples.
Now, why did we do that?
Why was logistic regression insufficient?
Why wasn't that line sufficient?
Why did we need a fat line?
Well, the reason is because of a problem that we're going to get into in a future episode
called overfitting.
Overfitting is basically, if I were to draw the line between the cats and the dogs,
I could draw it wrong, actually.
Let's say that I had a cat closer to the middle.
Well, if I wasn't smart, I might draw the line to accommodate that cat, what's called an outlier.
Something that doesn't really fit the bill of the majority of the data.
I might skew the line, maybe I'll tilt the line counterclockwise or clockwise a little bit
to accommodate for that one outlying cat.
In other words, I didn't make the most ideal line possible.
You and me were humans.
We look at a cluster of dots over here and a cluster of dots over there.
And in our minds, we can draw a vertical line right down the center.
Even if there is an outlying dot, we still have an intuition of where that line goes.
Where's the best line that separates the two classes so that in the future, if I were
to add a new object into the mix, it will go on the correct side.
But logistic regression is a little bit sensitive to outliers and things like this.
And this can cause a line that gets improperly drawn and this is called overfitting.
In an extreme example of overfitting, imagine a line that goes right up vertical and then
it squiggles out half circle to include that outlying cat and then keeps going.
Imagine that we created some wild function of polynomials that allowed that little squiggle
out.
That's an wild example of overfitting, but in our particular situation where we're using
logistic regression, which is a linear function, we only can work with a line, not a polynomial
function, just tilting the line, maybe counterclockwise or clockwise might cause some overfitting.
So what support vector machines do different than logistic regression in the case of coming
up with a decision boundary between the classes on the left and the class on the right is
it makes that decision boundary as fat as possible so that we can deal with these outliers
no problem.
Now the thickness of our line, it's called the margin.
We want as fat a line as possible, we call this a large margin classifier, large margin.
So that's the word for the thickness of the line is margin.
And then the word for the dots that are being bumped up against by this line that we're
drawing, they're called support vectors.
That's why this thing is called a support vector machine.
Support vectors.
It's kind of a weird word.
I don't know why we don't call this.
I think we should call this algorithm the fat line algorithm and we should call these
dots that the fat line bumps up against.
We should call them bumping dots.
But no, we call the algorithm a support vector machine or a large margin classifier and
that these dots that the fat line bump up against is, they're called support vectors.
A vector, so a dot on a Euclidean graph xy plane, this is what we're looking at, there's
a bunch of dots on a graph.
You can think of them as a dot or a point or you can think of them as an arrow pointing
from the origin to that dot.
You can graph that arrow with a function mathematically.
So you can represent these dots in another way and we call that a vector.
A vector is a line that points from the origin to a dot.
And that's why they're called support vectors.
They're the vectors that support drawing a fat line.
Okay, all finding good support vector machines seem pretty simple.
It's like logistic regression with a fat line instead of a skinny line.
That's the only difference, right?
Well, it's got one more little twist and this is where things start to get wild really
weird if you ask me support vector machines only handle linear classification.
So does logistic regression now you can throw some polynomials into the function and make
the situation non-linear.
But that's a little bit less than ideal, typically we typically if our situation is non-linear,
we move away from the linear classifiers into something more complex like a neural network
for example.
So both logistic regression and support vector machines are linear classifiers, but there's
a trick, a trick that can transform a support vector machine into a non-linear classifier.
And this trick is called the kernel trick, kernel, K, E, R, and E, L. You'll see kernels
used quite commonly in machine learning, they're very weird, they're very hard to understand.
I still haven't quite wrapped my head around them, but what I think of as a kernel, it teleports
you into another dimension or these rose colored goggles that you put on and they change
the way everything looks.
Okay.
So that sounds very strange.
Let me give you an example.
An example used from the machine learning with R book that I'll post in the resources section
is that if you're looking at a graph of dots, some are blue and some are red, okay.
And this is what it looks like.
You have a blue circle of dots in the center and surrounding that is a red circle of dots.
It's like a blue circle with a red border, but they're all dots, okay.
It's not drawn onto the graph, it's a bunch of dots.
So that is clearly a non-linear situation.
This isn't a bunch of cats on the left and dogs on the right, which is linearly separable
by our decision boundary.
No, this is a circle and a circle.
Those are not separable by lines.
However, if those dots represented something conceptually, for example, if we were looking
at latitude and longitude and those dots represented, say, snow on peaks of mountains versus
non-snow, latitude and longitude, well, that wouldn't really make sense as a way of looking
at this.
Would it?
What we really care about is altitude, how high up the mountain peak is and latitude,
how far north and south we are.
Those are the two characteristics that are more important rather than longitude.
Longitude doesn't help us at all.
So if we think about the problem different, we can actually transform our situation into
a new graph where dots are indeed linearly separable.
All the blue dots suddenly have become sort of a rectangle on the left or the right or
top or bottom.
Some sort of situation where we can actually draw a line between the two classes of dots,
okay.
So that's a little bit weird.
Let me think of another way of representing this.
If you have two circles, blue and red dots, maybe you can think of them instead of in
a Euclidean space of X and Y, you can think of them in a radian way.
So this is what a kernel does.
A kernel, what it does is
As it takes your data, the stuff you're looking at right now, which is non-linearly separable.
And it transforms it into a new set of dimensions.
So you're looking at the latitude and longitude representation of mountains in the world and
trying to decide whether or not to have snow on the peaks.
And you're scratching your chin and you're like, how am I going to separate this?
But I grab your hand and I'm like, no, no, no, come over here, come over here.
And I pull you around so you're looking at it from a different angle.
And you go, aha, okay, looking at it from this angle, things seem a little bit different.
So a kernel is something that you multiply your data by in order to transform it into a new
dimension.
So I think of it as it's looking at the problem from a different angle.
I think of it as like in Zelda, a link to the past.
I can't remember.
You blow, you blow on a flute or you do some mirror trick.
And it goes, and you're now in the dark world.
You do it again, and you're in the light world.
You're in the same place.
The whole world is really the same place.
Everything is the same, but you're looking at it different.
And it helps you to solve different puzzles.
So you can transform a circle world into a line world.
There's a whole bunch of kernels out there.
There's like radial basis function kernel, polynomial kernel, sigmoid kernel.
There's a whole bunch of kernels, a whole bunch of colored goggles that you can put on.
Imagine a drawer full of colored goggles that you could put on at any time.
But you have to know a little bit about your situation, yet to know whether the data
that you're dealing with is sort of could be transformed into a different world so that
it's easier to work with so that it is now linearly separable.
So support vector machines, very strange machine learning algorithm.
I really took me a while to kind of wrap my head around it, and I still don't know exactly
when it's preferred to be used under certain circumstances or not.
So let's hit it from the top one more time.
Let's reference prior algorithms that we've used.
Remember linear regression is a regression algorithm for coming up with a number output.
Okay, if we want to classify something in the past, we piped linear regression into
a new function called logistic regression.
Logistic regression is like using linear regression to classify things.
Is it a cat or a dog?
Is it go on the left of the line or the right of the line?
Now conceptually, logistic regression draws a line down the middle.
We call this the decision boundary, this line that separates the cats from the dogs.
Now with logistic regression, unfortunately, this line may be prone to overfitting based
on outliers.
If there's a lot of data that's bad, bad data or just noise or anything like this, it
could kind of screw up our line.
It might tilt it, tilt it down counterclockwise or clockwise may not be the best fit, the ideal
line straight down the center, separating the cats from the dogs.
So we have this new algorithm for classifying things called a support vector machine.
And it uses a decision boundary as well.
But it makes that decision boundary as fat as possible, a large margin.
It bumps up against the innermost dots on the left and the right classes.
We call those innermost dots support vectors.
And that large margin helps us prevent overfitting future examples.
That's step one of a support vector machine.
It is simply maybe a little bit more accurate, a little bit more efficient version of logistic
regression you might consider it.
Step two is this strange trick of the trade called the kernel trick.
And the kernel trick lets you take your data, which may be represented non-linearly if
you look at it like this, transform it by putting on some goggles into a new dimension.
And now suddenly it is linear.
So you take a non-linear data set, look at it a different way, and now it's linearly
separable.
Cool.
So a support vector machine is a classifier or can be used for regression.
And it has the ability to represent non-linear circumstances.
Now the problem is, like I said, you have this drawer of kernels, goggles that help you
look at situations from different angles.
Well, there's only so many of these, you know, like circle world or radial basis world.
There's only so many ways to represent a non-linear data set in a linear fashion.
And you have to know which one to use given the circumstance, unlike a neural network,
which is able to represent non-linear situations completely on its own.
It will learn the way to represent them non-linearly.
It can represent any number of complex situations.
So support vector machines and neural networks are often compared to each other because
they're both these black box methods and they can both handle non-linear situations.
But the difference is that a neural network in deep learning is more powerful.
It can represent more non-linear circumstances.
And you as the developer don't have to know in what way is this situation non-linear?
The neural network will learn that mapping for you.
Whereas with a support vector machine, you have to know sort of in what way is this circumstance
non-linear?
You don't necessarily have to know in advance.
You could just try throwing at it all the kernels in your drawer.
But if you don't have that sort of upfront information, it might be better to use a
neural network anyway.
So why wouldn't you use a neural network?
Well, if your situation can be handled with a linear support vector machine, okay, vanilla
support vector machine, or you do know about the situation and you can pop in one of those
kernels into your support vector machine, then support vector machines are a lot faster
than neural networks.
They're faster and they take up less memory.
And in fact, you're gonna see that this is a very common recurring theme in machine learning.
Like I said previously, machine learning engineers, they look at people who use deep learning
as a silver bullet for every situation and they say, you could do this faster with a dedicated
shallow learning algorithm for specific situations that call for the shallow learning algorithms.
So if your situation supports using a support vector machine, then you will get a lot more
speed and memory savings using that over a neural network.
But you'll have to know a little bit about your data set or your circumstances in advance
to help you determine whether using a support vector machine is for you or not.
So I kind of like to think of machine learning algorithms as you have this backpack, like
an enrol playing game, you have this backpack of tools that you can use.
You have a grappling hook for certain circumstances.
You have your sword and shield.
You have a magic wand, and so if you're presented with a puzzle, so you need to kill a bad guy,
you will use the sword and shield.
You need to open the entrance to a cave, use a bomb.
Well, neural networks and deep learning, they're kind of like a bazooka.
You can almost solve any situation with the bazooka, but maybe it's overkill and expensive
and can cause collateral damage.
So kill a bad guy, bazooka, open a treasure chest, bazooka, open a cave entrance, bazooka.
But why not use the cheap bomb in the case of the cave entrance?
Why not just use your hands and a key when it comes to opening a treasure chest?
The way I think of support vector machines is like a gun, and it kind of looks like a
space gun, like a plastic ray gun.
And it's kind of, for me, it's a little bit tough to know when to use this thing.
And you're trying to figure out what's the best tool for opening a door that's locked.
You can use a key.
You can use a bomb.
You can use your bazooka, or you can use this weird plastic ray gun.
And the proper approach is to try all of them, try all of them, and evaluate the performance
of all of them at the end of your script, determine which did the best, which took the least
amount of memory, the least amount of time, was the most accurate model, et cetera.
And it just so happens that it turns out a key in this particular case opened the door
the best.
Logistic regression handled situation A the best.
But I always think of support vector machines as this weird ray gun, and you point it at
the door, and you shoot, and the laser comes out, and nothing happens.
And you're like, huh, you turn it around in your hands and somebody behind you says,
oh, well, you're not using the radial basis kernel.
Of course, that's why it's an open.
So he hands you this little module, and you look at it, it's the radial basis kernel.
And you're like, uh, and you clip it into your ray gun, and you pointed at the door, and
you shoot and outcome these sonar circles, woo, woo, woo, and the door opens.
And he's like, see?
It was obvious.
And he's like, was it?
Support vector machines.
Now let's move on to naïve Bayes classifiers, naïve Bayes.
Bayesian inference is a very interesting and important component of machine learning
in general.
In fact, Bayesian inference really is a rung of the ladder of statistics.
And like I told you in a previous episode, statistics is the god math of machine learning.
Statistics is everything in machine learning.
The very basic principles of statistics like probability, joint probability, conditional
probability, et cetera, are used everywhere in machine learning, even if you don't know
it.
Many of the machine learning algorithms that we've been discussing so far, there are algorithms
that come straight out of a statistics textbook, linear and logistic regression.
That's statistics.
Statistics is really essentially boiled down into probability, probability and inference.
That inference is based off of probability.
And probability is sort of raw statistics.
So the algorithms that we've been learning so far are probability and raw statistics deep
down inside, deep down under the hood, they're just statistics.
the high level, the way that we've been looking at them, they kind of look like machine learning algorithms,
they kind of look like computer algorithms or complex mathematical equations. Yes, they are,
they are indeed, but they are truly fundamentally based on probability. Now, I'm not going to teach you
statistics in this podcast. I'm not going to teach you probability, but I am going to really quickly
run you through the basics of probability here in order to help you understand how naive-based classifiers
work, because to understand how Bayesian inference that is naive-based classifiers, how they work,
you have to understand the very basics of statistics. So like I said, all the algorithms that we've
been using thus far, they use statistics, but they use them under the hood, they use them conceptually
in principle. Well, Bayesian inference, which is a classifier, supervised learning algorithm,
but also can be used for regression, Bayesian inference is like raw statistics. It's like using
statistics in the raw to handle machine learning circumstances. Statistics in the raw, so Bayesian
inference is really just raw, true, pure statistics in order to make an inference or an estimate about
whether something is classified as this or the other thing. So let's try to understand probability
a little bit. Probability. Probability is very simple. It's the chances of something, the likelihood
of something, of an event we call it, an event. What is the probability of getting heads when I flip
a coin? Well, 50%, 50%, 50, 50, right? It's one half of the time it is heads, and one half of the
time it is tails. So the probability of this event of flipping a coin is 50%. Okay, so that's step one,
basic probability. Step two, joint probability. What are the chances of me getting heads first,
and then flipping the coin again and getting heads again? What are the chances? What's the probability
of A and B? Well, it is simply the probability of A times the probability of B. 50% times 50%.
That is 0.25. So the probability of heads, and then heads again, is the multiplication of the two,
which is 0.25. That is called joint probability, the probability of these things joined. Step three, conditional probability, and now we get into right proper statistics,
the good stuff, the meat. Conditional probability. What is the probability that my second flip gives me
heads? If the first flip was heads, that's an interesting question. I don't see how the first
flip has anything to do with the second flip. Exactly. If I flip a coin once, heads, tails,
okay, 50%. And I flip the coin again and get heads or tails. That second flip has nothing to do
with the first flip. The result of the first flip does not affect the result of the second flip.
Those are what's called conditionally independent events, independent because they do not depend
on each other. They do not affect each other, independent events. Well, there are some situations
out there which are not independent. They are dependent. So for example, what is the likelihood of
it raining today? Given it is cloudy outside. Ah, now there is an interesting question. If it is
cloudy outside, then it is, let's say, 40% likely to rain. The probability of it raining depends
on the probability of it being cloudy. We call these conditionally dependent events. And this is
all called conditional probability. Conditional probability is an interesting
thing. It is very useful and widely applicable in machine learning. And it has a mathematical formula.
Okay, probability, raw probability, step one was just probability. Joint probability, step two
is probability times probability. Just multiply the two. Conditional probability, step three,
is this mathematical equation. The probability of b given a, that is, the probability of rain,
given that it is cloudy outside, is the probability of a and b, the joint probability, a times b.
Over the probability of a. Okay, so the probability that it is rainy, given that it is cloudy,
is equal to the probability that it is rainy and cloudy, over the probability that it is cloudy.
Very strange. Very strange. This seems kind of non-intuitive. I mean, first off, it's a mathematical
formula and it's a little bit tough to kind of tease what everything is in this puzzle. But let's talk
a little bit more about probability, just general probability. Imagine we have a big giant circle
that represents weather. And inside that circle are a bunch of little circles. We have cloudy,
and we have sunny, we have rainy. These are built based on observations of the past. The number of
times that a day is rainy is the number of times that we've seen it rain in the last five years,
for example, over the total amount of times we've seen weather at all. Okay, so the way that we
build up probabilities, the way that we build up like what are the chances of it being cloudy at all,
is just that we look at days, day after day after day, and count the number of cloudy days. And then we
divide that by the total number of days we've observed. So that makes sense. Don't overthink it. It's
the number of times we've observed something over the total number of observations. So we have cloudy
days. We have rainy days. We have sunny days. Now, the joint probability of two events is the number
of times they overlap. That is in the case of non-independent events. The number of times they overlap.
So it's like a Venn diagram. We have cloudy days and rainy days and sunny days. Let's say that it's
cloudy 40% of the time and it's rainy 30% of the time. And there's a little sliver of overlap
between the two. Actually, not a little sliver. A very large chunk of the time, they kind of both fall
on the same day. We have both a cloudy and a rainy day at the same time. That's joint probability.
That's a times b. Joint probability. The number of the amount of overlap between the two. It's a
Venn diagram. And then conditional probability is a very interesting formula. It's very non-intuitive.
It doesn't make a whole lot of sense when we're trying to visualize this as a bunch of circles and
Venn diagrams. The conditional probability that remember the question that we're asking is,
what are the chances that it's going to rain today if I know that it is cloudy today?
And the formula says that the answer is the joint probability, the amount of overlap between
the two. The number of times it rains and is cloudy over the probability that being cloudy at all.
So the probability of a and b over the probability of a. So again, we have three steps so far. We have
basic probability and we build that up just by observing things over time. Okay, coin flips,
flip them a million times and you build up a database of 50/50. We have joint probability,
which is the probability of two things co-occurring. And then we have conditional probability.
And that is the probability of something if we know something else. And that, my friends,
sounds a lot like fundamental machine learning, right? What is the probability of it raining given
it is cloudy? Well, it is cloudy is a feature, a feature in our spreadsheet x, x1. And what we're
trying to determine is why, whether or not it will rain today, that looks a lot like the
logistic regression or linear regression or any other algorithm, any basic fundamental machine
learning algorithm that we've seen. This is kind of the skeleton form of machine learning. So
conditional probability is really core machine learning. So that's kind of the raw statistical
formulation of a machine learning algorithm, conditional probability. Now, the next and final
step is called Bayes theorem, Bayes, B-A-Y-E-S. That is the namesake for our algorithm here,
called a naive Bayes classifier. There was a man a long time ago named Reverend Thomas Bayes,
who was a statistician. And he learned a little trick of the trade when it comes to conditional
probability. Specifically, if you know the other thing, then the thing you want to know,
you can do a little reversey on our conditional probability formula. That's it. That's all Bayes
theorem is. It is using some statistics algebra, some probability algebra, and flipping stuff to
the other side of the equation. So if what we want to know is, is it cloudy? And we do know that
it is raining. So the opposite, the opposite of what we were asking before. Well, they're not the
same thing. Very obviously, they're not the same thing. How likely is it to rain if I know that
it is cloudy? Well, it is very likely to rain. Okay, maybe let's say 40%, maybe not that likely,
but 40%, 45%, likely to rain if it is cloudy outside. Well, how likely is it to be cloudy
if it is raining? Oh, totally different number. Now we're talking like 90%, 95%.
Have you seen rain on a sunny day? Yes, so have I.
the blue moon. It is substantially more likely to be cloudy if it is raining. Then it is to be
raining if it is cloudy. So conditional probabilities don't reverse. They're not the same thing, but
they're reversible. There is a way to reverse them. And that's called Bay's theorem. And Bay's
theorem looks like this. The probability of A given B, okay? So I want to know the opposite
order. Equals the probability of B given A times the probability of A all over the probability
of B. What the heck? I'm not going to explain where this comes from. You're going to have to learn
Bay's theorem and you're going to learn all this in statistics anyway. Bay's theorem is a very
fundamental component of statistics proper. You'll learn Bay's theorem in one of the early
chapters of your statistics textbook or the Khan Academy course. So it's not specific to machine
learning. It's a very raw fundamental core component of statistics in general. And all it does is
it gives you the ability to ask the question the other way around. So why is it so fundamental
then? It sounds like step three. We talked about regular probability, step one, joint probability,
step two, you know, joint depends on right vanilla probability. And then conditional probability,
which is step three, that depends on two and one. So they all, you learn them in sequence because
they depend on each other. It seems like conditional probability is the crux of what we need to use
statistics in the raw to solve probabilistic machine learning situations. Yes, that's true. But very
often the question isn't asked the way you wanted it to be asked. The question is the other way
around. So Bay's theorem is using conditional probability and doing a little reverse on the
equation so that you can ask the right question. Okay, so that was a little bit crazy. Let's talk
about an example using email spam. It's whether and spam are the two most commonly used examples
in understanding Bayesian inference. And in fact, whether and spam classification are two of the
most common applications of naive Bay's in the wild. At least up until now and to recent times when
I think deep learning principles are used a little bit more commonly in these spaces.
naive Bay's was the champion of whether prediction and spam classification for emails.
The way it works for emails is you break up your emails into all of the words of an email.
Let's say that we build up a dictionary of English words and we throw out all the very dumb words
and basic words like the is and we call these stop words. They are of course important in grammar
and understanding sentences. But they may not be as important in just classifying an email as spam.
So we throw out these stop words and we keep the essential words. We start to learn that certain
words are commonly co-occurring with spam emails versus non spam emails. So for example, the word
Viagra is very often seen in spam emails. But let's not get ahead of ourselves. First off,
what we want to do is just build up a database of how common every word is in an email in general.
And how common spam is in general. So we build up a probability of the word Viagra.
We build up a probability of the word friend, Saturday, weekend, every word under the sun. And
a probability of whether or not an email is spam. Let's say it's high. Let's say it's like 60%
of email is spam. Well, that's that. Okay, so we have a bunch of probabilities. That's step one.
Regular old probabilities. Step two, we're going to skip because we use joint probability in the
equation of conditional probability. But we don't really use it directly. So step three is conditional
probability. If I've got all these probabilities words and spam, what is the probability of an
email that I'm looking at right now being spam? Just straight up. Okay, well, that's 60%. We've
already said that. Well, what is the probability of that email being spam given it has the following
words because it does in the circumstance, it has the following words Viagra free act now,
etc. Okay, well, we will use the conditional probability formula. It'll give us a number.
It'll give us the probability of the thing being spam. And if we have to ask the question
a different way based on the information that we've provided, which is usually the case,
then we will use Bayes theorem to do a little reversey on the conditional probability formula.
And that'll give us our answer. Bayes theorem. Now, the specific algorithm for classification is
actually called naive Bayes, naive Bayes classifier. Why is it called naive? Well, there's a level to
which all the probabilities in our formula actually depend on each other. I kind of think of it as
like a Mexican standoff. It's like, what is the probability of this given this guy, this guy and the
other guy? Well, they all depend on each other. So think of like three guys pointing guns at each
other and they're all looking at each other. Well, what's the probability of this given that guy?
Well, the probability of this guy depends on the probability of that guy and the other guy. Well,
the probability of the other guy depends on this guy and this guy. So they're all kind of mutually
codependent. The naivety part of naive Bayes cuts off the dependence of events from each other.
It makes things not dependent on each other. And this makes the algorithm tractable able to be
computed within a reasonable amount of time. Without that naivety part, they call it the naive
assumption. The algorithm would be too computationally difficult for modern machines to perform. And so
in order to use Bayes theorem and conditional probability in the wild for machine learning
applications, you had to introduce this naivety assumption, which severs the dependence of events
from on each other. It assumes that they were all independent events. Now, as we did with support
vector machines, let's compare naive Bayes to deep learning. Naive Bayes is commonly used in text
based applications. Like I said, spam classification of emails. It's going to be based off of words
in the email. We call this a bag of words approach. It's called a bag of words because you're not
assessing grammar or how words relate to each other. You cannot with naive Bayes. The way words
relate to each other remember that would be dependent events that would not be independent.
Therefore, we would not be using the naive assumption. If words related to each other in a
grammatical structure, they would depend on each other and our approach could not be naive.
So we're going to assume that they don't depend on each other. Instead, we're just going to pull
out all the words of the email and we're just going to kind of keep our eye on trigger words,
like Viagra. What is the probability of an email is spam given the existence of the word Viagra.
So that's why it's called a bag of words. It's just all the words just throw them on a bag and hand
the bag to naive Bayes. A recurrent neural network, which is an algorithm that will get into in a
future episode, is a type of neural network. So it's a type of deep learning algorithm that is very
good at handling text-based applications as well. So naive Bayes and recurrent neural networks are
commonly pitted up against each other. But unlike naive Bayes, which uses a bag of words and the naive
assumption that there are no relations between the words, recurrent neural networks literally read
the email from left to right, top to bottom, and they keep grammar in mind, negating words,
modify the words they negate. I mean, I think of a recurrent neural network as taking an email,
printing it out, and it's a classy English gentleman who sits in his leather sofa and he has a pipe.
Well, I see the existence of Iagra, but let's not be too hasty because the use of some amount of
Antonins in this particular structural here, and I do find that they use abbreviations more often
than real words. Why would they use abbreviations? It's either that they're uneducated,
less versed in formal grammar, or that they're trying to save precious space so that they're
thinking you did a word edgewise. I believe through formal analysis of the documented hand,
we are indeed dealing with spam. But it took the guy almost a day to come to this conclusion,
by comparison to naive bays who sit in their folded his arms, he's got a cigar in his mouth. He's
like a "bub". It says Viagra. You don't need any other information. Recurrent neural network
looks up from the paper and he says, "Yes, of course it has Viagra, maybe increases the probability
of the thing being spam, of course, but let's not be hasty. Haste all these causes." And naive
bays snatches the paper out of recurrent neural networks hands and rips it up. It says,
"The goddamn thing spam, it has Viagra. I don't need to know anything else." So if time and memory
are crucial to your application, if things need to be fast and not consume a lot of memory,
the naive bays is a preferable machine learning algorithm to a more complex algorithm like a
recurrent neural network. But if you need more accuracy and complexity in the analysis of the
situation, then a recurrent neural network is more likely to be your guy. If time and memory are
less of an issue for your particular situation, and you'd rather have higher accuracy in a more
complex modeling of the situation, then deep learning is preferable to naive bays. But let's think
about email spam classification. You don't have all day to determine if an email is spam. When somebody
sends an email, the recipient expects to receive the email in very short order. Let's say no more than
one minute. Well, very powerful recurrent neural networks on very powerful machines could probably
do that in a minute. But I'm not so sure. Whereas an naive bays classifier could snap its
fingers and make a judgment, blink of an eye. So in the case of email spam classification,
it is very likely the case indeed that an naive bays classifier is preferred to a recurrent neural
network. And this is a prime example where
where we see a shallow learning algorithm
may be better for a particular purpose
than deep learning.
Even though deep learning is more accurate
and complex and magical.
In fact, in this particular situation
of using recurrent neural networks
for email classification,
they call this field natural language processing.
But using recurrent neural networks
with what's called word vectors,
we're gonna get into in another episode,
the way that it represents these documents
is as a point in vector space that can be compared
to other documents.
It's actually very magical.
So much so that the spin of natural language application
using this type of technology is called
natural language understanding,
which indicates if you might stretch your mind so far
that the machine may be understanding in a fundamental way,
the meaning behind what classifies a document
as spam or not spam.
Very interesting indeed.
So there you have it,
support vector machines and naive bays.
And I do want to admit,
I don't understand these algorithms
as much as the algorithms that I have presented to you thus far.
So this is probably one of my worst episodes.
I would encourage you to go learn these algorithms offline,
which brings us to the resources section.
Of course, the Andrew Ng Coursera course,
he has a week on support vector machines.
I will link to that in the show notes.
Andrew Ng does not cover naive bays classifiers.
I found that very interesting actually,
because naive bays classifiers,
that's one of the fundamental algorithms
of machine learning that you see brought up over
and over and over compared to more complex models
like neural networks and used in the wild today
with great success.
You'll see it in most introductory machine learning
textbooks and all these things.
So why didn't Andrew Ng cover the naive bays?
I actually found a video by Andrew Ng on YouTube
when I was trying to learn naive bays later on naive bays
and it clearly came from his course.
He took it out at some point.
I think that he didn't want to bog down newcomers
to machine learning with statistics.
Because like I said, to understand naive bays classifiers,
you have to understand base theorem.
To understand base theorem,
you have to understand conditional probability.
To understand conditional probability,
you have to understand statistics.
So the whole world of Bayesian methods,
it's the world of statistics, raw statistics,
and stats is hard stuff, my friends.
So it's important, it's essential,
but I have a hunch that Andrew Ng decided,
they'll get to that later.
I don't want to scare them away from the field yet
'cause you don't need it to succeed right away.
You can start doing linear and logistic regression
and you can deep dive right into deep learning
and neural networks and skip past all this statistics stuff.
But it is essential for you to know.
So I would encourage you to learn naive bays classifiers,
the machine learning with our book
that I'll put in the show notes
is it has a great chapter on naive bays classifiers.
It also has a great chapter on support vector machines as well.
And the mathematical decision making
great courses series that I've referenced from time to time
also has a whole episode, audio episode dedicated
to naive bays.
So I'll post that in the show notes
and I would encourage you to try to learn
the basics of these two algorithms offline
'cause like I said, I don't think I did a very good job
of presenting them unfortunately.
I prepared and I prepared,
but I was a little bit out of my element for this episode.
And finally, like I mentioned before,
how do you choose which algorithm to use?
When you know your situation,
whether it calls for supervised learning
or classification, regression, et cetera,
you'll be able to narrow down grossly
which algorithms to throw out.
And now you have in your hands 20 algorithms
that you could possibly use for classification.
And in order to decide which of these algorithms
you should use, you assess your data.
So for example, naive bays works well
with categorical data and missing data,
something many other machine learning algorithms
do not work well with is missing data.
naive bays works a-okay with missing data.
How many examples do you have in your training data set?
Are you working with text or numbers, et cetera?
Using these types of questions
will help you narrow down even further.
And in order to do that, I am going to link to a table
of pros and cons of various algorithms
under various situations.
And a decision tree put out by the Psychit Learn project
for choosing an algorithm,
giving various circumstances in your problem.
And from there, once you've got five algorithms to use
in hand, and you still don't know
which of these five is best to use, you just try them all.
You try them all and you see which one,
which one has the highest performance
based on some evaluation metrics.
In the next episode, I'm going to be talking
about some more miscellaneous machine learning algorithms,
things that are very dedicated.
So the last three algorithms that I talked about,
decision trees, support vector machines
and naive base classifiers,
these are all very general purpose,
power tool machine learning algorithms.
All three of these could basically be swapped with each other.
And knowing which one goes where is a little bit difficult.
But in the next episode, the machine learning algorithms
I'm going to be presenting to you
are very specifically tied to very specific use cases.
So it'll be a little bit easier.
It'll be one of those apples to oranges bits.
That'll make it easy for you to decide
that yes, you should use this algorithm
because the situation is A or B.
I'm going to be doing these episodes now every other weekend.
I've become quite busy recently, I apologize.
So rather than every weekend, I'll do every other weekend.
So I will see you two weekends from now.
Podcast Summary
Key Points:
Support vector machines (SVMs) and naive Bayes are powerful, general-purpose shallow learning algorithms primarily used for classification, with SVMs creating a wide-margin decision boundary to reduce overfitting.
SVMs use a "large margin" approach by maximizing the distance between classes, with support vectors being the data points closest to the boundary, and can handle non-linear data via the kernel trick.
Naive Bayes is based on Bayesian inference and conditional probability, making it fast and memory-efficient, especially for text classification like spam detection, relying on the "naive assumption" of feature independence.
Both algorithms have trade-offs
Choosing the right algorithm depends on data type, size, memory, and speed constraints—using a decision tree or pros-and-cons table helps narrow down options.
A practical approach is to test multiple algorithms on the same dataset, evaluate performance using metrics, and select the best-performing one based on accuracy, speed, and resource use.
These algorithms are less sensitive to missing data than many others, making them suitable for real-world datasets with incomplete features.
Deep learning models like neural networks are more powerful and flexible but often overkill for simple problems; shallow learning algorithms like SVMs and naive Bayes are preferable when speed, simplicity, and low memory are priorities.
Summary:
This episode explores two powerful shallow learning algorithms: support vector machines (SVMs) and naive Bayes classifiers. SVMs create a wide-margin decision boundary between classes to prevent overfitting, using support vectors as key data points. The kernel trick allows them to handle non-linear data by transforming it into higher dimensions.
Naive Bayes, rooted in Bayesian probability and conditional logic, is fast and memory-efficient, particularly effective in text classification like spam detection, due to its assumption of feature independence. Both algorithms excel with categorical data and handle missing values well. Choosing between them depends on practical constraints: data size, memory, speed, and the nature of input features.
A structured approach—narrowing options based on domain knowledge, evaluating data characteristics, and testing multiple algorithms—helps determine the best fit. While deep learning models offer greater flexibility, SVMs and naive Bayes remain valuable for specific use cases where speed, simplicity, and efficiency are critical. The episode emphasizes that no single algorithm is universally best; instead, a systematic evaluation and experimentation process is essential for successful model selection.
Tyler also highlights resources like Andrew Ng’s course and the scikit-learn decision tree to guide algorithm selection.
FAQs
Support vector machines (SVMs) create a 'fat' decision boundary between classes, making it thicker than logistic regression's thin line. This helps reduce overfitting by bumping against the closest data points, known as support vectors. Unlike logistic regression, SVMs are less sensitive to outliers and can be used with kernels to handle non-linear data.
The kernel trick transforms data into a higher-dimensional space where it may become linearly separable. This allows SVMs to handle non-linear classification by mapping data points using functions like radial basis or polynomial kernels, effectively changing how the data is viewed without explicitly computing the transformation.
Naive Bayes analyzes email content by counting word frequencies and using conditional probability to assess the likelihood of spam. It assumes words are independent, building a model based on the probability of spam given specific words like 'Viagra', making it fast and effective for initial spam filtering.
It's called 'naive' because it assumes all features (like words in an email) are independent of each other, ignoring real-world dependencies. This simplification makes the algorithm computationally efficient and tractable, though it may not reflect true relationships between features.
Use SVMs when you need faster processing, lower memory usage, and have a linearly separable or kernel-transformable dataset. SVMs are especially effective for small to medium datasets where speed and efficiency are priorities over maximum accuracy.
Start by identifying if your problem is supervised or unsupervised, and whether it's classification or regression. Then evaluate your data type, size, memory, and time constraints. Use decision trees or algorithm comparison tables to narrow down options, and finally test multiple algorithms to find the one with the best performance.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.