Supervised learning uses labeled examples (input-output pairs) to learn a function that predicts outputs for new data, including classification (categorical outputs like cat/dog) and regression (continuous outputs like temperature/price) tasks, with algorithms such as K-Nearest Neighbors, curve fitting, and deep neural networks. Unsupervised learning finds patterns and structures in unlabeled data, including clustering (grouping similar data points, e.g., K-means) and dimension reduction (simplifying data representation, e.g., PCA). Semi-supervised learning combines both approaches when labeled data is expensive. Reinforcement learning is a separate paradigm focused on learning control policies through experience, where an agent learns to maximize cumulative rewards by taking actions in an environment, using value-based or policy-based methods.
Supervised vs Unsupervised Learning: ML Foundations
Added:okay great so this is the explainer about supervise versus unsupervised learning um I mentioned this in the introduction to machine learning first explainer we're just going to dive into more detail here I've put supervised versus unsupervised learning but what I really want to highlight it isn't a competition so the verses really shouldn't come in there should be supervised and supervised learning but often it's classed as a sort of seen as either or it really just depends on the class of problem or tasks that you're trying to achieve machine learning so fundamentally we're going to go through those two examples give examples the types of algorithms that you would look at different classifications of supervised non-supervised learning so they're subsections and then we're going to look at semi-supervised learning which is a hybrid of the two just really at a high level and then finally we're going to go on to reinforcement learning which is a really nice area which is learning how to control systems through experience we'll get on to that so that's it okay to highlight again supervised versus unsupervised learning is often pitched you'll see LinkedIn posts you know which one's better which one's not it's not one's better than the other they're just fundamentally different tasks and you should yeah you should be able to understand given a challenge what what task what type of learning algorithm that you need to use so supervised learning at a simple level is really learning from labeled examples this means I show you examples of input X outputs Y and then I want to learn a machine learning algorithm that mimics that that process we supervise it by giving examples to learn the task at hand unsupervised learning is all there's no labeled data so it's just all inputs and what we're trying to do is find patterns or similarities in that data to extract knowledge or information about that data set in machine learning these patterns are often called features and you'll hear that used quite a lot um I personally think of machine learning as a pipeline of algorithms so this is moving you from real raw data all the way through deployment and validation of your machine learning algorithm and I would say that a lot of the time both supervisors and unsupervised algorithms are used in tandem in algorithms so often unsupervised learning is used as an initial step to simplify the patterns in our data and make it easier followed by a supervised learning task where we try and mimic examples of that data okay so first we're going to deal with supervised learning around learning from examples so don't worry about the mass here but we'll try and put a bit of math notation in there so you start feeling comfortable with some of the mass terminology that you'll see when you when you see more machine learning so what we say is we have a training set and in maths we denote a set by these curly back brackets just imagine it's a bucket right a bucket of examples in it um in this bucket of examples or set we then have examples which are paired together X's are the inputs y's at the output and they're given a number each to identify them so x0 y 0 x 1 y1 all the way up to x n y n so we have n samples in our data set now the task at hand really is to think well a simple way of thinking is right we have a machine where we can feed in inputs and it can spit out outputs and so we denote this output of the machine as f of x this basically says it's a function of the input now what we want to try and do is we want to supervise this machine in when we give it a particular X input the f of x is very close to the label that we want or the output or the um not necessarily a label get into that the output y that we've got paired in our data set and we want this to be true over all of our training data but we also want it to be true over data we haven't shown it at training this is a concept of generalization which we've touched on in the last explainer okay so under supervised learning there are broadly two classes classification and regression and what this really depends on if I jump back to the previous side it all depends on what Y is if it's a label then we call this a classification task if it's a number temperature for example or price for example then we could see it's a regression task they're both supervised learning algorithms but different methods some can do both but different methods are suitable for classification against regression problems so it's not either or just depends on what data you have you can also have more complicated methods which are combinations of this but we won't touch that in this explainer so to give you an example of classification example that I use a lot if we have inputs of images of either cats or dogs then the labeled outputs will be Associated outputs with that picture so here we've got cat cat dog we feed these in the machine to the example and then we want the machine to see an unseen picture here a picture of a dog and it's learned by example that it has similar features to other examples of dogs and it classifies that picture as a dog hopefully on the other hand and we have a a regression task so here we might have two sets of inputs here from naught to 100 just a made up example and then the data lives on this uh these black dots what we want to do is to generalize or learn to areas where we haven't yet seen data so in this regression task we've basically fitted a surface a curve to that data to try and recreate that data to give our best predictions to unseen data so where we don't have the inputs but to try and make good predictions this is a typical regression algorithm not the only way to do it there are many many but curve fitting in this case is a is a classic example of regression so here's a challenge for you uh you can stop the video but I would encourage you to go and find out about five supervised learning algorithm I don't mean in detail because that's what this course is about is diving into detail but if you can get on chat GPT say give me five supervised learning algorithms do it Google would give you the same and just come back and you know um just having found five uh and distinguish now if they give examples with it think is it a classification algorithm or is it a regression algorithm so hopefully you've done that and I'm going to give you three now so I didn't do the full five but I'll give you three the first one that you're gonna see in this um in this course actually after this video is K nearest neighbor um so it's like a more sophisticated version of lookup table so what we have so we're going to think about it for a classification problem but I also cover the regression problem in the next explainer so what we have is we have the space of possible input so imagine this area represents different X values with two values two coordinates in each these the labels identify as green triangles or yellow squares and these this is the classification so the the shape and color of the the um data point represents its label okay so now we ask the question well we've got a new instance here and we want to know if we had this as an input what would the output be in terms of label and what K nearest neighbor does so in this case the K stands for is a number in this case we're looking at three it will find the three nearest points and we'll discuss what near means but in this simple case it just means measure distance um the three nearest points and then we take a majority vote so in this case we say the three nearest points to our new instance two are green one a yellow and then okay so we will classify this point as um green uh green triangle rather than yellow Square nice simple algorithm we'll dive into the details more but an example of using labeled data to make predictions of unseen data so in this case it's a classification example and it's supervised learning so I've dutched on this so in general curve fitting can be seen as linear models and you'd be like well linear models just are straight lines we will cover in this course actually linear models can be more complicated than straight lines and we'll get on to why even you know fitting polynomials or more complicated functions it can be seen as as linear models as well but essentially what you're doing here is you define a way of defining a curve so this could be polynomials it could be straight lines it could be other types of functions that you put and then you find the best fit the least squared fit to the data so you try and minimize the error between the data and the prediction and so this gives you a predictive curve some people call it a response surface around making the predictions in this case supervised learning and a regression because the output is a continuous variable it's a number right and one of my favorites um a deep neural network so deep neural networks work are mimicked on the way we solve problems our brain is set up with neurons I won't get into lots of detail in this course but in the next course with Andy he starts to touch on deep neural networks and more complicated models what happens here is that the inputs I've gone back to the the classic is it a dog or is it a cat picture um what we have here is uh the picture goes in are the pixels form the inputs to the neural net it's then passed through a series of uh weights uh onto hidden nodes through activation functions we won't get into the detail um all the way through this network where it outputs a choice a classification of problems so here if the second one they've identified if the second one says one then essentially this is a this is a dog classification maybe one of these is a cat or other animal for example I'm just saying I don't know in this particular example what then happens is the weightings are of these um each of these components is then adjusted so that the behavior reconstructs the right label output given examples so this is an example of a classification problem and again supervised learning because we're giving labeled data neural Nets work for both classification and regression tasks so it's not just a classification with neural Nets but um nice example you'll learn more about them in the wider Ai and the World Series okay um so unsupervised learning is different we're trying to find patterns and I said oh it's your best friend in highly parameterized models what I mean by that is if you have lots of parameters and models like um the potential to make very rich models then essentially they can be very complicated to fit we'll see that they need lots of data we've already talked about the curse of dimensionality so trying to simplify the representation of these models to uh simpler representations but still capture the features in the input data is really important and this is often how unsupervised learning algorithms are used so again we're looking for patterns but we kind of consider two classes of unsupervised learning on the left here I can't remember what the data set is but this is an example of a clustering algorithm right so it's trying to group like inputs together so they are treated in a similar way this clustering k-means which we're going to go through in this section as a first unsupervised learning you'll meet is an example of a clustering algorithm uh the top right hand corner is an example or tries to represent Dimension reduction so Dimension reduction is where we have this representation of the full input which may have many numbers so for example a picture has many numbers to describe all the pixels the colors of the pixels but do we need all of that information to understand a picture is there as much smaller or concise representation in what we call feature space which allows us to have the same information in our data set but much more efficient in terms of number of parameters so examples of this could be Auto encoders or principal component analysis and I'll talk a bit more of that so this is a dimension reduction that we're trying to find the reduce representation of our parameterized model super helpful at the beginning of dealing with data okay so let's quickly cover what we mean by k-means so as I already said K means there's a clustering algorithm I'm going to put a strap line birds of a feather flock together right so the idea is that the means bit says okay what we're going to try and do is assign uh and um put the data into K clusters so here again we've used k equals three and what we're trying to do the um uh the groups to which the data is uh assigned to or the classification to which it's signed to um try is is chosen so that it minimizes its distance to the mean of that cluster so if you take the mean of all these points it sits somewhere in the middle it's called the centroid and actually all of these points uh are um so each of these clusters has a centroid and what you see is actually all these points the closest centroid to all these points are one um uh one cluster and the same here and the same here so this is where ring fencing uh clustering data together nice simple algorithm um pretty easy to implement can be really useful for sort of deciphering different uh types of behavior in your data set or patterns we'll get on to that more and then principal component analysis uh off widely used tool particularly at the early stages of data analysis what it's trying to do is trying to find the source of variability in your data so it's a dimension reduction technique so in this little sketch we have one input and another input and if this was your data set what we can actually notice that in this direction there is the most variability there's the biggest spread of data in this direction we're in this other direction there is much less spread so if we had to reduce the dimension from this 2D representation down to one um a good representation would be just along this axis because it includes the most variability and so we can choose uh to reduce the dimension we lose some information but we do it in an ordered way to retain the most variability in our data set so it still as diverse as possible our data set so we lose the smallest amount of information principal component analysis is a linear way of doing this there's more complicated ways autoencoder for example which we'll learn about but this is often used at the early days side to try and filter out noise out of data or only pick the most relevant features that drive variability in that data set really important always try at the beginning of the data analysis pipeline I nearly always do okay so I've covered supervised learning and unsupervised learning um so what is semi-supervised learning it sounds like we're making up also called weak learning um so one thing that often is the case is sometimes collecting label data is really expensive so you can collect inputs but actually the process of producing labels for that data is really expensive so just in one slide um trying to talk about how you can use a combination of the two methods together so on the left hand side what we have here is only labeled data so I've got two data points um and I want to build a supervised classification algorithm so I want to classify them as either green or blue for any general input in this space well the best we can do here is the decision boundary is the one that bisects these two points everything to the left here would be classified as green everything to the right would be classified as blue and that's the best you can do because that's the only label data you've got but then suppose we have other data which isn't labeled so these are other inputs and um and so how do you use those well if we use a clustering algorithm that for some reason these points can be clustered and identified as blue and these points can be clustered and identified as like the green one then this gives us our two labels right uh and I think it's a really interesting point and um and then if we can do that then we can use this clustering and labeled data together to give us a a more complicated decision boundary which is hopefully more accurate than just using only the label data itself so it's really about trying to exploit everything you've got from labeled data and unlabeled data together in a problem to try and build better models right we won't touch too much on this in this course but something really important to know about there's a whole set in class of semi-fires super learning algorithms I've touched on at a very high level the nuances of this can be a bit more involved but good to know that you know if you have limited label data and it's hard to collect more then this is something that you can look up and use right so reinforcement learning um I call it learning from experience it's really about how do you learn how to control systems um and I I love it as a topic I think it's a really exciting area of machine learning I think there's been growth in it there's lots of open questions but more and more we're seeing it used so I'm going to go back over my favorite example you've seen this before but just to recap what we've got here is a two-legged draft is the best I can describe it what the algorithm can do is it can control how the muscles contract and this causes this two-legged draft to walk fall over or completely lie down you'll see and the question is can a machine learning algorithm learn how to coordinate those muscles together to walk forward and also remain upright so you incentivize it it's objective its reward is essentially to remain upright and walk forward and the algorithms I'm going to run this now after different generations of this algorithm it's gradually learning better and better and more stable control you know you see in the last one that you know by generation 900 of this algorithm it's learned to walk off into the sunset you know really nice example quite fun but an example like how we learn from experience of trying different things about what we call a policy on how we decide to control this system how we get it to walk towards literally in this case something with a good outcome which we Define so in more General case um we think of reinforcement learning there's this sort of intricate pairing so we talk about an agent so the agent in the last case was the um was the the draft right it's learning how to act within its environment and that's the thing you're seeking to control so this agent interacts with an environment this environment can be uncertain it doesn't have to give back the same things often it can depend on other agents as well or these are called multi-agent systems really interesting problems but then just to go through some terminology and reinforcement learning we have state which is state is a representation of that system in a snapshot of time action is the range of possible options available in terms of actions it can take at a given time and then reward is a feedback mechanism to say how well it's done essentially once it's chosen an action and this reward is funded mental to how you incentivize it or reinforce it to do the behavior that you want think of it like training a dog so just to say in this example Pac-Man is like a classic example where people test out reinforcement learning algorithms because it's fun right and it's you know complicated to an extent so here the state representation is probably the representation of the system the location of the ghosts the direction they're traveling in because you can't tell if this one's going up or down so direction as well the number of counters that remain on where the big counters are would represent the the state of that system um the action so at any given time step that a player can choose to go up down left or right and so there are the actions available at any given step um and then the um the reward is defined by the game so I think with this I can't remember I haven't played Pac-Man in a while but essentially you get Five Points as if for eating these small ones more for the big ones but that then changes these into uh friendlies that you can eat uh once you go past the big ones and then if you eat a ghost you get a certain number of points and then ultimately you complete the game by eating all of the uh these crumbs should we say obviously if a ghost catches you you lose and that results in a negative reward or um it will terminates the game actually and your overall score is you're trying to maximize this you're trying to collect all the crumbs so every time you go through a step at this the state gets updated uh the agent has the ability to make an action including the null action no actions taken the environment increments it then says okay now based on the actions you've taken the state it passes back a reward and the state gets updated and we repeat this continuously so fundamentally in reinforcement learning there are I don't know there are two broad classes of methods uh and then of course it's not that simple because there are combinations and hybrids of the two um but generally we can think of value-based methods and so in this case we choose to take actions in this system which maximize the value so a given State at time uh time uh T then say that here we have four actions that we can possibly take we get a new state and we have a value associated with uh taking that action to the state that we're in so basically we can think of this as taking a state and an action as an input so given a state if I take an action what's the value of that action given my state this mb's comes a regression task actually these are inputs and the output of the machine is a value a number um basically a value is the estimated of all future possible rewards discounted uh but we won't get into that now um but that's um and so the value state which basically says okay what's the value in my state um if I take this action I will then choose the action which leads to the highest value position or the likely highest value position so just take talk about a classic example with noughts and Crosses what we have here is uh towards the end of a game we're going to pretend that we are we're putting down crosses so um we've started we've got cross in the middle um we're now this is our state so this representation three crosses uh three uh zeros and then we have three possibilities of actions we can take we can put an X here bottom right bottom left so the idea here is if we put it in the top right then there are two possible outcomes that can either the zeros are put in the bottom right and then we complete the diagonal or they put it rightly so in the bottom left and then they can't complete the diagonals and it's a zero so in this case it's um sorry a draw so we give it zero reward in the case where we win we give it one if it was in a losing position we could give it minus one for example uh we repeat this process so if you put it in the bottom right hand corner actually boast the possible outcomes come here as to be um a uh a draw so there's zero reward in this case uh in the opposite case it's almost the same as the first case just uh slightly differently um there's one possible win outcome and there's one possible draw outcome so under the assumption I mean that there's no discounting uh we won't get into that and our player plays at random then we can essentially assign the value of each of these positions as the average of these two rewards so the value of this action is essentially a half of zero plus one um this one is zero the value of this position the value of this is is a half so actually the best move I can take in this situation uh given what I know is either a or C which makes sense right and I I can't really choose between the two of them I'm indifferent I think of another area like so we've talked about value-based methods so making decisions to maximize the value of course it gets a lot more complicated than this there are lots of actions you can take there could be all sorts of you know almost an Infinity of combinations that could result think of chess for example much more complicated game than naughts and Crosses but the principles could be the same right um in policy-based methods the approach is different in this case we are trying to learn given a state what action should we take and we learn it's called a machine learn policy this this is a policy a mapping from state to action in the system and then we try and learn this machine to maximize the reward so we reinforce uh policies that lead to good rewards um and yeah this obviously involves kind of supervised learning algorithms um and um you know the object is instead of minimizing the loss or the fit we're maximizing our reward but in essence the same these are often very common they're often used in combination with value-based methods they're widely used in lots of exciting different areas okay so this was a flyby tour of supervised learning primarily and unsupervised learning we then touched on semi-supervised learning and then my favorite reinforcement learning associated with optimal control take away messages from this and things we're going to focus on know the difference between what a supervised learning task is and what unsupervised learning task understand when you might use them start building up your toolbox and methods that you know about which you can use in each case what are the benefits pros and cons in this explainer uh well sorry in this course next section and the remainder of the course and the work wider AI in the wild uh course you will have lots of examples of uh I would say predominantly supervised learning algorithms but some really important and useful unsupervised learning algorithms so this is the start of your journey there great
Up Next

Online DPO Fine-Tuning for LLMs: Hands-On Implementation Guide
@fahdmirza
768 views•2024-09-02

Introduction to Secure Multiparty Computation with Yehuda Lindell
@fhe_org
7.7K views•2021-02-04

HTTP Requests Explained: GET, POST, PUT, DELETE
@codecademy
103.1K views•2021-10-07

Enigma Machine Mechanics: WWII Encryption Explained
@JaredOwen
13.2M views•2021-12-11
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Computer Science
![麻省理工开放课程_线性代数[MIT][Strang]Lec01_方程组的几何解释](https://i.ytimg.com/vi/YeznlKTrpmU/hqdefault.jpg?sqp=-oaymwEmCOADEOgC8quKqQMa8AEB-AH-BIAC4AOKAgwIABABGFEgXihlMA8=&rs=AOn4CLBhKWvGGo7Wx9xfalcysEcZW89l7w)

![[線性代數] 第1-1 單元: Introduction to Linear Algebra](https://i.ytimg.com/vi/IG-EQUIk7P0/maxresdefault.jpg)






![統計学基礎(検定)[G検定・中級]](https://i.ytimg.com/vi/GrAw30LuYWs/hqdefault.jpg)









![[핵심 머신러닝] K-nearest neighbors & Distance Measures](https://i.ytimg.com/vi_webp/W-DNu8nardo/maxresdefault.webp)









![[Mì Python] Bài 4. Python với Keras (Phần 1)](https://i.ytimg.com/vi/hPhnqTtidnA/maxresdefault.jpg)









