The human brain encodes decision-making under uncertainty through specific neural mechanisms: the orbitofrontal cortex represents experienced utility (the hedonic value of outcomes we actually receive), while the ventromedial prefrontal cortex represents decision utility (the expected value of potential choices). These signals are learned through reinforcement learning algorithms where dopamine neurons encode prediction errors—the difference between expected and actual outcomes—which allows the brain to update its predictions and make optimal choices. Additionally, the brain explicitly represents risk signals, with individual differences in risk-seeking versus risk-averse behavior reflected in distinct neural activity patterns in these brain regions.
Risk and the Brain: Neural Basis of Decision Making Under Uncertainty
Added:well we're on time this week welcome to this third lecture in the Darwin College lecture series on on risk this isn't really commercial but there are copies of the publication's from previous lecture series on sale in the in the hall and there was someone there ready to take your money on your way out if you want so far we've looked at how notions of risk can be communicated and at how they can be misused and tonight we're going to look at how our brains appear to to manage them our subject is that the neural basis of decision making under uncertainty now speakers professor John a dotty who's Trinity College Dublin is the Thomas Mitchell professor of cognitive neuroscience he was educated in maths and psychology at Trinity College and and then did his doctorate at Oxford and he's worked at University College London and he's worked at Caltech his research has been concerned with how the brain deals with reward stimuli he uses magnetic imaging data to deduce where and how the functions of the brain such as decision-making and emotion and social cognition where they way the way the way they come from on what they do and he's currently developing computational models to refine these studies so we couldn't have someone better to tell us about what happens if you burrow inside your skulls and and and look at risk in the brain Jannetty [Applause] okay well and thank you very much and I'd like to thank the master and the organizers of this series for inviting me it's a huge privilege to be able to come here and talk to you and come to sunny Cambridge from from from Dublin here and this is a little bit of a misrepresentation the sunny sky here because it's I think about one day a year that we that we have this in Dublin but I think you're not that far behind here in Cambridge either so what I want to talk about today is really introduce you and to a relatively new field of investigation with in neuroscience called depending on who you talk to decision neuroscience or neuro economics and really what what this field is concerned with is trying to understand what goes on in the brain really what are the algorithms what are the rules that almost like a computer program what is that what's the programs that the brain uses to solve the problem of making decisions under situations of uncertainty and really if you think about it pretty much all decisions are most decisions that we make in our lives are made under conditions of uncertainty so we don't know all of the the pros and cons we don't know exactly what's going to happen there aren't certain outcomes and and that can be for the simplest decision whether it's deciding what would I like to have for dinner this evening or if I go to a restaurant you know what I should should I choose from the menu so even if you've been there before you don't know maybe the chef is off today you know maybe something Bad's gonna happen maybe the chef's had norovirus you know so you don't know what's gonna happen and so you even that is under situation of uncertainty up to kind of big life decisions you know what career should I choose and so the question is you know how does the brain actually do this and I think one thing I should say is the brains actually very good at doing this and we know that because humans are extremely successful right we've been very successful successful as a species so clearly were very good decision makers and so clearly something that's going on in our brains is doing very well okay so how does it do it so the first sort of intuition and as as neuroscientists that we had as to you know the functions are of the brain in decision-making came from a case and back in the the 19th century a very famous case now many of you may have heard of this and called Phineas Gage I think he's found his way into introductory neurology and introductory psychology textbooks by now and he was a railroad construction worker in the US and and what they would do is and he would kind of blast Rock every day and and that involved during a whole and then filling it with blasting powder and covering it with sand and then compacting all of that with a tamping iron okay and that tamping iron is that thing you can see him holding up there so it's just this big big iron rod and but one day outside the small town of Cavendish in Vermont and around 4:30 p.m. an explosion blasted an iron through poor gages skull and it went through under his eye socket into his brain and outside out the other side and a physician at the time Harlow actually and was he was gages physician and he and reported on the fact that there appeared to be a big change in gages behavior so he said you know the equilibria more balanced so to speak between his intellectual faculties an animal propensity seems to have been destroyed so after he had this and this brain injury he is fitful irreverent indulging at times and the grossest profanity which wasn't previously his custom manifesting but little deference for his fellows impatient of restraint or advice when it conflicts with his desires at times per tenaciously obstinacy a capricious and vacillating devising many plans of future operations which no sooner arranged than they are abandoned in turn for others appearing more feasible so what you can see there is apparently he seems to be having some difficulties making decisions he or you know he makes them and then he reverses them and he sort of comes back and sounds like how I make decisions but that's a different story and previous to his injury he possessed a well-balanced mind and was well regarded by those who knew him but afterwards his friends and acquaintances said he was no longer gauge okay so that's the first sort of observation and that you know having damage to your brain appears to impact on your personality and behavior but also hidden in there you can see it seems to be impacting on his ability to make decisions and more recently and Hannah DiMaggio and her colleagues used computer techniques to reconstruct where this tamping iron might have actually gone which parts of the brain might have been impacted so you can sort of see it coming in here onto the eyeball and it's coming through and this more anterior part of the brain so you can see there is see there's the eyes there and it's basically coming up through here and destroying that bit of cortex and the part of the brain that is actually destroying is called or mostly impacting is called the ventromedial prefrontal cortex so just to orient you this is the front of the brain this is the back and this area in blue here this going to certainly from here downwards would be part of this ventromedial prefrontal cortex and another part that often be thought of as part of this area is an area called the orbital frontal cortex and so called because it's just above the orbits of the eyes so though this region together with this collectively can be called the ventromedial prefrontal cortex and this was the area that seemed to be most at least according to and a DiMaggio and her colleagues most impacted in Phineas Gage now more recent work has been done by Anton Bashara and his colleagues and they are also as Hannah DiMaggio at Iowa at the University of Iowa at the time and Anton Bashar and colleagues had access to patients to human patients and that had had the misfortune of having this part of the brain this ventromedial prefrontal cortex damaged in some way so that can happen through it could happen through having a head injury or it could happen from having a tumor and our stroke there's all sorts of reasons why this piece of the brain can get damaged and and they took a group of these patients and they tested them in the laboratory on very simple decision-making tasks and the canonical example is they call the Iowa gambling task basically what it involves is a patient or a subject has to choose between four decks of cards okay and they choose a card on each trial they just make a choice and they either win money when they do so or they they win money and then they lose money okay so they either so basically their net gain can either be positive so they accumulate money or after a series of choices on some decks they actually start to lose more than they they gained okay so they identified some of these decks and they call them advantageous decks because you win more than you lose so you're kind of net gain is cumulative and whereas other decks are so-called disadvantageous decks in that if you choose those you tend to lose more than you win so there are disadvantageous decks and what they found was that the patients with damage to the ventromedial prefrontal cortex were actually very very bad learning to choose the unfair advantages deck so that's kind of new down here it's very hard to see that probably especially at the back but basically what it means is that they were sort of very very bad at converging on the advantageous decks wares and contest page or human subjects who didn't have damaged through controls or patients who had damage to other parts of the brain were actually able to do this over time they gradually learned through trial and error to choose the an for advantageous decks so that was the first report and that damaged this or at least formal study in the laboratory the damage to this area impacts on one's ability to make good decisions in an experimental context and and the other thing that sort of been reported about these patients is just like Phineas Gage is that often they're very bad after their injury sadly their lives fall apart they make very very bad financial decisions they make an appropriate social decisions sometimes their marriages collapse and they generally just become very fickle individuals okay so the question would be well what's actually going on so what's going on in this region of the brain or indeed other bits of the brain that are supporting the ability for people to make good decisions and what sort of concepts do we need to sort of understand what's going on so one concept that I think is very important is this notion of utility and we can go back to Jeremy Bentham the utilitarian philosopher and to sort of get it a handle on this so he said Nature has placed mankind on under the governance of two sovereign masters pain and pleasure it is for them alone to point out what we ought to do as well as to determine what we shall do they govern us and all we do in always saying all we think so clearly our behavior is shaped or by seeking out good things gaining pleasure and avoiding bad things avoiding pain and he said that utility is that property in any object whereby attends to produce benefit advantage pleasure good or happiness or to prevent the happening of mischief pain evil or unhappiness this notion of utility is some sort of abstract construct that refers to the ability or the propensity of an object to confer something pleasure or confirm pain or it not confer pain should I say and of course this idea is is very central in in modern economics so the idea is that a rational agent will make decisions will take choices in order to maximize his or her expected utility so they want to this concert if utility if you're irrational you should be doing everything you can to increase that and maximize that and avoid things that compromise or go against that third aim now of course according to this you know Bentham it was utility was just some property that was inherent in an object but it does have a more subjective property so I might much assign a higher utility to apples than oranges meaning I might prefer apples than oranges and might think that that confers more benefit to me whereas someone else may have the opposite utility and sort of difference is impression this can occur between individuals and even across time from one individual so I'm a you know like heavy metal music when I'm younger but later I will develop a penchant for for opera right so your preferences are in constant now another concept which kind of moves away from economics for a while is a notion from psychology from the psychology of learning and constructs called positive and negative reinforcer so a positive reinforcer could just be opera operationalized as really anything an animal will work to attain okay and whereas a negative reinforcer is anything an animal will work to avoid okay and so psychologists of learning such as Edward Thorndike here back in the turn of the the the late 19th century and study the mechanisms by which animals learn to shape their behavior to take actions in order to maximize their probability of obtaining positive reinforcers and minimize the probability of obtaining a negative reinforcer what you can see here is and this dog who is receiving a little treat from his master and maybe this dog in order to get that treat he went around and played dead or did some other acts and now he's receiving the treat that he's expected expecting because he performed this action or maybe the dog is thinking every time I play dead this this guy up here gives me a treat I've really got him under control and some reinforcers probably have innate values so they're probably certain things that were all pretty much programmed to find rewarding and to want to perform actions in order to obtain and you know it's actually still controversial as to what those might be but some candidates might be sweet taste might be a pop right a positive reinforcer bitter taste which sort of is often found in poisons sometimes might be a negative reinforcer they're clearly not always because some of us like tonic water occasionally and colic caloric content so just the post ingestion the changes in it that the cut the content of the of the food how much calories it has might also be something that's a positive reinforcer and maybe also sometimes a negative one water the smell of rotting foods painful touch so you know they're probably not that many examples but you can certainly think of more than than these and these might be things that were more or less sort of programmed to find reinforcing and because you know we need to be able to have our glucose or carbohydrates water and we want to avoid poisons and we want to avoid pain in order to survive and procreation and also these positive negative forces can produce stereotyped passions are responding so for example and a positive reinforcer like a food will engage you know cause us to salivate to evolve all heard of the Pavlov's dogs story and so clearly with salivation anticipation of food will approach food and there's also stereotyped responses and that we might do to bad things like we might try and withdraw we might sort of develop defensive reactions and so on ok so there are sort of stereotyped responses which indicates that these are probably sort of in some way sort of innate responses but other reinforces have learned value through association with other more primary enforcers so an example could be you know the size of food or the packaging on one's favorite chocolate bar right they're clearly not something that is you know you you you will have assigned value to from birth but it's something that you have learned through experience that when you see that packaging it means that you can get this this nice chocolate for example and perhaps the most ubiquitous secondary reinforcer this what I should say a secondary reinforcer is what you call a reinforcer that's being nor has value through learning as opposed to having an eighth value and the perhaps the ubiquitous one of these is money of course clearly we all use money it doesn't have any intrinsic value in its own right it's simply something that it's associated with our ability to obtain other kinds of more basic reinforcers and although they do say money doesn't buy you love but it'll certainly buy you pretty much everything else so we probably agree on we enforce your value or utility of certain things such as primary reinforcers I should say again you know some people may like pain so there are certainly some exceptions to that but pretty much we probably agree that certain things are we or have fundamental reinforcer value but there's plenty of room to disagree and other things ie develop complex preferences because our learning experiences are going to be all very different now let's get on to decision making so and here is a canonical decision problem so this is sort of a a very good example of the problems faced by any organism a human or even an artificial system let's say a robot trying to make good decisions in the face of uncertainty so it's called the three armed or unarmed bandit problem and what you can imagine is you're in Las Vegas and and you've got these three slot machines in front of you okay and you want to win as much money as possible because of course that's what economics says you should be doing should be maximizing your expected utility win as much money as you can and but you the problem is you don't know which of these machines pays out the most money so you know let's say that one of them probably is better than the others but you just don't know which one of these it's the best one so question is how do you solve that problem how do you actually learn to choose the one that gives you the most money well I guess probably what you do is he'd sample the machines and see what you get okay and you know let's say here you sample this and you're in Vegas so you win $1 okay and you might try these other machines and you might do that sequentially through trial and error now what's going on or what would you imagine might be going on in your brain that allows you to sort of eventually decide that you're going to choose one of these machines over the other well what might be going on is that you might build up an internal representation in your head about the average reward the average amount of money that you're winning on the different machines okay based on your past experience and you might use that to try and predict the future okay so you say well if I got a if reward on this this green guy and I didn't get much on these guys well probably that means this guy's the one that's giving me the most money so what you might build up is a representation of the expected utility the expected average reward or expected award that you're going to get on the different machines and then you might then decide to choose one machine over the others based on the fact that the expected utility of that is higher now the other thing that has to happen before even you have an expected utility is you need to be able to understand that some things are rewarding and other things are not so clearly what's rewarding in this case I should go back here is winning the money right so there could be other flashing things going on and slot machine that probably isn't that relevant and not rewarding what you really want to know is where to discriminate things that are rewarding from those things that are not those things that have X utilities and those things things that are not so what you might need to do that is to be able to process or it a process you experienced utility ok so in other words you need to be able to understand that certain things give you pleasure and that they feel good to have them when you get them and this is what happens when you actually get an outcome okay so that some part of your brain needs to be able to say AHA this is rewarding so how does that happen well how how can we study this in the brain well simplest thing we could do is just give stuff to somebody of various various kinds of materials for these kinds of stimuli and and and you know this person will report what their pleasure is what their objective pleasure or experience utility is for these items ok and then what you might do is look for areas of the brain where neural activity correlates with those subjective reports of experienced utility ok so how might we do that well and one way you might do it is if you're able to have access to somebody Hastur and skull open you can actually go directly in and peer inside their brain now that happens on under only very limited situations thankfully and and in particular if somebody were to be undergoing neurosurgery and occasionally as scientists we can have the privilege of actually measuring activity in sitting inside a patient like that and but typically of course we can't do that and so what we can do though is make use of indirect measures of neural activity that we can actually measure without having to open up somebody's skull we can actually measure it safely and painlessly with the person being fully intact and able to go home and have a cup of tea afterwards and what we can do is make use of the fact and that when parts of the brain are active when neurons are firing and they're sort of they're processing information and they need energy to do that and in order for for that energy to be produced you need glucose and oxygen and that has to be delivered via the blood stream okay so capillaries near the neural activity deliver that glucose and oxygen and for that to happen when a part of the brain is is particularly active there's an increase in blood flow there's an increased consumption of glucose as the increased consumption of oxygen so we can sort of use any one of those indices and and measure that indirectly without actually having to go inside the brain then we can get a handle on what's going on inside the brain and that's what this technique called functional MRI functional magnetic resonance imaging allows us to do and what fMRI does is it makes use of the fact as I just mentioned and that there is a increase or an initial decrease in oxygenation so if a part of the brain is active and that the neurons will use up a lot of the oxygen oxygenated blood that's around it'll take the oxygen oxygen away and you'll go from up see haemoglobin to deoxyhemoglobin and but there actually turns out to then be a compensation so there's a signaling process and more oxygenated blood is is delivered to that part of the brain and there's an increase in the relative amount of oxygenation and the what's interesting is that up see and deoxyhemoglobin happen to have different magnetic signatures okay they have different magnetic properties and what we can do is use a machine called an MRI machine magnetic resonance imaging machine which has a magazine finish that's about 30,000 times more powerful than the Earth's magnetic pole and we can put somebody in there and and measure changes in the magnetic signals due to changes in oxygenation so that gives us a handle on what's going on inside the brain because when there's an increase in oxygenation it generally correlates with an increase in neural activity okay and it's localized so let's get back to our attempts to look at experienced utility and let's take experienced utility for a pleasant food smell so we probably all when you go past a bakery in your hat and you have to you get the nice smell of fresh bread you know that generally is a positive thing so it generates positive experience utility for you so how can we study that and well we need to brain scanner we need somebody in it so we've got that and what we can actually do is deliver different odors okay different food smells so here we deliver they're sitting in there and we're pumping in these smells and it's either the smell of banana or the smell of vanilla okay and and the other thing I should say is that the person is hungry so clearly a bakery doesn't smell good if you're if you're full it only smells good when you're hungry okay so these people are hungry and they're smelling these food smells and we measure activity to these both of these things and the the subjects are selected - both - like bananas and like let's say vanilla ice cream okay so they like both of these odors and now what we do is we take them outside of the scanner and feed them with as many bananas as they can possibly eat until their fall okay so these are fun experiments both for me and for the experice objects' and and what happens then is that the pleasantness of the banana goes down okay it's after you've eaten it to satiety you've actually no longer particularly want banana and the smell of bananas not particularly pleasant but what's interesting is and that the smell of the other odor the vanilla odor actually hasn't changed much and so because we have this property called a sensory specific satiety and the way to think about this is almost like a restaurant phenomenon so if you go to a restaurant and it's an all-you-can-eat spaghetti place very very very upmarket or you can eat spaghetti restaurant and you eat and a big place is Getty and then you have the option of another big places of Getty okay and you probably say I really don't want any more spaghetti I'm stuffed but then the dessert tray comes around and you're offered the tiramisu you say well actually I've got a bit of room left and that's actually quite adaptive it's very good because it's it's an adaptive phenomenon because it enables us to sample different foods and actually maximize the variety of nutrients that we get okay so that's it's a very good thing but it's also good for an experiment because we can selectively change the experience utility of this banana and actually look at what areas in the brain track the changes to the banana but that don't change to the vanilla so we put them back in and we scan them again okay and then what we look for our areas of the brain that show a change a selective change in activity neural activity to the banana so they should decree it should decrease to banana but not change the vanilla and the area that we found and and what I should say just to orient you these light colors these kind of orange and red colors here what they mean is areas that shows statistical difference as to statistically significant effect and that we're looking for compared to other areas that are not colored in and these are five different individuals and across these individuals what's common is an area down here which is the orbitofrontal cortex just above the orbits of the ice that's coding for the value of this of these banana and vanilla foods and we know that because we can look at the activity there and basically from pre to priya's blue pre feeding post feeding is is red and there's a big change in the in the activity neural activity to the banana and if anything the activity to the vanilla has gone up okay which means it's actually it you might might indicate that subjects actually find the vanilla a little bit more pleasant afterwards because it's not the banana which is now no longer Pleasant okay so that's one indication that this area seems to code the experience utility so another thing we can do so you remember I mentioned that money is a is it's a secondary reinforcer or you can even think about it as an abstract reinforcer because it's sort of been associated with lots of different primary reinforcers so it has very generalized value and we can actually get people to play simple tasks in the scanner while we're scanning them where they make choices and they make those choices and they can win or lose money and here it's it's not here this person's won 175 pounds unfortunately we're not that rich as experimenters so we can't afford to pay them that so this was just play money but even if you use real money in smaller amounts I should add you get the same kinds of results and we then measure activity and what we look for in the brain are areas that correlates with the magnitude of the money that they win so for example if they 175 pounds we'd look for areas you know for that and interview if they won thirsty we'd look for slightly less activity or if they won 300 we'd look for more activities so we correlate the amount that they win with the brain activity and what we find is again a region of the brain correlating with the amount that they win so the activity here is bigger in this region here when they win more money and it's less when they win less money so you can see that here and so that sense suggests that this region this orbitofrontal area again is more medial it's called medial because it's in the middle orbital frontal cortex is correlation with the reward value of money with the experienced utility of money the other thing that's interesting because you remember Bentham referred to pleasure and pain there's also region more laterally than the orbitofrontal cortex that goes up the more that you lose money because not only can you win you can also lose and this area seems to track B go up the more that you lose that might be tracking that the punishment or the negative reinforcer value other kinds of rewards could include and and pretty faces so as humans we actually like looking at other people and and we actually like looking at people that are more attractive than we do looking at people that are less attractive and so what we did here is we put people in the scanner we didn't actually tell them we were looking at attractiveness and and we just showed them some images and the images were either attractive people or on a track or less attractive people should I say I'm only showing you the people that were in the attractive cash we here I think it's not right for me to do to show the people that were less attractive and and what we did is simply just measure activity to faces that were categorized as being attractive by a bunch of other people outside the scanner and by the subjects themselves when they came out and we compared it to times when they were looking at the lower tractive faces and we again see activity in this venture medial prefrontal cortex correlation with the attractiveness so it's more active when you see attractive faces and less active when you see low attractiveness faces and I'm not showing it here but we also found this more what's called lateral on the side area correlation where with with faces that were rated as less attractive okay so again it suggests that the experienced or hedonic value of this and of these faces is represented in this orbital frontal region ventromedial prefrontal region this is another example this is not from my own work it's worked from and blood and Robert Satori and in Montreal and and what they did was instead of looking at money or faces they actually looked at music and so they actually played a set of tones to volunteers which are either in harmony or not in harmony okay so they varied the level of dissonance between between the tones and and in a very you know in the level of dissonance is high it obviously sounds like loads of cats wailing or it can sound kind of relatively Pleasant when it's when it's very harmonious and they just looked at areas that correlated with the degree of harmony and again the more pleasant and the more harmonious the the music was the more activity was increased in this medial orbitofrontal region so again it sort of seems to be representing this valley or utility ok so here's another example so probably a lot of you like wine as I do and I would here's that here's a wine testing task I'm giving you two glasses of wine and I'm just going to ask you to tell me I'm not actually there's somebody going to come down with all with the wine now and that's a joke and I'm just gonna ask you to say well which one would you like sample it and tell me which one you like which one is better should I say better why probably many of you would fancy your fancy yourselves has been good at doing that and but what if I were to tell you that actually the wines have very different prices so this one is ninety pounds and this one is only five pounds okay now do you think that would influence your you know your experience pleasantness for the wine many of you probably think well actually no I'm just going to be very objective I'm just going to sample the wine and it's not really going to affect my experience utility at all because actually I've got quite a bit of experience with wine you might want to say that and what if I said well they're actually identical wines they're actually completely identical and would you detect that and many many of you might think you would so what we did is we did a study of this so we put people in a spraying scanner and they were actually getting squirts of wine into their mouths so we can't give them a glass because it's very kind of tight in there and and they see the amount of money that the wine costs the bottle of wine from which the the wine is being sampled so it's either five dollars here's this was done in the States five dollars and then they get to swallow the wine then they get a rinse they get to swallow and so on just like a professional wine taster would and and there's they get they get two wines is actually another third wine as well but that's I don't need that detail and and we tell them it called wine one it's the same line but we tell them as either cost five dollars or forty five dollars so the same people get to sample the same wine repeatedly but they're it's different wines and they think it's got different prices and similarly for wine - it's either labeled as costing ten dollars or $90 and the actual price of wine one was five dollars and the actual price of wine two was ninety dollars okay and then we looked at activity in this orbitofrontal ventromedial prefrontal cortex area and what we found was that the activity in this region didn't really track the you know objective value of the wine whether it was 90 or whether it was five it actually was modulated very heavily by what subject thought the wine cost so activity was much higher meaning it more rewarding for wine that they thought would cost $90 and was much less for this for the actual $90 costing wine when they thought it cost 10 and similarly for the five dollar wine it was much higher when they thought it cost $45 so what that suggests is that you know our experienced utility isn't just something that's objective it's not something that is immutable it's something that's very much influenceable by various factors such as how much something costs so and let's get back to decision making because that's what the what the talk is going to be about and so we know that there's this part of the brain that represents experience utility now let's start to think about how we can solve this and this problem this banded problem remember we had these representations here I'm calling them V VA VB VC which are just coding for the expected the expected utility for pursuing these different actions so then the question is how do we actually learn this you expected utility so we can actually and to understand this make reference to a field of computer science or artificial intelligence called reinforcement learning and reinforcement learning is concerned with how to build artificial systems like robots that are able to learn from their environment and behave adaptively ok so choose things that maximize rewards ok and algorithms are coming from this field have turned out to be quite relevant for understanding how brains and how people might make decisions so the question is how do we actually learn to build up a prediction of expected of a high expected utility for a particular action and well these models say is that what goes on is we encode the degree of surprise in our predictions so we're always making predictions and that's our expected utility or decision utility and our prediction is either confirmed or disconfirmed by what we buy the experienced utility we receive and the difference between our expectation of what we expect we're going to get and what we actually get is the degree of surprise that we have and that's called prediction error ok and this prediction error I'll explain it a bit more down here this prediction error has been linked by and a very distinguished neuroscientist here in Cambridge professor Wolfram Schultz and to the activity of a certain class of neurons called dopamine neurons that are deep in your brain ok and so that they've been shown to sort of encode this prediction error and so what would it look like so let's imagine you're coming along you're making your choosing the slot machine and you have kind of very this is just this the strength of your prediction ok so it's either you know this is very low here you're not expecting anything you choose the Machine you come along you win your dollar and what the prediction error does is goes up because there's a big surprise you weren't expecting to get anything and you just did so you won you dollar and so then you come along and you update your expectation because this was your prediction error was high so you update your predict your your prediction you come along again you make a choice again and again you win your dollar so this time your prediction error is lower because you've helped you're less surprised okay and you still update your your value again so you come along with third time you make a choice this time you actually don't get the dollar because remember this is an uncertain world you don't win all the time sometimes you lose now this time you're actually going to have a negative surprise a negative expectation so you're going to have a negative prediction error okay so this could be one mechanism and we think there's a lot of evidence suggesting that that's how the brain or at least learns about to make predictions about reward and utility it does this through the activity of these dopamine neurons so we want to do is see whether can we look in humans and see how people first of all represent their expectations of future award in the brain where does that happen and then also what do we see these kinds of prediction errors kicking in as people are learning do we find evidence of that and what we did is we had people play disbanded tasks except they're not they're not three here there are four bandits and they're playing this they're making choices and and they they win points okay now the different bandits they're not and paying off the same amount all the time but instead and theirs they're very variable in the amounts that they're paying off and so in red here is the amount that you get on average across 300 trials of playing this game for choosing the red stimulus the red bandit and it starts off high this is the amount of points that you win starts off high and it goes down and it's kind of middling from there this this blue one starts off low and it goes up a bit then it goes down again and so what you can see here is that you know to play this game well you kind of want to keep track of your expected utility on the different machines because they're changing over time and you want to keep choosing the one that you think is paying out the most but you have to be prepared to always you know to shift around and so we had people play this game what we did was we basically took this out these this algorithm this reinforcement learning algorithm simple mathematical model and we correlated it against people's choices so the choices that they were making and the the rewards that they experienced and the model made predictions about what it thinks given your experience if this were going on in your brain what your expectations should be for your you know your decision utility your expected reward what should that be and you have this for each trial as you make choices you have predictions for each of the machines and those predictions will change over time and so for example here we might predict that the subject should think that the Green Machine pays out the most this is the average reward on the green machine whereas the red machine is actually not doing so well it's only paying out 38 points okay and this will change over time and we can basically just correlate brain activity where we're measuring it in the in the fMRI scanner against the predictions of this model and that should tell us where this expected reward is coded when they're making choices or it decision utility and where it's coded is in the ventromedial prefrontal cortex surprise surprise so activity in this area not only codes for experienced utility is also codes for expectations of future Awards when we're making choices so activities higher the more we think something is valuable is going to give us something good and activity is lower the more we think it's going to give us less of something good okay and this is just another example from a different fMRI experiment looking at the similar kind of process again it's coming bringing out the same brain region and we can also look for prediction errors and we do find these in the brain as well so we find these in a different part of the brain and this is more if you kind of look go through the middle of your head down right into the middle of your brain actually and in an area called the striatum the part of the basal ganglia and and this area actually receives very strong input or signals from these dopamine neurons that Professor Schultz found - echoed this prediction error and we find in humans activity that correlates with this prediction error in this place called the stration so what that suggests is it provides evidence that these kinds of algorithms might be going on in the brain to allow us to learn and we the point is we learn our predictions of whether something is good or bad through experience and we learn by the degree of surprise we have in our predictions okay now the next thing I want to talk about is how does the brain make choices between different types of rewards so an economist would say that if you if if you're encoding utility for something that it doesn't have to be you know the same good it doesn't have to be just you know the utility for money or the utility for food you'll actually assign a utility to anything and we can make choices clearly between very different kinds of rewards we might decide whether we want to go out to the to the movies or you might decide that we've law that go to a restaurant or C go to the theater so we're always making decisions between things that are often not very comparable in an invert in any obvious sense and so presumably what we do is we sort of assign some kind of utility to these things and they get compared on some sort of common scale common utility scale and so here's an example whose three items here's an iPod here is a you two tickets in Madison Square Gardens from 2005 so maybe that's not that useful now and here's a bottle of of Middleton fine Irish whisky and I should say this I if there are any Scottish people in the audience this is this is real whisky proper whisky and and and so you what you do is probably computer utility and you'd say well actually the utility of this to me is is 14 utils these imaginary units of utility and this Madison Square Garden ticket maybe for sentimental reasons is worth 19 utils here but actually this bit whiskey here is worth 2,000 utils okay now that would be my experienced immediate experienced utility for whiskey but sadly I appear to developed an allergy to whiskey so it comes and bites me back afterwards so as a rational decision-maker I'd actually say my utility for this and it pains me to say so would be zero utils so I'll make a decision between this by ranking them so the question is how does the brain do this and you know obviously what you'll do is you'll take the maximum utility and so what we might imagine is that somewhere you kind of encode decode what the item is you recognize it it's a bottle of whisky or it's an iPod or it's a ticket and somewhere you'll value that item you'll say well this is worth a certain amount of utils this is worth a certain amount this is worth a certain amount somewhere in the brain these values will sort of mix together so you can then start to compare them and decide which one you want so somewhere there should be a common area of the brain where which represents the values of all these items that allows you to make a decision and so here's the task that we had so we did this and this is done in Caltech II in in Los Angeles and and what we had is we had P put people in the scanner and we had them make choices between different kinds of items so a food item food items and I should say these are all these subjects are also hungry they haven't eaten for four hours and these are basically junk food items okay so they're packets of crisps chocolate bars you know stuff that's really not good for you that you should be eating but clearly something that people are motivated to eat if they're hungry and and we also had trinkets so these could be DVDs and CDs a here it's a hash that says Caltech on us and you know and these things or it could be you know a pen or you know something slightly less used for a less interesting and so the point is that and people will have variable you know interest in these items so they'll have variable expected utilities or decision utility to these items and say well I don't really want a hat I really don't but I would like to have this DVD of you know of Star Wars say and and also monetary Gamble's so you can take you know would you like an eighty percent chance of winning four dollars okay so the bait the basic idea is that they see these items and they're given a certain amount of money on each trial they're given four dollars and they can use some of that money to purchase the item okay and basically what that purchase amount tells you is how much they like that item because they would be prepared to give more money for an item they like versus an item they don't like and it turns out at the end and we just randomly I'm not going to go into the details of we randomly choose a few sets of items and they get to if they decided they wanted that item and they if they paid enough for it they'll get us if not they won't and but I'm not going to go into that detail but it's set up in a certain way and that they should be motivated to pay the right amount of money that they think it's really worth so this is sort of based on economic theory so we have these monetary amounts to tell you how much they think this item should be worth on scan on each trial we can look at at an activity so we get this willingness to pay and I should say what we actually do in the scanner is we just get them to choose between either the item or a fixed monetary amount which is actually and the median amount that they're prepared to pay for all the items so basically on average they should choose this item the items 50% of the time they should choose this 50% of the time okay and what we can see is just I'm just plotting their behavior here is that they're willing to choose the item over the fixed amount in proportion that a proportion to which they're willing to choose that is proportion to the amount that they say they're willing to pay okay so it seems that this is a very good willingness to pay is a good measure of their actual choice behavior and we can look in the scanners areas that correlate with their willingness to pay and what we find is and again not a surprise here so err the the ventromedial prefrontal cortex activity there is gray sure for items that they are willing to pay more for and less for items that they're willing to pay less for okay so again we have this correlation with and this sort of willingness to pay measure which is an index of the decision or expected utility so that's that makes sense but what we want to look at is areas that are commonly active for the when you're making decisions over each of these three different items whether it's money trinkets or snacks and the area we do find an area of ventromedial prefrontal cortex that's active whether you're choosing between a trinket or a monetary item or food item it just represents the value of that item in the world okay so that's some evidence to suggest that and at least this particular bit of the brain seems to encode the utilities of many different kinds of items whether there are food money or or anything else which suggests that we do code things on a common utility scale and we use that to make our decisions okay so finally I think in the last 10 minutes of the talk I want to sort of get to this sort of risk component of the toe of the talk which of course is the title and now you know up to now everything we've been talking about has been involving decisions under uncertainty and part of uncertainty is risk and but let's look at sort of risk more directly so let's say I give you a choice between a 50% probability of winning at 2500 pounds or 100 percent probability of obtaining a thousand pounds okay it's almost like Deal or No Deal and now probably I would say a lot of you would say I'll take the thousand pounds thanks okay but obviously if you say if you were to play this game infinitely and the average amount of money you'd get for choosing this one would actually be more so you Wyn on average 1250 pounds whereas with B you'd only win a thousand man so you can basically compute the expected value that the probability times the magnitude okay so you know you might say well it actually why shouldn't you be choosing a okay well obviously one reason why you might not want to choose a is that you know that there's a chance that you could walk away with nothing and let's say if you were somebody who had no money and you wanted to buy dinner that night well then you know the chat there could be a chance that you have no money at all and therefore you'll just go you'll starve whereas you know that with the thousand pounds you're sure to be able to buy I think a pretty good meal for a thousand pounds I should imagine okay so and so why is it that we're sort of we're sensitive to this can a variable called risk this probability fifty percent chance of getting of winning us but this other 50% chance that we get nothing okay well of course we can go back to economics again and sort of there Bernoulli first hypothesized and that actually what we do is we have a thing called a utility which again we've you know we've been talking about since the beginning of the talk and and rather than expected value which is basically expected value keeps going up linearly with the amount of money multiplied by its probability of getting it whereas utility actually tapers off so it's basically encapsulate this notion of diminishing returns and so basically if I have the you know my utility for winning two thousand five hundred isn't twice or sorry my utility let's say for winning three thousand isn't twice my utility for winning fifteen hundred so you can take the amount that you win or that you're offered oops and you can basically go up and just look at what that is in actual subjective utility so let's say take our two thousand five hundred pounds and go up here and it's actually not all right the utility that is not two thousand five hundred it's actually just under twelve hundred pounds okay based on this hypothetical utility curve you might have where as utility for a thousand actually comes out as being sort of just under 700 or just over it's about 760 here and so basically what that means is that and with a concave utility curve that looks like this and you know the value of 2500 is less than twice that of the value of a thousand pounds so it's not the case that it's it's it's worth proportionally more and so basically utility of a thousand would be seven hundred sixty utils the utility of two thousand five hundred is eleven seventy and you take the probability multiply that by twenty five hundred that works out is less it's not more than the utility of a thousand pounds so that explains why according to economists if you have this if you're making choices and you you represent the utility of things in this way it would mean that you're risk averse and of course the other thing you could do is have instead of this curve being shaped like this it could be shaped like that and what that mean is you actually raise things more as even more valuable proportionately than smaller amounts so it actually you'd have you would utility for bigger things would be greater and that would then mean that you are risk seeking ok so if I mean you'd really like to take the gamble and of course some people are is seeking some people will like to take that gamble okay a lot about instead of representing things as utility explicitly representing the degree of risk or in certainty and if you think about this let's take a gamble a fixed gamble and which has different probabilities and and the probability 50% could be up here probably 25% of winning let's say 100 pounds will be here probability a 75% to be here and so on and basically what you can see and this is your risk here that your risk is maximal for a 50% probability so it's only if you either winner you don't so it's a binary gamble so the for you know 50% it's a maximum risk you know that you you know you're that risk is maximum either winning or losing whereas for a 25% probability gamble you know most the time you're not going to win if you're only gonna win 25% of the time so your risk is a bit less and whereas also 75% you're going to win most of the time so your risk is also around the same and you know for sure if it's a 0% probability you're not going to win and you know for sure if it's a hundred percent you are gonna win okay so again your risk is very low so your risk is really maximal around this point where you know at a 50% probability and what we can do is basically give people Gamble's in the scanner and look and see is there anywhere in the brain that just represents this kind of this risk signal as that as the probabilities of taking Gamble's changes and I didn't do this but and Hurston pro chef and Peter both certs I've contacted and they had people play again and it's also you recognize it it's sort of play your cards right sort of Bruce foresight I don't know if that's still on or who does that now but basically sort of you you have a you're given one you you make a bet you're given one card and then what the bet is is the next card that's turned over higher or lower than the than the first card okay so you can actually batch on higher or lower and we do that before you see the cards so it's a little bit unfair and you see the cards and then there's another card so you by the time you see this card if it's a six or a seven you know what the probability is roughly that you're you know exactly what the probability is that you're gonna win based on the next card assuming it's it's drawn Eve with even probability with equal probability okay so basically you can actually quantify what the risk is for this kind of decision and you can look at activity in the brain and what they found is again in this region this ventral striatum region I showed you earlier that seems to correlate with prediction error you see activity that's correlating with risk so it's basically showing this kind of u-shaped curve so it's maximum at 0.5 and a tapers off as you either certain to lose or certain to win okay so basically the brain does seem to represent this kind of risk signal again Wolfram Schultz has found evidence for that this signal may be present in in these dopamine neurons and that projects richly up to this this area ventral striatum okay and here's another example and and basically here and people and so this is a study that Philippe Tobler myself Rey Dolan and London and Wolfram Schill stood together and here subjects see a variety of different cues and that tell them something about both the amount that they can expect to win the amount of money they can expect to get and then two points and also and the probability that they're going to win okay so we can vary that independently so basically some choices are going to be riskier than others and we can again just like with the bow certs and pro shovin bow start study we can look at areas that's correlated with with this risk which is formally the variance in the distribution in the reward distribution okay and what we also did was we measured people's individual risk preferences so some people are more prone to being risk seeking than others they would like to take a chance they actually favored risky situations other people are going to be more risk averse so they would be less likely to want to take a risky option they'd rather take the safe option okay and so people vary in that we can measure that and what we can show is that when people are making these kinds of decisions where they're seeing these different kinds of predictions of the amount that they might they might win or lose and activity in again this ventromedial prefrontal cortex correlates with their uncertainty so this is risk signal and across individuals activity is stronger in this region the more risk seeking the individual is okay so the more is seeking they are the more likely they're going to want to take the risky option and the stronger the activity in this region correlates with with this risk signal whereas in this other area this more lateral area again remember you this was also we found a similar dissociation for rewards and punishments this more lateral region in the pre front and the orbitofrontal is stronger the more risk-averse people are okay so maybe what you could think is you know maybe they're thinking more about the negative consequences here the possibility that they might lose if they take that gamble so we're starting to understand areas of the brain that seems to be correlating not only with risk and expected value but also with individual preferences differences across people in how risk seeking or risk-averse they are so I'm going to finish up saying to understand how the brain makes decisions and what I hope you've seen is that as neuroscientists we can't just start from square one what we have to do is actually anchor ourselves in ideas and they have a long history in disciplines such as economics and psychology to really sort of grapple with and understand conceptually how the brain might be representing things like value utility and risk and and we can make a distinction between a very simple distinction distinction between decision utility which is the utility of what you expect you're going to get and you can use that to into you enter that into decision versus experience utility which is the thing the value of the thing you've actually got you've sampled the hedonic value of it and we do see regions of the brain and that represent these kinds of signals especially the orbitofrontal cortex involved an experienced utility and we saw evidence for ventromedial prefrontal cortex representing decision utility we also saw other bits of the brain I haven't but I by no means being comprehensive in saying all of the areas of the brain that are involved in this other parts of the brain that represents things like risk and we also saw that our utilities for things or decision utilities can be learned through trial and error as through a process such as reinforcement learning or it can be instructed you can actually tell people what the probabilities are and how much they're likely to win and and we also saw as I said that risk is explicitly represented in the brain as our individual risk attitude and sort of open questions are we now have a good idea where our different kinds of utility or value signals are located but we don't really know much about how these things the utilities for things actually get compared in order to make choices so we don't know is that done in one part of the brain or is it a more distributed process you know what part of the brain actually makes the decision well which part of the brain calls the shots where does the buck stop and the other question is you know why do we have separate representations of risk and utility because of course the kind economist might say you only need to have this nonlinear utility curve you to account for risk preference you don't need to have this risk signal and we don't know why that is one possibility is that you need this risk signal combined with other very elementary variables to actually build a representation of utility which you then use to make your choices and the other thing I didn't touch on but is a really interesting question it's not just one system for making decisions but there are several and how do they interact and what leads to decision biases or deviations from rationality so that's another question that's really interesting and sometimes we're not rational sometimes we do deviate from what should be optimal and what goes on in our brains to cause that that's a very ongoing area of research and so back to Phineas Gage so probably what was going on to underlie his difficulty at making decisions or indeed modern patients who have damage to this area tragically is that these representations these fine-tuned representations of utility experienced utility decision utility are compromised and if you don't have those it's very hard for you to actually make good or optimal decisions and I just like to finish it would be inappropriate of me to be giving a lecture at Darwin College without you know making the point that the reason why these kinds of representations are present in the brain why they are likely are is because through evolution these kinds of signals these kinds of computations where the optimal ones that allowed us to make good decisions so presumably through evolution these kinds of computations have become honed and improved upon in order to allow us to make good decisions and to procreate and survive and flourish okay thank you very much [Applause] well it all gives a new meaning to the notion of looking into people's minds I thought it was a fascinating approach to the idea of risk as to see how we how we make decisions and how we make decisions by looking at which parts of our brains bounce around for and we're doing so I think the unsung heroes of this talk for me of the I imagine students who engage in drinking gambling eating mating and spending while lying inside a but I guess he's a big magnetic drain pipe it raises questions whether their behavior was altered subsequently and there's a sort of economist myself I there is one crumb of comfort I thought our subject was going to have to hand over completely to psychologists because bankers and others behave like demented lemmings but it turns out that psychologists need us and that's a very refreshing thought it was a lovely talk as a doctor it's given us some very stimulating things to think about and really things that take is very profoundly to the to the frontiers of the notion of free will and that's really interesting [Applause] you
Up Next

Dr. Wolfram Schultz: Dopamine Neurons & Reward Processing Explained
@NaturedNurture
1.3K views•2024-01-24

Deciphering Neural Circuits Controlling Anorexia in Mice | Richard Palmiter
@AllenInstitute
2.3K views•2011-10-10

Vagus Nerve (CN X): Anatomy, Nuclei & Functions Explained
@Alilamedicalmedia
305.2K views•2022-10-31

How Exercise Benefits Your Brain: Science Explained
@TED
11.4M views•2018-03-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Neuroscience















![[강화학습기초및실습] 01. Introduction to Reinforcement Learning](https://i.ytimg.com/vi/RLTFzcGvHFg/maxresdefault.jpg)























![[알릴레오 북's 16회] 자유의지냐? 운명이냐? / 운명의 과학 - 한나 크리츨로우](https://i.ytimg.com/vi/hTYZF_6xt_M/maxresdefault.jpg)