In economic games, how players learn from experience versus description significantly affects outcomes; players tend to underweight rare events when learning from experience (unlike the overweighting observed in description-based experiments), and high variance in payoffs slows learning, which has important implications for designing economic institutions like school choice mechanisms and auctions that must account for how people actually learn rather than assuming rational optimization.
Learning in Games: Ultimatum, Public Goods, & Institutions
Added:going to uh talk to about work that uh Ido Arab and I have uh conducted together and separately over a number of years Ido is over there so So the plan is that I'll actually talk for both of us but that he'll answer questions uh and um this is work that uh of course my talk may be abbreviated well okay this is work that that we started together some time ago so I'll start by telling you about some some an older stream of work in which we looked at Learning in games so so I'm not going to be talking to you about Evolution but about learning but the thing about learning in games is that um players learn about the game at the same time as other players are learning about it and so they're what they learn about the game depends on what others learning are learning so there's a kind of co-evolution of learning going on and uh we'll tell you about some success we had in explaining some some things with simple reinforcement models of learning that that if you wanted you could think of as being like discreet time replicator Dynamics I won't I won't talk in detail about the models in fact Ken binmore has done some work on that years ago about using replicator Dynamics uh and then Ido and I went in different directions from that he uh decided to look at at radically simple environments in order to study individual learning in its clearest form with without complications and I got caught up in uh uh trying to design complicated Economic Institutions but our but our work has uh has has some convergence now in particular one of the things that Ido and his colleagues discovered is that how people learn depends a lot on on how they experience the environment so I'll talk to you about that and that'll bring us back to the question of of how do Economic Institutions uh work well and why do models of of Highly rational calculating actors often provide very good approximations when we don't think that that's what uh individuals are actually doing so to put it another way a lot of economic models are models of equilibrium again this fits with the discussion of replicator Dynamics but but mostly we don't have very good models of equilibration models of equilibration that we think actually describe what goes on in the world um but what we observe is that some games move quickly to equilibrium and some games move much more slowly if at all and so we'd like what we'd like a learning theory to do for us is is help us predict which of those will be which so let me tell you about two games uh which may already be familiar to many of you uh they're similar in various ways they're similar in the way they're play their structure and they're similar in their perfect equilibrium predictions so the first game is the ultimatum game and it's a two-player game where each player has one move player one proposes a division how much he should get and how much player two should get so typically you say to player one something like you have $10 divide it between you and player two player two can either accept the division proposed by player one in which case they each get the payoffs that player one proposed or he can reject it in which case they each get zero so there's a simple argument that says if they only care about how much they get then player two is going to look at this and he could get zero or X2 as long as X2 is not smaller than zero he'll accept therefore player one knowing that whatever offer he makes will be accepted uh can give a very small amount to player two and uh take the rest for himself but that's not what we observe in the laboratory on the contrary what we observe is that player twos who are offered very small amounts tend to reject them and of course it's not very costly for them to reject them because they're being offered a very small amount now here's another game which is a public goods game we've heard some public goods games discussed already but this is a very extreme public goods game but it's similar to the ultimatum game not not just in that it's equilibrium will be extreme but that it's a two player game each player has one move player one chooses a quantity of the public good that he is prepared to supply player two learns this quantity and chooses a quantity himself and then there's a total public good that's supplied I'll come back to that in a minute but the payoffs to the players they get some reward for the total public good minus the cost of how much they provided but but how much is the what is the total public good what the reason this is called a Best Shot game is the total public good is the maximum that player one or player two provides the best shot and what that means is that the contribution of one of the players will be always totally wasted because it won't contribute to the total quantity that exists but it will be costly to supply so there aren't going to be equilibria in this game where both not going to be perfectly equili where both players uh provide a positive amount only one player will provide a positive amount whoever provides the larger amount determines what the rewards are and the smaller amount is completely wasted but it's costly to provide so in the parameters that we're going to be using cost 82 cents I think is that right uh yeah per unit so when you look at the Redemption values of the units you what what will make sense player one goes first remember so he has the option of of being the one who provides nothing he can provide zero which is what he does at the perfect equilibrium um and player two is left to decide how much to provide if he's simply calculating how much to provide he should provide four units because the margin cost of the fourth unit is the benefit is above 82 cents the cost is 82 cents and so the equilibrium is a very unequal equilibrium player one will provide nothing he'll get the benefit of four units um so uh where am I four units so he's going to earn $3.7 player two will get that same benefit but he'll have to pay for the four units he's only going to earn 42 cents that's a 9 to1 uh ratio when we see ratios like that in the ultimatum game they get turned down they get rejected so of course in the best shot game one thing that could happen is player one will provide zero and player two will provide Zero by and large that's what happens in the ultimatum game when player one proposes that player two should get something very small but that's not what we see in these two similar games here's the ultimatum game results and the equilibrium prediction remember is that player two will give will get something very small and accept it but what we see is in the first periods player ones propose on average almost half you know not quite 40 not not not a little more than than uh 41% uh and in the 10th period there they are still on average giving around half so player ones for whatever reason in this ultimatum game will talk about it they are not proposing that player two gets nothing they're proposing that player two gets almost as much as they do now here I'm going to look at the best shot game under what's called partial information this this means that the players do not know each other's payoff they only know their own PS okay and I'll show you under full information in a moment so under the best shot game in partial information player one the prediction remembers that player one will provide a quantity of zero but in Period one he starts out providing quite a bit on average but by period 10 the average quantity is less than one right so many player ones are providing zero you have to provide discret integer units of of uh quantity one um and these are are uh well let's take a look at what happens but but I'll show you the player two's decisions in a minute but one thing you might wonder is what will happen in full information in other words these are pretty different player one is providing very little as predicted by equilibrium in this game he's providing a whole lot contrary to what's predicted by equilibrium in this game but in but in this game they they don't know the other players payoff so if if the motivation is considerations or fairness those just can't operate in the partial information game but if I if I start to inform the players of the other players payoff then maybe if if it's fairness that makes payoffs equal maybe they'll become equal in the best shot game but that's not what happens when the players know each other's payoff player ones start off offering a bit less but but they're still offering a positive quantity but they quickly get down to everyone offering zero zero with you know zero error so they go the the player ones learn to provide zero units of the public good when they know the payoffs of both players okay so let's quickly take a look at at sort of scatter diagrams of the data here's round one of the partial information game and what you see is there's plenty of people providing both player one and player two are providing some units of public good that's that's wasteful okay it's inefficient but very few player ones are providing zero in the first period but by The End by round 10 nobody is no no pairs are are both providing positive quantities and lots of player ones are providing zero units and many player twos are responding by providing four units and here's the full information condition what you see is that even at at period one lots of player ones are providing zero starting initially and by the end again no one is is there there are no joint provision of of wasteful amounts of quantity 1 and quantity 2 and lots of player 2os are not responding with zero but they're responding with the perfect equilibrium response which is they provide four units so these are games that that are similar to each other but go to different equilibrium here's the ultimatum game data uh in round one player one is offering a lot of fives out of 10 a lot of fours out of 10 some small numbers the the lighter numbers indicat some of those offers are rejected by the end very small offers are not being offered anymore offers less than five are still sometimes being rejected and what we see at the end is not that we've gone to equilibrium of offering zero but that uh in fact the the mode is is a pretty fair offer between player one and player two so these these two games are pretty different from each other in their outcomes but they look pretty different in their format and in their um in their equilibrium so what's going on well one thing that's going going on that that this summarizes a lot of data is in the ultimatum game if if you were giving advice to player one and you looked at the revenue that accured to player one by offering different amounts X2 what you'd see is that the maximum payoff to players one come when they offer 50% when they offer five out of the $10 to player two that's something that very seldom gets rejected everything else well I don't think offers of six got rejected but but when you offer six you only get four yourself so you so you make less and if you offer uh if you offer only one to uh player two when he accepts you make a lot of money you'd make nine but he rejects often enough that your expected return is less on the other hand in in the best shot game the story is very different when you order the equilibri when you offer the equilibrium amount you get your highest payoff because as soon as you offer a positive amount it becomes profitable for player two to offer to offer zero remember it's wasteful for and but cost ly to provide this public good so player ones who uh provided positive quantity found that player twos didn't provide a positive quantity and that therefore their payoff was very low and that eventually their largest payoff came from offering zero so that's a learned response these games don't look that different in uh in the uh the first period and but the difference we propose it has to do with the feedback you get off the equilibrium path player ones who make close to the equilibrium offer of they offer very little to player two in the ultimatum game they get rejected and they when they deviate from the equilibrium offer they get a higher payoff player ones in the best shot game when they deviate from equilibrium don't get rewarded they get a very low payoff uh but they get a high payoff when they uh when they stick to the equilibrium um so reinforcement learning models do a pretty good job of explaining this and again I won't go into the the details of different reinforcement learning models but we look at three time models that that that track the game but you can think of them uh as if they were uh replicator Dynamic models and uh what goes on in the ultimatum game is that proposers learn not to make small offers much faster than responders learn not to reject them so if you go back to the data what you see is there are no more small offers but there are still offers being rejected if they're less than 50/50 and it makes sense that that should happen because when you reject a small offer it doesn't cost you very much so so you don't learn very quickly not to reject small offers but when you make a small offer and have it rejected that costs you a lot so you could learn quickly in these models to uh to not make a small offer on the other hand in the best shot game player ones don't get much reinforcement from deviating from equilibrium so they they learn to Free Ride they get higher payoff on average when they provide zero and then player tw's learn to provide four so so we converge to equilibrium there so we need a theory of learning that would maybe be more robust than the ones we had been using uh in those days and I want to tell you about those um uh and what we'd like a theory of learning to do for for games is tell us what players will learn from the feedback and also incidentally how the information about the game changes something like the best shot game right so we see a different kind of convergence from a different initial distribution in the best shot game than in the with full information than with than without information about the other's payoff so these reinforcement learning models like replicator Dynamics don't have a place to put that information that that difference the way it shows up is different distribution of first period behavior that appears through some kind of cognition that players have when we give them more information and then the reinforcement learning Dynamics do a good job of picking up from this new different initial distribution so one question is how how do these initial different distributions come up and so what Ido and his colleagues did is they started um looking at very simple choices over lotteries but instead of doing what became famous through the approach of con and tki instead of describing the lotteries to their to the participants in their experiment verbally telling them what the probabilities were they just let them experience the lotteries gain experience by playing the lotteries and seeing the outcome so they have a quote in one of their papers that says uh that describing the lotteries saying here's a lottery that gives you a 10% chance of this payoff and a 90% chance of that payoff that became the fruit fly of choice research and it it sort of drove out other kinds of experience that we might be interested in studying so what they started to do is they simplified a great deal and they said let's just look at experiments where where all you're going to do is push two Keys one of two keys and your task is to click one of these two keys and when you do that you will see the payoffs that you got from the key that you clicked and also the payoff that you would have gotten had you clicked the other key so your your feedback after a trial of this experiment which you can of course do very many trials would say you pressed the right hand button you got to pay off of a a dollar or maybe it's a sheel um had you pressed the left hand button you would have gotten zero that's what you learn and now rather than being told what the probabilities are that generated those numbers you're just going to be invited to push them lots of times and we'll see how your behavior evolves so what they find is that while in many circumstances this leads to behavior that looks like maximization there are regular circumstances in which it does not so for example here's a pair of buttons one of them gives you zero with certainty and one of them gives you a 10% chance to lose uh1 and a 90% chance to win $1 so it has a negative expected value but it's being picked more than 50% of the time over uh these are blocks of maybe these are 500 periods is that right uh okay so over 100 periods uh you are not learning that this is a a bad gamble uh on the other hand here's a good gamble here's a a similar one where you can get zero from one button and this Lottery from the other 10% chance you win $10 90% chance you lose one it has a good expected value but they only get up to about 25% a little higher than 25% of choosing that so they're not maximizing expected value and what it looks like is that they're under waiting rare events that is you're you're learning from experience you never see this number 10% you just only win $10 10% of the time and what it looks like you're doing is you're saying you know I don't win very often I mostly lose a dollar whereas I could get zero for certainty if I push the other button and here you're saying you know I rarely uh lose $10 I'm mostly getting a dollar where I'd get zero if I got the other so I'll I'll stick to this one so these are non-expected Maxim value maximizing decisions and they look like you underweight rare events and one reason that's important is Conor and tki developed a whole theory of of how you would choose these lotteries among these lotteries if they would describe to you and what they found is you overweighted rare events right they found just the opposite so they found that if they describe to you uh you know a small probability of losing $5,000 uh you know in in a hypothetical choice you you would choose losing $5 with certainty and if they gave you a small chance of winning $5,000 you would jump at the chance uh even when you were were facing things with with the same uh expected value and they attributed this and they built into prospect theory the idea is that you look at small probabilities and you overestimate them but when you're but what what Theon as colleagues found is that when you're uh looking at uh these from experience that is no one shows you the probabilities you just play the game and and learn what to do you do the opposite you you um you you seem to uh underestimate the low probabilities and they refer to this as as the experience description Gap so you can start to see where I'm going in the best shot game we're doing some description and some experience when we tell you what the other players payoffs are that's the description of the game and then there's the experience of playing the game uh so you can see these things in the in in very controlled ways in this simple environment with with the same decision so here's they describe to you the lottery and then you get to play it many times you know 100 trials and what you see is here's a lottery that has a a 120th probability of giving you a reward so it's a good bet it's a 3.2 expected value compared to getting three all the time and um and you over you know you react to that it looks like a big probability and you choose it more often than not in the first trial but less Often by the time time you've played 100 times even though it has a high expect higher expected value you after all aren't very often getting the big reward you're mostly getting zero whereas you could have been getting three with certainty so that's a hard choice to make it's hard to be that that disciplined um here's another uh here's a study by uh Greg Baron and Steve lighter and Jen stack uh one of those people was was Ido student and the other two were mine and what they do they do an interesting experiment with similar lotteries they say a certain gain of 10 cents versus 13 cents with a very high probability and a big loss $15 with probability one in a, a small enough probability that people are not going to experience it in the 100 trials that they're going to have but they have two conditions in one they describe the game to you at the beginning that's they give you a warning right at the beginning and then they say now you can play you can play 100 trials and and get the the payoffs and uh here's the curve the the probability of taking the risky action starts off lower than if they don't tell you uh the about that probability but so that's not surprising but then after 50 trials they tell the other group they say you've been playing this game for a while you haven't experienced that $15 loss but you know there's a risk that you will we we didn't tell you about this but now we're telling you so starting in trial 50 the two groups have exactly the same information they've seen the same sequence of random events and they've gotten the same warning but they got it in different order these guys got the warning at the beginning and these guys got it at the end and they don't converge they stay different and the way these the the one of the things they were thinking about when they did this experiment was uh there was a pain medication called viox that uh was discovered to increase the risk of heart attack and new prescriptions just fell to the floor you know no one wanted to take a pain medication that increased the risk of heart attack but people who were already taking via who' found that aspirin didn't work for them and that this worked pretty well they like to keep taking it right they hadn't had a heart attack it worked for them it was it was a good medicine and so they were so their behavior you know the the guys who got the experience before the warning were different than the guys who got the warning before the experience so there's something going on there there there's you know people are complex and even after they have the same information how that information has accured to them uh changes the way they learn and what they conclude now one of the old uh things about um one of the old observations about learning was variability sometimes slows learning but but here's another look at that which says many of the things that in the conent derski framework looked like risk aversion you know you switch from uh picking the action you could get zero for sure or you could get an action that gives you one well everyone uh learns to to choose to maximize if you instead of getting one for sure you get this gamble that has expected value of one it's a lot harder to learn to choose that gamble um compared to getting zero so only 58% people are choosing it over over time uh but you could interpret that as risk aversion or loss aversion but here you see it the other way if you were risk averse or loss averse uh then then when you go from minus one to this minus one gamble that has much more variance you uh you know you should drop away but in fact these numbers are pretty similar and what seems to be going on is that uh variability it it's harder to learn about a gamble that that has high variance so let me show you something that my colleagues and I did um in games again in a more complicated environment because because I want to talk to you a little bit about this co-evolution of learning so I want to tell you very briefly about 10 period prisoners dilemma games so I'm going to guess that everyone knows what the prisoners dilemma is what the 10 period prisoners dilemma is you're going to play 10 rep repetitions of the prisoners dilemma with the same person and then you're going to be matched with a different person and play 10 repetitions of the prisoners dilemma with that person and so forth okay you're going to do that 20 times so you're going to play 200 experiences in groups of 10 and uh as I say the the hypothesis we want to look into is that variance slows learning and what you have to learn in a 10 period prisoners dilemma is that you'll make more money over the 10 periods if you start off by cooperating and cooperate for a while okay so here's what so one one condition of the experiment was because these payoffs are you're you're going to acue monetary payoffs uh we did these for historical reasons we did them in numbers that related to roulette probabilities but these aren't probabilities these are if you choose a and the other place chooses a you will each get 10 and a half cents credited to you and after you've played 200 rounds you'll get paid your crude amount and in this you're going to see not just what you get paid but you're going to see that you chose a and the other chose a you're going to see what what everyone did but the other condition the next condition of the experiment is instead of getting paid 10. half cents you're going to have a 0.
105 probability of getting a dollar okay and that'll be independent of the other guy's probability you'll still see what he did you'll you you might not get the good payoff but you'll see that he cooperated if he cooperated you'll see that he defected if he defected okay so in this condition the the they see the same Matrix matrices but the payoffs are probabilities the these payoffs are probabilities and what they see here are zeros and ones you either get a dollar or you get zero so sometimes you will both have cooperated but you won't get the cooperation payoff okay so here are the deterministic some some deterministic data and I'm showing you these by trials rather than an aggregate to show you that we're seeing a very robust phenomenon okay in blue I have rounds one to five so in rounds one to five actually a lot happens in those one to five so there's already learning in the first five trials but what you're seeing is pretty high cooperation in round one trailing off to lower cooperation in round 10 and then the whole curve shifts it tilts so that cooperation gets probability of cooperation gets really high in round one as people learn that if you cooperate in round one the other guy will cooperate in round two and you can keep that up for a while but in fact cooperation drops off a bit by round 10 by round 10 the other guy is going to be defecting so that whole curve shifts you're learning to cooperate as the other player you're playing with learns to cooperate and that happens every time we do it these are the conditions with the probabilistic payoff and there's just no learning right there's somewhat declining cooperation as you go along but it's the same in the first five games you play as in the last five games you play they just fail to learn to cooperate they they don't learn those High cooperation so this looks like like a win for the hypothesis that uh variant slows learning to learn to cooperate in the prisoners dilemma the repeated prisoners dilemma the person you're playing with has to also learn to cooperate because what makes your cooperation valuable to you is that it elicits their cooperation if they are not learning to cooperate then it won't be valuable to cooperate and you shouldn't learn to cooperate okay but there's another hypothesis that you could go on which is maybe variance just makes cooperation less desirable maybe it's less fun to cooperate if sometimes I get a dollar and you get zero or vice versa but it turns out you can you can control for that by looking at by rematching every period so you playing one period prisoners dilemas and when you do that deterministically what you see is that you get a reasonably High rate of cooperation but it deteriorates pretty quick to some base rate of you know 10% or so of provision of public goods just just out of good feeling this now if random payoffs made cooperation less desirable then we should see less cooperation when we give them random payoffs but if it just slows learning then we should see more cooperation because what they're learning in this game with deterministic payoffs they're not learning to coop they're learning not to cooperate and in fact we see more cooperation they learn more slowly not to cooperate they get there but it takes them longer so so this is co-evolution of behavior in the game through learning and when we slow The Learning in this game we make it go away that is instead of slowly learning to cooperate there's actually no point in learning to cooperate if if your partner is not learning to cooperate the benefit of cooperation comes only from cooperation okay so uh there are other turns out the eo's observation idon and his colleagues observation that that there's this big experience description Gap has has uh a lot of consequences that that that follow from it here uh just I I won't talk more about this but this just shows in in some very simple artificial Commodities that index funds it's hard to learn to love index funds they have you know lower variance than the underlying stock but they it's very tempting to follow the underlying stock to to you know Chase big returns and uh flee from past losses uh but what it what turns out to model this Behavior well this this description Gap is uh predicting from small samples so here's here's the list of the gambles I've showed you so far and their expected value uh you can always get zero by choosing the status quo and here's you know what people have learned but here are the predictions from a model which says you get a sample you look at a sample of past actions and and from that small sample that might have four or five observations you pick the best one and and go from there so uh so so just incidentally one of the things that that I've collaborated with EO on is is U we've run competitions open competitions where where we present some experimental data and invite people to model it and to predict a a a prediction set a different set of experimental data drawn from the the same distribution of games in the same population and we ask people to submit programmed models that will try to predict these data and the ones that work well tend to have um small samples so this seems to be there's some reason to think that that may be what's going on in in some of our data so why might people choose to use small samples well one of course is it's bounded rationality right it it might be really hard to remember 500 periods of play and figure out the best thing to do on the basis of them another might be you know in in in these very simple things where you just learn from experience what's going on EO certainly doesn't promise you that you are dealing with a stationary distribution and has run some experiments where you're not dealing with a stationary distribution and if you at all suspect or want to explore the possibility that you're not dealing with a stationary distribution then then small samples of similar things might might be very relevant to to make your decision so um so one of the things you can look at is how do people choose those samples if they're using small samples what kind of pattern recognition might they be using to try to find good samples so for example here's a series of you can push two buttons you push this button or this button on the first trial this one yielded minus one and this one yielded zero and the second and the third but on the fourth this one yielded two and that yielded zero three more minus ones and then a plus two three more minus 1es and then a plus two three more minus ones what do you choose now well a lot of us and a lot of the experimental subjects would say looks like we're due for a two right and pick a two uh even without the colors right that's a pattern that that looks regular enough and it doesn't have to look that regular even even after just one experience that after three minus 1es are a two a lot of people will will go for the two after three minus 1es so and and depending on what the underlying generation function is that might not be crazy so among the experiments that have been done are experiments where the numbers are generated by in a non-stationary way by a Markov process that changes States and that as you get minus ones it changes States and makes a A plus two more likely so here is some data from such a thing here's a lottery with independent Trials of uh minus 10 or plus one and you know people don't take that very often but here's one where there's some predictability you get a minus 10 every uh eight or 10 trial based on this underlying unobservable Mark of process and people get pretty good at picking this and they get pretty good at picking it when it's good to pick right when when the deck gets too rich they go away from it and come back to it when it looks like U it'll it'll pay off so you know a little bit like counting cards in blackjack you can you know that the distribution isn't entirely stationary you could learn to do something with that okay and um again you know so so this is actually the the experiment I was just talking to about this is a three-state Markov process and uh subjects learn to play it pretty well in the sense that when when there's going to be a plus one they choose it a lot more than uh than when there's going to be a minus 10 uh but what you're what you're seeing here is depending on the probabilities the whether or not they learn to maximize whether or not they learn to choose the high expected value gamble changes a lot and one of the things going on there is the uh underestimation of of uh from experience of uh small probabilities okay so right now we can't tell you very much about how people choose similar situations how they choose a sample of similar situations but it seems likely current hypothesis from from this kind of work is that is that people do that not in a crazy way and then and then react based on past experience so um so the question of of you know what can we say even without knowing the the exact nature of this is is how I want to come back to the larger economic questions that that motivate this which is how can we think about different Economic Institutions that that will help people learn to behave in sensible ways um so you know rare and unpredictable events are important right so so one of the things that that you can study for instance is things like you know flood insurance do we underinvested flood insurance because year after year there's no flood and then there's a flood um and and if we think that's what's going on then how can we change the information we give to people about insurance and the risk they face similarly wearing seat belts or or investing in bubbles um and and so my work has has been sidetracked in recent years into trying to design institutions of various sorts that uh that that will be the that'll promote welfare and so right now the way the way our our thoughts are coming together on this is you know often the institutions we design are well modeled as if we think it's a good approximation to treat people as if they were rational and the question is why is this seems to be that that learning and R and rational models will will say something similar if the experience if the thing you get that maximizes your welfare also often gives you a good outcome right so that doesn't have to be the case but what you're good at learning are things that happen a lot what you're what you're bad at learning are things that happen seldom and so the question is how to how to create the experiences that that will give you the right kind of feedback so for example one of the things my colleagues and I have been doing over the years is designing school choice mechanisms in American cities uh and the school choice mechanisms that that we've been designing in for New York high schools or for Boston Schools or in Denver and New Orleans have the property that that families are asked to give preferences over schools and it turns out that in contrary to past ways of doing this the the new mechanisms that we use make it a dominant strategy to reveal your true preferences they make it safe to confide to the school system um which schools you would really like even if you have only a small chance of getting into them now when we introduced this in Boston some years ago we were replacing a system in which it was in which you came to Great harm if you didn't get the school that you listed as your first first choice and therefore you had to be very careful about what school you claimed was your first choice so um these are subtle differences and you could hardly expect people to learn them from experience so the description is going to matter too that is part of the the design job is not just to design something that will over time give good incentives so that when people talk to their neighbors they'll get reinforced to do the right thing but to give a description a communication package that will allow them to know that the world has changed and that the advice they used to get uh is no longer accurate because there's a a new mechanism we want to start them off you know properly in a new period one so when we first rolled out the the Boston school choice mechanism there was a brochure that went out in backpack mail you know you you send it home with kids and things like that and on some page in the middle of a big booklet it said it explained what I just said to you and then it said you know if you have any questions here's a telephone number you can call and the telephone number rang on the desk of a man named Colton Jones who was an employee of Boston public schools and who understood very very well the differences between the old and the new system and the the good properties it had and so he started getting calls that that he said were were were a little tentative people would call him up and they'd say I'm calling because on on page six of the bore if I understand it correctly it seems to say that that it's safe for me to list the true choices of schools even if uh if we have only a low chance of getting them and he'd say that's right you you've understood correctly and they'd say ah you know cuz the reason I asked is we went to the Family Resource Center and then the nice ladies at the Family Resource Center explained to us that we had to be really really careful what we put as our first choice which had been true the previous year okay so so what what had the glitch that had happened in the first year of the roll out of the of the Boston school choice was that the description didn't match the new environment and now when we talk to American cities we tell them not just we'll help you design an algorithm but we'll help you design an algorithm and a communication package so so one thing that we're trying to learn from this is is how do those things go together if you know they they used to tell people in Boston that they should put down their true preferences but people pretty quickly discovered through the experiences of friends over the years that that wasn't a good idea that that was a bad advice so you so people learn you can't expect them to put down their true preferences when you're going to harm them if they do but on the other hand it's hard to learn that the system has changed and it's now safe okay that's a very hard lesson to learn let me mention that how you how you learn things like that can also depend on the kind of feedback you get so here's this is my last slide here's um a graph of some experimental results from some different kinds of auctions in an in a laboratory experiment and the and what I'm graphing here sorry the the the axis aren't labeled what I'm graphing here is time uh repetitions and here I'm graphing what percentage of people's True Value they are bidding in the auction and it's 100% up here and the different auctions are a bunch of auctions motivated by eBay and other iterative mechanisms where you bid if if if you you're not the high bid if anyone wants to bid more you can bid more you see whether you're winning bidd you were the winning bid or not and so forth and they're comparing it with the second price sealed bid auction now the second price sealed bid auction the way that works is you write down or enter on a screen uh the bid you want to make the auctioneer opens them up the high bidder wins and pays the second highest price and it's not hard to show that in a private value auction of the kind we were talking about here it's a dominant strategy to bid you True Value because of course you don't pay your bid your bid just determines whether you win or not what you pay is the second highest bid and that together with with some other things makes it a a dominant strategy but it's hard to learn that so here you see that so it's a dominant strategy this is the only of these auctions that where you have a dominant strategy it's a dominant strategy to bid 100 % of your true value but it takes a long time to learn that why is that supposing you think that maybe it's smart to bid less than your value in many kinds of auctions that is smart well you bid a low value you lose the auction but you don't discover that until the auction is over now you have a second auction you bid higher and so forth whereas in these auctions if you start off with a low bid then you're not the high bidder and the auctioneer says does do you want does anyone want to bid again you say I do you know that that object is worth a lot to me I started with an initial high bid but I immediately understand in the C before the auction is over that that's a mistake I get feedback that allows me to see that I can't win a valuable object for a low price I have to I have to pay more so in these auctions although it's not a dominant strategy to be bidding your true value they learn much quicker to bid their true value there there's a related result here that that says you should keep going until you get to your uh at least higher than the second highest price so these two mechanisms if you wanted to analyze them as if people were very smart you would not predict that there was a difference between them and if anything if people were really really smart they would the dominant strategy argument should convince them but in fact people form Impressions about what to do and then they adjust them based on their feedback so both the description and the experience matters and other people's experience matters when you're playing with other people so when you do Market design when you when you try to design New Economic Institutions whether they're school choice or labor market clearing houses or my colleagues and I have been building kidney exchange mechanisms uh over the last couple of years uh whenever you start the mechanism is new so there's no reason to think that people will start immediately at equilibrium they're going to have to learn what to do by experiencing the mechanism and one of the things you want is in in Practical Market design is not just to have a good equilibrium but you want to make sure that people don't die while getting to it that it has good properties on the path to equilibrium stop there
Up Next

NRMP Match Algorithm Explained: How Residency Matching Works
@nationalresidentmatchingpr5234
17.7K views•2025-09-12

Triumph of Orthodoxy Icon: Byzantine Art & History Explained
@BenCallan
2.1K views•2024-08-06

FastAPI vs Flask vs Django: Choosing the Right Python Web Framework
@TechWithTim
302.5K views•2024-05-26

Game of Thrones Opening Credits: A Cinematic Analysis
@gameofthrones
46.3M views•2011-04-18
Related Study Plans & Knowledge Roadmaps
Structured learning paths in General & Interdisciplinary Studies



















![Plenary Talk by Ariel Rubinstein on 17 12 2015, ISI Delhi [2/3]](https://i.ytimg.com/vi/Re1njz8mAnk/sddefault.jpg)



















