In infinitely repeated games, the One Deviation Property allows verification of Subgame Perfect Nash Equilibrium by checking only one-period deviations rather than all possible deviations; grim trigger strategies can sustain cooperative outcomes (like mutual cooperation in the Prisoner's Dilemma) as SPNE when players are sufficiently patient (discount factor δ ≥ 2/3), as the threat of permanent punishment deters deviation.
Infinitely Repeated Games & Grim Trigger Strategies | Game Theory
Added:in this episode i'm going to talk about infinitely repeated games well what does that mean that means the horizon of the game is infinite so there is no final stage game there is no end to the game the game will continue forever once again it doesn't sound realistic it isn't but we just use those games to picture environments where players do not foresee any immediate end or sort of foreseeable end to the game all right well obviously this brings some technical issues for example what is going to be happening to the payoffs well if you think of a payoff for example a t from zero to infinity uh u i a t so delta to the power t well here well the question is are is this going to be a bounded payoff right i mean today i'm gonna get payoff two tomorrow three and then three periods later four or five so it's like you know a bunch of positive numbers when i add them up may i actually ending up some sort of infinite number because i mean a huge number well probably it's going to be huge right because it's an infinite horizon but thanks god because this delta is very i mean in between zero one the payoffs will be bounded all right meaning it is not going to conver diverge to infinity all right so therefore this delta is actually very very important parameter in the payoff calculation so that the uh you know the total payoff in this repeated game does not blow uh uh to infinity all right so it's going to be bounded well the other problem is how am i going to use a backward induction because there is no final stage or final period how can i build up my spne well remember the uh one deviation property this is exactly what we are going to use uh so in games well the the the version of the theorem one deviation property uh we described was true for finite horizon extensive form games but this is infinite horizon but the a similar version of this one deviation property we we described in our earlier videos holds if the payoff in an infinite horizon game is bounded so therefore for our infinite horizon repeated games we can actually use one deviation property well again for the details please go back to this video but in a nutshell the idea was the following you do not have to you know build up spme from the final sub game because there is no such uh sub game in this game infinite horizon repeated game but what you can use is that given a strategy profile you know strategy of each player if you want to check whether this is an sp e or not what you have to do for any player and after any history right remember every history defines a sub game what you should check is the following so for any history look at the the player uh players uh so history that follows uh player one's move all right so look at player one's move does he have incentive to deviate well but there there might be a bunch of different deviations right he can deviate in in this period and then uh in the next period or maybe 10 more periods and then he may actually not deviate sort of play according to the strategy well the one deviation property says look only at one deviation in this period deviations or periods t plus one so given a history of length t look at the deviation in periods t plus one all right um well what about for the rest of the game for the rest of the game uh let's assume these players are going to stick to their strategies all right so the player deviates only in periods t plus one he does something else uh something different than what he's supposed to do but for the rest of the game all the players are going to play according to their strategies so given that there is no profitable such deviation uh and if this is true for each player and for every history well then you know what that strategy profile forms and sp e all right so that's basically the rough idea well normally obviously there might be infinitely many possible strategies but the trigger strategies are normally and usually sp and e for high enough delta right for high enough delta so patience is important for our sort of sp needs if the players are not a patient well unfortunately not a lot of strategies will be sp because they you know again look at the the most extreme case where delta is equal to zero so the players are completely impatient the only thing they care is today so therefore the only sp and e in this game is going to be the stage game nash equilibrium right okay so but the thing is if delta is high enough uh we can support a lot of strategies and a lot of payoffs as an sp and and and and the trigger strategies are widely used strategies so what do they look like well usually the trigger strategies consist of several action profiles usually two is enough but sometimes more than two is might be necessary one profile is usually called cooperative profile all right so they cooperate to achieve that profile usually that profile is not nash equilibrium all right so we call it cooperative profile probably because although it's not nash it gives both players a higher payoff while then there's that the other the second group of profiles is called punishment profile well what is the role of those punishment profile is basically if they do not play cooperative profile strategies i'm sorry well they're going to switch to punishment profile and they're going to keep playing this forever and usually the punishment profile strategies are nash equilibrium off the stage game and the the trigger strategies usually work in the following way uh the game starts with players playing cooperative profile actions all right and every period they're going to play the cooperative profile unless they see that one or maybe more than one player deviates so if anybody deviates all right well then all the players move to punishment profile meaning they play the punishment profile strategies for the rest of the game forever all right um so that's that's basically the idea of uh trigger strategies or you know the the strategy that i just described you is usually called grim trigger strategy so those strategies are most like uh you know a carrot and stick type of behaviors that we observe in real life it's like uh so let's let me give a more concrete example so the prisoner's dilemma once again because it has a unique nash equilibrium but the nash equilibrium is inefficient so corporate deviate corporate deviate so corporation is going to bring each player two and two payoff but the thing is it's not a nash equilibrium because given that one party cooperates the other party has a huge incentive to deviate so zero four four zero and they actually end up playing dd and by the way in this game d and d is a strictly dominant strategy all right so the c the action the strategy you see is is strictly dominated so the only nash equilibrium of this stage game is d and d where both player gets payoff of uh one but the question is can these guys achieve a payoff two and two at each period meaning can they play cc forever can this be a sub game perfect nash equilibrium outcome well the answer is yes it can be sub game perfect nash equilibrium outcome but how well here is the grim trigger strategy players start playing cc well by which i mean in period at the very beginning of the game player one is going to play c player two is going to play c all right for the rest of the game for the rest of the game they will play oops cc unless if well let's let's put it this way if nobody plays uh d all right d before so in period zero if they both play c and c in period one they will play c and c well because they played c and c in both periods zero and one in period two they will play c and c alright so in period hundred they will play c and c if nobody before all right if nobody before played d alright so basically this is the cooperative profile the cooperation stage however right somebody may have deviated if one or two players ever play d so maybe player one plays d at some round or maybe player two or maybe both play d right well if that happens well then uh play the players are going to play d and d forever so here's one thing i would like to say is that we cannot really be too formal in terms of describing strategies in an infinite horizon repeated game because the strategy remember the strat the definition of strategy was at every um decision note we should tell what each player is going to do well here in an infinite horizon game there are infinitely many possible decision note for each player so a strategy profile is going to be a a huge very complicated animal so trying to be too formal about it doesn't really make much sense so what do we do we verbally define it um well but nevertheless one thing that is very very critical you have to cover all possible histories of the game and so tell us what the players are going to do all right so that that description should tell me what players are going to do after any possible history so here the thing is you know they may play c c c c c c all right so c c so in that case we know what the game what the next period will be it's going to be c c but the thing is it might be for example c c in period one and then c d in period two right this is one possible history so the the thing is oops sorry cc cd what's going to happen here well i know that according to those strategies it's going to be dd forever right good what if it was c c c c c c c c and then all of a sudden d t well i know that for the rest of the game it's going to be d d right so therefore the the the next stage move is going to be c and c for both players if you observe nothing but c however in any period if you observe d by any player well then you know what the rest of the game they will the rest of the game forever they're going to play d and d alright so that's that's the green so this play dd is a punishment profile and as you see by the way dd is in nash equilibrium of this game so therefore playing dd forever all right so regardless of the history playing dd forever is basically a nash equilibrium in that sub game remember our earlier result repeating the nash equilibrium of a game is an sp e or therefore it's a nash equilibrium in any subgame so this is dd is is a nash equilibrium very good but what we have to check is is playing cc can we support it as a as a nash equilibrium of this entire game so therefore this strategy profile is sp any well we can if the discount factor is large enough well what is the thing what is i mean how can we uh guarantee this how do we check this well simple you first check no deviation payoff oh well you have to do it for player one and obviously for player two but here i'm going to do it only for player one because the payoffs are symmetric strategies are symmetric and therefore once i get some inequality about delta it will be exactly the same inequality for player two because everything is symmetric all right so but you have to normally you have to do it for all the players so for player one i'm going to check calculate his no deviation payoff and then i'm going to calculate his deviation payoff so the nice thing is that here uh there's only one deviation right i mean normally your strategy tells me to play c and so if you want to deviate well again one deviation properties like we allow players to deviate only in one stage while there's only one thing that you can do it's like deviating to d so therefore uh the deviation payoff is going to be well what if uh player one deviates to d all right so and then compare these two well no deviation payoff should be greater than or equal to the deviation payoff so therefore uh the no deviation of the strategies that described here actually best responds the other guy's uh strategy so what is the no deviation payoff here well in which period are we talking about well that's the nice thing about repeated games infinite horizon repeated games and those strategies it doesn't matter because here remember um the the nice thing is that as long as players played cc before uh well whether this is period 10 or 100 or 1000 doesn't matter because the rest of the game is exactly the same right i mean they will keep playing this game again and again forever and so therefore you can think of this we do this analysis for any period right for any period where the we we observed no deviation before well you may ask what if the same analysis for history is where deviation occurred well remember in those sub games according to those green trigger strategies players are going to play dd forever and we know that this is a nash equilibrium of these sub games because repeating the stage game nash is always a nash of the sub game or you see what i mean all right so therefore if uh player one uh does not deviate well his entire game payoff is going to be what well he's gonna get remember two in period zero well or in period that we started but let's suppose without loss of generality this is period zero and then he gets another two a period later and so it's two delta and then another two in in two periods later so two delta square and so on well the nice thing about it we can actually simplify this as two parentheses one plus delta plus delta square etc all the way to sort of forever well the nice thing about this this is a geometric sum because delta is less than one and greater than zero this is actually equal to two divided by one minus delta well what about the deviation path all right so again one deviation so player one deviates to d well what about player two player two remember we fix his strategy he's playing c in this period so therefore by this deviation player one is gonna get four payoff good so there's no delta because this is what exactly he's gonna get in this period next period well so this is the trick of the one deviation property we allow just one deviation one period deviation and then assume that players are going to play according to their strategies for the rest of the game so is it going to be 2 delta plus 2 delta square etc no why well look at those strategy profiles it says look you these guys are gonna play cc if nobody before played d but here player one actually plays the and so he changes the history of this game and therefore the strategies are going to react to that how well they're going to play d for the rest of the game all right so by deviating you trigger punishment this is why we call it trigger strategies so by deviating player is going to player 1 is triggering the punishment so that means they're going to play dd forever which means 1 times delta plus 1 times delta squared plus 1 times delta cube and so on it goes like forever so one thing uh you may wonder well i'm to player one i deviate today and i know my opponent is going to start punishing me tomorrow but why don't i deviate next round as well well i mean you can try remember if your opponent if you know that your opponent is going to start punishing you and therefore playing d the best action for you is to play d because d d is the nash equilibrium so therefore if you try to deviate in the second period as well the best you're gonna get is zero which is worse than one so there's no point in deviating in next period and next period this is actually why the one deviation property works also by the given the fact that playing dd is in nash equilibrium and this is exactly why we use the punishment profile as the nash equilibrium strategy profile all right so so we know that deviating more than one period is not going to be profitable for you so you can deviate only just one period all right so this is the deviation payoff what is it equal to well that's simple this is four plus taking into delta parenthesis it's going to be one plus delta plus delta square plus forever and so this is equal to 1 divided by 1 minus delta so this path is equal to 4 plus delta divided by 1 minus delta the question is when this guy is greater than or equal to this guy so the no deviation is greater than or equal to the deviation pair so if deviation payoff is greater that means these guys actually are going to deviate even at the very beginning of the game so therefore these strategies cannot be sub game perfect nash so when two plus one minus delta greater than or equal to four plus um sorry this is let me erase this part this is 4 plus delta divided by 1 minus delta right well multiply both sides by 1 minus delta i know it's positive so it's it's not going to change the inequality so i'm going to get 2 greater than or equal to 4 minus 4 delta plus delta so i have here i have 2 on the right hand side i have minus 3 delta but i send it to the other side as 3 delta so you know what as delta greater than or equal to two over three these uh the no deviation payoff is going to be greater than or equal to deviation payoff hence this strategy is sp and e if delta is greater than or equal to two divided by three that's it i don't really need to check anything else once again we know this is a very complicated extensive form games but they did the trick that the punishment profile is in ash equilibrium and the punishment is playing dd forever and so any sub game where any sub game that follows some player playing d is actually i mean those strategies are going to form a nash equilibrium and then so the only the history that we should check is that a history where no deviation ever occurred and then in this history we already know that deviation is not going to benefit this agent because he's patient enough so therefore as long as delta is greater than or equal to 2 3 those strategies this grim trigger strategy is the sp and e of this game all right question is are there any other sp e in this game yes there are many well here according to this we actually support on the equilibrium path what payoff we support two divided one other two divided by one minus delta if delta is very close to 1 by the way this is a very large number right as delta goes to 1 this thing goes to infinity but remember we assume delta is less than 1. so this may be very huge but the thing is what we normally look at is the average payoff average payoff what does that mean that means we basically uh multiply the payoff by one minus delta all right and so this is the average so on i mean on average what is the payoff they're gonna get here it's two right there each player is gonna get two forever the question is maybe they're going to play cc sometimes uh and and and cd sometimes can this be a nash equilibrium a subgame perfect now not according to these strategies according to some other strategy yes it is possible and so question is in this game where in the nash equilibrium the the unique nash equilibrium is strictly dominant strategy and it's inefficient but what we saw is that an efficient outcome can be supported as a subgame perfect nash equilibrium outcome of this game if the players are patient enough and if this game is repeated infinite horizon the question is is that it i mean can we support any other average payoff average here again i mean maybe sometimes uh they're gonna play cc but sometimes maybe they're gonna get dc or or cd so you see what i mean they're gonna get something more than two is it possible so that's what we are going to answer next
Up Next

Stackelberg Oligopoly: Quantity Leadership Model Explained
@AshleyHodgson
34.7K views•2015-04-11

Mundell-Fleming Model: Negative Goods Market Shock Explained
@Inlecture
831 views•2020-05-07

Mixed Strategy Nash Equilibrium Explained (Game Theory 4) | Matching Pennies
@selcukozyurt
69K views•2020-10-21

The Age of Easy Money: Fed & Inflation | Full Documentary
@frontline
21.2M views•2023-03-15
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Economics




















![(AGT3E12) [Game Theory] Folk Theorem](https://i.ytimg.com/vi/dUEsI0qwMQs/maxresdefault.jpg)










![[공기업경제학 기본강의 19] 제5장 시장이론 (6)](https://i.ytimg.com/vi/JyvevU7w4Pw/maxresdefault.jpg)





 | Tacit Collusion | Infinitely Repeated Game | 49 |](https://i.ytimg.com/vi/ErT4KANRvaE/maxresdefault.jpg)

