The Prefrontal Cortex Basal Ganglia Working Memory (PBWM) model explains how the prefrontal cortex maintains task-relevant information through intrinsic excitatory loops, while the basal ganglia acts as a gatekeeper that toggles between maintaining existing states and updating them with new information based on dopamine-mediated reinforcement learning; this mechanism resolves the 'homunculus problem' by showing how neural circuits learn appropriate task representations through trial-and-error without requiring an internal decision-maker.
Executive Function: PBWM Model & SIR Task | Cognitive Neuroscience
Added:so in that strip model we just simply kind of clamped the prefrontal cortex neurons on and even though we kind of adjusted the strength of those prefrontal cortex connections we didn't really answer the question this homunculus question how does the prefrontal cortex know what task neurons to activate it in the first place one of those come from how that how do they how do they get activated that's really the key question for making the prefrontal cortex smart and our answer is that there's some kind of interaction with the basal ganglia that enables selection of appropriate task states and so this is this idea that we can go from basal ganglia action gating where you pick the most kind of you know appropriate rewarding response instead we can think about that in a cognitive context as selecting the most appropriate rewarding goal yeah so these five different loops through the prefrontal cortex and basal ganglia and you can have a gating signal kind of process in prefrontal cortex in addition to each of these other levels and so the basal ganglia at every level is really playing the same function of helping choose the right thing to do that's going to be most rewarding and our particular idea about how this works is in terms of toggling or essentially latching the active maintenance and so the idea is that if the basal ganglia is not firing then the prefrontal cortex will have intrinsic kind of loops of excitatory connections that will cause individual neurons to fire and activate other neurons and then they reciprocally kind of Co activate each other and you get this kind of positive feedback loop that sustains the neural activity over time and this is really a classic idea that's been supported in a number of models and empirical studies and so the basal ganglia essentially when it's not firing when it's in its default mode if you recall the output of the basal ganglia is this snr also in known as the GPI and these neurons are firing away and inhibiting the thalamus and sort of preventing any disruption of this ongoing maintenance of activity in the frontal cortex and then if the basal ganglia does get a kind of input that causes the go pathway neurons to out-compete the no-go pathway and inhibit that tonic activity okay then that dis inhibits the firing of these stuff lamech neurons and you get essentially this big burst or wave of activity between the frontal cortex and the thalamus and we think that's really critical for essentially turning on updating and latching a new pattern of activity in the frontal cortex it's on this diagram you can see here with yellow being the active neurons that we've switched from this you know one pattern here so now we're maintaining this new pattern reflecting the kind of new inputs coming in from the posterior cortex all very kind of cartoon representations of these things and there are computational models that we'll go through that kind of show how this can actually result in stable active maintenance and selection of appropriate tasks representations and so our version of this we call the prefrontal cortex basal ganglia working memory model PWM and it's essentially that same diagram I showed you earlier where we have the prefrontal cortex holding on to this kind of task context holds task states etc and the basal ganglia playing this role of essentially updating toggling on and off those actively maintained States according to a history of learned reinforcements from the dopamine pathways based on what has worked in the past so let's look at this system in the context of a very simple working memory task so our simple task we call the stored ignore recall task and it really boils working memory down to its bare essence you store something you hold on to it while you ignore other things that might be distracting and that's actually a critical aspect of working memory as compared to simple short-term memory and then you have to retrieve that information that you previously stored and so if we look at how the model works here I just click that step trial we see in the input that we have some kind of signal in the environment that's telling us it's time to store this information and we're not you know they were really simplifying everything so this is like who cares what this is it's just some cue could be context could be all kinds of different things that tells us this is something that's important we should hold on to it we get also one of four different arbitrary stimuli ABCD and that's the thing we're supposed to hold on to so this is what we get is an S and a D and now we go to our next trial and in this case it was an immediate recall trial so we just have to retrieve that D that we got and you'll notice that we don't get any new input here we're just trying to retrieve here at the output layer the item that we had previously maintained randomly generated so we just get what we get and so now finally we got a store trial here where we got store a and then the next trial is an ignore trial and here critically the system gets this kind of distracting input this D that's coming in and somewhere in the prefrontal cortex over here if the models behave incorrectly it should have held on to the a and B holding on to that in this kind of active memory form this delay period maintained activity and then when we get the retrieval cue here we get another ignore trial so they're random so we can get fairly long periods here's another one this is a good test finally we get the retrieval cue the recall trial and it's supposed to output a and I'll note that we're seeing the a here because we were looking at the plus phase if we look at the mine Faye's we can see that the model did not get this correct as we would expect to model really is starting out with kind of random initial weights and has no idea how to hold on to this information what information should be held on to and it learns that purely through this kind of dopamine based reinforcement learning driving the go and the no-go pathways in the basal ganglia of the model we use the word matrix to refer to the MS NS the medium spiny neurons in the particular matrix matrix sub compartment of the striatum also because we love the word matrix and the movie matrix so that's why we call it matrix so this is the straight I'm gonna go the go in the no-go pathways going into the GPE which is the GP no-go pathway and then we kind of combine the effects of the GPI which is the output and the thalamus so this kind of disinhibition turns into a net excitation it's easier to see it's easier to deal with in this model and so now you can see kind of which pathways are activated in firing essentially I go and then that drives updating of the working memory in prefrontal cortex and we can train this model and then look at the trained model to see exactly how that updating works as it kind of goes through the whole system and now we can see here the model training did it is it going through all these different trials of experience doing this task over and over again over here we can see the dopamine pathway we have a kind of raw reward based on whether we got the right output answer and then there's a prediction of whether we're gonna get reward and this is the kind of key case here we're looking at the risk or low Waggoner version of the dopamine pathway turns out we don't need the full power of the TV or the Pavlov biologically-based model in this case a simple Rescorla Wagner model works and so that the SNC here this final dopamine pathway is actually reflecting the difference between when we get reward relative to the prediction of and I can just run that again and you can see these bursts and dips of dopamine happening as the model is learning the task sometimes it starts to get an expectation that drives a dip and so when it gets a purse to build mean that reinforces the go pathway and when it gets a dipper dopamine that reinforces the no-go pathway you can see then the balance between the go and the no-go is reflected in whether the model is gating here and in the end pathway of the GPI thalamus and then that gating signals those gating signals are driving the extent to which the prefrontal cortex is updating its memory representations and actually as we see the model train up here one of the key things that you can start to see is when some of the maintenance stripes back here in the PFC are starting to hold on to information over multiple trials you can see them kind of slow down and stop here this one actually seems like it's getting a little bit stuck it's holding on to the ignore information which is not so adaptive and it's not doing such a great job of holding on to the store information as the model starts to latch on and try different strategies of gating it gets this pattern of reinforcement and Punishment eventually it will seize upon this goal of holding on to the kind of store trials and then that's when it finally starts to succeed on the task and here we're tracking over time the learning performance of the model you can see the kind of percent correct a percent error in this graph we can look at the absolute level of dopamine in this Green Line that tells us how much kind of dopamine overall is getting generated and we can see its predictions of overall level reward how well it thinks it's doing and so it gives you a nice tracking of the performance of the model over time and now we can step trial by trial here let's get a store trial so we got a store and you can see here where we're activating these patterns and these stripes these would call these different kind sub-groups prefrontal cortex neurons stripes they're independently datable we have so we have four different stripes of prefrontal cortex that we can independently gate part of the learning process exploring in parallel each one kind of trying out different strategies and that greatly speeds up the learning because it can search through kind of different strategies in parallel instead of doing it sequentially over time and so here we can see that you know it has aided in information into these two stripes and how we can tell that is that these neurons in the maintenance layer which we associate with the deep layers of the cortex is it's those guys are activated and that's a result of these two neurons here or firing and so these kind of are in correspondence it's these two rightmost stripes of the maintenance layer it also did this kind of output gating over here which we'll get to in a second that's what these other two stripes are showing so overall it's now encoded this kind of store D signal into two separate kind of parts of this maintenance area in prefrontal cortex interestingly you can see why it was actually kind of hedging its bets here this one of those guys has now gated again and that's resulting in the updating and holding on to this ignore trial information in this first right but critically you can see back here that one of those stripes that the back one is still holding on to the original store signal this has the key information that it was the D that we were supposed to hold onto and now as we go again we get the retrieval cue that is able to read out from output gaining and so now this is the other key part that the the model has to learn to do is in this other part of the GPI thalamus which is mapped over here to the output part of from cortex we have this additional idea that when it's time to use information that you're holding on to and working memory you have a separate gating decision to output keep that information and use that to control behavior and so that is actually what's responsible for transferring this kind of activity from the maintenance layers over here back over to the output deep layers and that all happened in time to generate as we can see over here in the minus phase the correct answer of the D that was originally stored so this is how you go from sort of preparation holding on maintaining preparing for something and then deciding now it's time to actually act and take that information that I've been preparing and holding on to and use it to control money behavior this same framework has been used to simulate many many different types of working memory tasks as we'll go through in a little bit so it's very general and you can see quite quite interestingly just through this kind of trial and error learning process with these dopamine rewards it eventually does learn the kind of right strategy and and so that essentially helps us get rid of this homunculus this infinite regress and we can see that there really are kind of concrete mechanisms involving what we know is actually going on in these different brain areas and the prefrontal cortex and basal ganglia that enable the system over time through learning to kind of be smart
Up Next

Dopamine's Role in Working Memory | Huberman Lab Clips
@HubermanLabClips
38.7K views•2024-04-26

Bessel van der Kolk on How Trauma Affects the Body and Brain
@bigthink
226.3K views•2025-10-03

Vagus Nerve (CN X): Anatomy, Nuclei & Functions Explained
@Alilamedicalmedia
305.2K views•2022-10-31

How Exercise Benefits Your Brain: Science Explained
@TED
11.4M views•2018-03-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Neuroscience


































![D. Vernon - Cognitive Architectures, pt. 2/3 - iCog Talk [13/01/2021]](https://i.ytimg.com/vi/IinI6Zj6HnM/maxresdefault.jpg)



