Neural networks can achieve symbolic processing through an architectural principle called the 'relational bottleneck,' which uses external memory as a binding mechanism that forces the network to learn abstract relational structures rather than memorizing specific data. This architecture, demonstrated through the Emergent Symbols Through Binding (ESB) model, enables neural networks to perform symbolic tasks like Raven's Progressive Matrices with radical generalization from minimal exemplars (as few as 5% of symbols), achieving sample efficiency orders of magnitude better than standard transformers. The key insight is that by creating a bottleneck that only allows relational information to pass through, the network is forced to discover abstract symbols and functions, enabling it to generalize to novel data it has never seen before.
Neural Networks & Symbol Processing in Cognitive Architectures
Added:So, it's um my great pleasure to introduce Jonathan Cohen. I of all the things I was supposed to do to prepare for this meeting, I should have created u an introduction. So, um I'm going to let him introduce himself. Uh he has a long history of research in uh cognitive architectures from different perspectives, often from the connectionist viewpoint. Um, I heard him give a talk and I said, 'Wow, it would be really great to have that perspective for the people in our community to hear.
And so, uh, take it away, Jonathan.
Um, thanks, Don. Uh, yeah, I'm actually thank you for letting me introduce myself. It both spares me the embarrassment and lets me say the things that really matter. Um, the first of which is that, um, I am incredibly honored and touched to be here. Um, as John said, I've spent most of my career building what, you know, I'm not sure I would fancy as cognitive architectures in the sense that you guys mean. I've always admired that effort and aspired to it. Um, but only recently felt like I have anything close to something to say about that. So, um, that's aspirational for me mostly. Um, I'll also say a little bit about history. Um, I actually originally started my career as in with medical training. I won't tell you the whole story about that, but what I will say and then went on to get training as a psychiatrist. And basically what I learned in psychiatry is that we're like, you know, we're like uh car mechanics that don't know about combustion and so beat a path back to basic science. Um and uh maybe the best summary of my experience in my medical training was that I often say the best book I read in residency was um the C programming language by Kernahan Richie.
Um I got bit by the bug. I started reading everything I could about computing. It seemed like a really good way to think about the brain. Um, some of you may know Terry Winrad or know of Terry Winterrad. I had the fortune of being at Stanford at this time and somehow found my way to his doorstep and he was very very um welcoming um uh you know I said I want to understand how the brain works and I want to understand from a computational perspective. He says, "Yeah, boy. We all do and someday, but we're going to start simple." And he basically set me off in other directions. Um, given that CS wasn't just at quite at the point at that time, I was incredibly fortunate though, not long after that, to stumble into a talk at Berkeley that somebody had told me to go. People were basically trying to get me out of their hair, and so they were sending me anywhere they could to get get rid of me. Um and somebody sent me up to Berkeley to hear about some guy that um uh actually it was I think it was um Gordon Bower who told me about this and he said there was a guy giving a talk up at Berkeley. Um his name is Dave Rumbleart and uh you might want to go and listen to what he's spewing because it sounds like the kind of thing you're interested in though you know we're not really sure where it's headed.
And so I went and heard Dave Rumlhart give a talk about backrop and I the scales sort of fell from my eyes about at least how I could come at it from a neural perspective but ask questions about cognition.
Um, I beat a path from there to his colleagueu's doorstep, Jay Mcclulland, and spent my time retraining as a basic scientist under Jay's toutelage. Um, and arrived in 1986 at a time when the Holy Wars were in full fury. Um, and I hope somebody has transcribed the bulletin boards of the day because some of those discussions really belong in the historical record. I mean, this was when Herb Simon and Alan New were there. Jay had just been hired. Jeff Hinton was still there and you can just imagine the flames that were I I I suspect some of you might have even been party to those discussions. Um and you know during the day um I was you know a beautiful um soldier in my camp but I had one it turns out there were only two of us in in the grad school. This is at the psychology department at the CMU and it turns out that there were only two of us in the psychology department graduate.
Well we started with three but one bailed. I don't even remember why. And so we landed with two people in our program, me and a guy named Ken Kadinger, who many of you may know. And so Ken and I were both sort of and he was John Anderson's student. And so he and I sort of were beautiful soldiers in, you know, in in our respective legions um during the day, but at night we'd sort of crawl out of our tents and bring our candles to some secret spot and talk about how we could find some kind of middle ground, how we could find some sort of meaning in what it what it would look like to have an architecture that could do the kinds of things that people do and that SOAR and ACTR were were trying to um formalize, but in a way that actually happened in the brain And um as I say that was an aspiration and it remained an aspiration for many decades. Um always sort of close to my heart especially since I sort of picked a part of the modeling space and connectionism that was the least occupied and the most problematic which is how do you get neural networks to do something interesting like planning or problem solving. And I ended up having to reduce that to the simplest possible question which is how can you implement some form of control in a network? how do how do you get some sense of, you know, intention or goals? And so I spent a lot of time sort of working that through. And I think, you know, we made some progress. Um, not that it's ever been sort of viewed that way, nor would I expect it to, but but in some sense, we sort of discovered attention and how it could be built in neural networks that I would like to think anticipated, you know, the paper that says attention is all you need. Um, in this paper, I will argue that attention is not all you need. It's something you need, but not all you need.
Um anyway, so that that's sort of the background for me. Um and I feel like, you know, we're in a remarkable point now. And I don't have to say this to any of you. You've just heard talks already, you know, sore perspective, how to integrate re reinforcement learning or LLMs and and and I think that's just absolutely the way we need to go. The question is how do we do it? And there's lots of ways. and I'm going to articulate a few um and then offer what I'd like to think is an insight about a point of contact that I think might make a difference. So that's sort of the background and a very high level statement. Um let me just start let's see if I can get going here. Yeah. Um by just a little summary. I know I'm preaching to the choir here but I think it's worth stating nonetheless. Um you know we all well most of us respect the miracle of symbolic computing. I mean, how could it not be um an amazing thing?
It's computationally general. It is, you know, got maximum flexibility. I mean, that's just restating the same point.
Um, you know, all these words that everybody's in search of now, interpretability and reliability, um, sort of come for free. Um, and you know, I mean, if anybody thinks the brain doesn't do symbolic computing, I'd like them to explain to me how it is that they the brain came up with symbolic computing, right? Right? I mean it's almost almost a tautology that the brain is capable of symbolic computing. Um and in some sense is one expression of the miracle of symbolic computing. On the other hand, I think we all agree that it has its challenges. Um it's inefficient at least to program or to configure especially in hyper complex domains exactly domains where it struggled such as language and visual processing for a long time. Um and you know I also don't think I need to make this argument at great length but again it's worth noting um there is a equally miraculous development namely deep learning or connectionism or back prop really to be specific um uh that complements symbolic computing. It's computationally efficient. Um some refer to it as automated function approximation which is just a really fancy way for saying you program it with without really knowing how to program it but you give it an algorithm and a lot of data and it programs itself. you don't know exactly what the program is, but at least it accomplishes something close to what you want. Um, and yet it is sample inefficient, domain specific. Um, and then as I intimated a moment ago, you know, it it challenges our ability to interpret um what it's doing and as a consequence, we have concerns about its reliability, at least in in untested domains. So, this is sort of, you know, the the standard story. Um, you know, so where are we? Well, you know, it's been like this for 50 years, and I'm not sure we're much past that, sort of. Um, what we'd really like is this, right? We'd like some sort of marriage of these two, some understanding, much as physicists understand that, you know, molecules are stochastic and gases are, you know, you can't say where every molecule is, but we can say something about an ideal gas law, and we understand what the relationship between those is. The same thing with Newtonian mechanics and statistical mechanics, etc. We like what I think philosophers refer to as reduction principles. How do we go from the symbolic level to the lower finer scale level of description um and and that you know that's one way of putting it but I also think we would like to understand at a mechanistic level you know how symbols arise like what are the right approximations and what are the architectural features that give rise to it in something like a neural network.
Um and so there have been a few approaches right one um sorry so the challenges are to to integrate flexibility of symbolic processing in traditional architectures with in effect the efficiency of function approximation and and the configurability of neural networks for achieving that. Um and so there's sort of two kinds of efforts as I see it. Um there are neuros symbolic approaches and and this by the way I think excludes the sore community at least from what I've heard about it from John's talks and then the few that I heard um this morning. I think that's taking a slightly different approach which isn't fully articulated here but I think it falls within the sort of general um sort of scope of what I consider to be these two approaches which is one you come at it from the symbolic perspective. You say well there's this deep learning stuff right and obviously it's really good at doing certain things. So can we start with the symbolic primitives and then get it to do the programming for us? That's basically what it comes down to. Um and you know in in in one world this is referred to as program induction. This is sort of the tenonbound camp. You know we're going to assume that there are these symbolic primitives and we're just we're going to sort of get the neural network to do the program induction for us but it's still going to be symbolic processing because we built that in. And I think there's a certain spirit of that in the kinds of architectures that were talked about this morning. You have the sore architecture where you specify certain kinds of operations that we can call are symbolic and then you're going to get deep learning to sort of augment that by doing some of the programming for you. Um, and this stands in contrast to and oh sorry this is often referred to as the neurosyolic approach. Um, I I'm not sure exactly whether that's the best term, but that's the one I keep hearing referenced. And so in response to that, I'm going to coin a new term called the neoconnectionist approach. So the traditional connectionist approach is that you start with undifferentiated stuff, all these little processing units that have nothing but weak weights connected between them and you pick a statistical algorithm that trains them and you give it lots of data and it trains itself, right? But there's never any it's it's often referred to as endto-end training. There's never any prespecification of any kind of structure whether it's representational or architectural other than maybe the feed forward stack of of layers in in the network. And I don't know how many of you have ever heard the the light bulb joke about connectionists. Anybody heard this? How many connectionists does it take to screw in a light bulb?
If you laugh, you'll understand connectionism. None. If you wire it right, you don't need the bulb.
At least some people understand connectionism or at least the historical connectionist approach. But I will say this is one of the things I never bought. And I think that was always sort of in, you know, it made my sort of stomach a little queasy was just to think that it was just going to be a bunch of undifferiated units and a learning algorithm. There has to be something that evolution learned that makes us different, for example, than apes, right? That have a lot of neurons and are able to learn. And as far as we know, there's really nothing qualitatively different about their brains from ours. There's not a cell type or a particular circuit that anybody's found that makes the ape different from us that we know of. And yet we do something qualitatively different. So there's got to be something. Is it just more neurons? That may be part of it. It seems unlikely.
And so when I refer to neoconnectionist, I refer to the following. The selection as parsimmonious as we can be about that of inductive biases that favor abstraction effectively that favor the discovery of symbols. And there are two broad forms of inductive bias we can imagine. Um one is architectural and that's what I'm going to spend most of my time talking about. But I want to mention the other one as well which is training. And you know we don't learn to do math just by being in the world. Our brains are capable of doing symbolic processing but they certainly don't exhibit all the features of symbolic processing if they aren't subject to some kind of curriculum. If you don't go to school you don't learn math. Maybe you pick up a book, but if you don't if you don't have some kind of structured introduction, you don't dis discover math simply by your experience of the statistics of the natural environment.
You have to have some sort of shaping of the procedure. And I'm not going to say much about that in my talk. If there's time, I'd be happy to say more about that. We've begun to explore this and I think we've made a few more discoveries about the kinds of things at least qualitatively that that help um take advantage of and augment the kinds of architectural biases that I am going to talk about and the second one is exactly that architectural structures to the system that though they themselves don't ensure that there is symbolic processing make it possible under the right curriculum.
So where does this leave us with regard to the state of you know of these ideas in the context of LLMs? Well, in some sense they're sort of you know we we could debate this right but the first approximation they're doing a pretty good job of faking symbolic processing if they're not doing it per se. At the end of the talk I'm going to argue that they actually are doing it in the the the most you know formal rigorous sense of what one might mean by symbolic processing. But for the moment, let's not assume that and just say there's growing evidence of abstract reasoning in large neural networks. This is a paper by um one of my colleagues, Taylor Webb, who together with Keith Holio, went through and tested the then sort of current model GPT3. This is several years ago on a bunch of analogical reasoning tasks and and various other measures of abstraction like Ravensburg as a matrices. And I'm not going to go through all the data here, but suffice it to say that GPT3 was pretty much at human level performance. And as far as we can tell, GPT4 in many ways exceeds it. It can do anal analogies that at least undergraduates can't solve.
So there's at least primaasy evidence that they're doing something symbolic like um but and again I think you all know this, they're wildly data inefficient and also energy inefficient. I'll say a word about that, not too much, but I'll mention in just a moment. Um, and they're largely domainspecific, right? You train up, you know, chatbt to answer any question about history, but it can't drive a car, no less fry an egg or or raise a kid. So, they're impressive, but they're not truly what we're looking for. And there is an existence proof of what we're looking for. It's ourselves, right? And so, this is just a little madeup slide to make the point. Again, I suspect I'm preaching to the choir, but I think it's sort of fun to reflect on. This is just a a a sort of fictitious graph of different tasks. Those are the different sets of bars. And then different agents in the different colored bars um trying to perform these tasks. So, one is building a web. Well, spiders do that pretty damn well. Um building a dam, you know, beavers can do that, build a nest, robins, etc., etc. out to chess and go and navigate. And you know in that case we have um artificial agents alpha go alpha chess castle autopilot um that do pretty well at or in some cases um radically exceeding human performance.
So if we look around us we're populated in a world now by agents that pretty much beat us on anything that we might want to take on. Right? And yet if you connect the dots and you ask well how well does any agent do all of these things it's not even close right? the area under that pink line which is connecting the human dots is vastly orders of magnitude um more area than any other connection of any other lines.
And that just makes the point about our flexibility. So even if I had to build a web, I wouldn't do it as well as a spider, but I could catch a fly, right?
I could get some string and sp, you know, spray some tacky glue on it and I might get lucky if I put it in my pantry. In fact, I'm sure I get lucky.
Not flies, but moths. Um so, you know, what does this say? Okay. Well, it says that there's some kind of sweet spot between flexibility and efficiency. Um, we have a near limitless range of tasks um at adequate performance. That's our flexibility. And we do it with reasonable amounts, often little or no training. Um, I think one of my colleagues, Tom Griffith, has estimated that GPT3 took something like 5,000 years worth of reading or what the equivalent of a person would have to do if they wanted to get exposed to the same amount of information. And we do it in what maybe 10 years. So like three orders of magnitude less data. Um and I'm not going to emphasize this point but it's hard not to mention it. Um we do it on 20 watts. And I did a quick estimate of what GPT3 required in energy both in terms of training and then operating costs advertising it over whatever the three or four years that it was useful until it got replaced. So the training was I think um something like uh 10 megawatts and operating when it was when it was being used was like about a tenth of a megawatt. So if you sort of average it out let's call it a megawatt a megawatt to do a fraction of what we can do on 20 watts. So there is a lot of room for improvement. Some of that may be biological like literally physical like you know chemistry for better living but there's got to be a lot of it that's algorithmic. It can't just be the efficiency of chemical computing.
So, how does it accomplish this? All right, so that gets me to my talk.
Here's the outline. I'm going to propose an architectural principle. I'm going to show you that it works in at least a simple toy setting. Um, I'm going to show you how it works in that toy setting. And then I'm going to discuss in much more sort of handwavy broader form, although there are papers that substantiate much of what I'll refer to.
um uh its generalizability across different architectures, its relationship to symbolic in particular production system architectures. I'll be eager to hear what you all think of that whether I'm doing it justice or not. Um for those who are interested, I'll say a word about biological plausibility.
That's something close to my heart, maybe not everybody's. Um and then um if there's time though I doubt there will be um some variations on the theme that start to get um to applications at least to human cognitive science where we can use these architectures and start to provide pretty um sort of detailed accounts of the the space of human cognition at least in the domain of semantics where with these architectures we can start to account for sort of all of the historical semantic idiosyncrasy like the triangle inequality and the reversal of similarity effects. And I say that as a tease because I doubt I'll get to it, but but we have a couple papers on this if anybody's interested and if there's time afterwards, I'll be happy to say more about it. All right, so let's start with um the architectural principle. Um it was inspired by a paper um about 10 years ago um by uh some people at Deep Mind um called the neural touring machine about an architecture that they proposed called the neural touring machine. Anybody here know about this architecture? people familiar with this?
I I'm surprised on many counts um because I think this was one of the deepest insights that Deep Mind had. In fact, I think it was probably the only insight that I'm aware of that they put on the table that was a real insight that was a real innovation. They did amazing work, don't get me wrong, of course, you know, who am I to say, but it's for sure true um in in building out neural networks, but most of the ideas were borrowed from others. This one was in many ways original and it sort of scratched an itch I'd sort of had for a long time and caught our attention. The innovation was as follows. Let's take the standard LSTM which is what they used to solve the Atari games and for the most part was used to solve um Go and chess and most of the other things they've done.
Variance thereof. LSTMs you usually think of as using backrop but you can do it with the reinforcement top end or loss function. Um there's a talk this morning I think Bob Ray talked about stuff sort of like that. DQNS are a version of that. Um and uh add to it an external memory.
basically what we would as psychologists call an episodic memory, a system wherein that LSTM in addition to encoding information into working memory and then sort of using that integrating that information and then spitting something out and deciding when to encode stuff into in into um working memory using its gating mechanisms.
Also decide when it would write essentially to a touring tape. It would take its memory state or some part of its memory state and simply put it in a cache RAM stored vertidically but retrieved in a content addressable way. So when it decided to retrieve something it would do it through a standard dotproduct match. And I'll sort of hit one of the punchlines here by saying essentially the matching procedure that production systems use when they match conditions in the production rule with the state of working memory. This is not working memory by the way. This is external memory which is interesting and I'll come back to that later as well. But in the sense of being able to sort of match to states and then retrieve new states that might affect what's going to happen next that's going to get brought back into the LSDM. here's the working memory of this system if you have to label it.
Um and then decide what to do. Um would give it the chance to sort of basically implement a touring um general computational device. This is and that's why they called it the touring machine.
And they went on to show that if you train this thing and and all of this is trainable and I'll say a word later about how one might train this part of it, but all of this is trainable. And so it starts with sort of no representations here and empty memory and it just learns representations here.
It not only learns what kinds of mappings to do from input to output but also what mapping what what information it needs to be able to store in a way that it can later retrieve it. Right?
And what they showed is that this machine could be trained on genuinely symbolic tasks. Things like copy operations where you give it x arbitrary number of y's and then a signal to say retrieve x. Right? And if you're familiar with LSTMs or any kind of recurrent neural network, you'll know that they suffer horribly from gradient loss. And they if if you try and improve that, then they don't really have their integrative properties. And so there's a sort of silent sheribdis that you have to walk. And they're very hard to tune to accomplish symbolic like computation at the same time that it preserves the properties of sort of standard neural networks of integration and and sort of averaging.
So okay, so this thing can get around that by storing here. And as I say, they did a pretty good job of showing that at least um on some tasks achieved what one might argue is symbolic computation. The problem is much like its successors, the transformers, it and a transformer is a variant of this in certain ways. I won't have time to go into that in any detail, but but again, if anybody's interested, I can come back and say it's this is basically I wouldn't say it's a notational variant of a transformer, but it's a slightly rotated and tweaked and reparameterized transformer or the other way around. transformers are reparameterized version of this one could couldn't think of it. Um so uh the problem is like transformers it's insanely data inefficient. It took hundreds and hundreds of thousands of trials just to get it to learn to do a copy operation. So it was sort of a proof of principle but it didn't achieve this efficiency that at least we know that the human brain is capable of. If I give you a copy operation that, you know, you might take a, you know, a little bit of of of exposure to figure out what I mean, but but even if I don't tell you exactly what I'm doing, if I go a bb a and then a bb a and then q lq, you'll get it right at some point.
You'll get when I signal, tell me what the first thing was, you'll know how to do it.
So we looked at this and we realized it's suffering from the same problem that all neural networks suffer which is that it monges everything together over here even though it can cache stuff the loss is coming through here and that loss is getting infected by the data much as whatever it stores here right and the problem that neural networks suffer if any if there's any one thing that for me way that they that that they suffer from it's from the inability to tell when it's useful to abstract and when it's useful to use the data and this thing doesn't have any advantage over any others. It it in principle it could abstract but it doesn't know how to do it without getting hit over the head with the data much as transformers.
So we thought well what if instead of just thinking about this as a tape we thought of it as a binding mechanism which from a neuroscientific perspective of course is way it's always been thought of this this in so far as you think of this as episodic memory in neuroscience we think of this as a model of the hippocampus which is doing binding. Um, I say this for those who appreciate what I'm referring to. I don't want to take time to to slide track on that, though. Again, I'm happy to come back later if people are interested in that. But here's what I mean in the present context. All we have to do, sorry, went the wrong direction, is add another LSTM, another controller, if you will, on this side. And the trick is going to be that whenever this thing writes here, this thing will write. And whenever this reads, this will read. And vice versa.
And when they write together, those elements will be bound. That's what this orange line means here. So if I write here, it's just like a dictionary in in in Python, right? If I write a key, I write the value here. It it's labeled as keys and values for the machine learning community, but but really you could just call these fields, right, or entries and they're getting bound and and I can retrieve either by matching this and then I'd get back this plus this or I could match this and I would get back this. If I match this, I get this plus this. Okay, but critically this side never sees the data. These are binding only in the sense that when you retrieve this, you retrieve this. But this side never sees this and this side never sees this. That is there's a firewall.
critically by again the ingenuity of of um of deep mind folks in developing this and in so doing making this differentiable and I'll show you later how you can do this you can differentiate through all this so I can train a network just on this with the loss coming here or even the loss coming here and that's the example I'll show you in a couple of minutes okay where the loss comes from here and it will benefit by the gradient errors both here and anything that happens here you can differentiate through this. So you never get the data corrupting this or these representations corrupting the data but the errors can pass through okay mathematic as John Anderson liked to say and again I'll show show you how that works in a bit. So what can we do with this? Well, we can train it ah and and critically because of the similarity based retrieval then basically when this thing is retrieving something based on its similarity to these this is getting whatever it's being forced to sort of respect any relationship among these that this might observe between these or better still to put down symbols here that will be useful in responding to that relation and I'll show you exactly how that works in a minute. So the critical thing is that when you train this what can this possibly learn? It doesn't know anything about the data but because of the similarity based retrieval mechanism it can only learn relational structure that is abstractions. It's forced to abstract. It has no choice to abstract.
Either it either learns nothing or it learns abstractions. It never sees the data. Okay. And so this is an inductive bias that we refer to as the relational bottleneck because in effect it's a bottleneck. It's not allowing the data through. it's only allowing sort of relation information through. Okay. So, first I just want to show you that it works and then I'll unpack how it works and and then but both sort of operationally and then under the hood in terms of the neural network aspects of it and then come back up and say a little bit more about how it might apply more.
Um I should say if anybody has any questions or concerns along the way, please jump in um and interrupt me. Um I'm happy to answer questions as I go.
Okay. Um, we refer to this, by the way, as the emergent symbols through binding network. And I'll make that clear why we call it that in just a moment. So, we trained it up on simple relational tasks drawn from Raven's progressive matrices.
So, we simply asked the network to see two arbitrary symbols and say whether the same or the different here. Of course, it would be different. If they were both the same, they'd have to say the same. Um, relation match the sample.
Two things are the same here. Now, pick which of this set goes goes with this.
It would have to pick this one because these are the same. If these were different, it would have to pick this um distribution of three, which I'll I'll I'll show you in just in a moment. Um and then abstract sequence rules, which I'll come back to at the end. So, we trained it on all these kinds of tasks, and we did it in a space of Unicode characters where in this case, we had a 100 different arbitrary characters that had no relationship to each other. The network wasn't trained on these other than to simply embed them arbitrarily.
And we trained it um to learn to do these tasks first just to see if it could do the tasks having given all the stimula. And then we throttled it back and we trained it only on 85% of the stimula to see how it generalized to these other held out stimula. So we played the typical machine learning generalization game. And then we throttled it back to 50% and we throttled it all the way down to 5%. 5% which is um I won't explain why fi three is 5% of a 100. It has to do with how many you have to show. But the bottom line is we trained it in the limit on just the number of symbols in this case just three symbols that were sufficient to to to to perform the task. No more no less. So, just like if I was to say to you, um, uh, cup um, uh, uh, uh, uh, tape sky or sorry, cup, tape, cup, and then I go sky, lamp, you say sky, right? I only needed three items plus two more as exemplars. That's why it's 5% of the 100. Okay, so we train it on exactly that number and ask how it does. And here are the results. We did it with um our our model is in red. Uh the SPN is red. We did it with a transformer, the neural touring machine itself, a straightup LSTM. Um ironically, a prior candidate um from Deep Mind called relational net, which did the worst. That's the green one. It can't do anything. Um but all these other models do pretty well up to about 50%, but when you get to radical generalization, they're at chance. And the model is perfect. And not only is it perfect at radical generalization, genuinely extrapolating from like the smallest number of exemplars needed to to illustrate the the the rule, but it does it with insane sample efficiency.
Instead of hundreds or thousands of epics, it's like 10 or 20 depending upon the rule. Some are 50, but but like orders of magnitude faster than others.
Okay. So, what is the claim? How does it work? Sorry, let's come back to the claim. So, here's how it works. Here's the distribution of three tests. The distribution of three, you're given three items and then you're given two more and you have to pick what the third the completion is of the of of the second set. And it's called distribution of three because there's a distribution of three unique items here. This in some sense is the simplest, you know, most basic symbolic processing task. So, I have to basically recognize that I have one of these, one of these, and one of these. And then I get two of these and I've got to pick the answer which is of course going to be the square, right?
That completes the set. So here's how it works. Oh, and I should say um in this particular model, you know, our hands, our thumbs were on the scale. We didn't train it to do to decide when it was going to encode from one side and then retrieve from the other. Um we had an external controller that synchronized both when these things wrote and when they retrieved. We're now in the process of building models that know how to learn learn how to do this. And at the end I'll show you LLMs have figured out how to do this. I mean we know that now.
Okay. And I'll show you examples of that in just a moment. Um but for the moment just we'll allow that there's a controller that says okay this thing needs to look oh and and also there's no sequential attention in this. Uh we're hand the model itself doesn't have it.
We're hand sort of sequencing the stimula. We do have models that know how to do the sequential attention now. Um so we give it um the circle and it encodes a circle and then it lays that down. on its memory trace. And this thing starts out by just having to write something because it's commanded to write whenever this thing writes. So it writes some arbitrary thing. I'm showing you what it looks like once it's learned. And I'm just picking the letter A to designate sort of what it learned to represent there. Okay? So it represented some thing called A. Could have been anything. Okay. And now it sees the square, right? So encodes the square and it writes down square. And this thing then decides to write B. And then it sees the triangle. And so it writes the triangle here and this thing decides to write C. That's what it's learned to do. Okay. And so now we've encoded the stimula and now we're going to start um encoding and retrieving.
Right? So here we've been writing to memory and now we're going to retrieve.
So it sees the triangle. It encodes the triangle and it retrieves. It does a match and triangle matches something that's in memory, right? And so it retrieves triangle. And remember what I said before when this retrieves this has to retrieve. So this gets back a C.
Okay. And that goes into the LSTM LSTM and it does two things. It encodes that it's at seen a C. Remember LSTMs have the capacity both to serve as a memory and sort of an integrator. And they're really good at learning sequences. So this thing that one of the sequences has learned, you've already seen, is to write down ABC in these slots, right?
And then when it sees the C, it learns to encode that in working memory.
And then we see it looks at the circle.
it retrieves circle on the on the sort of the the data side that's associated with the the the encoded A. It retrieves A. So now it says it's A that's in its working memory. And now it knows it's looking for B. It's learned this, right?
This is all this stuff over here has been learned. We didn't code any of this. It's just that the LSTM learned to do when it was trained on 20 to 30 epics of this. Okay, so it look it's now it knows that it's seen A and C and so it's looking for B and so it encodes the pentagon. Now now we're looking at the options and and and so it's going to look at the pentagon and it's not going to find anything. So it does a question mark and now it's keeping track of positions. So position one doesn't know what's there.
Position two, it goes with square. It it retrieves square, retrieves B. Ah, it was looking for B and it got B. That's a match to its internal working memory state. Okay. And so now it encodes that it saw B in position two. You know, rinse and repeat.
A it it's a circle. There is a circle there, but that's in position three. And now at the end, it knows that it's looking for um position uh for B. And so reports out whatever is associated in working memory with um B. So I know I've sort of put this in words. You'll have to trust me that this thing trained and trained as efficiently as it did. What I'll argue and and well what I'll tell you is that when you analyze the representations here, these are literally symbols in the sense that whether I give it circle, square or triangle, asterisk, you know, um pi sign, plus sign, it doesn't matter. All it has to do is get three different things here and every time it just writes ABC and then it just goes through this routine. And because it only is using symbols here, it's basically binding those symbols to the to the to the fillers the the datim right as they come in the order. And because it knows about the order of those symbols, it's learned a function which is the distribution of three and it can perform the task perfectly over any data that it can encode here. It doesn't matter what those data are. It never has to have been trained on them past one set to learn this function. So what is this thing learned or how can we describe it?
It's um using external memory not just as RAM but as a bottleneck as a way of arbitrarily binding two separate items without one affecting the other freeing up one of them to learn to be entirely abstract. So it's learning a function over here. learning an abstract function one two three keep them all different right and this is serving as the variable binding apparatus that one needs in any sort of self-respecting you know venoyman architecture and then together you get the emergence of symbolic computing it learns the symbols here super simple example but I think it passes muster in terms of the formal claims that are being made here okay what's really cool is that once you see this you can start implementing it in a variety of ways So one of our colleagues at MA, a guy named John Carlo Kirk in um Yosua Benjio's lab took a look at this. This is the ESPN model and he said, you know what's this thing really doing? Well, in the abstract, it's building up over the set of data a correlation matrix, a similarity matrix of each vector with every other vector. And then it's just reading out from that what the data, you know, what what the data are or what the right response is at each trial. And so he simply reconfigured this to be simply a correlation matrix that gets populated by the data and then a um decoder that reads out of the similarity matrix um the answer. And this is in effect acting as the relational bottleneck. This this correlation matrix it's not as plausible because he had to construct that similarity matrix or he had to sort of lay the rows down manually with CC code.
This is doing it with an LSTM um and learning how to do that with sorry with this external memory and and the LSTM but in principle it's doing the same thing and it when you run it it learns all the same rules in fact it's a little cleaner because it doesn't have the gradient loss of the LSTM for learning the function you can simply articulate the function in the form of this similarity matrix and so it's it's a little bit more efficient but it's the same idea okay and notice that the data never make it past this right this thing is only operating over the similarity matrix which has lost all the data themselves right it's just saying Is the first thing similar to the second thing?
Okay. And in fact, you can put this in a transformer. Transformers have attention heads as I'm sure you all know. Those attention heads are operating over queries, keys, and values, right? So, it learns a set of queries to query its data, which is basically a humongous context window, which is basically a hack for an episodic memory. That's the reason they're having to make them bigger and bigger. So they can take more take kind of more and more context and achieve essentially what you know what what the brain has in the form of episodic memory and what production system models have in terms of of of of working memory, right? Um and it uh and it operates sort of by querying taking a query and then using that to match to a key that gets a value and then all of that information that whole vector is passed through to the next step. So it can all coingle. So it's a lot like the neural touring machine, right? It doesn't have this bottleneck.
But if you simply isolate the values that are being retrieved from the keys and queries. So you do a key query match and then you simply retrieve a value that that's associated with the vector that the key query the query query key match found you basically have the relational bottleneck. And so we've built transformers called abstractors.
This is work done by a colleague of mine John Lafery at Yale with a student of his Anie Alaba and show that it has all the properties of the ESPN. And because now you can stack them, right? You can actually do more complicated things. And so we've shown that this works not only in visual reasoning tests of the sort that I've showed you here, but in sorting tasks. Here it was a transformer. So Q sort can be learned by this thing in in again about three orders of magnitude more efficiently than a standard transformer. Um and sequence learning. And here we've begun to map this into um uh or used it to to simulate tasks that we can have people do. and it shows many of the properties um that people exhibit. For example, advantages of block versus interleaf training. Um one of the curricular points that that that I sort of alluded to at the beginning that I think is central to how these things work.
So, okay. So, that sort of shows you how it works or that it works and how it works. Um and I want just to take a quick moment here. Um, yeah. So, I I should be done in at least a couple the five minutes. Is that okay, John, if I take another five minutes?
Sorry.
Yes. Okay, great. Thanks. Um, so, um, the question now is, you know, are transformers doing anything like this or are they really just memorizing stuff?
And I they maybe they're certainly memorizing stuff and it may be that some are only memorizing but we've looked in ones to which we have access to the inner representations um and asked um have have they implemented anything like this relational cross attention or abstractor like um relational bottleneck mechanism.
And to do that, we picked a simple set of sequence tests, the sort that I've been discussing all along, ABA or ABB.
Um, and we got a bunch of LLMs, not the biggest ones, but ones that were on the one hand capable of performing this task with with basically oneshot learning where you basic or two-shot learning.
You give it two examples of the sequence like this, and then you ask it to solve it on the third. All quote unquote in context learning. One of my least liked terms. It's reinventing the idea of simply activation based processing, but so be it. Um, and so we picked models that could solve this task in twoshot in context learning. Um, but that had open architectures that we could go and probe the actual representations. And I don't have time to go into all the details about exactly how we looked, but I'll just give you a quick sense of it. Um, first we postulated the the architecture in a transformer would look something like this. There would be a first level of heads of attention heads that were doing the symbolic abstraction pretty much exactly as I showed you in the ESPN model where it sees a a set of tokens and it simply puts the symbol A down for the first B down for the second um and then it retrieves based on the first and the second or one that isn't like it whether it's one of these two symbols or something different. Right? So this is exactly just rotated the ESPN architecture and then another layer that actually learns based on the symbolic representations to perform the ABA task or ABB task and then finally a stage um that dreferences the symbols back to their original token. So there are these res connections in transformers you may know about that are bypasses that would allow this to sort of map the symbol back to the um to the original token. So this was the hypothesis arch hypothesized architecture that we hoped to find in a transformer. And we did a bunch of work to lesion these transformers um using what we call causal mediation analysis where we basically either knocked out heads one by one or we took heads and we took the state they were in for one rule and then replaced it with the state that they should have been in for a different rule. having measured the the system in that in that in that configuration and asked if we could basically force it to perform the first rule. And and if we could do that by by basically TMSing, right, electrically sort of stimulating that attention head to be in the state that represented one rule versus the other, it should respond as if it thought that that was the rule that it was performing. And we looked, we did a massive search for these to see if there were any heads like this. And we didn't find many, but we found them.
So these are the symbol abstraction heads. Um and up is higher layers. So these are the lower layers. These are later layers. And then beyond on average below above that were these symbolic induction heads. The one that's that were actually learning the function. And then finally retrieval heads. Unless you think this is a multiple comparisons. We were dutiful in doing careful permutation analysis. All these are statistically significant at insanely high levels like 0.00001.
Okay. So these are there and you can go in and look at them and watch the representations as they move. How they're learned we don't have anything to say about. But I would argue that this is evidence um that symbol abstraction heads um and induction heads live in transformers at least in some instances and it's across all of the transformers. This is the case for all the different transformers. They vary in the specifics um obviously but the same principles apply.
Okay. So um back to the ESPN architecture. You can think about this probably you've already seen it um as a production system. Right? Here's the matching mechanism.
Here are states of the world, the conditions. They match. They give you an output that's going to be read as an action. Rinse and repeat. Right? And in fact, I've taken one model and literally taken the EM and taken productions out of a production system model and plot them in here and have it perform exactly the same task. So this works as a production system model. Basically, you have to you have to program this. Right?
Now, there is a difference.
here in a production system model, at least as I understand it, here's where the semantics are, right? There's there's spreading activation among these things. And so which of these you match is determined not just by the match itself, but by which things have been prior activated because it's the state of these that matters. It's not just the match. Um, and so the semantics lives in here. And so you have to program all that into it. And if I have this wrong, please forgive me. And I'm I stand to be enlightened. Um, but at least that's my rough understanding of how actar works, which is what I'm more familiar with. I don't know all the details of Soore. In contrast, in neural networks, all the semantics lives here. This is assemantic. This is genuinely episodic, right? This is declarative memory, but only in the episodic sense, not the semantic sense. Semantic memory lives in here in what I would call well, you'll see in a second. So that's sort of a mapping to production systems.
It also happens to comport beautifully with the architecture of the brain. Not looking like this, but looking like this. Okay. So if you think about this as the endrinal cortex, the portal to the hippocampus, which is where we think binding and episodic memory occurs, and then these systems as where all the semantics is in the neoortex and all you stipulate is that posterior neoortex projects to different part of the entrinal cortex than prefrontal cortex or at least there's some separation between these, right? And then the binding among them occurs in here. you have an abstractor in your prefrontal cortex and we've begun to sort of build models now that are looking at using this to make predictions of human brain activity and we have a few provisional um hits.
So I promised the last thing I'll do here is to just show you how this works in backdrop. How do you backrop through one of these things? So think of these two as sort of those parts of entoal cortex or of the data and the abstract side. And what's going to happen is every time that I get a new piece of I write to this, I'm simply going to add weights here that are proportional to the strengths of these activities. I'm basically going to play take this activity pattern and plop it into my weight matrix. Okay? And we call it a modern hopfield network because that's essentially what a hot field network does. Okay? And it's recurrent in the sense that I'm going to plop the same weights. I'm going to plop these weights here and plop these weights here and then I'm going to use this for retrieval. So it's as if I'm this whole thing is a hotfield network where I where I get to I guess I should have said hopfield network in the context of it's used to model hippocampus where it's assumed that weights are are formed by rapid synaptic plasticity sort of single shot or two or three shots. So here I'm rapidly programming these weights by simply plopping these activity patterns down into those weights. That's what LTP does, right? And I'm doing it for this side here and this side here. And then these are connected through the hippocampus. The these two are bound. So this might be a unit that's sort of a conjunctive encoding of this representation. And this one is a conjunctive coding of this representation. And then these are then bound. And so now I can just keep going about doing this. And as long as I have enough units here, each of which serves as sort of a different different entry, then I can encode my entries in my on my touring tape. Hippocampus in this view is sort of the touring tape. Okay. And now when I go to retrieve I place the pattern of activity here. And now I ask what is this best match? Well these weights are going to give me the natural dot product. Right? That's exactly what neural networks do. They take dotproducts of activities and weights.
So this will do a dot productduct which is exactly my memory b content addressable memory retrieval operation and it will retrieve this will get the biggest dot product that will activate this and then that will through these weights retrieve the pattern that was originally associated with this. And because all of these weights are just, you know, you know, continuous values and and and and all of these are linear or continuous nonlinear functions, um I can flip it around basically copy these weights and flip it up. It's a way of dealing with recurrence in neural networks and now train through that. So I know that was quick, but at least it gives you a sense of how it can be done.
Again, I'm happy to elaborate if anybody's interested.
Preserves relational bottleneck. And now we can imagine sort of how this fits into a bigger scheme. Well, we have our encoders and decoders here. We have RNN's over here. And essentially we've got ability to represent context that can be used for attention. The RNN is doing integration. And I think those are the critical ingredients. And so I'll just end on this summary slide saying abstraction requires inductive biases in the architecture as well as curriculum.
I haven't said much about the curriculum. The TE's on that is block versus interle turns out to be super super important. Um integration is critical. That is the formation of context representations that represent your past. It's the only way you can deal with a non-Marovian world with Marovian processes. You store representations of the past in the form of your knowledge presumably somehow integrated. Um you can do that through time averaging which is lossy. That's the way it's typically done with LSDMs.
You can do it with sort of windowed integration. And this is essentially what um transformers learn to do or you can do it with episodic memory which is essentially lossless um in so far as you have enough space to encode everything.
And then you need attention. You do need attention just not just just not just attention. Um and attention is just simply the use of context representations that have been learned um with these biases to do selections of mappings um or the shaping of statistical structure in more sort of continuous um higher dimensional spaces.
And uh as I said I could do a deeper dive on this last part but since I'm out of time I will just skip to the acknowledgements. Um these are all the people who did the hard work and gave me the privilege of sort of representing it. Um, thank you.
That was fabulous. Um, thank you so much. And so what we're going to do now is I've invited uh three of my colleagues from around the country to comment or discuss or ask questions of Jonathan on it. Um, we hadn't talked about what order they were going to go into. So, I'm going to guess I'm going to start from the bottom. Uh hopefully this doesn't uh go against the way they see themselves. I'm going to start with um Andrea Stoko. Andrea, you go first and introduce yourself and then uh we'll have Christian and then Paul. Okay. Can I just interrupt to say that I I I'm both honored again and frightened.
These are all the luminaries that I've heard about my entire career growing up.
Only one of whom I've had much contact with, that being Chris John. So well this is uh I think this is thrilling to have you here for us I think in terms of information transfer and novelty uh this is just great at least for me to get some ideas that I had not had before. So I I love it. So Andrea hi uh thank you Jonathan that was a fantastic talk and I'm absolutely scared to be the first to speak after you.
So I I'm going to take the role of of the I'm going to ask you a question. It is probably the closest to neuroscience uh in this panel. So I really love your point about the curriculum. This is something that has been like um on the back of my mind for many years now. And one of the things that I've been wondering a lot is the fact that if you look developmental, you talk about evolution, but even developmentally the brain doesn't mature at the same paces.
And even if you divide up frontal, interrinal and posterior parietital cortices, there is a developmental trajectory. Some parts will develop early, some parts will develop much later even. And this seems to play an important role. You don't want children to start binding at age one. And you don't want them to start planning at age four. You want them to develop first layers of representations that are more basic, more content based, and then develop episodic memory and then develop abstractions. At least this seems to be like the the general rule that nature has learned and to me this has always been like a natural way to identify modules in an architecture by developmental times. I was wondering whether something to be learned about in your architecture about these distinctions or in training.
Yeah, I mean I agree with everything you said. Um John is the format I respond or did you want everybody? Okay for 35 minutes and uh what whatever makes sense. Um but we're going to start with the four of you and then we'll open it up after that. Okay. Okay. Yeah. So um yeah, I think developmental trajectories are critical. Um, I will be honest and say I don't have too much to say about that yet except to share your intuitions and some work that Jay did many years ago showing that you know um you can account for stage like stages of development by the the the the extent to which the network can exploit the structure it's learned at one level to construct new structures at the next level. And the closest we've come to that in recent work is a collaboration with another one of Jay students, Tim Rogers, who I sort of consider like the the authority on um neural network models of semantic cognition with Jay, but but sort of really staked that out as his space. And together we've teamed up and sort of exploited the idea that um control, which is what I work on and semantics is what he works on, are really just an artificial distinction of convenience and that they're really just two sides of the same coin. And if you had to sort of label them differently, what you'd say is semantics is about lower order statistical structure of the features of the world. And control is about the higher order structure of those features and the semantics of the spaces that develop around those features. For example, category structure and how that maps onto affordance, right? Which ones are useful for which kinds of actions, right? I think about an elephant differently if I want to ride it than if I want to get away from it. Um and and and so if you then sort of train up a system that has a layered like structure but with a bit more indirection to the point that's supposed to be responsible for modulating the lower order structure in the service of action, you get sort of this natural progression of learning.
And there's one paper that we just had in um published in psych review um and it's called the semantics an integrated model of semantics and control the ISC model. Okay. And that makes the sort of as close as we get to the developmental point. I mean it's super crude it's simpler soft but I think it gets the principle across but I think as important to as and I think the the evolution has figured that out then you know myelination of the prefrontal cortex happens much later than milination of posterior structures possibly for that very reason. you don't want to sort of commit to structure at higher levels, right, until you've got the structure on which that needs to supervene at lower levels. And so this model sort of doesn't have any mileation processes, but it it sort of shows that kind of that kind of trajectory. Um what we learn from that model though or what what that model doesn't explain is how you infer the right context in which you're going to need the context you're going to need to operate. How do I infer that this is the moment where I need to run away from an elephant as opposed to step on top of it? Well, that seems sort of an obvious question. If it's coming at me, I run away from it. But if it's just standing there and I want to think, well, what am I going to do? I need to evoke a context and how do I know which is the right one, right? And so in these models, including all the predecessors and all the ones that Jay did and all the one and and Tim and and the Rumlart network from which all of these originally emanated, if any of you are familiar with this, the the Rumlart semantic network, the context is presupposed. the the the the the designer picked what relevant context might be useful um and then put those one hot inputs along with features of the world and then it was trained to sort of get the right context information based on but it didn't do context inference and so we in the second line of work which addresses I'm mentioning this because it has to do with curriculum we simply allow that the context had to be inferred not by prespecifying it but based on the prior stimula So we simply took a model that was exactly like the semantic networks but we added one extra layer that integrated the last n stimula that it had seen and this was square wave integration. That's why I had that on my slide where it's windowed right so would do it would it would take the last three or five or 10 or whatever representations very much like in context learning in a transformer. um and it would simply sum them and it would take the sum of those representations and use that as the attention signal as the context vector.
Okay? And when you train a network up like that from scratch, not only does it learn context representations, but it learns to invoke them based on the stimula that it's just seen.
And that translates into this idea that when you're learning about something, the world has temporal autocorrelation that gives you that block-like structure. So when I go to the zoo, I care about the sounds and the shapes of animals, maybe less than their color because they're all sort of gray and brown. And then I go home to my coloring book and I'm given a whole Crayola set of 57 colors. And I want every animal to look different. I care about the colors less than the shapes and the sounds, right? And so I'm doing a lot of drawing and coloring about colors and those that's all summing up and creating a context vector. And that gives you a maybe an intuition for why we think blocked versus interled. If you if you randomly interled, you'd never get the categorical structure that makes things similar along one dimension at the zoo in a different dimension in your coloring book, right? It's all still the same stuff. And when you build a model like that, you get all these things that I mentioned at the beginning of the talk. get things like asymmetry of of similarity judgments, triangle inequality, which are all context effects for free. It comes out of the training of the model. Nothing was hand there's no the only thumb we put on it was this architectural bias that we we set the context layer up above the the the feature initial feature encoding layer and then we block the structure of the curriculum and then that context layer was integrating over that block.
That integral is essentially a soft form of relational bottleneck. It's data lossy, right? When you when you add up 10 representations, you don't know it could have you could have added 10 things in 10 different ways to get the same number 10, right? And so the integration is like a soft relational bottleneck and it learns the right abstractions over category structure that are useful for making the right inferences in different settings. So that that I know is just a promisatory note, but that part of the curriculum we're all over right now. This this is maybe the most interesting line of research I think in the lab is is trying to understand what are the temporal dynamics. It aligns I'll say one last thing. Um it aligns with empirical data about how kids learn from Linda Smith. I don't know if you know her work about burstiness. Yeah. Yeah. And and the burstiness is exactly what we didn't know about that when we did this. And when we heard about that we oh my god and one of the things we want to do is get the exact functional form of the burstiness that she measures empirically and make that the window rather than just this hard-coded you know sort of discrete window. Yeah, the brutess is exactly what I was was thinking about when you were describing this machine.
This this is like this is like wonderful. Thank you.
Great. Uh Christian, uh I have a question and a meta question. I guess I'll ask the question first. Uh so you mentioned the relational bottleneck uh prominently. Uh are there any other bottlenecks that potentially would be of interest? And for example, I think here of the fact that various cognitive architectures discovered sort of in different ways and at different times that grain scale matter. And so for example in Actar we broke down the original scale of of the architecture into sort of what you call the atomic level of a simpler scale. And that gave us a lot of nice properties and I think saw to some extent with their commitment to the 50 millisecond uh cycle time discovered that in in different ways and that would be an example of sort of a different general lesson that has been learned and I I I don't know that it's quite right to think of it in term of a bottleneck but in term of sort of a a way of scaling and organizing computation.
Can you think of other ones that are fundamental? Or do do you think that the relational bottleneck is just the one fundamental one?
No, I would not be so presumptuous to say it's the only fundamental one. I would say it's a fundamental one that needs to be there for certain kinds of efficiency of abstraction. Um I presented it in a rather sort of cartoon form and an extreme form as a firewall, but then I pointed out that there's ways that you can get softer versions. So integration there's data loss there's that's a bias towards abstraction but some of the data are preserved right if I integrate all colors I'll get a different mean representation if I integrate all shapes for example right and so that's a data sensitive form of relational bottleneck because the colors are all related to each other in ways that shapes are not and vice versa. Um in the broader sense there is this notion if for if you're familiar with it of the information bottleneck in information theory that that simply says you know you get you want to compress as much as possible. It's really just a variant of rate distortion theory that says you know you want to compress your representations as much as possible. Um, and that aligns really nicely with the idea that compression if it's done right is a form of can can be a form of generalization. And I think we're just sort of reexpressing that idea. So if you're familiar with some of the work by like Chris Sims and others and there are other lines of work like this that that show you can sort of retrieve uh uh um Shepherd's universal law of generalization from um applications of the of the information bottleneck. And this you can think of this as sort of one example of that. So at that level I would say it's a pretty broad principle and and I don't know Christian whether all the examples you gave would fit under that but I would offer it as a working hypothesis that they might. So the specific forms of integration the data the structure of you know what's going to be most useful I mean that that is up for grabs. So you know the the the visual system seems to like you know linear transformations and rotations um as a basis set and so genes have learned that that's a good bottleneck. Um the ear seemed to think forier descriptors were better and so the tempanic membrane learned that bottleneck. Those are things that are learned over evolution.
similar things may apply in the space of online you know of of learning you know the agents learning and um I think there the devil is in the details but but they all sort of had this flavor of being a form of compression that's structured in particular ways right so I gave you one example of how that can happen great no that's great uh so so my meta question uh we're 50 years into the cognitive architecture program and John Paul uh Andrea and I have tried to sort of recognize and condense and formalize what's been learned and and part of the intention of that was to reach out to other communities uh neuroscience AI robotics uh uh to to try and express the lesson that have been learned and uh we've tried I don't know that we've been terribly successful And can can you think do you think that there is value in what has been learned in the program of research of cognitive architectures for the kind of efforts that you're doing? And if so, what what's the best way to try and extract and communicate that value?
I have two really strong responses absolutely of value. Don't give up. Keep doing it please. You guys have been an inspiration for me anyway and I know some of my colleagues and certainly the junior ones as I expose them to that space of work. On the other hand, if you know what's good for your mental health, you'll stop right now because machine learning people who make the money live in industry and are basically ruling the world right now don't give one hoot about principles or science. They just want to make it work. And no matter how much you talk to them about principles, unless you figured out how to get another tenth of a percent over the greatest latest benchmark, right, on how to build a rocket, I mean, and I'm and yeah, I see you guys have not surprisingly the same experience as us when we submit papers to Nurips, you know, it's just unbelievable the the the reviews we get. And so, you know, I for one, I'm going to keep beating my head against the wall. Oh, and I I will say, by the way, as you probably also know, that very little if anything that they're doing wasn't anticipated and discovered 30 years ago, right? In on the connection side, no less all the things that you guys have discovered, right? And so, if they're not even willing to listen to us on the connection side and say, "Look, we we know something about episodic memory and how right, what's the hope?" I don't know. I But I don't think it's it's I think you can't stop trying. And and my hope is that with us training our students to talk to each other um and interact and and I hear that happening from your community again through hearing John's talks and going to ONR meetings that we go together and hearing all the people that are trying to reach out and and us doing the same. Maybe we'll build a stronger integrated cognitive science, cognitive neuroscience, I don't know what to call it, but you know, principled science of cognition and computation that at some point they'll start realizing is useful.
You know, can I just say one last thing?
The the place where I see this where they're going to hit the wall and if I if I if I'm trying to be strategic and think where can we set ourselves up to have been able to say in an active way I told you so where they'll hear it right is in compilation automatization. So all of their agents they decide what this what should be done efficiently and what shouldn't what they're going to train.
But autonomous agents as you well know through the process of compilation and as I've always been interested in through the process of automization um can know when to do that. I decide whether to touch type to to continue to hunt and peck or learn to touch type right. I decide that there's an intertemporal choice between the time and effort you know spent doing this versus the time and effort spent learning to do this. And if we could formalize that and come to some sort of unified understanding about how that decision is made informed by the the way it's going to play out in a neural network architecture which I think we understand quite well now um and all of the work that you guys have done in facing that you know in every architecture you build what are you going to compile how is it going to compile what's going to be going through working memory what's going to be directly sort of you know learned as as as as a condensed production I think we could get ahead of them because they're not thinking about that at all. But once they have to start putting all these LLMs on edge chips and sending them to space where they don't have access to the mother ship, they're going to start realizing that they need to know how to do that, right? They can't just, you know, call up Google central and find out what the next state of the LLM is, right? They're going to have to put it on something small. And so it's going to have to make make these trade-off decisions. So I think that's one place that we could look to developing principles and maybe even showing, you know, ideally we show some successes. um that they need to pay attention to.
Yeah. No, I think that's exactly right.
So So I'll I'll pass it on to Paul to give time for everyone. Yeah. Sometimes somebody should ask a question he's passionate about, but um Paul, it's up to you. Okay. Hi. Um, yeah, I was fascinated by your talk and I to admit I I cheated some in preparing for this panel by reading some of your papers and that inspired me in in various ways.
Um, one of them was a followup on some of what Christian was talking about having to do with what's called the common model of cognition. I don't know if you've had a chance to look at that.
I have briefly, not as much as I feel like I should. Okay. In in reading your papers, I came across at least at least three of what I consider to be challenges to the way we think about the common model. And I want to focus on one, but the three are um whether working memory is actually a memory um the relationship between procedural and semantic memory and the one I want to focus on which is the relationship between uh automatic and control processing. Now in some sense your talk touched on all three but it didn't really focus on those on those questions. uh but I want to focus on the last one which showed up in one of your later slides when you mapped it over the brain and you showed the abstract and the task specific uh portions of the model and if I understand right from one of the papers I read by you you're proposing that is the means for autoomatization that essentially the controlled processing involves abstract or I think what you referred to there as shared representations and the uh u autom automatized processing were these task specific representations.
In contrast, in the kinds of architectures we've worked on, this goes back all the way to my thesis in 83, but in the common model, we talked about uh essentially procedural composition. So in terms so we go from deliberate um uh Christian mentioned the cognitive cycle.
So a process that takes many cognitive cycles to the compilation of that process to where say a single rule can yield the same result and that's generally the common model model of the transition from controlled automatic. If I understand your proposal it's it's a shift from these abstract shared representations that require seriality to task specific representations that can go on in parallel. Now, one could think of those as being compatible or as alternative as alternative stories stories. I was wondering if you had a sense yourself as to how you think about that. Are these alternatives that are incompatible or is there a way you can view them as u two sides of the same coin? Completely the latter um modulo the details, right? And that may turn out to be important. I I'm sure it will.
But in in in the sort of broader way of thinking about it, I think they're the same point. Um, and in fact, I took inspiration from the stuff that you guys have done and how you think about, you know, composition or compilation. Um, in thinking about what might be going on in a neural network. Um, to be fair, most of the work that we've done on that is in what I'll call the spatial domain rather than the temporal domain. Um, and by that I mean semantic space, not physical space. And um the idea there is that um when you first learn a task, you want to draw upon representations you already have. And the more generally useful they are, the more you can draw on them and the more risk you are of drawing on them at the same time. That's the idea of a shared representation, right? and and then eventually if you need to do that mapping a lot and you don't want to risk interference of that that representation being used in other mappings at the same time you sort of you know penocytose off a new representation that's dedicated to that mapping and so when I say spatial I mean exactly that you start with lowdimensional representations for things like colors and words for the shroop task right and then if it turns out that you want to use the color representation at the same time as the word representation but the word representation you don't want it to interfere with what you're going to say about color. You develop a new dedicated orthographic representation that's specifically, let's say, for pointing in a direction that the word tells you as opposed to sort of reading it out loud, right? And I'm sort of summarizing some work that we've done that maybe you've maybe you've seen. So that's in the spatial domain. That's sort of the hidden layer like do you have these lowdimensional, we used to call them minimal basis set, but now we're going to call them compositional representations shared that are useful for mapping lots of different things but not at once versus dedicated ones.
Right? So and that that story I think is rich. I mean there's a paper that remains to be published but it's available in archive form that's 200 pages and covers every single phenomenon we could think of in the task in the the flat mapping task space right task switching experiments the psychological refractory periods similarity all of them that don't involve critically any sequences now in act or sorry in in so in production you're interested in sequences and the compilation you talk about often is in the multiple steps and so I think you just rotate space into time and you get the same problem.
Though the specifics may be different in terms of the tools you need to do it, but it's still the same idea when you're doing multiple steps. Why are you doing that? Well, because each of these is compositional. It's useful in lots of different ways and you have a toolbox, a basis set in which to use them. And they're beautifully compositional, not only in terms of different purposes, but maybe even in the order in which they can be applied, right? Maybe there's some syntax, there's some constraints, but they're exactly the same character as these lowdimensional representations we talk about as being shared representations in multitasking space.
Right? They're they're both notions of specialization as being the focus of autonomization, though one seems to focus on specializing the representation and the other one specializing the processing.
I'm wondering if there's a a joint story that could be told that sort of combines those two in some interesting way.
Did we lose Jonathan?
So, so I I and that's unfair. I And as and and I think I sort of allowed that by saying the devil may be in the details for how you treat time differently than space. That has to do with integration, right? Um but but I think in the end it is about the specialization of the representations.
Yes, I I would stake a strong commitment to that that when you're doing it sequentially, you're using different representations and at some point you realize that the first one you want is not the one that's sort of the entry of the in the sequence but has exactly the mapping you need to get to the end. You've dedicated so you know if I see ABC D ABC D ABCD eventually I don't bother with the B and the C. I see A and I think D. Okay. And there's a representation, a dedicated representation that A evokes that gives me a mapping to D as opposed to B which then can get me to C which then gets me to D. So maybe the way to think about it is so in the original steps you have very general rules that have maybe a small number of conditions to each. Um in the compiled thing you tend to have these very long sets of conditions.
That's the representation essentially is a specialized representation. That's exactly right. the condition, the articulation of the condition is exactly what's in that episotic, right? That's that's what you're mapping that you're going to retrieve. That's exactly right.
Okay. Thank you.
one other it's more of a clarification question from ear early part of your talk when you when you talk about the um ABC in those three-part problems and you go through one one problem and you get a b and c for the for the three different shapes and then you go to the second problem or the third whichever is the test one and somehow you know that the first one is still a I wasn't quite sure how that happened because that's the seems to be the focus of the abst raction totally the key is that the abstract side learns one two things. It learns that one, two and three form a set and it always has to satisfy itself by having in memory all three of those.
Okay? And that it's going to lay down those symbols that it's learned always in exactly the same order, the same representations in the same order. One, two, three or ABC as I had it. Okay. who learns all that and and on the abstract side that's all it learns that's all it can learn and it's the timing of the task which of course is you know central to how it's performed that you give first one stimulus then a second and then a third so those ABCs perform exactly their role those are the roles and whatever you put in are the fillers it doesn't matter because they can't they happen in that order and once you have literally a pointer to each of those three different pointers you could call it a pointer too right Now it's just a matter of when you see one of the things you retrieve that one of the pointers.
When you see the next one, you retrieve whatever the pointer is that was associated with that and you're left with the third pointer. And those pointers are just the ABCs. And so now it just doesn't matter what the data is because it's learned the abstract rule and you have this variable binding mechanism. It's literally a variable binding mechanism, right? the variables are ABC and the data the fillers are you know whatever you happen to put in the order you put them in in the in the in the problem that does that make sense thank you yeah thanks Paul I'm going to go back to Andreas do you have any follow-up questions besides I do I do it's going to be quick and I'm going to riff off Paul's question that was also my other backup question was about the fact that in this newer architect aritecture that there is a whole and of course like my my my personal training bias. Let me see like that immediately like because my favorite part of the brain is the basog ganglia and you talk about procedural knowledge but you have production rules but you were putting that in the hypoc campus and it occurred to me that there are two ways to think of binding and one is like directional like you have a key and you eat a value and this leads to temporal dynamics and the other is like pure superposition which is another way to think about binding in the hypocampus and those seems to be two different uses of binding and also two different structures in the brain.
So I I I apologize for not having been more careful about that. Um uh the RNN is is the basil ganglia or contract.
Okay. Okay. There are loops. Randy O'Reilly a close colleague of mine has spent his entire career building out that idea. Um they're subject to exactly the kinds of in fact we proposed LSTMs as a model of that well before they were popular in machine learning. Hawk Rider and Schmidt Hoover published their first paper and literally we saw that paper and said ah that's the B the the basil ganglia PFC loop um and it's a gating mechanism the basil ganglia gates what one part of the cortex wants to get into another right the prefrontal cortex acts to maintain at least some persistence and and as well as integration and so that that yeah I should have labeled it I'm sorry but that that's where that is and you're right that does that's the production part of the system yeah all right perfect thank Christian, you have a small point or a question to follow up.
All right, John, if you have Are you asking if I have a follow-up question?
Sure. Um, actually, not. Great.
Christian. Um, could I ask the people on uh the Zoom to raise your hand if you want to ask a question? And I'll just let Christian fill in the blanks for a few minutes.
Christian, you have anything? Uh, sure.
Oh no, that sounds fill in the blank sounds like a really important role. Uh so so to to to to go back to the the the last thing that uh in in Jonathan's answer to my question about the the importance of sort of that he talked about compilation and automatization. I thought uh I I I thought of it as metacognition which is something that we we're talking uh we're discussing very much right now in the context of the common model and something that we we've in a couple of instances use cognitive architectures as metacognitive overseer of more automated neural net type type system which seems to be a natural sort of separation of of of specialization there and then really the best way of leveraging the generality of cognitive architectures.
Uh do you think of that kind of thing as a separate system? Do you think as something that could should ultimately be integrated into into the the the sort of the lower level system itself?
So I I'll swap out the um the light bulb joke for a different metaphor. I am a connectionist in the following sense.
it's turtles all the way up. Um, and so I think that it's true there's a level of analysis I made a point to pointing that out and I know that you agree of when you decide or how you decide what's worth automatizing and what's not, right? But I think that we can invoke the same kinds of learning mechanisms one layer up. Now at some point you have a halting problem, right? I mean you it's it's you can't go forever, right?
Horowitz and Russell noted this in terms of the cost of computation analyses they did many years ago and you know obviously the brain faces that um so how many layers up you get to go I don't know but we have some examples where you can invoke for example reinforcement learning mechanisms to make that choice.
So there's a paper unfortunately not published but it's also available in archive a deep learning neural network in the multitasking space. It's again not sequence learning. It's about whether it's going to learn separated or shared representations for simply simple IO mappings, right? But it it's I think as I in my response to Paul I you know I think it's a similar principle um where an RL mechanism can learn when it's worth doing that and we actually that was based on a simpler sort of basian analysis um where you can sort of write this down in the simplest most abstract form as an intertemporal choice that is the cost of serialization against the cost of training and if the system has some way of estimating its learning rate and you put a cost on serialization that is the time it takes to do something you can find out, you know, what the optimal thing to do is and then you can parameterize it to learn that. Um, and and so there's a couple papers I can point you to if you're interested in in that. They're very simple models, but I think they point out how one could take this on with the same kinds of approaches and apparatus that we've that we've taken on um these other problems.
So, and is it specialized? I mean, you know, is it another layer specialized?
It's different than the lower layers, but it's sort of doing the same stuff.
Thanks. Uh, she hi. Hi, I'm Zil and I'm a postto from UC Santa Cruz and I I really like I mean your talks and all your thoughts especially on how you talk about like I mean these I mean the meta cognitive process of these models on things like sample efficiency and their performance on pattern recognition. So my question is I I want to dig a little bit further into this meta cognitive process especially on the like on the architectured model and their influence on the sample efficiency. So to make it very simple, do you think this sample efficiency comes from they are either trained differently or the model architecture being differently or do you think that human just have uh like I mean I would say better maybe that's not a better word for that like mean a better way to like I mean look into patterns learn on these patterns or just having overall like I mean a matter of cognitive process that the neural networks or other AI models are not able to master yet.
I can't answer the last question. That's sort of where the research lies. I I don't know what makes people special. Um either with respect to other animals or I do I think I have some ideas about what makes them special with regard to to the current generation of large um neural networks and it is the inductive bias. So so the ESPN model benefited not by curriculum. There was nothing about the curriculum. We trained it on the same curriculum that the transformer got and that you know that all the other comparison models got had nothing to do with the curriculum. It was entirely the inductive bias. That said, the model that I was mentioning in response I think it was to um to Andrea's question um this integrated semantics and control context inference model does leverage curriculum to do even better. Um honestly we didn't really look at the sample efficiency because we don't know what to compare it against. I mean we did compare it to transform existing transformers on its ability to perform these semantic tasks but we didn't train those transformer models. So I don't know what the sample efficiency is there of the curriculum but in other work where we've looked explicitly at blocked versus interle training and there it was in sequence learning. So there actually now I think about I should have mentioned this there is one example um where we have these models doing sequential performance tasks learn learning sort of finite state um grammarss if you will and um there we didn't do any work on compilation there but but that's a place we could look um in any event um where uh where we where we there where we look though the model does benefit in terms of its efficiency based on the curriculum um but that wasn't really very richly explored. So, I don't want to say that with too much confidence, but but I have I'm have a very strong intuition that it'll be the interaction between the two, right? Um, you know, you learn math with better with a good teacher than a bad teacher.
Maybe that's a soft statement, but I think there's some truth to it. All right, our final question is from Brian.
Hi. And uh, thanks for the the talk.
That was wonderful and fascinating to listen to. I just had a really quick question. If one wanted to um like download and try out some of the models you talked about in the various iterations, uh do you have a recommendation of like where to start as far as like um easiest to start learning and diving into requiring least amount of like expertise and how to get it to function well, that sort of a thing.
Great. I'm really glad to you're interested. um you know to the the best the most responsible advice I could give you is PyTorch and then I can point you to some implementations of our models.
That said um as Christian knows and I'm a little bit embarrassed to admit I've been working for many years on a a basically an API. It was didn't start as one for PyTorch but it can be thought of that way right now. It's called Sulink and it's in Python, but it's an attempt to try and distill away all the under the hood, you know, machinery and magic of, you know, atom optimizers and, you know, all of the crazy stuff of neural networks and just give you tools that are pretty much well enough formed to do interesting stuff with in a way that you'd recognize. So, I have an episodic memory mechanism. I have an RNN. Um, and then other standard stuff like decision-making modules. It's sort of like a plugandplay Lego, you know, environment where you can put together models. And um if you're interested, I would love nothing more than for people to be trying it. We've never advertised it or or or promulgated because I I don't want to do that until it's totally rock solid and I feel like I have enough models that are useful.
And so that's why it's taken so long, but I think it's getting close now. And some of these models, we just sort of crossed threshold in getting them implemented. So the ESPN model is now implemented and there's a few others that I've referred to that are implemented and it also has the virtue of having built-in optimization not back prop but like standard grid search op for for hyperparameterization. Those are tools that are built into the environment. Um so it it does make it a little bit easier to get models not just built but to work. Um, and so if you're interested, just reach out and I'd be happy to point you and and any time you're willing to spend playing around with it would be tremendously appreciated, especially if it comes with feedback. Awesome. Thank you so much.
This was fabulous. Um, I just want to thank uh Jonathan, Andrea, Christian, and Paul. Um, this really was uh informative and I think a lot of us learn new new things and uh I think it'll affect us in in our research. So, but can't ask for anything more, Jonathan. So, thank you so much. Thank you.
And I look forward to continuing the conversation and a boy, your comments about publication and transfer of information. I I just don't know how we I think there is a community of people that are like-minded um and that we have to find a way um and not just that on our program reviews to uh get them together and I and and there should be cross talk across these levels and and across the people doing this.
So, you know, maybe COGSAI somebody, you know, it' be great to have a COGSAI symposium on this that Yeah. Well, there's I mean there's also um ICCM, but u maybe that has too much of a reputation of uh of symbolic and not I don't know
Up Next

End-to-End Encryption Explained: How E2EE Protects Data Privacy
@SimplilearnOfficial
4K views•2024-07-25

Building Real-Time ML Pipelines with Feature Stores and MLOps Frameworks
@ODSCAI
5.1K views•2022-02-20

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence






![He Co-Invented the Transformer. Now: Continuous Thought Machines [Llion Jones / Luke Darlow]](https://i.ytimg.com/vi_webp/DtePicx_kFY/maxresdefault.webp)
































