Artificial intelligence and machine learning algorithms, particularly generative models like transformers and graph neural networks, can accelerate the traditional drug discovery process by generating new molecule candidates from a vast chemical space (estimated at 10^30 molecules), using the same computational approaches that power natural language processing to design molecules that are effective, safe, and manufacturable for treating diseases.
AI-Powered Molecule Design for Faster Drug Discovery
Added:We are talking about computer science and the search for new medicine.
I'd like to introduce you right now to Lei Li, Language Technologies Institute, Carnegie Mellon University School of Computer Science. Lei.
Thank you so much for making the time to speak with me today.
Thank you for hosting me.
Tell me a little bit about your job and what you're doing at CMU.
I'm a computer scientist working on artificial intelligence and machine learning.
The type of job I do is mostly research on developing novel algorithms using data to develop autonomous software that can develop intelligence to smart tasks like our humans.
The particular problems I'm interested in, generative algorithm, generative models that can generate text, generate image, generate new content including generating new molecules for drug purpose.
It must be very rewarding to be working on solving problems that can benefit humanity.
Yes, exactly.
The reason I started working on molecule generation and molecule design for drug purpose was I started around 2020, so I was attending a workshop, a conference in MIT that was right before Covid.
So in February where at that time there was Covid in China, but in US there wasn't too many cases yet.
So I was able to attend the conference in person and the conference was talking about molecule design and drug discovery and manufacturing using AI.
I was writing there talking about all this and I wasn't working on anything yet, but I was listening to all those advanced research at that time and right after the conference I received the email saying, oh, someone attended the conference, got Covid, so are you affected?
Fortunately enough I didn't got affected by that, but then that was a moment I realized, oh, it's very important.
It's a very important topic and also AI can play a bigger role in this area that can be really beneficial for all humans.
Well, I mean look, that's when mRNA leapt into the basic public space and as something that we realized was recently programmable.
This is effectively what you're working on in a different capacity.
So take me through this. What is it that you're doing in the space right now.
Talking about biomedicine. This is a huge space.
If we talk about pharmaceutical research and the path or the effort to develop effective drug takes a lot of effort from basic research, trying to identify the target for disease and then trying to identify some molecule that can cure this particular target, bite to the target and be effective in our human body.
We're not being toxic to our human body and then after that we need to do a clinical trial like phase one, phase two, phase three and then manufacture it to be cost effective and deliver it to the market.
So the whole process can take over 10 years with 1 billion to even 10 billion US dollars to develop one effective drug.
So that's a very long duration and our goal is to accelerate certain parts, certain steps in the procedure. In particular, we want to accelerate the design and search of this molecules both before clinical trial and after clinical trial. We have to do the clinical, we have to have the real patient to try these molecules, but before that it also takes several years to find an effective drug candidate.
So our goal is to use AI to accelerate this process so that instead of taking several years, we can just use our high performance computer GPUs, use the same AI algorithms that we are using for natural regulatory processing computer vision to find those molecules that can be potentially effective for certain disease target while satisfy certain chemical property we want. It should be saludable in our human body.
It's not toxic to our human body.
It should bide well to a particular target for the disease.
It should be easy to manufacture and synthesize this molecule in the lab and in factory.
This is a very interesting point that you bring up.
I was a health reporter in the year 2000 and they were talking about the concerns over R and D for drug development.
This is what 24 years ago potentially bankrupting our system for all of the reasons you just outlined and it's still a factor and now we have this interesting work going on in AI and ML that is able to very much more quickly, this is in its infancy, start pulling apart and finding these new molecules.
But what I don't understand is how NLP works.
Is it because of how it looks in sort of space and time?
Why is natural language processing a part of something like this to help discover these new molecules? Can you help me understand that?
That's great.
So there are different kinds of drug molecules.
Usually when we are talking about target-based drug design, we are referring to small molecule.
So are the molecules with below let's say 50 or 100 heavy atoms other than hydrogen?
So like carbon for example.
So if we think about the potential space of all possible molecule with just 50 heavy atoms, so it can be, it's a very large number.
It's maybe more than the number of say sand on earth.
So the rough estimate is 10 to 30, that's the one. After that we have like 30 zeros.
So this is a huge space and these are all possible molecules, but it doesn't mean we can all synthesize all this molecule and it doesn't mean all these molecules can be effective.
We want to find from this very large space, very tiny portion, those are the promising candidates that can be potentially effective for certain disease.
And the way we represent these molecules, those are small molecules. So the way we represent them, we can use two ways to represent them.
First we can write them as a chemical formula, it's just like a string. So we can write it just like a English sentence, it's like letter.
Instead we have a different letter system in this chemical language, chemical string.
So we can use the same way we are modeling natural language modeling our English sentences.
So we are modeling this English character sequences.
We can use the same algorithm to model these chemical atom sequences and we can develop the same generative AI algorithm such as this variational auto encoder and the transformer which is behind ChatGPT.
So we can use exactly the same large language model to model the molecules.
The slight difference is for molecule, it's not only sequence.
We also have graph because these atoms that they are abiding to each other, they have some bond, they form a graph, they have 3D shapes as well. They have some geometry as well.
So we want to model their geometry.
So for that purpose we need to use a second type of model called graph neural network.
So right now many models including our own models are based on a combination of transformer which is modeling the sequence and graph neural network, which is a modeling the geometry combination of the two.
We can better model this molecule sequence and their geometry then we can find from this very large space and to generate those effective molecule candidates.
My mind is blown. Let me start with that.
When you first realized that this was successful, that you were able to effectively discover new molecules by combining these two particular types of machine learning and AI and potentially reduce the amount of time.
Was this an aha moment or was this a slow burn?
You realized it was working and had potential?
Was there a moment when you knew we're really onto something here?
We got a lot of insight from earlier research, both from our own group and also from other teams like deep mite.
They have work on the AlphaGo.
It's also very similar problem.
They want to search best gaming strategy in the game of Go.
If you consider the possible play strategies in the game, it's 360.
That's also very large space.
So we are searching in the similar space in the chemical space and in their problem they can use reinforcement learning, generative AI to find effectively find the winning solution even better than the best human champion.
So that's actually the moment we got some inspiration.
We think it's possible or we do have a very large search space and a lot of the space even containing those molecules that never experiment, never testing the lab environment, we don't know yet.
We never synthesize those, but we can possibly use this algorithm, AI algorithm from this large space to effective fender path to search for the most effective molecule.
I love that you mentioned scientific collaboration effectively collaboration has been a theme across every one of these interviews that I've done and I can see how attacking the problem both from a structure of these molecules as well as a function of these molecules is what will ultimately, hopefully deliver better therapies in the end for people.
How difficult is it to build these models?
I feel like this particular field of study is relatively new in the big scheme of things. How challenging is it to build these?
It is very challenging, the most challenging part.
First we need to find the right model architecture to represent this molecule. So I mentioned a small molecule, but there are also a larger molecule we need to model as well like peptide, protein, DNA as well.
Usually for drug for disease target, target is usually a protein, so we need to model this protein, the target as well.
So the problem for small molecule drug design is given target, let's say for example Alzheimer's disease.
So Alzheimer's is a particularly difficult disease because we don't know too much about its target.
We have some suspect candidate for the targets but not confirmed yet. For those target, usually those target are protein.
We want to find some small molecule.
This molecule can be searched found from a molecule database.
This is what this pharmaceutical company are doing.
Usually they have a very large database with say 100 million molecules, but this database is only a small portion of the much larger hope possible molecule space we can search for.
Then we want to find from those large spaces, small molecule that can fit into the target.
This is called binding. For binding, usually there should be a pocket on the protein.
This pocket has certain shape, which means our small molecule should have some certain shape as well so that it can fit into that pocket so that it can bind together to behave with the ideal desired function in our human body. And if we want to develop effective molecule using AI, we need to have data to train our AI model for new disease target like Alzheimer disease or many other cancers.
We don't have much collected data from the lab environment.
So possibly we have some measured molecule and protein binding affinity like how effective they are and the strength of the binding, but we only have a few with very small data or even no data. We need to train a AI model that can effectively generate those molecule. So that's a key challenge.
That means we need to have the right model architecture to represent the molecule, represent the protein, represent their binding as well. We need to have sample efficient training algorithm so that we can work well with small data.
We want to develop models that can effectively tell the function between the small molecule and disease target tell their chemical properties and we need to effectively synthesize those molecule to tell whether a particular molecule can be synthesized or not. If it's from the data already no database, we know how to synthesize them, but if we generate a new kind of molecule that has never been synthesized before, we also need to develop and to generate the synthesize steps so that we can synthesize it in the chemical lab.
Do you think, given everything you've described across the process and that process you just described, is the new drug discovery that can take 10 years and a billion dollars from discovery to execution, do you think in your crystal ball, which is not how scientists tend to, to think that this could be a process that could almost to completion be run by AI and algorithms?
Ultimately?
I would say it's this process is it's always human, and AI interaction process like collaborative process.
Of course we need an advanced AI exam, we need computer scientists in this procedure, but we also need a domain expert, which has those chemistry pharmaceutical expert, and also lab scientists who really have the domain knowledge about the particular disease and the target and then they can use these AI tools in their daily drug development life cycle.
You mentioned the key challenge obviously in getting good data.
Again a common theme across all of the AI ML discussion that we have had here. How do you get ahead of that challenge?
How are you hoping to tackle that and grow these data sets to train these models?
Yeah, this is a great challenge.
So first we need to collaborate.
So we are calling for open science like collaborative research so that we can team up with other universities like MIT or other places and also industry, these pharmaceutical companies so that we can work collaboratively to the extent that maybe they can share some data, not really for commercial purpose, but for research purpose.
We can push forward the research and we can potentially collaborate with, for example, CDC and NIH to build a common consortium to, so each of this party will contribute a little bit of data to the data repositories so that collaboratively we can use the data to train a bigger model.
I can see that's already happening now for some domain. For example, there's a very large protein database that's kept being collected for over several decades, protein data bank, PDB data.
So we are using that to develop our protein models as well.
So it's highly effective. Yes, they are, right.
So this data is very important and I'm looking forward to collaborate with more, especially from industry partners as well.
Lei, what's the best part of your job?
Oh, the best part of my job. I think the excitement is from the moment.
I think that there are two I feel really excited about.
One is I can turn my algorithm into some real findings.
So recently we just published two papers on protein design and we are pushing this into synthesize the proteins that found using our algorithm and test that in the lab.
We're working with our partners in chemistry department of the universities, California, in Santa Barbara, and we're very close to synthesize this new protein and we are very excited about to see this algorithm actually work and can be effective in real biomedical purpose. So that's one moment in time.
I'm very excited. The other moment is I'm always enthusiastic to work with students.
I have two students graduating this year, and at the moment they defend their thesis and they become real PhD.
I feel really proud of them.
I think that's our contribution as an educator and as a researcher at university.
Is there anything I'm not asking you, there's a million things, but is there something big that I should be asking you that I didn't.
Under this small molecule, we are also working on protein design, which I also find very interesting and exciting.
In particular, recently Google, I just published their Alpha 43 paper predicting the structures for protein and molecule complex. We are working on slightly different problem.
We are working on protein design, which is given a particular function.
We want to design this protein so that it'll have this desired function. For example, it can be a protein that meets this green for recent light when it's attached to a certain biomolecule target.
Another example is designing enzymes that can be used in biochemical reactions so that it can accelerate the reaction and manufacture the biochemical materials. So again, we are using AI algorithm, generating AI algorithm to design these protein.
Our key idea here is we are designing not just their sequence, we're designing their sequence and their structure protein structures together.
So we're using this core design approach that we can design better proteins with to function.
I want to talk a little bit about the, I guess, luminescence because this would be a situation where you could in real time determine if a therapy is working, correct.
I mean that is when you're talking about cells lighting up, I've seen this in mouse models where they're trying to determine if course A or course B is effective in treatment.
Tell me a little bit more about that.
Yes, exactly.
So this can be used to tell whether a certain biological pathway is activated or deactivated so that we can identify what goes wrong and where it goes wrong.
We can identify the biological target and then we can design certain drug molecule for the purpose.
Lei. I really appreciate your time, Lei.
Thank you very much for your time and perspective.
Thank you.
You've been listening to Does Compute the podcast that examines how computer science is building useful stuff. That works. Thanks so much for joining us.
I'm Steph Stricklen.
Up Next

Reverse Engineering a Patient Monitor's Binary Protocol for Vulnerabilities
@mattbrwn
46.5K views•2025-02-10

Building Real-Time ML Pipelines with Feature Stores and MLOps Frameworks
@ODSCAI
5.1K views•2022-02-20

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence





![Machine Learning And Deep Learning - Fundamentals And Applications [Introduction Video]](https://i.ytimg.com/vi_webp/bLHqHRWUUWg/maxresdefault.webp)











































