This research demonstrates how protein language models can analyze biosynthetic gene clusters (BGCs) to predict whether two BGCs belong to the same biosynthetic class or predict chemical similarity between natural products, even with limited training data; the approach uses contrastive learning networks trained on protein sequences to compute vector representations that capture the relationship between gene cluster organization and natural product diversity.
Protein Language Models for Biosynthetic Gene Clusters
Added:hi my name is Tanya and this is the presentation of my poster on B synthetic gen clusters and protein language models many organisms such as bacteria fungi and plants produce intricate chemicals that are not needed for the graph and Inter production and th are called secondary metabolites or natural products or NPS NPS are a rich source of drugs with most antibiotics being D of NPS in a producer organism and P are synthetized by set of enzymes inced by genes that often each each other on the chromosome and the C by synthetic gen clusters despite the clinical importance that some NPS have only a small number of natural cure in is explicitly described in cases like this the world based Solutions and calistic models which don't require a lot of data work best we don't know anything except those rules it's hard to generalize it's hard to predict newes den and characteriz their products the big is publicly available as a set which has Rich annotations for B and corresponding natural products is called M the version we currently use in this work contains information of 2 and a half thousands of B synthetic TR clusters as you can see at the upset plot which describes it some of the bgcs in miik are notated with several different B classes we exclude them from the photo analysis does having 2000 of BS left for experiments the total number of samples we have is quite small but we can still use itong in there and the protein language models might help us here typically these models are trained on General tasks like mask language modeling task where the model is trained on to predict am amino acids in a protein sequence and typically the set for the training is huge this models produce Vector representation soal impedance on am Mino acid level or on a whole sequence level by by averaging the amino acids in bedance the table illustrates some examples of state of the art models and with their performance after F tuning for several down stream tasks looks nice hey for our data set of 2000 bgcs we use PR language model as a backbone of CMS Network the scheme of the typical CMS network is shown on the slide our input is a pair of protein sequences are corresponding to a pair of bgcs the the network computes a pair of for the two sequences and comput C similarity between them we can f tune the network to mph the impeding space to make the cosign similarity between edings close to some user defined metric one of the cool features of TM network uh is that they capable to provide nice results even in cases where the join data is very limited that's exactly our case uh in in this work we'll try to solve two different tasks with Sam Networks let's take a closer look in the first task we try to predict if the two bgcs belong to the same B synthetic class or not try to solve this task with two different data set split schemes the one simply tries to keep the percentage of samples on in which class similar in all splits and another simply considers one of the seven best entic classes to be in our test data set and uh uses for validation everything else first we wanted to make sure that our model actually on something the first PCA plot shows the meing for bgcs which were computed by pre-rain model with no fing after that we trained the model in two different settings uh here you can see how our light and space have changed note that the both pictures looked similar despite the fact that the model which produced the Bings in the further picture below didn't know anything about NP class we also tried to look at the cosign similarity distribution the model which uh we used to comp similarities for the that plot was trained with not information about re yet it can separate them from their classes embedding on R plot computed by the model which had no information about NP class it doesn't separate NP fromti which is okay because these two classes are similar in terms of domain organization to evaluate the results of our experiment we computed the S scores of f models as a baseline to compare we took the big skate distance Matrix we filtered the bgcs to be is both B results run and in our model test data sets what can we say about this experiment first the find workor and improv by synthetic class separation the F model perform well for the classes they haven't seen and might be used for the NOA prediction as a preliminary work we try to compute ano similarities between all pairs of fingerprints belonging to the pair of natural products produced by BJC within the same B synthetic class in the plot you can see that theot similar between NPS differ within each class this means that the data set in the terms of NPS is quite diverse in addition the B synthetic classes themselves have the similar Dom oranization we decided to simplify the task by focusing B belong to the same to the only one class inp as a result we are dealing with an extremely low resource data set here 228 samples in total despite that after contining our model learned something that could predict anot similarity between the natural products given the sequences of the corresponding pgcs in this task we also use the same bcap distance Matrix as a Baseline and compute the correlation coefficient between the metric predicted by our model and the correspondent distance the negative signs are fine since we compute the similarity the metric here when the similarity is high the distance is slow and VI Versa in this experiment the model was capable to predict chemical similarity within one in of class uh we plan to try to use other multimodal data as attached metric in the future thank you for your attention future project updates will appear at our GitHub rle
Up Next

The Future of Egg Production: Precision Fermentation with Arturo Elizondo
@TheProofWithSimonHill
1.5K views•2022-08-01

Algae Biofuels: Harnessing Microalgae for Renewable Energy
@LosAlamosNationalLab
623 views•2020-12-03

Microbial Degradation of Plastics: Biodegradation Pathways & Sustainability
@majeedhammad
2.9K views•2021-04-11

CRISPR and Genetic Engineering: How Gene Editing Works and Why It Matters
@kurzgesagt
30.5M views•2016-08-10
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biotechnology






































