Directed enzyme evolution is a powerful protein engineering technique that uses random mutagenesis combined with high-throughput screening to improve enzyme properties, particularly when rational design fails due to incomplete understanding of enzyme structure-function relationships; this method involves generating diverse mutant libraries through error-prone PCR or saturation mutagenesis, then screening thousands to millions of variants to identify improved enzymes, as demonstrated by achieving 900-fold activity increases in aryl malonate decarboxylase through iterative rounds of mutagenesis and selection.
Directed Enzyme Evolution: Biocatalysis Explained
Added:Welcome to the MOOC “Tailored materials and enzymes for industrial processes”. My name is Robert Kourist and this is the fifth unit of the basic module and deals with directed enzyme evolution. So the previous unit I explained that the catalytic properties of enzymes are stored in the primary sequences and we want to improve and we use rational protein design or directed evolution in order to improve these properties for applications. And in rational protein design we can bring in mutations at selected sites that for instance we identified based on structure elucidation or molecular modeling and then we predict that the effect of a certain amino acid exchange we incorporate this mutation alter the structure and then alter the function.
The problem is that it is exceedingly difficult to make an accurate prediction of a change of an amino acid sequence, but in some cases and I explained one the last unit this works reasonably well. The case was the one of the aryl malonate decarboxylase which produces optical pure alpha substituted carboxylic acids. In this case the enzyme decarboxylates malonates, it cleaves carbon dioxide and in the wild type the enzyme produces the R- enantiomer by protonation from one side and exchange of this Cysteine into a serine, which does not protonate anymore and the introduction of the cysteine at the opposite side of the active site completely inverted the enantioselectivity.
This is a classical case of rational design, because a rational prediction was introduced was, investigated experiment and could be confirmed. Unfortunately, the activity of the resulting mutant was here G74C/C188S was by 18,000 fold reduced. So somehow this cysteine or maybe the serine or maybe both they interfered with activity and at this stage the researchers had no clue, how they could recover the activity of this mutant and therefore they went for directed evolution. In rational protein design we only need to introduce very few mutations and this can be done in a very controlled way, it's very simple. In direct Evolution the generation of the mutants means it influences what kind of amino acids we can access and how they're distributed. So this is a crucial element of the experiment and there are basically two approaches. One is purchase chemically synthesized libraries, the other is to do it yourself in the lab. Chemical synthesis is very expensive, cost several thousands of euros but of course it's not work intensive and mutagenesis in the lab can be done with very simple ingredients but it needs a skilled coworker. So there's personal cost and then the decision is just a question of priorities.
However, also and chemically synthesized like libraries and PCR generated libraries are a bit different. So it's not exactly the same, so it's a mixture of priorities and also what kind of libraries we want to have. In mutagenesis we have PCR based methods and we have other methods and for the sake of simplicity I will focus here on polymerase Chain Reaction based mutagenesis, because this is now the most frequently used method. There are others there are chemical reagents that induce mutations and we don't like to work with them, because they're cancerogenic, or there’s UV mutagenesis could be also used and there are mutator strains that are deficient of repair systems, so that each cell doubling they would cause a certain number of mutations, also a bit difficult to control. So I would think I think polymerase chain a reaction is the most frequently used method right now to generate libraries for directed evolution. Error prone PCR was one of the very first methods used for mutagenesis for directed evolution. It is quite a simple method and can be done in every lab that has the ingredients and machines for a regular PCR. So in the regular PCR a double strand the DNA is our template, then there are specific primers that bind and the polymerase amplifies the two strands and then the product of the PCR serves as a template for the next cycle. So the product of the first cycle is also the template of the next so in every step in every cycle the concentration of the DNA doubles and because of this, this is called a chain reaction. So the polymerase has a very high accuracy, so it has an error rate of one in 100,000 or one in a million and then polymerases have usually a second unit at attached to them which slide along the DNA and check the double strand for errors and this also has an error rate of one and a thousand. So this gives us totally an error rate in the billions. With this high accuracy it would be difficult to insert mutations. However, there are ways to increase the error rate of a polymerase and this is done in so-called error prone PCR and these are quite easy quite simple ways. So we can use a polymerase that does not have the proof reading capacity so this already increases the error rate a lot and we can add manganese instead of magnesium so both are bivalent ions and magnesium is in the active site of the polymerase to coordinate the nucleotides and if there is manganese, somehow the enzyme has a higher error rate. We can also spike a nucleotide concentration this means if we have more for instance of dCTP, then the likelihood that dCTP comes into the active site of the enzyme is higher leading also to a higher error rate and if we combine all this this can increase the error rate to 1 to 2%.
So in this case we would still have our template, primers bind and then the amplification causes errors which I have here denoted with these red crosses and these errors would of course be then again in the next cycle as a template and the polymerase then if that amplifies conserves these mistakes but also introduces new errors in every cycle. So with the number of cycle and the conditions of this we can also adjust the mutation frequency and this is a widely used method. It has some limitation and usually it is done for the generation of very large libraries between 50,000 and typically 100,000 clones. Sometimes we have already concrete idea for an amino acid position and we want to insert all 20 amino acids into one position and in the last unit I presented site directed mutagenesis where a certain nucleotide exchange is introduced by a mutagenic primer, which is a primer which has one different nucleotide or several different nucleotides, but still binds because the large part of the primer is still complementary. For saturation mutagenesis we can use primers with a so-called N and N stands for four bases G A T and C and so a primer which is possess it has an N is actually a mixture of four primers 25% of each base and if this is then Incorporated in the amplification and the template is degraded we have now the N in our Gene which means we have a mixture of genes with each 25% of different base and by using one or several bases with an N we can put in every of the 64 codons, which means we can put in every amino acid at a certain position provided of course we have a certain idea which of our 200 or 300 positions to set to mutate. The last component for a directed evolution experiment is the high throughput screen. So here we need a simple way to measure a reaction in a microtiter plate and the activity of the the AMDase can be determined by the quantification of substrate and product for instance in a high HPLC. But this would mean for each mutant we need to make at least one HPLC measurement which would be a large number of HPLC measurements already, but in this case we can if we decarboxylate a malonate and produce a mono acid we have a pH shift. And because here we have two carboxylic groups only one so the medium gets more basic and then we can use a buffer with a low buffer capacity and you see this here. In these wells, a reaction took place and the shift to the basic a pH was visualized with addition of a pH indicator and here in these wells no reaction took place. Then there was the wild type enzyme was added at certain points to show to see for of all these four measurements all the six measurements have the same rate and then we get a So-called mutant landscape. So this would be here inactive variants or lower active variants, this will be here our wild type. You see we have quite of a spread so this essay is not so accurate and then we might have some where we have a higher activity so with such as green. We can now efficiently screen thousand of variant 100 per plate and there also microplates that have 364 wells which even give us many more clones to screen. So coming back to the experiment so the mutant G74C C188S has the cysteine on the other side is inverse but has lower activity and the determinants of the activity were unknown but it is known that the enzyme always cleaves the same carboxylic group, which is this one which one pointing out towards us and this is accommodated in a hydrophobic pocket, which results in unfavorable interactions and this leads to the cleavage of carbon dioxide.
So it was thought that a variation of this hydrophobic pocket might change the activity and these are here residues in this active site. So we have a methionine, we have a valine 156, we have valine 43 we have a tyrosine 48, a leucine 40 and a leucine 77. Here is the catalytic cysteine. This should not be altered of course and here we have the serine which is the cysteine in the wild type. They constitute this hydrophobic pocket and now by saturation mutagenesis each of these residues was saturated so first of all this serine which is now not anymore needed was altered and it was shown that the glycine in this position gives a five times higher activity and this was then used as a template and next here different residues were saturated and here for instance if leucine 77 is is altered methionine was the most active variant and interestingly here the substitution of methionine 159 by leucine led to a 100 fold almost 100 fold activity increase and then when this was used as a template again the hit Y for the F led to a 900 fold high activity so by one to three cycles of mutagenesis the activity could be increased 900 fold. So this is a typical experiment of a randomization protein engineer. So the the effect is striking we have a 900 fold activity increase but it is very difficult to understand why substitution of M959 with a leucine which is not much bigger leads to such a striking activity increase. One thing we can see here, we can see here that all the hits all the improved variants have hydrophobic amino acids, which makes a lot of sense because all of these amino acids are hydrophobic. So there was no arginine, lysine, serine or glutamate found and actually in hindsight it is also possible to do this experiment now with the restriction of the amino acids put and putting in only hydrophobic amino acids. So there are lot of degrees of freedom but it is also not known if this here is the best variant of the enzyme, but directed evolution brings produces improved variants without an accurate understanding why here for instance these changes of these hydrophobic amino acids lead to this activity increase. This makes it such a powerful method. So if you compare the rational design and the randomization we have several important factors. One factor is how well do we understand our enzyme. If we have a very good hypothesis making a very good prediction one mutant can be sufficient and indeed in the AMDase example a double mutant inverted selectivity completely. If we have very good knowledge then we need to screen only few. If we have not so much knowledge understand our enzyme rather poorly then we need a tremendous screening effort and this is an equation of experimental capacities. So nowadays we have a lot of ways to get structures without experimental structure information. In many cases we can make we have similar enzymes proteins that have been crystallized so we can make homology models and we can also make very plausible de novo predictions and also often we know something about the mechanism even if it's a very new enzyme. On the other hand if it's an interesting enzyme often it's quite new so we do not have so much experience. So we are often here we have some knowledge but not enough knowledge to make accurate predictions. This is a typical example of the hydrophobic pocket so the researchers doing this Miyamoto and Ohta they identified this hydrophobic pocket as activity driver but they could not pinpoint and predict mutations so they could use this for randomization. We have a similar situation with a screening effort. It is quite easy to screen a few thousand variants. We can do this with an HPLC, with a pH screen so this is quite simple.
To screen millions or even billions this would mean we need a selection assay or fluorescent assay and they are available for only few enzymes. So therefore the quite usual situation is that we have some idea for instance we can predict the active side residues and we have a medium through screen and this means we can screen 100 to 200 variants or maybe 5,000 variants and this is now a way how enzyme engineering is done often nowadays. So with this I introduced you to a few basic concepts of molecular biotechnology, biocatalysis so the way to produce enzymes and optimize them by rational protein design and directed enzyme evolution and the following the following units will then discuss how enzymes can be improved for the use of multi cascade reactions by immobilization to carriers and with this I would like to thank you for your attention.
Up Next

Directed Evolution of Enzymes: Golden Age Insights
@iqf_csic
129 views•2024-06-20

Quantitative Real-Time PCR (qPCR): Principle, Method & Data Analysis
@animatedbiologywitharpan
129.2K views•2023-09-06

Microbial Degradation of Plastics: Biodegradation Pathways & Sustainability
@majeedhammad
2.9K views•2021-04-11

CRISPR and Genetic Engineering: How Gene Editing Works and Why It Matters
@kurzgesagt
30.5M views•2016-08-10
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biotechnology










