Recombinant protein purification involves a systematic workflow: first, selecting an appropriate host organism (bacteria, insect cells, or mammalian cells) based on protein type, required posttranslational modifications, yield needs, and available resources; second, designing expression vectors with suitable promoters (constitutive or inducible) and affinity tags (such as His-tags or GST fusion partners) for purification; third, optimizing expression conditions including induction timing and temperature; fourth, lysing cells using methods ranging from gentle (osmotic shock) to vigorous (ultrasonication); fifth, applying chromatographic techniques including affinity chromatography (for initial purification), ion-exchange chromatography (for charge-based separation), and size-exclusion chromatography (for size-based separation); and finally, verifying protein concentration through absorbance measurements or densitometry and confirming purity and identity via SDS-PAGE, immunoblotting, or mass spectrometry.
How to Purify Recombinant Proteins: Methods and Techniques
Added:[Carrie Iwema]: I'd like to welcome you all to our How-To Talks by Post Docs series.
Today, we have a new doctor, Prerna Grover, who just completed her Ph.D. this past spring in integrated molecular biology with Dr. Tom Smithgall.
And today, she will be talking about how to purify recombinant proteins.
[Prerna Grover]: Thank you, Carrie, for the introduction, and thanks for organizing this wonderful forum for those post-docs.
So, welcome, everyone. Thanks for coming. I'll be talking to you today about how to purify recombinant proteins. Before we get started, just by a show of hands, how many of you are familiar with recombinant protein expression or production at all?
Okay. So, um, to get started, the first question is, why do we care?
Why do we want to purify recombinant proteins at all? And it can have several functions.
So, to examine the biological activity of new proteins that haven't been characterized, or to get their enzymatic properties, specifically for enzymes. It's also helpful to analyze the structure adopted by the proteins, so proteins can be crystallized.
And you can study the structure-function relationships of recombinant proteins. There are also application-related needs where you can actually raise specific antibodies towards these proteins by using them as antigens.
Or you can develop grants to specifically target a protein, which is what I did for my thesis project.
And again, of course, produce therapeutic proteins such as vaccines and, you know, peptides that are being currently tested in the clinic for diseases.
So, out - to outline my talk, this talk will be divided into three parts.
So the - in the first part, we will be talking about the host selection and purification strategy. We will then proceed to purification techniques that can be used.
And finally, I'll talk about how you can analyze the concentration and purity of protein you've purified.
So, to go into a little bit more detail, whenn we talk about in the first part, we'll first see what - how to identify target protein where, like, what boundaries should be used for that. Determine the amino acid sequence.
Decide on an ideal host organism to express this protein. Prepare expression vector constructs and eventually, test and optimize the expression conditions. In the second part of the talk, I'll be talking about how you can decide on a purification strategy and then purify these proteins and then in the final part, we'll talk about how you can determine the concentration and then confirm that it's the correct protein and look at its security. So, as with every research project, I'm sure you're familiar.
The first step is, you know, a literature search where you basically need, you know, your protein.
So the first thing to look for would be has a protein ever been expressed in any host organism?
So if that protein has been expressed, you can look at the specific part of the sequence that's being used, and it's also useful to check for any specific mutations that have been introduced in earlier studies.
The reason I specifically mention this is because sometimes mutating or deleting a single residue can enhance the solubility of the protein significantly. In case the protein has not been expressed, ever, you're working with a totally novel protein, then it's helpful to do a Blast search to actually look for homologues, if the protein from any other - let's say if you want to express human neighbors, so has the human - has the able protein from Drosophila ever been expressed?
And you can also do a protein sequence alignment to where you can identify what boundaries should you use.
So, if you just want to identify two domains, it's hard to, you know, just ascertain that from the sequence, but if you align multiple protein sequences, you can use that to identify domain boundaries.
It's also helpful to see what posttranslational modifications may be required and I'll be talking about that quite a lot in the later slides as well.
And if that - if the protein has been expressed in the past, then what is the most common host that has been used for expression?
So, as we try to identify the host for protein expression, one big question is the protein prokaryotic or eukaryotic. It's much easier to express proteins in bacteria.
So, if it's prokaryotic, then you're in luck. But, a lot of eukaryotic proteins are also expressed in bacteria.
So, that's one question.
And then, what is the normal localization of the protein, whether it's cytoplasmic, meaning whether it's secreted or targeted to specific organelle, because these will decide what are the posttranslational modifications that are introduced into the protein.
And again, it can be important tool as you are trying to pick a host. So, what downstream applications are you looking for?
If you are looking for, you know, analytical, functional, biological assays, you may only need a small amount of protein in the nanograms or micrograms range.
However, to produce antibodies, you need more protein and structural biology needs a lot of protein in the milligram scale.
So, again, different hosts are able to simp- like, the yield that you get from different host varieties, so, if you know how much protein you need, that can be another helpful factor.
Then, another question is how important are posttranslational modifications for folding our activity.
Because bacteria usually don't introduce any posttranslational modifications, whereas if you are expressing a human protein for which posttranslational modifications are very important, you would prefer to choose mammalian cells.
And finally, what resources are available for protein expression, in terms of vectors and expression systems?
So, if you work in a bacteria lab that does not have access to tissue culture, then it would be hard for you to work with insect cells or mammalian cells and they're also more expensive to work with.
So, you know, just factoring in all of these things. So, I'd like you to take maybe a minute and, you know, in an ideal world, based on your research, pick a favorite protein that you would want to purify and then answer these three questions if you know the protein is prokaryotic or eukaryotic. Do you know if there are post-translational modifications that are required for folding or activity?
In an ideal world, if you get the protein, what experiments would you want to do with this protein? Wait for a minute.
Is everyone ready? So now together, we'll answer the question, what is the best tools to express this protein.
Does anyone want to volunteer with their protein information? Or we can just - OK. OK. So, there are several hosts that can be used for protein expression and we'll be talking about three of those here, bacteria, and E.coli's the predominant bacteria that's used. Insect cells and mammalian cells.
Other than this, there is yeast. There are Drosophila cells. There are experimental sponges.
There's a host of organisms that you can use and optimize for expression. But we'll focus on these three and compare their characteristics.
So in bacteria, the protein can be either cytoplasmic or it can be exported to the periplasmic space or secreted out into the media.
In insect and mammalian cells, the protein can be either intracellular or secreted.
If you are expressing - so, if you want the protein to be secreted, you can actually add a sequence at the end of the protein that will target it for secretion.
That secretion requires posttranslational modifications that may affect your protein if that's - it's not, it's not normally secreted. The expression from bacteria, like the relative process, is much quicker.
It takes only a few days where you clone it out and then express it, whereas it takes longer time to express proteins and insects, as well as mammalian cells. The cost of expression is cheap for bacteria because you essentially need just flasks and your media is really cheap, whereas if you're culturing insect cells or mammalian cells, they require tissue culture conditions. They required their own incubator and the media is high.
You need tissue culture equipment. So, the relative cost of expression is high for these two systems.
The protein yield is high for bacteria. You still get a moderate amount of protein from insect cells, but mammalian cells, the protein yield is low.
And, of course, there can be exceptions to all these rules where, you know, protein may work fantastically in mammalian cells and you might get a ton, but it may be insoluble in bacteria or something.
And as I mentioned earlier, posttranslational modifications, so, for example, if your protein myristoylated naturally and it's important for its activity or function, then bacteria may not be the way to go because bacteria is not going to myristoylate your protein.
Whereas in insect cells, some posttranslational modifications I've heard, not all of them, but if you are trying to express a human mammalian protein, then you would get most accurate posttranslational modifications in mammalian cells.
I'd like to point out there now that now there are vectors where you can express two proteins from the same vector.
So, you could actually encode an enzyme that, you know, let's say, that adds the myrisitc acid moiety and express your protein on the same vector and then you could get, like, some posttranslational modifications.
There's a lot of recent studies which are, you know, playing around with the conditions for that now. So, to summarize for protein expression, these are like the various factors to compare hosts.
In terms of speed, mammalian cells are the slowest. Insect cells are in the middle and bacteria are the fastest. Bacterial cell are the cheapest to culture.
Insects are again in the middle and mammalian cells are most expensive. The typical yield is highest for bacteria, lowest for mammalian cells, and insects are again in the middle. For posttranslational modifications, mammalian cells are best for human proteins.
And if you're trying to produce a protein that is going to be used for therapeutic purposes, then if you are making that in mammalian cells, it would be the most likely to get approval.
So, moving on to expression vectors, again, there's a host of expression that are there, but I'm just providing you examples of bacteria, insect and mammalian expression vectors.
So, there are some important features for expression vectors. One of them is, of course, the origin of replication and antibiotic resistance, so you can propagate the vector in bacterial cells.
For expression, these proteins would have a multiple clone inside where you would clone your gene in, which is indicated by this MCS here.
And it would have a promoter that precedes the MCS and a terminator that follows the MCS.
So, in buc- this bacteria right now, there is T7 promotor, which is an unusable system and I'll be talking about that. In insect cells, there is a polyhedrin promoter, which is right here, which is a constitutive promoter, and in mammalian cells, there is a CMV promoter, which is again a constitutive promoter.
So, the promoter may be constitutive, which means that it's always expressed or invisible, which means that you can selectively turn it on. Some proteins may be toxic to cells or sometimes they may be expressed at such a high level that you want to limit the amount that they are expressed.
And that's why it's useful to have inducible promoters.
Also, you can have weak and strong promoters that would give high and low expression levels.
So again, if the protein is toxic, then maybe it's easier to express it at a lower level and so you would use a weaker promoter for that. You can - some vectors also have shot peptide tags that may be vector encoded, so they might be included in the vector or you can even introduce them into the prime of. Some vectors, also from the fusion partners, which are larger proteins, such as GST, that, again, maybe vector encoded.
Additionally, if there is a fusion partner or even a shot tag, there may be a protease cleavage side in between your protein and this tag so that you can actually clean this off.
So, to talk a little bit about protein diets. Why do we care about them?
So, one factor is that they aid in direction, during expression and purification because, let's say, if you have a stretch of six, histidine residues, which is known as a histag.
Then, you can easily probe for that histag and you can direct your protein. And they also enhance the ease of purification, because then you can have an antibody, or, like a resin, that's modified to basically attach to this tag and it's only going to bind your target protein because only your target protein is going to have that tag.
And that's the principle of affinity chromatography that I'll be talking about later.
So, as I mentioned, there are two types of tags: peptide and fusion partner tags.
These are some examples. So, peptide tags are usually small.
They're between 6 and 10 amino acid residues.
So, it would be a stretch of histidine residues, or Arg, or FLAG, or c-Myc, whereas fusion partner tags are large size.
So, these are essentially proteins, so Glutathione S-transferase, Maltose-binding protein, or Thioredoxin.
The peptide tags don't affect the solubility of the protein, mostly, however, fusion partner tag, some of them may enhance solubility.
So, for example, MBP is known to enhance the solubility of protein.
So, if you're struggling with an insoluble protein, it might be useful to add a tag to it.
The peptide tags are small, so they may not need to be removed. They may not interfere with biochemical activity or structure of your protein, so you can, like, leave them in. Whereas if you have a fusion partner tag, it's, you know, it's a full sized floating protein, so it's probably going to be removed for most applications. It might even interfere with your protein activity.
For example, um, GST is a dimer.
So, if you have GST fused to your protein, it's going to dimerize and, you know, it's not going to behave the way it would on its own. So, you'll have to remove that. Tags can be cleaved either chemically or by enzymes, just thrombin and, um, TEV. And these enzymes have been optimized for purification conditions that they are able to cleave these guys.
And there's a bunch of factors that play into selecting an enzyme that you would use.
An important point to note here is that sometimes, these tags can actually, once you remove them, they produce solubility.
Not all of them do, but sometimes that happens. So, that's another thing that goes into the decision.
If the solubility is going to go back down, is it worth even, you know, going through the extra steps?
And if you do you cleave by enzymes, then you will have to add an additional step because you have to remove the tag, as well as the enzyme you used from the protein you've purified.
So, I'm going to talk about sample expression strategy, so we'll start with bacterial cells.
So, we talked about hosts and we talked about vectors. Now, let's say you have an expression vector with your favorite gene.
Anyone to transform, you know, you first will transform equalized cells, a strain that has been specifically engineered for expression of protein.
You will grow these cells to a mid-log phase, so you have enough amount of cells and they're metabolically active.
Then, you would induce protein expression using - I'm going to give an example of IPTG here - for a certain amount of time, which could be between four hours to, you know, twelve hours or longer.
And finally, you will harvest the media or the cells depending on whether your protein is secreted or intracellular.
So, I talked about briefly about the T7 expression system earlier, so I'd like to explain how that works. This is basically based on the lac operon.
So, I'm not going to go into details of the lac operon, but essentially how it works is that LacI is a protein. So, lac operon essentially senses the presence of sugars in the media and in the presence of lactose, these proteins get activated. The lac operon gets on and genes are synthesized that are useful for utilizing lactose, whereas in the absence of lactose, this LacI protein, which is a repressor, represses the lac operon and there is no expression of downstream genes.
Now, this is a strain of E. coli that's called BA21 and it's been optimized for protein expression. So, it has a mutated promotor upstream of the LacI, so that the LacI protein is always expressed and it's expressed at high levels. Then, under the regulation of the lac operator, it has a T7 RNA polymerase.
So T7 is essentially a bacteriophage that infects E. coli cells and it has its own RNA polymerase.
So, this T7 RNA polymerase, G, is encoded downstream of the lac operator and then in our expression vector, we have our target gene, that's encoded under the control of the LAC operator as well as the T7 promotor.
So, what happens is in a wildtype cell, when there's no lactose present, a large amount of LacI, which is this repressor protein is expressed. This repressor protein binds at the lac operator and there is no expression of the T7 RNA polymerase.
So essentially, there's no RNA polymerase in the cell.
And then, because there is no RNA polymerase, it can't bind T-cell and amino target gene is not expressed until it's induced.
When you induce with lactose or IPTG, which is a non hydrolyzed stable form of lactose.. This IPTG molecule can bind this LacI green protein and it causes a conformational change, so that another LacI protein is unable to bind to the stack operater.
As a result, the T7 RNA polymerase is expressed. This T7 RNA polymerase can then bind to the T7 promoter on your expression vector and cause expresision of target gene.
So, does anyone have questions about that? So, um, this is a sample expression strategy for bacteria. Moving on, if you are - if you want to express your protein in insect cells, in that case, before you - after you have the expression vector, you first need to generate a baculovirus and I'm going to talk about - details about that in the next slide.
But once you generate a high titer of baculovirus, you can plate your insect cells. You can infect these cells with the baculovirus and you can harvest the media after 72 hours.
So, how the system works is basically baculoviruses are viruses that infect insect cells.
And they - so we, - so we encode our foreign gene, which is, you know, our target gene in a donut plasmid that's recombinant and it's encoded under the polyhedrin promoter.
Now we need to - so this is, the baculovirus has a huge DNA and we need to include this gene in the bac - which DNA, which is the baculovirus DNA.
So, for that, we - there's this system that has been developed by Nitrogen where you have on complement DH10 bac E. coli cells that have this DNA and helper virus and you can transpose your donor DNA into this virus.
So, you essentially, um, transformed your plasmid into these E. coli cells and you can pick four.
You can pick four bacteria. Their transposition has taken place by the simple blue white selection.
You can then do a mini prep of the molecule weight DNA and you can confirm that you have your protein by PCR.
Once you have your recombinant DNA, which is essentially the baculoviral DNA, you can transfect that into insect cells and that will result in the recombinant baculovirus virus particles.
So, once you have this recombinant baculovirus particles, they're usually relatively low titer to begin with.
But, you can serially infect insect cells and you can get a high titer baculovirus particle that you can use for protein expression.
So, this was the insect cell system. And then for mammalian cells, you can essentially just plate the cells.
Transfect them with your expression vector. Host, we - now, we transiently tranfect.
We can then - you can have either constitutive protein expression or inducible.
And if it's inducible, then you can add your inducer. And finally, after about 72 hours, you can harvest the media or the cells.
So, for most of the talk, I'm going to assume that we are working with intracellular proteins.
And after you have these intracellular proteins expressed in your cells, one question is, how do you lyse these cells? So there are several methods that are based on how stringent - sorry, how harsh they are. So, there are gentle extraction processes such as osmotic shock where the product yield is low, but there's also a lower proteases released. Then, there's enzymatic digestion with lysozyme, which can usually be done only on a small scale and may still need to be combined with mechanical disruption. And mechanical disruption, you know, under these moderate extraction processes where you can grind either with abrasives or you can actually freeze, thaw, for several cycles and ice crystals are formed, which leads to lysis of the cells.
Finally, there are some vigorous extraction processes, such as ultrasonication, which uses sound waves, high frequency sound waves to essentially lyse the cells open.
And this may result in the release of nucleic acids, which may cause viscosity problems.
So, you can use DNAse and also, it's helpful to use protease inhibitors for all of them, so that as soon as you lyse the cell open, it's not that, you know, the protease just use up your protein.
And then, um, these three, Microfluidizer, French Press, and the Manton-Gaulin Homogenizer, these usually use pressure to lyse the cells open.
So, once you have the small scale expression done, I'm just going to show you a sample, because once you have everything ready, you would first want to test these on a small scale, whether it's working or not, before you express protein on a large scale.
So, the example I'm going to show you is it's an Abl protein.
It's the recombinant - regulatory domains of Abl that have a histidine tag, so six histidine residues, upstream of this Abl protein and they're expressed, they're cloned into this pET21a vector which neutralizes the T7 proter that I described earlier.
And these are expressed E. coli Rosetta cells, so Rosetta cells basically are a modified version of the Beta 21 cell.
So, they have the T7 expression system engineered into them and they also have some rare tRNA codons for tRNAs that are rare in E. coli, but commonly used in humans.
I induce these cells with 0.4 mM IPTG for four hours at twenty five degrees Celsius. So, there are gonna be four fractions that I'm going to show you.
One is uninduced cell, which is, but, total. Then, induced cell - induced cell fraction of total cells and then from the induced samples, we separate them into the pellet and soluble. So, this is Coomassie Stained Gel.
You can basically run these fractions on a gel and you can stain with Coomassie, and then when you destain, you can look at all the proteins that are expressed in the cell.
Hopefully, you'll be able to see bands that are relatively stronger and there's a difference between uninduced and induced, which is what you see here.
And then, this is the pellet. So this would be protein that's included in the inclusion bodies and that would be the insoluble fraction. And finally, this is the soluble fraction, which is present in the superheated.
So, I was lucky that I got some soluble protein.
If you are not totally sure, you can also immunoblot. In this case, because I had alpha-his, I could immunoblot with an alpa-his antibody and you can confirm that the protein you think is your target protein and it has the his-tag attached.
So, worked for me, but it doesn't work for all proteins and I have come across conditions when it doesn't work and I have insoluble protein. So, how do you work around that? So, there are conditions that improve solubility.
So, one of the things that leads to in solubility sometimes is that the bacteria is producing the protein at a very high rate and the protein folding machinery cannot keep up.
So, it ends up - the proteins end up getting aggregated and misfolded and they get incorporated into inclusion bodies.
So, one thing that you can do is actually lower the rate of synthesis.
You can try a lower induction temperatures. If you do that, then the protein would be synthesized relatively slowly and you might get more solubility. Or you can try alternative E. coli strains that are again, engineered to express proteins relatively slowly.
Another alternative is to try other growth and expression medias.
So, some of them modify how sugar is utilized in the system, and again, the protein is expressed at a slower rate. And I tried that for one of my systems and it worked.
So, Enpresso is like one of these medias that's available commercially.
In case none of these factors, you can also purify protein from inclusion bodies.
It's more complicated because you need to use denaturing conditions.
So, you need to use these harsh solvents to actually dissolve protein that's precipitated.
And once you do that, you can purify. And then, you can post-purification refold the protein, which can be done on or off the column.
So, one of the common starting points here is on his-purification because it can stand these denaturing conditions.
Other chromatography techniques may not be able to do that.
So, this was the first part of the talk where we talked about data mining getting us sequence for expression, deciding on a host organism, and preparing the expression vector construct and then testing and optimizing conditions. Does anyone have questions so far?
So, moving onto purification techniques. So, again, assuming that we have our protein in the lysate, you first need to - after you lyse open the cells, you need to spin them down and get rid of all the cell debris, so you only have a clarified lysate sample.
And then, based on the - based on the protein properties, you can choose the number and type of chromatographic techniques that you can use.
So, if you have a bad protein, you can use affinity chromatography straightaway.
And then, there are other chromatography techniques, such as ion-exchange, size-exclusion, reverse-phase, etc. So, you can use - once you use one method of chromatography, we need to ask the question whether the protein you gain from that is pure.
So, if it is, one way to test that is to basically run them on a SDS phase gel and then stain with Coomassie. If you will see only one prominent band, that means that your protein is pure and you can confirm the identity of the protein under the right concentration.
If it's not pure, then you need to go back and you need to use an additional chromatographic technique to separate out the contaminants from your protein.
So to explain the concept of these chromatography techniques a little bit, the first one that I'm going to talk about is affinity chromatography.
So, it's based on specific binding to tags, such as His, GST, FLAG, etc. I talked about these earlier.
So, you can actually obtain resins that are specifically modified to be able to bind to these tags.
So, um, shown here is an example. So, the sphere in beige is actually a polymer bead that's attached to your ligand which is in black.
And your protein of interest is this red crescent shape. So, you have this mixture of proteins that you load onto the column and the column has these protein - the polymer beads that are bound to your specific ligand.
So, only your protein of interest is going to be able to bind to these black ligands, while the rest of them are going to get washed out, as you flow buffer through the column.
Once you've washed out all the contaminants, you can actually elute your target protein in different ways.
One of them is shown here where you can just add a solution of the ligand at high concentration.
So, that's going to elute out your protein bound to the ligand, but you have to find another way to separate your ligand from this protein. Other ways of eluting the target are using the competing molecule that binds to the ligand.
So let's say you have another molecule that binds to the ligand and then your protein is going to competed off from the ligand and you'll get your protein.
Or you can also use chemical conditions that disrupt the ligand-protein binding.
So, that was shown on the more, you know, explanatory format, but there's automated purification systems that you can use. So, one example is this.
This is available from GE Health Care. It's the ACTA Explorer and there's a bunch of different quantifications and versions of these ACTA purification systems that are available.
So, this here column that you see is actually a purification column and it's a size-exclusion column that I was talking about earlier. And you can actually obtain columns that can be used for His- purification and other purifications in one or five mL sizes.
And this is connected to a fraction collector that collects samples in a ninety six well plate.
So, in the previous version, you had to manually collect samples and manually add the buffers.
But, this is an automated system that's run by a software.
You can collect one to two mL fractions in these ninety six well plates and the software generates this purification curve for you.
So, this is a purification probe from His- purification and in blue is the absorbance at 280 nanometers.
So, aromatic residues absorb UV light at 280 nanometers and there is an increase in the absorbance, which is correlated with concentration.
So, for example, when you load mixture of these proteins in the beginning, there is an increase in the absorbance and as the loading finishes, it comes down. As you're washing, you see this baseline absorbance.
And then, so green actually shows the buffer conditions. I should've mentioned that.
So, here, you're changing the buffer and you add your elution buffer.
And so this protein comes off the column and you'll see a peak that correlates the high protein concentration again.
So, once you have this curve, you can use that and you can run fractions that correspond to this peak on an SDS page gel.
So, here's a sample Coomassie stain gel. This is the preload, which is essentially what you're loading here.
So, you can see that your protein of interest is expressed at high levels, but there's also a lot of other proteins. Um FT 1 and FT 2 are essentially the flow-throughs which are when loading.
So, in case your protein doesn't bind to the column at all, it's going to come off on the flow-through.
And then, this is your - so E4-F4 is this part where you are - the peak where you are eluting your protein of interest and you can see that there are some contaminating bands in E4 and E6.
So, as you are pooling these fractions, you can maybe pool starting from E8 to F4, or maybe even E7-F4, and your protein is going to be pretty pure for most applications going from here.
If you don't get protein, there's, you know, several contaminating bands that you need to separate out.
You can add on chromatographic techniques to this. So, another chromatic technique that you can use is ion-exchange.
So, this is based on the charge of the protein. So, every protein has an isoelectric point where it's neutral and which is related to the pH, sorry, which is related to its amino acid content. So, once you have - know that the ion of the protein and you have a buffer, you can predict whether your protein is going to be positively or negatively charged at that pH.
So, to understand the concept of ion- exchange chromatography, these grey beads here are polymer beads that are negatively charged, as you can see by the minus sign here, and then you have proteins that are either positively or negatively charged.
So, proteins that are positively charged at the specific pH are going to be able to bind the negatively charged polymer beads whereas the negatively charged proteins are going to flow through.
So, you can have a, of course, an ion- exchange where your bead is positively charged and your proteins - and you are accumulating negatively charged proteins.
Or cation exchangers, which is the other way around.
To elute these, so, um, for example, for this example, the negative proteins - charged proteins are going to flow through, but to elute the positively charged proteins, you can either add an NaCl gradient, so these sodium and chloride ions are going to compete with the protein to bind on the bead. Or you can use a pH gradient because then the charge of your protein is going to vary based on the pH.
I'm going to talk about one more technique that you can use for chromatography which is size exclusion, also known as gel filtration.
And this is based on the size or the molecular weight of globular protein molecules.
So, um, important point of caution. This works well for globular protein molecules but not so well for proteins that are elongated or lodge-shaped. How this works is that there are porous polymer beads and the pore size can be regulated, like there are different pore sizes that are available. Um, and when you load a mixture of proteins, these proteins are separated by their size.
So, the largest proteins are not able to enter the - these orange spheres that are not able to enter these beads and they flow through and they are collected first of all. And then, the smaller proteins are able to actually enter these porous beads and they travel through the column at different speeds and they elute at different times.
So, you can actually separate out proteins based on their size and smallest proteins elute at the end.
So, um, basically, proteins elute as a function of buffer walling that has flowed through the column or time.
So, as you realize, we're using buffers for purification for all of these different systems.
So, how do you decide on what buffer to use? It's helpful to look at the pI of your protein and decide your pH based on that.
It should be separated enough so that if there's some change in the pH that your protein just doesn't crash out of solution.
And also, you can cite the pH based on the downstream applications because they may work for some assays and, you know, it might interfere with others.
The common ingredients for buffers usually include a buffering agent, of course, also sodium chloride, because it helps stabilize the proteins.
And some people like to add glycerol. And then most people, especially for intracellular proteins, add a reducing agent to the buffer, such as beta-mercaptoethanol, DDT, or TCEP. OK.
So, let's say, you know, things worked well for you. You'll get soluble protein and you're able to purify it and it aggregates or precipitates in solution which happened to me and it took me a long time to troubleshoot that. Sometimes, it can be simple.
Sometimes, maybe the protein concentration is too high, which happened for one of my proteins, that it was produced in bacteria.
It was a lot of protein and it was coming off in a sharp peak, so it just crashed out after purification.
And for me, if I diluted it, it worked. If the protein concentration was less than two mL/mL, it was usually fine.
You can also test increasing sodium chloride concentrations.
So, anywhere between a hundred and fifty what's - on an average in all buffers. So, anywhere between a hundred and fifty to five hundred mM NaCl is fine. You can - if that doesn't work, you can also test the effect of additives, such as glycerol or detergents that tend to stabilize these proteins.
But, you know, worst case scenario, maybe your protein requires nonspecific cofactors for proper refolding, and that's why it's crashing out in solution.
So, that summarizes the second part of the talk for purification techniques and does anyone have questions at all?
So, I want to talk about determining the protein concentration and then checking its purity and confirming the identity.
So, one of the ways to determine protein concentration is using densitometry.
So, any example is shown here. This is a Coomassie stained SDS-page gel.
So, you can run known concentrations of a standard protein and I used BSA where I have it from 0.5 mg/mL to 8 mg/mL. And you can run different volumes of your test protein that you've purified and you can see that - so essentially, you can create a standard curve from your standard protein and then you can look at the intensity of these bands and plot them on that standard plot and find out what your protein concentration is going to be. Another method to determine protein concentration is to measure the absorbance at 280.
For that you need an Extinction Coefficient which can be - um, you can calculate the theoretical Extinction Coefficient on Expasy, like it's the first thing that comes up when you Google how to do this.
Um, the formula that's used here is absorbance was equal to ECL. So, E is the Extinction coefficient. C is the concentration of your protein that you're trying to determine and L is the pathway.
So, for this, you can either use a whole style spectrophotometer or you can use a Nanodrop, and I've used both and it works. Of course, for the spectrophotometer, you're going to dilute off your protein because you need one mL.
And you measure the absorbance and it's helpful to test at least three to five samples and make sure that you get reproducible readings, and, you know, you can average those out then.
And then, there are also a bunch of colorimetric biochemical assays that are available. Essentially, for all of them, the concept is that you have a reagent that you add your protein to.
And then it shifts the absorbance. You measured the absorbance and you get a read out and - and, again, you need a standard curve like BSA for this as well.
This is the only one that's solely based on your protein, but the theoretical Extinction Coefficients are based on assuming that all the aromatic residues are equally exposed to the solvent and are absorbing.
So, I've tried all these methods and they're usually - the concentration that I determine is, you know, within the range of error and reproducible.
So, to confirm that the protein you have is pure, one of the things that you can again do is analyze the protein on a SDS page gel and stain with Coomassie.
So, if you are using this method and that works out, because here I can see that my protein is pure. There is no contaminating bands at all present.
Another thing you can do is, um, you know, immunoblot or ELISA analysis. So, you can probe with antibodies against either the tag, if you are using a tag, or the target protein and that can confirm that, you know, your protein, but it's the correct protein.
And you can also do mass spectrometry analysis. So, that essentially determines the molecular weight within an error range of one dalton.
So, what's the exact molecular weight, you know that, you know, it's your target protein.
And I think there's another How to Talk for Mass SPec analysis a few weeks later.
And, um, finally, you can also - once, you know, you confirm it. You have a concentration, you can do biochemical assays to assess your protein activity and function, which depend on what your protein is.
Thank you. Thanks for coming in for your attention and happy to take any questions. Does anyone have any questions?
[Colleague]: You said you could- to purify using column, I've looked into His-tag purified columns. We were thinking about trying to use those, you can put the soluble fraction through. They seem really easy to use.
[Prerna Grover]: I haven't used them like, we usually use the Explorer, but we have to purify a FLAG type protein, so I might be using that in the near future. I know people have done that and it works.
It essentially works like a immunoprecipitation, right, that you have a column. I think they work, but I never use them personally.
[Carrie Iwema]: Alright, well thank you very much.
Up Next

Histone Code Read by Reader-Writer Complexes | Chromatin Biology
@biotechnologyonline
10.6K views•2020-06-11

Algae Biofuels: Harnessing Microalgae for Renewable Energy
@LosAlamosNationalLab
623 views•2020-12-03

Microbial Degradation of Plastics: Biodegradation Pathways & Sustainability
@majeedhammad
2.9K views•2021-04-11

CRISPR and Genetic Engineering: How Gene Editing Works and Why It Matters
@kurzgesagt
30.5M views•2016-08-10
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biotechnology






















![[TALK 5] What is its Structure? Dom Bellini, Jane Wagstaff and Shaoxia Chen](https://i.ytimg.com/vi/tjasuJmkueM/maxresdefault.jpg)
















