Short Tandem Repeat (STR) loci are non-coding DNA regions containing variable numbers of 2-4 base pair repeats that differ among individuals, making them ideal for DNA identification. PCR amplifies these loci using specific primers, producing products of varying lengths based on repeat number. Capillary electrophoresis with fluorescently labeled primers and allelic ladders enables precise allele calling, allowing forensic identification, paternity testing, and criminal investigations through standardized loci like those in the FBI's CODIS system.
STR Loci and PCR Genotyping: DNA Identification Explained
Added:okay today we're going to talk about the str which is to say short tandem repeat loci in human dna and how pcr is used to genotype people using these loci for identification purposes now we sort of ended last time discussing pcr but if you think about it you do a pcr reaction using somebody's dna what you end up with is a bunch of dna in the test tube you need to analyze it in some way in order to learn anything to determine in this case genotypes but for whatever other purpose the pcr reaction might have had i want to remind you a bit about primers because that's going to become important shortly and i'm just showing a short stretch of i believe this is the hemoglobin c dna the first 61 to 60 nucleotides of it so to amplify anything we need two primers one the forward and one the reverse the forward primer will bind to the left region has this three prime end pointing to the right the reverse primer binds to the right and has this three prime end pointing to the left now remember when you're given cdna sequences you're only given one strand and we have to infer the existence of the other strand and we need to take that into consideration when we're figuring out what primers we're going to be using so here's the same dna sequence the first 60 nucleotides and i simply wrote the complement i just took the upper strand and just complemented it and then spaced it appropriately remember the units of groups of 10 base pairs don't mean anything they're just for our convenience and i can represent that as a double-stranded molecule running five prime to three prime left to right on the top and then three prime to five prime left to right on the bottom so now we have to figure out what the primers are going to look like to amplify the stretch of dna okay so conceptualizing this we need two primers one running with this both with the three prime hydroxyls pointing in towards the insert and what that means is that the forward primer will actually equal the upper strand sequence because it needs to bind to the lower strand sequence and that will be a little bit more clear in a second okay so the sequence that you're given what you see is what you get as far as your primer because that's the sequence that's going to match the lower strand and prime synthesis to the right okay now the upper strand primer the reverse primer is the opposite and what you think about you have to consider in that case the complement because what's going to succeed is our reverse primer is actually the complement of the sequence that you're given in your cdna sequence and the reason of course is you need to have your three prime oh going to the left and so that's primer strand has to be writing 5 prime to 3 prime right to left not left to right so what that works out to be is that primary sequence equals the lower strand sequence which is the complement of when what you're given so you're always going to have to remember this when you're creating primers for any kind of pcr reaction the invisible strand has to be taken into consideration okay so that's shown in this in this slide where we're actually using the sequence and writing the primers that we diagrammed in the previous one we can see that if we have a primer which is going to bind to the lower strand and use that as a template to create new upper strands that that primer sequence properly pointing in its through prime end towards the right is equal to the sequence that you're given in the cdna it binds to complementary strand which are not given in the sequence and will prime synthesis in the direction towards the region we want to copy okay the reverse primer is kind of the opposite because we don't want to use the actual sequence because that would have the wrong polarity it would be pointing away from what we're trying to copy instead we're going to use the complement of that which will properly have its 3 prime n pointing to the left and that complement of course matches the complementary sequence in the cdna so when considering primers you always have to consider the polarity and the fact that you have a strand that you're not being given and that you have to imagine or copy out in terms of designing your primer okay so in this example we are going to use pcr to do dna testing using these str loci so a little bit of background about str loci um they're also um called microsatellites sometimes but in the fbi sense using the word the term str is generally used so what it means is short tan and repeat okay what you'll see in some examples in a second what the short tandem repeat means is that these regions in the genome or loci consist of certain repeats of usually the ones that are used or four base pairs which occur over and over and over and over again and what makes them important or useful is that the number of repeats is variable among different alleles in the population so for example one person's locus might have 10 copies of the repeat another would have 12 another would have eight or whatever so people are very heterogeneous in that way and also there are many different alleles out there and many people are heterozygous for these low side now these are not coding regions these are usually located either between genes or within introns of coding genes and the fact that they're so variable in the population and there's so many of them they're very useful that they've been chosen as a way of identifying people basically based on their pcr so the fbi spent years and years studying these and they've chosen a group of these regions for common dna testing purposes and so these have been sequenced in thousands of individuals and their population frequencies are well understood and they've found very many important uses in terms of identification of people um paternity etc so these are short-term repeats they're non-coding they occur between genes so they're sort of junk dna because they don't have any actual function okay so importantly each locus is contains a tetranucleotide repeat region and which is variable in length between alleles and that what that means will be a little bit clearer when you see some examples i think so here's a representation of the commonly used str loci they're represented as being scattered all over the chromosomes right each little picture represents one of our chromosomes so you can see that they occur or they're the ones that are used occur on most of the chromosomes and they have sort of stupid non-informative names that you just have to get used to most of them are not gene names they're just designations for example we've got d3s1358 that refers to the short arm of chromosome 3 for example and the 1358 is just a numerical designation of that locus so they're not normally gene names they're just locus names you just kind of have to get used to that but the important thing is we all have these loci and the actual number of repeats present at each one differs from one allele to the other okay so the codis combined dna index system has been around for quite a long time and these are the original 13 that was chosen by the fbi for identification purposes now 13 is the original one now 20 or sometimes even 27 different loci are used for genotyping purposes now importantly these are very standardized that primers have already been chosen that are going to amplify the regions and most if not all people and those are standard and the reagents for the pcr are also purchased through suppliers that are approved so you don't have to design and purchase your own primers you just buy the kit okay also in this type of pcr you will you are going to need some way of observing the pcr product after it's formed and generally one of the primers is labeled with a fluorescent compound for each of the loci under under test so you just buy the stuff okay so the names are stupid we've got d3s one three five eight and it's shown on the chromosome diagram so these are mostly in non-transcribed regions and the repeat sequences are going to differ among the different loci okay so for example they all have a short repeat for example agat might be one particular tetranucleotide repeat and it's the number of repeats that defines that allele and they're usually discussed simply as the number of agat alleles as at that locus okay so here's an example this is str locus d8s1179 and it's showing the genbank entry for this particular locus and the entry tells us where the primers are going to be binding the 17 to 41 means the primers that are the forward primer that was chosen starts at nucleotide 17 in the sequence and extends to nucleotide 41 and the reverse primer is 169 to 193 and they're being nice to you they're telling you make sure you complement it because it's the complement of that sequence that you would need to make the primer and you can count that off in this sequence and see where it's located so that's just the genbank entry okay so here i've put in boldface the actual primer that's been used or the complementary sequence of the primer that's been used the forward and reverse fibers okay but immediately adjacent to the forward primer occurs the repeat region and that's what's variable from locus to locus well allele to allele all right so it's t atc t-a-t-c t-a-t-c t-a-tc etc extending for several um usually at least seven and often as many as twenty different repeats so the important thing here is that everybody has two copies of this locus because you've got one from mom and one from dad and everybody has some number of tatc's but it might be seven or it might be 20 depending on what particular allele that is so the one you get from mom may or ma may not be probably it's not the same length as the one you got from dad so when i say allele i'm talking about the number of repeats at that particular allele where 7 or 20 or 15 or whatever would be the way you designate that particular allele now there are some variants which are like plus or minus the nucleotides so we don't need to worry about those particularly right now so that's what it looks like and you can imagine if you do a pcr reaction on it what you're going to do is get a product that starts at that first nucleotide of the forward primer and extends to the last nucleotide of the reverse primer including the titc region which lies between them so the primers are set in unique regions which are going to be the same in every individual but the important thing is that what they're going to do is amplify a region between them which is going to enable us to determine what the allele actually is at that locus in that allele okay so how this works if you think about it and here's a very simple example where we're going to have a forward and reverse primer for this locus and those are going to be the same in everybody and then we have our repeat region which in this case is aatg so critically here the number of aatg differs among alleles so for example the one on the top we have eight copies of aatg there so that's the wheel number eight for this particular locus the other allele in this person has only seven copies of the repeat but they still have the same primer binding size so the important thing that's going to happen is if we do the pcr as shown the product of the top allele is going to be four nucleotides longer than the one from the bottom allele even though the primers are found in equivalent locations so the length of the product is going to vary with alleles because the number of repeats between the primers varies with the alleles so this is easy to see this one has eight nucleates eight copies of the rupee each repeat is four nucleotides so this product is going to be um one repeat more or four nucleotides longer than the other one okay so here's an analysis here of one that you can basically do some arithmetic on this to try to decide how long the prime how long the pcr product is going to be so we do this simply by looking at the numbers on the trace on the um sequence where the first nucleotide of the forward primer is 17 and the last nucleotide of the reverse primer is 193. right so subtract these to get the length of this particular allele and we get um well it's 177 because it's inclusive right so the final length of the product um is going to be 177 right because what we did was subtract those two numbers okay so in terms of the allele you've got to figure out how much of that is unique sequence versus how much of of it is the actual repeat sequence okay so you count the tatc's i get 13 maybe 12 because the other's a t-8tt um sometimes these things are a little bit ambiguous but we're going to go with that all right so thinking about this i'm going to consider not this allele but other alleles how long would the pcr product be if this was allele 15 instead of 13. so if this is allele 13 we'll say and it's 177 nucleotides long we know that allele 15 has two extra copies of the repeat so each repeat is four nucleotides so two extra copies means your eight nucleotides more so the length of that product for allele 15 would be my 177 add eight that gives me a length of 185 right right okay so when we analyze these by pcr importantly everybody has the same sequence as adjacent so that we can bind the same primers to everybody and the variation that we see is among the repeat region which is simply the number of repeats for that particular allele importantly pcr produces different sized products depending on the number of repeats present so everybody gets two alleles one for mom one from dad or you could be homozygous so when they you give these genotypes they don't go into the length of the pcr product they simply say what the locus is in this case is d8s1179 and the allele number is the between 7 and 20 because that's the range of alleles present in the population so it would just show d8s1179 allele 10 comma 11 if there are two of them so there's lots of different possible genotypes for the for this locus and then many more if you consider that we're using 13 at least or maybe 20 different loci together in order to do a genotyping on somebody okay here's an example a particular str locus has 10 repeats of four nucleotides it's amplified by a particular pair of primers to produce a pcr product that's 200 base pairs long another allele of this locus has 15 repeats of the same four base pair the length of the pcr product from this locus is therefore okay right the first one had 10 repeats of four nucleotides then one were asked about as 15 repeats of the four nucleotides so it has five more repeats of four nucleotides so it's 20 nucleotides longer than the first one so the answer is d right right okay so one consequence of the fact that we know what the population looks like that the shortest possible allele that we see has seven repeats and the longest has 20 right along with a few variant alleles that are there and these alleles will differ in the length of the pcr product by four base pairs each right and the sizes themselves and we'll get to this in a little while differ according to the placement of the primers right the repeat region varies but one has one's choice about where the primers can be placed and that will contribute to the determination of the total length of the pcr product which includes the repeat region and the part of the flanking dna that's included all right so now we have to analyze these right and it's going to be a little bit difficult because the products differ from each other only by four in four nucleotide increments basically so you can't really run it on a plain old agarose gel and expect any to see anything that you can see because the resolution of something like an agarose gel is not adequate but fortunately we do have capillary electrophoresis systems now which can be used to resolve single nucleotide differences for example as in sequencing so it's a simple matter to use these capillary systems to distinguish the products of our pcr in individuals in order to call what alleles they have so systems are set up so that we can easily read our genotype for each of these loci by comparison to the known alleles okay so the pcr products will be labeled with fluorescent compounds on one strand and the sizes will be determined as they migrate past the laser and detector which is represented here this represents what's going to happen in a capillary electrophoresis system although what's represented is actually what looks like a good old acrylamide gel the idea is completely the same though so what's happening here is we've loaded our sample on the top of the gel into the well turn on the power and we're going to let the fragments migrate but instead of stopping it at some point and staining the gel like we would for a regular old odd roast gel we allow the fragments to continue to migrate until they actually run off the bottom now so what happens during the migration is that the bands or fragments or whatever are migrating past a laser which is going to excite the fluorescence in whatever the fluorescent molecule is that's attached to the primer so the laser excites the fluorescent compound and what's being emitted is detected by a system which can detect and also quantitate which color is being sent as it goes by so what's happening is they're migrating so we're going to produce what's essentially a trace of like a sequencing type trace as as the band approaches the laser beam gets its peak and migrates past the laser beam so instead of seeing a bunch of bands on your gel what you're going to see is a pattern of peaks where each peak represents one allele of one particular locus that's labeled with a particular color that makes sense all right so calling genotypes we have two alleles each one from mom and one from dad they can be the same as each other or they can be different so the gel analysis has to be able to say what's what now the best way in which these are calibrated is to run the people samples along with what are called allelic ladders but sometimes just size standards to determine the size of those products so the software is extremely sophisticated and it does the allele calling for us okay so here's a simple way of looking at the allelic ladders for these str loci represented as if there were bands on the gel which they used to be they're not anymore okay so in this particular situation the allelic ladder is basically a um collection of all possible sequence lengths or fragment lengths that would occur in a population so in this instance the shortest one is contains five repeats and the longest one contains 11 repeats with the bands appearing at four base pair increments in between so the five is the smallest the next one up is six seven eight nine ten eleven et cetera so it's a simple matter to compare the dna pcr products from different individuals to these ladders in order to call their genotypes so for example the individual on the left they have two bands one from one from dad and their shorter band co-migrates with allele number seven they have a large another large band and that matches up with allele number nine so we call their gene attack to be seven nine the person on the right they have two bands one matches in migration or size to repeat number six and the other one matches number eight so we can call their genotypes based on that so that's what the allelic ladder is really it's a way of calibrating your capillary or gel in order to be able to determine the actual length of the pcr product and therefore determine what the actual genotype is of the person okay so normally these are not bands under gel they're sequencing traces because the sigma zig like traces so what this shows is a very complicated and kind of scary looking way um the allelic ladders for a whole bunch of these loci so for example on the top row we have one seven d8s1179 and there are twelve peaks there meaning there are twelve possible alleles in the population so any person would have at least one and at most two alleles that match one or two of those peaks and that would enable us to tell what their genotype is and the same for these other loci they have different numbers of alleles but there's one peak that would enable us to match the dna of any person to that particular locus okay so we'll get to the why are there so many rows here um later but the idea at initially at least is that these are colored different colors like the top row is blue the next one is green next one is yellow the next one is red okay so the idea here is that um we're going to label the different loci the products of different loci different colors and that's going to facilitate the analysis and that enables us to run them together all in one on one system more later okay so for example all we're going to do is call the alleles by comparing to an allelic ladder so we have with woman norma she has one band in the region that matches d3s one three five eight and we call that allele 15.
so she's actually homozygous for those two that that allele at that locus okay but along the top we have two other loci ones vwa one's fga and again norma has two alleles 14 and 16 for the vwa locus and two alleles 24 and 25 for the fga locus so her alleles are called by comparing the size of her products to this allelic ladder right so her genotype you just call it 15 15 14 16 24 and 25 all right so this gets complicated because they want to get the most information out of any um capillary run or any pcr run so what's actually done is many reactions are done simultaneously in the same tube so the primers for different loci can be labeled with different colors fluorescent compounds so even if they overlap each other on the sequencing trace they can be distinguished according to the color so the color of the peak tells us which locus it is the other thing that's done is that the primers are chosen so that the products from different loci end up with different size ranges because you can decide yourself ahead of time how big you want your product to be according to where you put your primers relative to the repeat region it sounds complicated and it sort of is right so the idea is you multiplex them and then the machine calls the alleles fortunately it's a good thing you can you can have one allele or two of each locus but you can't have more than that unless something weird is going on okay so here's just some examples of these um well alleles of different individuals at different low side and illustrates the fact that the top um set of loci those are primers products from those loci um those are all labeled blue so they can easily be just be distinguished from something that might overlap in length because those other ones are green or yellow as the case maybe right so the software fortunately calls and repeats the alleles at each locus separately to get the genotypes i think i think i just said this so here's again traces with allelic ladders for a lot of these and again they're labeling different colors and so for example the um green ones and the blue ones right this csf1po products overlap in length with the d2s1338 products for example but the machine can tell who's who because one set is blue and the other set is green so you can actually do all these simultaneously with the machine being able to tell the difference partly because of the size but also because of the color to get the complete genotype out of a minimum number of runs to save money so i think i said this already okay so here's some different commercially available primers right people don't make their own they just they buy them and they do differ according to what particular um purveyor of primers might have so you can adjust the length of the product especially if they're going to be run together according to where the primers are put that's the point okay so for example there's three different primer sets here and they give different sizes according to what the primer's location actually is so for and there's more of this table i chopped off at the bottom so for example for um allele a a little seven for example one particular primer set gives you a product of 157 a different primer set gives you 123.
yet another primer set gives you 203 now the repeat region that's being amplified is the same it's just a matter of the position of the forward and reverse primers that will then be used to amplify that right okay so the gruesome reality here is this is what it really looks like when they do it um and these are all runs that will not always run simultaneously all the time but they're they're set up so that they have different colors and also come to different sizes and therefore one single run can resolve all of these loci all together at once right so there's four different colors 27 multiflex one run and this somehow able the instruments in the software are able to actually resolve all these and call these alleles now this does take it's not all by machine we need a forensics expert to actually look at these and make sure that everything's actually okay with these okay so you're not going to have to analyze anything like this but it illustrates some of the points okay so here's a case of a rape case where there were dna profiles of the victim and two suspects okay so what's shown in the top the sample one is the victim and then we're genotyping for all these different loci's together all loci together and the numbers so for example one for d3 s1358 she has allele 16 and 17. now in parenthesis is the um crime scene dna which does not match the victim but if it's it will match one of the suspects so that those extra low sides that she has are basically rape evidence and that what they're going to do is match the rape evidence to the different suspects and you can see that on suspect number two his dna matches the rape evidence and absolutely all of these loci so this can be used to exclude or not exclude particular individuals as criminals in that case now here's paternity and we're not going to analyze this now we're going to save that for genetics but the same idea is true that you can tell if somebody's the father what alleles they will have that will match the child and so it's you don't need to analyze this but it's another very very common use of this particular technology okay there's something new called the anti-rapid dna instrument where amazingly this is done automatically with no human intervention at all the usual situation um they can generate dna profiles within two hours with no human hands-on time beyond placing the samples in the instrument and these are a little bit encouraging in the sense that you can actually get somebody's genotype when they're still in custody and compare their genotype to databases for criminals it's also a little scary that you can genotype anybody that happens to get arrested it's kind of scary and whether the information was going to be used properly okay so without human intervention there are sample tubes and reagents inside the machine basically they take a dna sample from somebody's saliva put it in the machine and the machine spits out a profile right this is in contrast to when you send dna for traditional dna analysis it can only be done by certified labs and it can be a relatively time consuming process now in both cases you have to get a dna sample you have to pure um purify it somewhat then you have to do your amplification with all your primer sets and then you have to run it in the instrument the capillary type system and then evaluate the profiles that you get so it's a complicated series of events that can be done apparently automatically in these machines okay this is what it looks like so you just stick the samples in the machine and it spits out a trace and anybody that's been arrested for a crime because they can ask for a dna sample if they want to does it automatically
Up Next

How to Read an Electropherogram for DNA Fingerprinting
@daniellemyers1706
932 views•2019-05-14

Circadian Metabolomics: Sleep, Food Timing & Human Clocks
@tscnlab
359 views•2022-11-10

Enteric Nervous System Explained: The Gut's Brain | Neurobiology Lecture
@alumniu6029
438 views•2018-09-12

Bacteriophages: Earth's Deadliest Killers and Future Antibiotics
@kurzgesagt
34.6M views•2018-05-13
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biology







































